Tillbaka till svenska Fidonet
English   Information   Debug  
FILM   0/18
FNEWS_PUBLISH   5090
FN_SYSOP   42250
FN_SYSOP_OLD1   71952
FTP_FIDO   0/2
FTSC_PUBLIC   0/13992
FUNNY   0/4886
GENEALOGY.EUR   0/71
GET_INFO   105
GOLDED   0/408
HAM   16823/16961
HOLYSMOKE   0/6791
HOT_SITES   0/1
HTMLEDIT   0/71
HUB203   466
HUB_100   264
HUB_400   39
HUMOR   0/29
IC   0/2851
INTERNET   0/424
INTERUSER   0/3
IP_CONNECT   719
JAMNNTPD   0/233
JAMTLAND   0/47
KATTY_KORNER   0/41
LAN   0/16
LINUX-USER   0/19
LINUXHELP   0/1155
LINUX   0/22471
LINUX_BBS   369/957
mail   18.68
mail_fore_ok   249
MENSA   0/341
MODERATOR   0/102
MONTE   0/992
MOSCOW_OKLAHOMA   0/1245
MUFFIN   0/783
MUSIC   0/321
N203_STAT   940
N203_SYSCHAT   313
NET203   321
NET204   69
NET_DEV   0/10
NORD.ADMIN   0/101
NORD.CHAT   0/2572
NORD.FIDONET   189
NORD.HARDWARE   0/28
NORD.KULTUR   0/114
NORD.PROG   0/32
NORD.SOFTWARE   0/88
NORD.TEKNIK   0/58
NORD   0/453
OCCULT_CHAT   0/93
OS2BBS   0/787
OS2DOSBBS   0/580
OS2HW   0/42
OS2INET   0/37
OS2LAN   0/134
OS2PROG   0/36
OS2REXX   0/113
OS2USER-L   207
OS2   0/4837
OSDEBATE   0/18996
PASCAL   0/490
PERL   0/457
PHP   0/45
POINTS   0/405
POLITICS   2178/29554
POL_INC   0/14731
PSION   103
R20_ADMIN   1134
R20_AMATORRADIO   0/2
R20_BEST_OF_FIDONET   17
R20_CHAT   0/899
R20_DEPP   0/3
R20_DEV   400
R20_ECHO2   1989
R20_ECHOPRES   0/35
R20_ESTAT   0/719
R20_FIDONETPROG...
...RAM.MYPOINT
  0/2
R20_FIDONETPROGRAM   0/22
R20_FIDONET   0/248
R20_FILEFIND   0/24
R20_FILEFOUND   0/22
R20_HIFI   0/3
R20_INFO2   3882
R20_INTERNET   0/12940
R20_INTRESSE   0/60
R20_INTR_KOM   0/99
R20_KANDIDAT.CHAT   42
R20_KANDIDAT   28
R20_KOM_DEV   112
R20_KONTROLL   0/13377
R20_KORSET   0/18
R20_LOKALTRAFIK   0/24
R20_MODERATOR   0/1852
R20_NC   76
R20_NET200   245
R20_NETWORK.OTH...
...ERNETS
  0/13
R20_OPERATIVSYS...
...TEM.LINUX
  0/45
R20_PROGRAMVAROR   0/1
R20_REC2NEC   534
R20_SFOSM   0/341
R20_SF   0/108
R20_SPRAK.ENGLISH   0/1
R20_SQUISH   107
R20_TEST   2
R20_WORST_OF_FIDONET   21
RAR   0/9
RA_MULTI   106
RA_UTIL   0/162
REGCON.EUR   0/2072
REGCON   0/13
SCIENCE   1/1206
SF   0/239
SHAREWARE_SUPPORT   0/5146
SHAREWRE   0/14
SIMPSONS   0/169
STATS_OLD1   0/2539.065
STATS_OLD2   0/2530
STATS_OLD3   0/2395.095
STATS_OLD4   0/1692.25
SURVIVOR   0/495
SYSOPS_CORNER   0/3
SYSOP   0/84
TAGLINES   0/112
TEAMOS2   0/4530
TECH   0/2617
TEST.444   0/105
TRAPDOOR   0/19
TREK   0/755
TUB   0/290
UFO   0/40
UNIX   1237/1316
USA_EURLINK   0/102
USR_MODEMS   0/1
VATICAN   0/2740
VIETNAM_VETS   0/14
VIRUS   0/378
VIRUS_INFO   0/201
VISUAL_BASIC   0/473
WHITEHOUSE   0/5187
WIN2000   0/101
WIN32   0/30
WIN95   0/4291
WIN95_OLD1   0/70272
WINDOWS   0/1517
WWB_SYSOP   0/419
WWB_TECH   0/810
ZCC-PUBLIC   0/1
ZEC   4

 
4DOS   0/134
ABORTION   0/7
ALASKA_CHAT   0/506
ALLFIX_FILE   0/1313
ALLFIX_FILE_OLD1   0/7997
ALT_DOS   0/152
AMATEUR_RADIO   0/1039
AMIGASALE   0/14
AMIGA   0/331
AMIGA_INT   0/1
AMIGA_PROG   0/20
AMIGA_SYSOP   0/26
ANIME   0/15
ARGUS   0/924
ASCII_ART   0/340
ASIAN_LINK   0/651
ASTRONOMY   0/417
AUDIO   0/92
AUTOMOBILE_RACING   0/105
BABYLON5   0/17862
BAG   135
BATPOWER   0/361
BBBS.ENGLISH   0/382
BBSLAW   0/109
BBS_ADS   0/5290
BBS_INTERNET   0/507
BIBLE   0/3563
BINKD   0/1119
BINKLEY   0/215
BLUEWAVE   0/2173
CABLE_MODEMS   0/25
CBM   0/46
CDRECORD   0/66
CDROM   0/20
CLASSIC_COMPUTER   0/378
COMICS   0/15
CONSPRCY   0/899
COOKING   42825
COOKING_OLD1   0/24719
COOKING_OLD2   23948/40862
COOKING_OLD3   6830/37489
COOKING_OLD4   0/35496
COOKING_OLD5   9370
C_ECHO   0/189
C_PLUSPLUS   0/31
DIRTY_DOZEN   0/201
DOORGAMES   2139/2396
DOS_INTERNET   0/196
duplikat   6103
ECHOLIST   0/18295
EC_SUPPORT   0/318
ELECTRONICS   0/359
ELEKTRONIK.GER   1534
ENET.LINGUISTIC   0/13
ENET.POLITICS   0/4
ENET.SOFT   0/11701
ENET.SYSOP   34568
ENET.TALKS   0/32
ENGLISH_TUTOR   0/2000
EVOLUTION   0/1335
FDECHO   0/217
FDN_ANNOUNCE   0/7068
FIDONEWS   25343
FIDONEWS_OLD1   0/49742
FIDONEWS_OLD2   0/35949
FIDONEWS_OLD3   0/30874
FIDONEWS_OLD4   36498/37224
FIDO_SYSOP   12979
FIDO_UTIL   0/180
FILEFIND   0/209
FILEGATE   0/212
Möte HAM, 16961 texter
 lista första sista föregående nästa
Text 16818, 236 rader
Skriven 2026-08-17 22:09:56 av Sean Dennis (1:18/200)
  Kommentar till text 16815 av Mortar M. (4157.fido_ham)
Ärende: The ARRL Letter
=======================
Hello Mortar,

17 Aug 26 09:40, you wrote to me:

 MM> Is this OK with ARRL?  Site owners can get rather touchy about
 MM> scraping.

From the bottom of each ARRL Letter:

"Copyright (C) 2026 American Radio Relay League, Incorporated. Use and 
distribution of this publication, or any portion thereof, is permitted for 
non-commercial or educational purposes, with attribution. All other purposes 
require written permission."

I make sure that's left in every "letter" I scrape.

 MM> Can you post the script?  I'm interested in how that's done.

=== Cut ===
#!/usr/bin/env python3

### ARRL Letter Posting Script
###
### By Sean Dennis with assistance
### from Microsoft Copilot
###
### (C) Sean Dennis KS4TD
###
### Released under the MIT License.

import requests
import re
import textwrap
import unicodedata
import os
import sys
import time

BASE = "https://www.arrl.org"
LETTER_LIST = "https://www.arrl.org/arrlletter"
CACHE_FILE = os.path.expanduser("~/.arrlletter_cache")

# ------------------------------------------------------------
# Retry-capable fetcher with browser headers
# ------------------------------------------------------------
def fetch_with_retry(url, retries=5, delay=3):
    headers = {
        "User-Agent": (
            "Mozilla/5.0 (X11; Linux x86_64) "
            "AppleWebKit/537.36 (KHTML, like Gecko) "
            "Chrome/124.0 Safari/537.36"
        )
    }

    for attempt in range(1, retries + 1):
        try:
            print("Fetching {} (attempt {}/{})".format(url, attempt, retries))
            r = requests.get(url, headers=headers, timeout=20)
            r.raise_for_status()
            return r.text
        except Exception as e:
            print("Attempt {} failed: {}".format(attempt, e))
            time.sleep(delay)

    print("ERROR: All attempts failed.")
    return None

# ------------------------------------------------------------
# Utility functions
# ------------------------------------------------------------
def log(msg):
    print(msg)

def to_ascii(s):
    return unicodedata.normalize("NFKD", s).encode("ascii", 
"ignore").decode("ascii")

def load_cached_issue():
    if not os.path.exists(CACHE_FILE):
        return None
    try:
        with open(CACHE_FILE, "r") as f:
            return f.read().strip()
    except:
        return None

def save_cached_issue(issue):
    try:
        with open(CACHE_FILE, "w") as f:
            f.write(issue)
    except Exception as e:
        log("Warning: could not save cache: {}".format(e))

# ------------------------------------------------------------
# 1. Fetch ARRL Letter index page (with retry)
# ------------------------------------------------------------
index_html = fetch_with_retry(LETTER_LIST)
if index_html is None:
    sys.exit(1)

# ------------------------------------------------------------
# 2. Find newest issue link (single or double quotes)
# ------------------------------------------------------------
issues = 
re.findall(r"href=['\"](/arrlletterissue\?issue=\d{4}-\d{2}-\d{2})['\"]", 
index_html)
if not issues:
    log("ERROR: No ARRL Letter issues found.")
    sys.exit(1)

latest_issue_path = issues[0]
latest_issue_url = BASE + latest_issue_path
latest_issue_id = latest_issue_path.split("=")[-1]

log("Latest ARRL Letter issue: {}".format(latest_issue_id))

# ------------------------------------------------------------
# 3. Check cache
# ------------------------------------------------------------
cached = load_cached_issue()
if cached == latest_issue_id:
    log("Cached issue matches latest. Nothing new to download.")
    sys.exit(0)

# ------------------------------------------------------------
# 4. Download the issue HTML (with retry)
# ------------------------------------------------------------
page_html = fetch_with_retry(latest_issue_url)
if page_html is None:
    sys.exit(1)

# ------------------------------------------------------------
# 5. Improved HTML cleanup
# ------------------------------------------------------------
import html

# Remove scripts and styles
clean = re.sub(r"<script.*?>.*?</script>", "", page_html, flags=re.DOTALL)
clean = re.sub(r"<style.*?>.*?</style>", "", clean, flags=re.DOTALL)

# Remove all HTML tags
clean = re.sub(r"<[^>]+>", "", clean)

# Decode HTML entities (&nbsp;, &#39;, etc.)
clean = html.unescape(clean)

# Remove "undefined" junk lines
clean = clean.replace("undefined", "")

# Collapse multiple spaces
clean = re.sub(r"[ \t]+", " ", clean)

# Remove repeated blank lines
clean = re.sub(r"\n\s*\n\s*\n+", "\n\n", clean)

# Remove repeated photo captions (ARRL duplicates them)
clean = re.sub(r" \[Photo.*?\] ", "", clean)

# Remove ARRL navigation/header lines
clean = clean.replace("ARRL Home Page", "")
clean = clean.replace("ARRL Audio News", "")
clean = clean.replace("ARRL Letter Archive", "")

# Remove unsubscribe footer
clean = clean.replace("Unsubscribe from this list.", "")

# Clean up leftover blank lines
clean = re.sub(r"\n\s*\n+", "\n\n", clean)

# Normalize line endings
clean = clean.strip()

# Normalize line endings
clean = clean.strip()

# ------------------------------------------------------------
# 6. Convert to ASCII
# ------------------------------------------------------------
ascii_text = to_ascii(clean)

1# ------------------------------------------------------------
# 7. Normalize whitespace and wrap at 72 columns
# ------------------------------------------------------------
out_lines = []
for line in ascii_text.splitlines():
    line = line.strip()
    if not line:
        out_lines.append("")
        continue
    wrapped = textwrap.wrap(line, width=72)
    out_lines.extend(wrapped)

# ------------------------------------------------------------
# 8. Write final file
# ------------------------------------------------------------
output_file = "/tmp/arrlletter-latest.asc"
try:
    with open(output_file, "w") as f:
        f.write("\n".join(out_lines))
    log("Created {}".format(output_file))
    save_cached_issue(latest_issue_id)
except Exception as e:
    log("ERROR: Could not write output file: {}".format(e))
    sys.exit(1)
# ------------------------------------------------------------
=== Cut ===

Here's the actual script called by MBSE (note that the "mbmsg" lines are 
wrapped):

=== Cut ===
#!/bin/bash

# Environment
# PATH needed since cron doesn't inherit bash's environment
export PATH=/usr/local/sbin:/usr/local/bin:/sbin:/bin:/usr/sbin:/usr/bin
# Required by MBSE
export MBSE_ROOT=/opt/mbse

$MBSE_ROOT/bin/arrl_letter.py

# Post in MIN_HAM
mbmsg post "ARS KS4TD" "All" 61 "The ARRL Letter" "/tmp/arrlletter-latest.txt" 

# Post in Fidonet's HAM
mbmsg post "ARS KS4TD" "All" 117 "The ARRL Letter" 
"/tmp/arrlletter-latest.txt" -
=== Cut ===

I believe the Python script will run under Windows with minor modifications.

-- Sean

... "I have a love interest in every one of my films: a gun." - Schwarzenegger
--- GoldED+/LNX 1.1.5-b20260304
 * Origin: Outpost BBS * Johnson City, TN (1:18/200)