Skip to content

Blog

I miei computer: cronologia hardware e audio (1988–2012)

Ho sistemato dei vecchi appunti, ed ho raccolto la successione delle macchine che mi hanno accompagnato nel tempo, tra floppy disk, megahertz ed audio digitale, dalle prime Sound Blaster fino alle interfacce da studio.

1988 — IBM PS/2 Modello 30

  • CPU: Intel 8086 (8 MHz)
  • RAM: 640 KB
  • Disco: Hard Disk 20 MB
  • Supporti: Floppy drive 3.5" da 720 KB (DD)
  • Grafica: MCGA (Multi-Color Graphics Array: il precursore del VGA, capace di 320x200 a 256 colori)
  • Monitor: IBM monocromatico da 12" (B/N)
  • Audio: PC Speaker interno (cicalino a 1 bit)
  • Software: MS-DOS 3.30, Windows 2.11 e Word 1.0

1992 — Il clone AMD 386/40

  • CPU: AMD 386DX 40 MHz
  • RAM: 4 MB
  • Disco: HD 120 MB
  • Supporti: Floppy 1.44 MB
  • Grafica: OAK VGA
  • Software: DOS 5.0, Windows 3.1
  • Audio: DAC R-2R artigianale su porta parallela (il leggendario "Covox" autocostruito con le resistenze: 8 bit mono per far suonare i MOD a 4 canali con ModPlay, ModEdit + ModRes ed infine con FastTracker II)

1996 — L'era Pentium e l'audio a 16-bit

  • CPU: Intel Pentium 150 MHz
  • RAM: 32 MB, portata a 64 MB a fine 1998 (upgrade da Jolly Computer, per gestire i banchi SoundFont con Vienna e le tracce audio su Cakewalk Pro Audio senza dropout)
  • Ottica: CD-ROM Asus 40x (aprile 1999, da Silicon Valley a Padova)
  • OS: Windows 95 OSR2
  • Audio: Creative Sound Blaster AWE64
  • Software Audio & MIDI:
  • Sequencing & DAW: Cakewalk Pro Audio e il bundle Creative (Voyetra MIDI Orchestrator Plus)
  • SoundFont: Creative Vienna SoundFont Studio
  • Tracker: FastTracker II in ambiente DOS

2000 — Il cambio di millennio con AMD

  • Data acquisto: 13 dicembre 2000
  • CPU: AMD Athlon "Thunderbird" 800 MHz (Socket A)
  • RAM: 128 MB SDRAM PC133 (o 256 MB, da verificare)
  • Monitor: CRT 17"
  • Audio: Creative Sound Blaster Live! (chip EMU10K1, l'era dell'EAX e del riverbero ambientale hardware)
  • OS: Windows 98 SE / Windows 2000
  • Note: Insieme al PC arrivò la scatola di Fuga da Monkey Island (Monkey Island 4), fresco di uscita a novembre 2000: il primo capitolo in 3D, che sul vecchio Pentium 150 sarebbe stato impensabile far girare!

2004 — Il salto ai 64 bit

  • Data acquisto: 12 novembre 2004
  • CPU: AMD Athlon 64 3200+ (Socket 754, core ClawHammer)
  • RAM: 1 GB
  • Monitor: CRT 17" riciclato, successivamente un LCD
  • Disco: HD 160 GB
  • Audio: Creative Sound Blaster Audigy2 Platinum (con il pannello frontale da 5.25" pieno di ingressi jack, ottici e MIDI)

2012 — Dell XPS 8300 e l'home studio USB

  • CPU: Intel Core i7-2600 (Sandy Bridge, 3.40 GHz, 8 MB cache)
  • RAM: 8 GB DDR3 (4x2 GB) a 1333 MHz
  • Disco: HD 1 TB SATA 7200 rpm, successivamente SSD512GB
  • Monitor: LCD 17" riciclato, successivamente LCD 27"
  • Grafica: NVIDIA GeForce GT 530 (1 GB)
  • OS: Windows 7 Home Premium SP1 (64 bit)
  • Audio: Roland Quad-Capture (il salto di qualità: stop alle schede audio PCI interne, ma senza gestione SF2 nativa)

MODRES: reverse engineering di un TSR audio negli anni '90

MODRES era la libreria audio di MODEDIT, un tracker per MOD molto diffuso in ambiente DOS. Non aveva documentazione pubblica, l'ho quindi ricostruita a suo tempo con il Turbo Debugger ed inviata a Ralf Brown, ed oggi nel 2026 tali voci sono ancora presenti nella Interrupt List.

Cos'è MODRES

MODRES era un programma per DOS di tipo TSR (Terminate and Stay Resident) che permetteva di riprodurre in background i moduli musicali in formato MOD quattro tracce dei tracker Amiga.

Il suo uso principale era all'interno di MODEDIT: MODRES era la libreria audio che permetteva al tracker di ascoltare i moduli ed i campioni mentre li si editava, e di continuare a suonarli anche quando il tracker era in background.

Supportava:

  • PC speaker
  • D/A converter su porte parallele LPT1-LPT4
  • Sound Blaster (porta 02x0h)
  • Disney Sound Source
  • Configurazioni stereo su due porte parallele
  • Stereo-on-1

Esponeva le sue funzioni attraverso l'interrupt INT 2F, con AX=8220h come base.

Il reverse engineering

MODRES non aveva documentazione pubblica. Tramite il Turbo Debugger della Borland ne ho seguito l'esecuzione istruzione per istruzione, mettendo breakpoint ed esaminando registri e memoria. Con molta pazienza, ho ricostruito le sue logiche interne di funzionamento.

Le funzioni che avevo documentato sono sette. Le strutture dati due: MODPARM e SAMPARM, più una tabella di output device e una dei pitch.

La documentazione

MODRES - PLAY MODULE

AX = 8220h
DX:CX -> MODPARM structure (see #2646)

Return:
AX = status
  5722h successful
  2000h parameters out of range
  other MODRES not installed

See Also: AX=8221h - AX=8223h - AX=8225h - AX=8227h - AX=8200h"RESPLAY"

Format of MODPARM Structure (Table 2646)

Offset  Size    Description
00h     WORD    signature 504Dh ("MP" = Modparm)
02h     BYTE    output device (see #2648 at INT 2F/AX=8221h)
03h     WORD    segment of start of main module (pattern) data
05h  31 WORDs   segment of start of sample numbers 1-31
43h     BYTE    pattern at which to start playing (00h to 7Fh)
44h     BYTE    function
                00h play from pattern [offset 43h] until end of the song
                01h play indicated pattern [offset 43h] only
45h     BYTE    Machine speed
                00h 10-12Mhz
                01h 12-25Mhz (default)
                02h 25Mhz+
                03h mix speed 10kHz (fast 8Mhz machines)
                04h mix speed 12kHz (10Mhz machines)
                05h mix speed 13kHz
                06h mix speed 8kHz (test for 8Mhz machines)
46h     BYTE    allow >64k sample playing
                80h MOD has samples >64k in it
                else all samples in MOD are <64k

Notes: Main module data and all samples must start on segment
boundaries. In version 2.00 (ONLY) this function carries on
playing (works in the background).

See Also: #2647

MODRES - INSTALLATION CHECK

AX = 8221h

Return:
AX = status
  5722h successful
  other MODRES not installed
BX = BCD version number (BH = major, BL = minor)
DX:CX -> Output Device structure (read-only) (see #2647)

See Also: AX=8220h - AX=8222h - AX=8225h - AX=8227h

Format of Output Device structure [array] (Table 2647)

Offset  Size    Description
00h 20 BYTEs   ASCIZ name of the output device
               (end of list if first char is FFh)
14h    WORD    apparently always FFFFh
16h    WORD    0000h if output device not available
               else first I/O port for the output device
18h    WORD    second I/O port for the output device (for example
               if it is stereo)
               0000h if only one port used or device is not available
1Ah  7 BYTEs   ???

See Also: #2646 - #2648

Values for MODRES v1.52 output device index (Table 2648)

00h    PC speaker
01h    D/A Converter on LPT1
02h    D/A Converter on LPT2
03h    D/A Converter on LPT3
04h    D/A Converter on LPT4
05h    D/A Converter on LPT1&LPT2 (stereo)
06h    D/A Converter on LPT1&LPT2 (mono)
07h    Sound Blaster (port 02x0h)
08h    User Defined D/A (mono)
09h    User Defined D/A (stereo)
0Ah    Stereo-on-1
0Bh    Disney SS su LPT1
0Ch    Disney SS su LPT2
0Dh    Disney SS su LPT3
0Eh    Disney SS su LPT4

Note: This list may vary between versions of MODRES

MODRES - UNINSTALL

AX = 8222h

Return:
AX = code segment of the program

Note: This function does not release the TSRs memory; the caller
must do so

See Also: AX=8220h - AX=8221h - AX=8223h

MODRES - PLAY SAMPLE

AX = 8223h
DX:CX -> SAMPARM structure (see #2649)

Return:
AX = status
  5722h successful
  2000h parameters out of range
  other MODRES not installed

See Also: AX=8221h - AX=8224h - AX=8225h - AX=8226h

Format of SAMPARM Structure (Table 2649)

Offset  Size    Description
00h     WORD    signature 5053h ("SP" = SAMPARM)
02h     WORD    segment of start of sample to play
04h     WORD    length of sample (IN WORD)
06h     BYTE    output device (see #2648 at INT 2F/AX=8221h)
07h     WORD    pitch to play (see #2650)
09h     BYTE    volume (from 00h to 40h)
0Ah     WORD    loop start
0Ch     WORD    loop length
0Eh     BYTE    machine speed (see INT 2F/AX=8220h)

See Also: #2646

Values for Pitch to play (Table 2650)

C 0 is 06B0h
C#0 is 06B0h / 2^(1/12)
D 0 is (06B0h / 2^(1/12)) / 2^(1/12)
...

Note: C 1 is 06B0h / 2. C 2 is 06B0h / 4. Etc.

See Also: #2649

MODRES - ???

AX = 8224h
DX:CX -> ???

Return:
???

See Also: AX=8221h - AX=8223h - AX=8224h

MODRES v2.00+ - GET LOCATION IN MOD

AX = 8225h

Return:
AL = status
  00h playing
  01h reached end or stopped
AH = speed of MOD
BX = position within pattern 0000h-0400h
CL = position within the song (track number)

See Also: AX=8220h - AX=8221h - AX=8223h - AX=8226h

MODRES v2.00+ - STOP PLAYING

AX = 8226h

Return:
AX = status
  5722h successful
  other MODRES not installed

Desc: Stops playing the MOD file before performing critical
operations such as disk accesses

See Also: AX=8220h - AX=8221h - AX=8223h - AX=8225h - AX=8227h

MODRES - CONFIGURE

AX = 8227h
BX = function
  0001h set default playing speed (06h)
  0002h select output device
    CL = output device (see #2648 at INT 2F/AX=8221h)

Return:
AX = status
  5722h successful
  2000h parameters out of range
  other MODRES not installed

Note: Function 0001h should be called every time a new module
is loaded

See Also: AX=8220h - AX=8221h - AX=8222h - AX=8223h

Note sulla documentazione

Alcune cose che vale la pena notare, rileggendola oggi:

  • La MODPARM inizia con una signature (504Dh = "MP"), come quasi tutte le strutture dati di quei tempi. Serviva a verificare che il puntatore passato alla funzione fosse davvero una struttura valida.
  • Il campo machine speed non è un valore in MHz, ma un indice che il TSR usa per scegliere la frequenza di mixaggio. La voce 06h è "test for 8MHz machines": il programma si adattava alla macchina su cui girava.
  • La tabella dei pitch parte da C 0 = 06B0h e calcola le note successive come divisioni per 2^(1/12). È il temperamento equabile applicato ai registri del timer. Quasi certamente però il programma utilizzava una tabella precompilata per evitare di calcolare i valori a runtime, che però non sono riuscito a recuperare.
  • Non sono riuscito a capire il significato della funzione AX=8224h che è rimasta non documentata.

Dove è finita

La documentazione è nella Ralf Brown's Interrupt List, alla voce INT 2F/AX=8220h e seguenti. È ancora consultabile su ctyme.com/rbrown.htm.

Il riconoscimento

Nella sezione CREDITS della Release 55 (28 settembre 1997) della Ralf Brown's Interrupt List, c'è questa riga:

12/96 A Federico Thiella fthiella@stud32.math.unipd.it  MODRES

Dicembre 1996. Non ho più quella casella di posta ovviamente!

filtersql v1.2.6: il bug banale che ha rivelato un abisso architetturale

Quando ho pubblicato la versione 1.2.6 di filtersql ero convinto di aver raggiunto un buon livello di sicurezza. Poi, puntuale come un ordigno svizzero, è saltato fuori un nuovo bug.

Era la classica svista minuscola, la riga di codice che guardi e pensi di risolvere in pochi minuti.

Ma quando ho messo le mani nella correzione mi sono reso conto che la patch banale risolveva in effetti il bug, ma non il problema che si è rivelato essere più grande.

Il bug non era nella logica: era architetturale.


Premessa 1: La giungla dei nomi di colonna nei database reali

Vero che dipende dal DBMS, dalla versione e dalla configurazione, ma in linea generale i database consentono di definire i nomi delle colonne con grande ed eroica anarchia. Una colonna può contenere spazi, accenti, caratteri speciali e simboli Unicode. Si può addirittura quotare il carattere di quoting.

Ed in produzione in effetti non c'è limite all'ineleganza: da colonne come "qtà totale" o "n° ordine" che pur non essendo una idea grandiosa sono tutto sommato legittime, fino a casi limite ma inaspettatamente validi come "totale " (con due spazi in coda, o magari con tab, newlines, ecc.).

Per garantire una quotazione avanzata e sicura senza impazzire, ho scelto di applicare una semplificazione: rimuovere eventuali spazi bianchi accidentali agli estremi (trimming). La colonna "totale " veniva ripulita in "totale". Questo avrebbe ridotto la copertura dei casi più assurdi, ma i casi più dignitosi sarebbero rimasti salvi.


Premessa 2: L'incubo dei conflitti di scope (Multi-Tenancy)

filtersql gestisce il concetto di scope: una mappa di condizioni fisse (come tenant_id = 42) applicata lato server a tutte le operazioni (select, insert, update, delete). Non sostituisce davvero i permessi del database, ma offre comunque una sua utilità nel blindare le applicazioni web multi-tenant.

Ma cosa succede se un client invia un payload che prova a modificare proprio la colonna riservata allo scope?

A parte la rottura di un equilibrio estetico, nelle letture e nelle cancellazioni non si pone davvero un problema, dato che le condizioni finiscono tutte in AND e al massimo non esce e non si cancella nulla, ma in scrittura la faccenda si fa interessante:

Caso insert:

insert into users (tenant_id, tenant_id, name) values (400, 500, 'Mario');
A seconda del DBMS, questa istruzione può assegnare il primo valore (400), prendere l'ultimo (500), comportarsi in modo imprevedibile o fallire miseramente.

Caso update (molto peggio):

update users set tenant_id = 500 where tenant_id = 400;
Questa singola riga di codice trasferisce uno o più record dal proprio tenant a un altro. Ed a parte perdere i propri dati che spariscono dal controllo, un utente malevolo potrebbe sfruttare questo comportamento per spostare un account admin dal proprio tenant al tenant di un nemico per prenderne il controllo.

La soluzione logica: il problema sembrava risolvibile con banale controllo di collisione che ho implementato facilmente.

Se una colonna appartiene allo scope, deve essere vietata nei campi di insert o update.

Facile come bere una bottiglia d'acqua.


Il Bug della v1.2.6

Rileggi le due premesse, inserisci un minuscolo errore logico, ed ecco che l'ingranaggio salta.

Se lo scope protegge la colonna "tenant_id" e un client invia in scrittura la chiave "tenant_id " (con uno spazio in fondo), cosa succede?

  1. purtroppo il controllo dei conflitti confrontava immediatamente le due stringhe grezze: "tenant_id" contro "tenant_id ";
  2. le stringhe risultavano diverse ed il controllo dava il via libera;
  3. solo successivamente, la funzione di quoting sanificava l'identificatore applicando il trim().
  4. risultato: la query veniva compilata ed eseguita modificando proprio tenant_id!

Risolvere questo specifico passaggio non era poi complicato, bastava anticipare il trim() al momento giusto, prima del controllo di collisione. Il bug sarebbe stato risolto senza tanta fatica. Ma è stato lì che mi si è aperto l'abisso.


Il problema architetturale: Uguaglianza Stringa vs Equivalenza Semantica

Due stringhe identiche sono per forza uguali, il che non è particolarmente sorprendente.

Ma due stringhe diverse possono avere lo stesso identico significato logico per il database.

Cosa succede se il server imposta scope: {"tenant_id": 400} e il client invia un payload con values: {"users.tenant_id": 500}?

Un controllo basato su stringhe non rileva alcun conflitto ("tenant_id" != "users.tenant_id"). Il sistema genera allegramente questo codice SQL:

update users
set
  users.tenant_id = 500
where
  tenant_id = 400;

Mentre diversi motori, quali PostgreSQL o SQLite, trovavano questa sintassi indigesta e la rifiutavano giustamente con un errore, MySQL la accettava con grande entusiasmo, eseguendo lo spostamento temuto.

Ma non solo!

Cosa succede se le colonne sono case-insensitive? Cosa succede se a livello di database è definito qualche alias? Cosa sarebbe successo con le decomposizioni Unicode (NFC vs NFD), dove la lettera à può essere rappresentata da un singolo codice o da due codici separati (a + accento)?

Le stringhe in byte erano diverse, ma per il database puntavano alla stessa identica colonna.

Riconoscere tutte le possibili varianti semantiche che un utente o un database potevano inventarsi per identificare la stessa colonna era una corsa alle armi persa in partenza. Chi sa davvero se due nomi sono equivalenti è il motore del database; ma filtersql genera solo stringhe e non è, per filosofia, collegato al database.


La Svolta: L'Asimmetria Consapevole

Per uscirne vittorioso ho capito che non dovevo rendere più "intelligente" il confronto tra stringhe, ma dovevo cambiare il contratto all'ingresso.

Ho introdotto un'architettura asimmetrica:

  1. in lettura (select, where): Il sistema rimane flessibile. Può accettare identificatori qualificati (users.tenant_id) o percorsi JSONB (data->>key), perché quotati e sanificati separatamente.
  2. in scrittura (insert, update, delete): nuova regola: tutte le chiavi in values o id devono superare la validazione _validate_bare_identifier.

Un identificatore di scrittura valido dev'essere un bare identifier puro:

  • nessuno spazio iniziale o finale (il trim() non pulisce più il dato: se c'è uno spazio, la query viene rifiutata con InvalidIdentifierError).
  • nessun punto (vietata la qualificazione di tabella come users.tenant_id).
  • nessuna sintassi JSONB o carattere di controllo.
  • forma canonica Unicode NFC obbligatoria.

Normalizzando sia lo scope che i campi di scrittura in NFC e applicando casefold(), l'area di confronto per le collisioni si è ridotta a uno spazio unidimensionale in cui nessuna ambiguità semantica può più nascondersi.

Nota: nella 1.2.8 ho poi esteso la stessa regola anche a _quote(), così che l'incoerenza non possa ripresentarsi in nessun altro punto del codice.


E la Allowed_Columns?

Ci pensavo da tempo, ma avrei dovuto implementarla prima.

Il collision check risolve il caso in cui il client prova a scrivere su una colonna di scope. Ma non risolve il caso in cui il client prova a scrivere o leggere su colonne pericolose. Per quello serve una whitelist.

In un'applicazione reale (e a maggior ragione in una pipeline con gli LLM), lasciare che un client o un'IA interroghino o filtrino su qualsiasi colonna esista a DB senza filtro è un rischio enorme: rischi di esporre password_hash, credit_card_token o filtrare su stipendio_netto.

Il prossimo passo sarà quindi introdurre la possibilità di limitare le colonne coinvolte a quanto previsto nella allowed_columns.

In questo modo i ruoli rimangono ben distinti e puliti:

  • il developer definisce a mano la whitelist applicativa dei campi esposti per quello specifico endpoint o utente.

  • filtersql fa il lavoro sporco da compilatore: sanifica l'input, quota gli identificatori, applica lo scope multi-tenant senza possibilità di bypass e difende il database dai DoS.

Tutto questo sarà pubblicato in una delle prossime versioni.

Tanti saluti ed al prossimo bug!

filtersql.org — github.com/fthiella/filtersql

L'IA sostituirà anche i vertici aziendali?

Commento del 2025 su un blog, in risposta a un post sull'impatto dell'IA nel mondo del lavoro.

Ai vertici aziendali l'IA piace molto perché promette di poter sostituire molto personale, in particolare chi svolge mansioni di livello medio/basso. Ed in effetti questa prospettiva è molto fondata, dato che molti dipendenti sono intrappolati in lavori superflui, ripetitivi o a basso valore aggiunto, a volte nemmeno per colpa loro ma spesso per inefficienze organizzative stratificate nel tempo.

La vera ironia, che cerco sempre di far presente, è che l'IA non minaccia soltanto la base della piramide, ma potrebbe sostituire anche gli stessi vertici aziendali che oggi si stanno impegnando molto per introdurla!

Chi decide di introdurre l'IA per ridurre il personale dovrebbe chiedersi, prima, se la stessa logica non si applica anche al suo ruolo.

L'IA come collaboratore, non come scorciatoia

Commento del 2025 su un blog. Lo ripubblico perché la distinzione tra "AI come collaboratore" e "AI come copia-incolla" è ancora la più importante da fare.

Tempo fa avevo scritto una libreria che estende un progetto open source, si tratta di un prodotto di nicchia che ha pochi download, è abbastanza semplice ma mi serviva e non esisteva, e svilupparlo è stato molto interessante. La qualità del codice che avevo scritto era discreta, ed il prodotto funzionava abbastanza bene anche se con diverse limitazioni di cui ero ben conscio.

Di recente ho provato con AI a vedere dove poter intervenire: ho chiesto di analizzare la qualità del codice, se ci fossero incoerenze, codice duplicato etc. ed ho avuto diversi suggerimenti, alcuni molto utili, altri che invece non ho ritenuto interessanti. Mi è stato segnalato un bug che, in realtà, non esisteva. Ma effettivamente un caso particolare era gestito, a livello di codice, in maniera non ottimale. Ho sistemato il sistema di inizializzazione, l'ho reso configurabile, ho scritto la documentazione, ho scritto una suite di test per verificare il corretto funzionamento della procedura e per essere certo che, quando correggo il codice o aggiungo funzionalità, il prodotto continua a funzionare.

Tutte cose che avrei saputo fare anche da solo, ma che portavano via tempo e concentrazione. Sostanzialmente ho usato l'AI come motivatore, o più che altro come collaboratore, non ho fatto nessun copia ed incolla a caso.

Da un lato credo che uno sviluppatore discreto, utilizzando opportunamente l'AI, possa arrivare a risultati eccellenti. D'altra parte anche uno sviluppatore scarso, utilizzando impropriamente l'AI, può ottenere risultati che sembrano buoni - ma non necessariamente lo sono. Diventa quindi più difficile discriminare chi è bravo da chi non lo è.

Joel Spolsky sosteneva che per assumere uno sviluppatore devi vedere come scrive codice, non come parla. È una delle 12 domande del suo Joel Test: "I nuovi candidati scrivono codice durante il colloquio?".

I "take home test" sono un'evoluzione di questo principio. Per dire, non so in Italia ma all'estero le ditte che dovevano assumere sviluppatori facevano una selezione fin da subito tramite "take home test", dei piccoli progetti da sviluppare in qualche mezza giornata ma che davano una idea della qualità del codice che si sapeva produrre. Oggi come si fa? Con un po' di AI, ed evitando nel contempo di fare troppi copia+incolla indiscriminati, è molto facile far sembrare di essere molto bravi.

Probabilmente l'AI spazzerà via diverse professioni, sono cose che sono sempre successe con le varie innovazioni tecnologiche ma forse siamo di fronte ad eventi più estremi.

Grafici? Traduttori? Interpreti? Speaker? E non sono tranquillo nemmeno per il futuro degli sviluppatori, sì, la qualità dell'AI non è sempre ottimale, ci sono molti errori, aspetti da migliorare, e molte attività per le quali l'intervento umano è sempre indispensabile, ma sarà sempre così? D'altra parte anche molti sviluppatori pagati come tali non è che siano sempre all'altezza del loro ruolo! Forse ne serviranno sempre meno? Pochi ma di livello alto? Staremo a vedere!

Appunti su Pandoc: stili personalizzati per Word

Brevi appunti pratici sull'integrazione tra Markdown, Pandoc e Microsoft Word tramite file di stile personalizzati.

Documenti Word

Generare un documento con gli stili di riferimento:

pandoc -o custom-reference.docx --print-default-data-file reference.docx

Si possono quindi modificare gli stili all'interno del file "custom-reference" con le normali modalità di Microsoft Word.

Introdurre nuovi stili:

<div custom-style="Super big">My super big text</div>
Normal text. <span custom-style="Highlighted text">This is highlighted</span>

(devono essere esistenti nel file custom-reference.docx)

Si può quindi compilare il documento con:

pandoc documento.md -o documento.docx --reference-doc=custom-reference.docx

Per quanto riguarda le tabelle non è purtroppo possibile modificare lo stile standard. Si può però effettuare un passaggio con python:

import docx
document = docx.Document('/tmp/xxx.docx')
for table in document.tables:
    table.style = document.styles['custom_style']
document.save("target.docx")

Fun Facts about the 8088

Note: A classic compilation of x86 low-level optimization tricks originally attributed to Microsoft engineer Chris Peters in the 1980s. Preserved here as a historical document on the craft of shaving off single bytes and clock cycles on the Intel 8088 prefetch queue and bus architecture.

1. Comparing a register

The fastest and smallest way to compare a 16 bit register to zero is to OR it with itself, e.g.

OR      BX,BX           ; 2 bytes, 3 clocks
JGE     BXisPositive

this is much better than comparing it with zero, e.g.

CMP     BX,0    ; 3 bytes, 4 clocks (bush league)

For the ultimate in comparing with zero, try to use the CX register. The 8088 contains the single instruction:

JCXZ    CXisZero        ; jump if CX is zero

This instruction makes a short jump if CX is zero.

To destructivly test for 1 or -1, use the DEC or INC instructions:

DEC     DX              ; 1 byte, 3 clocks
JZ      DXisOne         ; if zero, DX was 1

or

INC     DX              ; 1 byte, 3 clocks
JZ      DXisMinusOne    ; if zero, DX was -1

The LOOP instruction is just a fancy way of writing:

DEC     CX
JNZ     CXisNotZero

The difference is that LOOP is 1 byte smaller and 2 clocks faster. The LOOP instruction can be used to compare CX with multiple values as follows:

LOOP    CXisNotOne       ; If CX = 1 then...
...                      ; ...do this code, else...

CXisNotOne: LOOP CXisNotTwo ; if CX = 2 then... ... ; ...do this code, else...

CXisNotTwo: LOOP CXisNotThree ; if CX = 3 then... ... ; ...do this code, else etc.

Its possible to check if a signed number is in the range 0-n with a single comparison to n:

CMP     DX,639          ; if <0 or >639...
JA      OutOfRange      ; ...its out of range

This is smaller and faster than:

OR      DX,DX           ; Never CMP DX,0!
JL      OutOfRange      ; If negative, its out of range
CMP     DX,639
JG      OutOfRange      ; if greater, its out of range

This can be generalized to any signed compare within a specific range. If you wanted to make sure the AX register contained a number from -5 to 17:

SUB     AX,-5           ; Subtract off lower bound
CMP     AX,17-(-5)      ; Compare upper - lower bound
JA      OutOfRange

You cannot compare a segment register. To do so copy it to a register or memory location, then compare it.

2. Setting a register to zero

To set a register to zero, the smallest, fastest way is to XOR it with itself, e.g.

    XOR     BX,BX           ; 2 bytes, 2 clocks

is smaller and faster than:

    MOV     DX,0            ; 3 bytes, 4 clocks

There is one side affect: MOVing does not affect the flags, but XORing does. The 8088 aficianado will only move a zero into a register in the rare cases where the flags must be preserved.

3. Incrementing, Decrementing

It is smaller to increment or decrement a 16 bit register then an 8 bit register, so if it doesnt matter, use a 16 bit register, e.g.:

    INC     DX              ; 1 byte, 3 clocks

is smaller than:

    INC     DL              ; 2 bytes, 3 clocks

Same thing goes for decrementing, its smaller to use a 16 bit register.

Its smaller (but not faster) to increment a register twice then to add 2 to it, e.g.

INC     DX              ; 1 byte,  3 clocks (total)
INC     DX              ; 2 bytes, 6 clocks

is smaller and slower than:

ADD     DX,2            ; 3 bytes, 4 clocks

One side affect: INC and DEC do not affect the carry flag, ADD does.

4. If Then Else

When confronted with an If, Then, Else problem the assembly language programmer will often write it as Else, If, Then. For example, a sample problem might be to return 123H in the DX register if AL<5, otherwise return zero in DX. Using If, Then, Else produces:

 CMP     AL,5                ; Is AL less than 5?
 JB      ALisBelow5          ; yes...
 XOR     DX,DX               ; Set DX to zero if AL>=5
 JMP     SHORT Continue      ; proceed

ALisBelow5: MOV DX,123H ; Set DX to 80H if AL<5 Continue:

Using Else, If, Then:

 MOV     DX,123H             ; Set DX to 80H if AL<5
 CMP     AL,5                ; Is AL less than 5?
 JB      ALisBelow5          ; yes...
 XOR     DX,DX               ; Set DX to zero if AL>=5

ALisBelow5:

The idea is to do the work for the most likely case, then do the comparision. If you were right you're all done, if not then do the other case.

5. Copying strings

One of the simplest ways to copy a null terminated string is as follows:

; ; DS:SI points at source ; ES:DI points at destination ; CopyString: LODSB ; read character into AL, inc SI STOSB ; store charcter, increment DI OR AL,AL ; was it the null JNZ CopyString ; no, repeat

A fast way to copy a string where the count is known is

SHR     CX,1    ; divide count by two (2 bytes = 1 word)
REP     MOVSW   ; move the first part fast
JNC     Even    ; No carry if count was even
MOVSB           ; move the odd byte

Even:

Slightly faster but slightly bigger is:

SHR     CX,1    ; Get the word count
REP     MOVSW   ; move the words, leaving CX = 0
RCL     CX,1    ; if count was odd make CX = 1 
REP     MOVSB   ; (possibly) move last byte

This is slightly faster because there is no jump to empty the prefetch queue.

The REP instruction checks to see if the loop count in CX is zero before starting, there is no need to check it beforehand.

6. Testing long pointer for null

Sometimes its necessary to check if a long pointer is null before using it. A sequence of code that works well is:

 LES     DI,[LongPointer]        ; get the long pointer
 MOV     CX,ES                   ; copy ES to CX
 JCXZ    PointerIsNull           ; if zero, dont use it

This assumes segment zero is an invalid pointer.

7. Exchanging

To move a register to or from AX takes 2 bytes and 2 clocks

MOV     AX,DX           ; 2 bytes, 2 clocks

However, to exchange a register with the AX register takes 1 byte and 3 clocks:

XCHG    AX,DX           ; 1 byte, 3 clocks

A clear case where its possible to optimize for size or speed. This optimization is only available with the AX register. A full understanding is actually more complex. Although XCHG takes 3 clocks it is often faster because it only uses one byte of the prefetch queue.

8. Testing bits

The 8088 contains an instruction that allows you to test various bits. It works by doing a non destructive AND operation.

TEST    AX,0800H        ; 4 bytes, 5 clocks
JZ      BitIsOff

If one of the bytes is zero (a common occurence) this can be optimized to:

TEST    AH,08H          ; 3 bytes, 4 clocks
JZ      BitIsOff

The best way to test the hi bit is:

OR      AX,AX           ; 2 bytes, 2 clocks
JNS     HiBitIsOff      ; jump not signed (hi bit zero)

A destructive way to test the low order bit is to shift it to the right into the carry flag:

SHR     AX,1            ; 2 bytes, 2 clocks
JNC     LoBitIsOff

You can test multiple bits with the TEST instruction:

TEST    BL,11000000b    ; check 2 highest bits
JZ      BothAreZero     ; if zero, both are zero

Sometimes you want to jump if either are zero instead of both zero, this can usually be accomplished by the NOT instruction and reversing the sense of the jump instruction:

NOT     BL              ; reverse the bits   
TEST    BL,11000000b    ; check 2 highest bits
JNZ     EitherAreZero   ; if either were zero, then this 
                        ; is non zero

9. Absolute value

A fascinating way to get the absolute value of the AX register was discovered by Marlin Eller:

CWD             ; replicate hi order bit of AX into DX
XOR     AX,DX   ; do a 1's complement or do nothing
SUB     AX,DX   ; add 1 to get a 2's complement

The boring method does not affect the DX register and can be used on any register:

OR      BX,BX   ; never CMP BX,0!
JGE     NotNeg  ; if negative...
NEG     BX      ; ...make it positive

NotNeg:

The boring method empties the prefetch queue with the JGE instruction, making it much slower.

10. Length of null terminated string

To get the length of a null terminated string, scan for the null at the end with a starting count of -1

; ; ES:DI points at null terminated string ; XOR AL,AL ; look for null 2 bytes (total) MOV CX,-1 ; CX = -1 5 bytes REPNE SCASB ; CX = -len-2 7 bytes NOT CX ; CX = len+1 9 bytes DEC CX ; CX = len 10 bytes

This count does not include the null at the end. If you want it to include the null, just delete the final DEC CX. The use of the NOT instruction is quite interesting here.

11. Returning flags

The 8088 contains instructions for setting (STC) and clearing (CLC) the carry flag. To set the zero flag, simply compare some register with itself:

CMP     DX,DX            ; set zero flag

To clear the zero flag, OR the stack pointer with itself:

OR      SP,SP           ; clear zero flag

This is making the safe assumption that the stack pointer is not zero.

12. Shifting

Variable count shifting is slow on the 8088. Its faster to shift twice then to set a count of 2:

SHR     AX,1    ; 2 bytes, 2 clocks  (total)
SHR     AX,1    ; 4 bytes, 4 clocks

is much faster than:

MOV     CL,2    ; 2 bytes,  4 clocks  (total)
SHR     AX,CL   ; 4 bytes, 20 clocks!

Variable count shifting is slower when shifting less than 5 bits, after that the prefetch queue makes variable shift counts faster.

13. Multiply and Divide

The multiply and divide instruction are some of the slowest instructions on the 8088. To give some perspective, a register to register MOV instruction takes 2 clocks, while a signed divide (IDIV) using registers can take 184 clocks.

If your goal is to write fast 8088 code, multiplying by constants can usually be done as a series of shifts and adds:

;
;  Multiply the AX register by 10
;
SHL     AX,1            ; AX = AX *  2    ( 2 clocks)
MOV     BX,AX           ; BX = AX *  2    ( 4 clocks)
SHL     AX,1            ; AX = AX *  4    ( 6 clocks)
SHL     AX,1            ; AX = AX *  8    ( 8 clocks)
ADD     AX,BX           ; AX = AX * 10    (11 clocks)

A multiply would be more than 10 times slower, but would take fewer bytes. Multiply is useful when neither argument is constant or you need to save bytes.

Mark Zbikowski uses this method to divide the AX register by 512:

SHR     AX,1         ; divide AX by 2
XCHG    AL,AH        ; divide AX by 512 (AX is unsigned)
CBW                  ; AL < 128, so this sets AH to 0

14. Converting bytes to segments

To convert a byte count to a paragraph count try:

;
;  DX contains a byte count
;
ADD     DX,15       ; round up to next paragraph
MOV     CL,4        ; 2^4 = 16 bytes per paragraph
SHR     DX,CL       ; divide by 16 by shifting 4 times

DX now contains a paragraph count. This assumes the value in DX is less than 0FFF1H. To cover the extended case:

ADD  DX,15        ; round up to next paragraph
RCR  DX,1         ; divide by 2, including carry 
MOV  CL,3         ; 2^3 = 8
SHR  DX,CL        ; divide by a total of 16

15. Call, Return, Jump

A near call followed by a near return can always be replaced with a near jump:

JMP NearProc  ; 3 bytes, 15 clocks

is smaller and much faster than:

CALL  NearProc  ; 3 bytes, 19 clocks (total)
RET     ; 4 bytes, 35 clocks

Its often possible to eliminate the JMP entirely by moving the subroutines adjacent to each other.

Conditional jumps on the 8088 are always short, i.e. the destination must be within -128 to 127 of the instruction pointer. It seems every time a single line of new code is added some conditional jump becomes out of range. One technique to get around this is to find a similiar conditional jump to jump to:

JC      OutOfRange     ; I want to jump to disk error...
...                    ; ...but its too far away, so...

OutOfRange:

JC      DiskError      ; I jump to this test for carry

Although this is not in the scope of this document, out of range jumps are usually the 8088 telling you that your subroutines have grown too large and should be broken up.

16. Multiple Entry Points

An old 8080 trick involving multiple entry points can be adapted to the 8088. Instead of doing this:

Entry1:
    MOV     AL,1                    ; 2 bytes (total)
    JMP     SHORT EntryCommon       ; 4 bytes
Entry2: 
    MOV     AL,2                    ; 6 bytes
    JMP     SHORT EntryCommon       ; 8 bytes
Entry3:
    MOV     AL,3                    ;10 bytes 
EntryCommon:

The hearty and brave will do this:

Entry1:
    MOV     AL,1                    ; 2 bytes (total)
    DB      03DH                    ; 3 bytes 
Entry2:
    MOV     AL,2                    ; 5 bytes
    DB      03DH                    ; 6 bytes
Entry3:
    MOV     AL,3                    ; 8 bytes
EntryCommon:                        ; flags are modified

The DB 03DH is the opcode for a CMP AX,xx. In this case the bogus CMP AX's are used to swallow up the MOV AL,x that follow.

A special case of this discovered by Pat Tharp optimizes for dual entry points:

TrueEntry:      ; come here to set AX = TRUE
  DB  0B8H      ; (opcode for MOV AX,...)
FalseEntry:     ; come here to set AX = FALSE
  XOR     AX,AX   ; AX = 0 = FALSE

Whats happening here is that the 0B8H is the opcode for MOV AX,### which will put the opcode for the XOR AX,AX (nonzero C031H) in AX. This is a clear win of 3 bytes over the boring method:

TrueEntry:
  MOV     AL,1    ; set AX nonzero
  JMP     SHORT EntryCommon
FalseEntry:
  XOR     AX,AX   ; set AX zero
EntryCommon:

17. Assertion Macros

When using these advanced techniques its important not to expose yourself to bugs caused by changing constants in your program. For instance, in the Multiply and Divide section of this document there is a code to quickly multiply by ten. If the constant should later change from ten to twelve this code would no longer work. An assertion macro would flag this code as being in error, saving many hours of needless debugging. In cannot be stressed to strongly that advanced 8088 programming requires liberal use of assertion macros and extra documentation.

Dynamic Markdown

Markdown è un sistema di markup che può essere utilizzato per aggiungere elementi di formattazione a dei normali file di testo. È molto apprezzato dagli sviluppatori, e non a caso è il sistema di riferimento per GitHub oppure per StackOverflow ma può essere utilizzato per qualsiasi tipo di documento, compresi report, libri, manuali, testi, etc.

Uno strumento indispensabile per lavorare, tra le altre cose, su documenti markdown è pandoc che viene definito come il coltellino svizzero per i documenti di testo. Può essere infatti utilizzato per convertire documenti markdown in documenti Word, Open Office, LaTeX, PDF, etc.

Ad esempio:

pandoc test.md -s -o test.odt

oppure:

pandoc test.md -s -o test.pdf

(necessario aver installato pdflatex, come ad es. portable MikTex, e configurato il PATH in maniera adeguata).

È possibile inoltre applicare dei template di alta qualità, come il template Eisvogel

pandoc test.md -s -o test.pdf --template eisvogel

Markdown Dinamico

Ho pensato che sarebbe estremamente comodo poter disporre di un documento markdown "dinamico". Ad esempio, per inserire delle tabelle in un documento Markdown posso utilizzare il plugin MarkDown per Adminer che ho scritto qualche tempo fa, in modo molto semplice e con i risultati spesso migliori di un copia-incolla da Excel su documento Word!

Tuttavia sarebbe meglio ancora poter pescare i dati direttamente dal database di origine (così come da un CSV, XLSX, o quant'altro).

Un sistema che potrebbe funzionare è Jinga2 ovvero uno dei più utilizzati framework per template di Python. Questo sistema però, per precisa scelta architetturale, non prevede la possibilità di mescolare codice di programmazione con condice markup. Avendo vissuto in prima persona gli anni '90 dello scorso millennio, dove il codice HTML si mischiava al codice di programmazione senza capire bene il limiti di dove finiva uno ed iniziava l'altro, posso dire che tale scelta è perfettamente condivisibile. Tuttavia, pur essendo una opzione sempre valida, non è quello che stavo cercando!

Linguaggi di scripting?

Si potrebbe usare un Makefile con qualche script bash, magari integrato con dei sistemi che generano output in formato markdown come il mio perl-Sql-Textify.

Oppure PHP che, nel bene e nel male, permette di inserire del codice all'interno di un documento e può quindi essere utilizzato per rendere un documento Markdown dinamico.

Oppure ancora il buon vecchio HTML::Mason, che mi ha dato grandi soddisfazioni all'inizio del millennio permettendo di integrare codice HTML con il linguaggio perl. Purtoppo né HTML::Mason né Mason2 sono attivamente sviluppati da anni, anche se proprio nel momento in cui ho iniziato a scrivere questo post (ovvero il giorno 11 febbraio 2023 - scrivo molto lentamente!) è stato pubblicato HTML::Mason versione 1.60 che contiene un piccolo bugfix. Non ci sono poi aggiornamenti successivi.

Script Python?

Ho pensato di risolvere il problema con dei piccoli script Python. Un documento Markdown sarà composto come segue, con del codice Python inserito all'interno del documento:

---
title: "Report Dinamico"
author: [Federico Thiella]
date: 2023-09-11
subject: "Report Dinamico"
keywords: [codice, python, esempio]
lang: "it"
table-use-row-colors: True
book: True
...
^ import mktools as mk
^ c={'host':'myhost.mshome.net', 'database':'mydatabase', 'user':'myuser', 'password':'mypassword'}
# Report Dinamico creato con Markdown

^ q=mk.query_db("select * from sezioni", c)
^ for r in q['rows']:

## Titolo: <& r[2] &>

<& r[3] &>

Inserisci una tabella:

^   w=mk.query_db("""
^      select
^        col1,
^        col2,
^        col3
^      from
^        progetti
^      where
^        progetto_id={}
^      order by
^        col1, col2, col3
^      """.format(str(r[0])), c)
^   w['head'][1]['name']='Colonna 1'
^   w['head'][1]['align']='right'
^   w['head'][2]['name']='Colonna 2'
^   w['head'][2]['align']='right'
^   w['head'][3]['name']='Colonna 3'
^   w['head'][3]['align']='left'
^   mk.markdown_table(w['head'], w['rows'])
^^
\pagebreak
  1. questo codice Markdown dinamico verra compilato dallo script pp.py in uno script puro python, che può essere eseguito e che genererà in output il codice Markdown statico finale, pronto da poter essere utilizzato e/o convertito in altri formati;
  2. la libreria mktools.py contiene delle funzioni utili per eseguire velocemente delle operazioni su database:
  3. mk.query_db esegue una query e restituisce intestazione w['head'] e righe w['rows'];
  4. mk.markdown_table che converte la tabella descritta da intestazione e righe in formato markdown.

La compilazione del documento avviene in questo modo:

python pp.py -i documento.md | python | pandoc -o documento.pdf --template eisvogel.latex -s

(lo script dinamico e la libreria di utilities saranno pubblicate e descritte su GitHub quanto prima. Non si tratta di un sistema "professionale" ).

Mako Templates

Un'ultima opzione può essere il sistema Mako Templates, la cui filosofia "Don't reinvent the wheel...your templates can handle it!" contrasta con il mio paragrafo precedente. Lo approfondirò appena mi sara possibile!

Numeri di telefono riciclati: una falla di sicurezza?

Commento del febbraio 2023 su un blog, in risposta a un articolo che raccontava il caso di un utente che aveva ereditato per caso l'account WhatsApp di un'altra persona a causa del riutilizzo del numero di telefono.

Quando mi è stato assegnato il primo cellulare di lavoro, molti anni fa, ho cominciato a ricevere SMS e telefonate da amici che non sapevo di avere, che mi invitavano a feste ed eventi vari.

Una signora, che si spacciava per mia nonna, mi chiamava spesso piangendo per dirmi che, anche se mi rifiutavo di riconoscerla, per lo meno avevo ripreso a risponderle al telefono.

Infine un ragazzo mi accusava di avergli rubato cellulare e numero di telefono.

Gli operatori telefonici riutilizzano sempre i numeri di telefono, nel mio caso devono averlo fatto un po' troppo in fretta.

Il problema

Per quanto riguarda questa falla di sicurezza, mi sembra talmente scontata che mi stupisco che sia stata sottovalutata: nella maggior parte dei servizi di messaggistica il numero di telefono viene verificato solo in fase di attivazione (tramite chiamata o SMS) e poi l'applicazione viaggia su binari diversi rispetto alla SIM.

Puoi cambiare SIM, numero di telefono, puoi anche togliere la SIM, etc. ma a WhatsApp la cosa non interessa. Se un numero viene riassegnato, WhatsApp come fa a sapere se sei il vecchio utente di prima che ha cambiato telefono o un nuovo assegnatario del numero?

Le policy di cancellazione automatica per inattività (come i canonici 45 giorni previsti da alcune piattaforme) mitigano solo in parte il fenomeno, ma restano una toppa parziale.

Lo stesso problema con i domini

Ah, il problema succede anche con i domini internet. Se un dominio viene riassegnato, cosa che avviene per diverse ragioni non sempre volontarie, il nuovo proprietario può non solo leggere tutta la tua (nuova) corrispondenza, ma anche prendere accesso a tutti gli account esterni se registrati con una casella di posta legata al dominio "perduto".

Il pattern è sempre lo stesso: una risorsa viene riassegnata, ma i servizi che ci si appoggiano non lo sanno.

Windows Home Profile

Percorso Profilo Utente di Windows

La variabile %HOMEPROFILE% contiene il percorso del profilo corrente, tipicamente:

c:\Users\username

mentre la cartella dei documenti è di default la sua sottocartella Documents quindi:

c:\Users\username\Documents

tuttavia il nome della cartella Documents è configurabile e potrebbe essere stato impostato a piacere, per cui non si può dare per scontato che il percorso %HOMEPROFILE%\Documents sia sempre quello corretto.

Il percorso corretto è registrato nella seguente posizione del registro di Windows:

λ reg query "HKEY_CURRENT_USER\Software\Microsoft\Windows\CurrentVersion\Explorer\Shell Folders" /v Personal

HKEY_CURRENT_USER\Software\Microsoft\Windows\CurrentVersion\Explorer\Shell Folders
    Personal    REG_SZ    C:\Users\username\Documents

Alcuni linguaggi di programmazione mettono a disposizione librerie e funzioni apposite per interrogare il registro, mentre con la shell è necessario utilizzare uno script apposito.

Script di lettura Percorso

Il seguente script legge il percorso richiesto, e lo memorizza nella variabile %value%:

@echo OFF
setlocal ENABLEEXTENSIONS

set KEY_NAME="HKEY_CURRENT_USER\Software\Microsoft\Windows\CurrentVersion\Explorer\Shell Folders"
set VALUE_NAME="Personal"

for /F "tokens=2*" %%A IN ('reg query %KEY_NAME% /v %VALUE_NAME%') do (
    set value=%%B
)
echo %value%

questo script sembra funzionare sia con le versioni più recenti di Windows che con le precedenti, e supporta percorsi con gli spazi.