Skip to content

2018

L'IP era vero, il mittente no

Commenti del 26 ottobre 2018 su un blog, in una discussione sul tracciamento degli utenti. Li ho riorganizzati in un unico testo, tagliando le ripetizioni.

Il discorso è complesso, e forse si fa confusione perché si mescolano due problemi diversi. Io li vedo così:

  • molte aziende su internet trattano nostri dati personali che noi non sappiamo di aver dato loro (e con questi dati ci campano, ecco perché la professione di data scientist è sempre più richiesta e sempre più pagata...)
  • molte aziende non sono nemmeno in grado di trattare e proteggere i dati che invece abbiamo affidato loro

Vale anche per i dati sanitari, ad esempio... purtroppo privacy, GDPR etc. vengono visti come adempimenti burocratici da applicare alla lettera senza capire che cosa si sta facendo e perché, col risultato che si possono anche avere sistemi che sulla carta sono compliant ma in pratica fanno acqua da tutte le parti...


Un esempio che mi è capitato in prima persona molti anni fa, e che forse può far capire meglio il senso del mio discorso. Era il 1999, e via mail arrivava un virus. A me e a molti partecipanti di un newsgroup. Il mittente era anonimo, ma l'IP era vero.

A quel tempo non si discuteva su Disqus o sui forum, ma sui newsgroup. Un newsgroup era un elenco di post, cioè nuove discussioni o risposte ad altre discussioni. Ad ogni post era associato il suo autore - anche fittizio, non serviva registrazione - e di ogni post era possibile ottenere l'indirizzo IP da cui era stato inviato.

Volevo capire chi era il vero mittente a partire dall'IP, ma come fare senza scomodare il provider?

Ho pensato di scaricare qualche migliaio di discussioni dai newsgroup, e le ho caricate su di un database con IP, data ed ora. Ho confrontato gli IP delle mail anonime con gli IP nel database, e la cosa ha funzionato perfettamente!

E ho trovato diverse cose interessanti: utenti diversi che condividevano lo stesso IP, un utente che si spacciava per utenti diversi, e l'autore inconsapevole della mail - il vero autore non sapeva di aver mandato il virus, perché il suo PC era infetto.

Corollario: discutere sui newsgroup rendeva pubblicamente associabile il tuo IP al tuo nickname!

Oggi i newsgroup non vanno più di moda, l'IP potrebbe non voler dire molto e gli IP delle discussioni in ogni caso non sono più pubblici (ma il gestore del forum li ha eccome) e solo una piccola percentuale di utenti partecipa alle discussioni - ci sono però altre soluzioni più moderne.

Non sto dicendo che tutti commettano illeciti, sto solo dicendo di fare attenzione che i fornitori di servizi dispongono di molti più dati su di noi rispetto a quelli che noi pensiamo di dare, e i dati sono come i maiali, non si butta via niente!

Questo lavoro è probabilmente la cosa più importante che ho fatto in quegli anni, perché dimostra di aver capito come funziona davvero il mondo: le tracce digitali raccontano una verità che spesso non coincide con l'intenzione di chi le ha lasciate.


Il punto però non è solo Facebook o Google presi singolarmente. Il punto è come funziona il web nel suo insieme.

Oggi i newsgroup sono stati sostituiti, ma il meccanismo di correlazione non è scomparso: si è semplicemente spostato sui circuiti pubblicitari e sui widget di terze parti.

Se un utente visita in sequenza siti diversi che condividono lo stesso sistema di commenti, lo stesso fornitore di statistiche o la medesima rete di banner pubblicitari, il fornitore terzo dispone di tutti gli elementi per ricostruire la cronologia di navigazione, anche in assenza di un account registrato o di un'autenticazione esplicita.

Le tecnologie cambiano, ma la dinamica resta la stessa: le informazioni che condividiamo sono molte più di quelle che pensiamo di trasmettere.

Show text files width

I am often working with fixed width text files. Sometimes text files aren't produced properly, so I have to make sure that all rows share the same width.

This is one liner perl thay I use quite often, on a Windows shell:

c:\> perl -lne "$h{length($_)}=1; END{print join \"\n\", sort keys %h}" source.txt

this is how it looks like in Linux, with a simpler way of quoting:

$ perl -lne '$h{length($_)}=1; END{print join "\n", sort keys %h}' source.txt
  • the -l flag handles newlines
  • the -n flag adds a while loop

for each line we calculate its length with length($_), we set the hash element $h{length($_)} to 1, which means that we have at least one row with the calculated length. When we are finished scanning the file, we print the list of hash keys in sorted order, which is the list of calculated lengths.

If the one liner returns only one length then we are fine, all rows share the same, hopefully correct, width. If it returns more lengths then I'll have to further investigate the problem.

Cut a fixed width text file in CMD

I recently had the need to cut a fixed width text file at a certain column, and I had to do it in the Windows 10 world, without any particular utility.

In linux it would be rather simple, thanks to the cut command:

$ cut -c -10 example.txt

On my windows machines there's always Perl available, and one alternative solution would be using one liner substitution:

perl -pe s/(.{10}).*/$1/ example.txt

but on my colleagues machines there are no particular utilities installed, so I had to create the following script:

@echo off
setlocal EnableDelayedExpansion
for /f "delims=" %%r in (example.txt) do (
    set s=%%r
    echo !s:~0,10!
)

it's still possible to use a single line command, without the need of a .cmd script, but firts we have to launch the prompt with a special parameter cmd /v which indicates that delayed expansion is enabled (it is disabled by default to be compatible with MS-DOS 2.0 batch files!). After we launch cmd /v the command becomes:

for /f "delims=" %r in (example.txt) do @(set s=%r && echo !s:~0,10!)

not really intuitive, is it? I think I am missing my teenager days when I used to program Assembly x86: that was complex, but for good reasons. Cmd scripting is still complicated, but for no knowns reasons! Also, please note that this solution is painfully slow.

Pratical use of Sql::Textify

I know that perl is considered out of fashion nowadays, but for some tasks it's still a good and handy choice. Here's a practical use of my module Sql::Textify. Every morning I need to check the status of some tasks, like the number of times my web services have been called the day before, with some performance analysis and the number of errors. At the same time I want to know if my pentaho tasks are all finished. And I want to know if all my backups are updated.

Luckily I managed to write all the info I need into some sql tables, updated automatically. So basically I just need to login to my Adminer.php instance and run some queries or some views. But how if I need to share those info to my colleagues? I can quickly export a dataseto to a Markdown text file, ready to be beautified with Markdown Here plugin.

But I wanted to make things more practical and faster. My initial idea was to make a Mason2 plugin that calls Sql::Textify and this might still be a good idea to handle complex contents, but since my content is often simple and Mason2 is out of fashion anyways, I wrote a simple script that handles everyting.

Here's an example. First we need a very simple template html page:

<html>
<header>
<style>
..insert a good style..
</style>
<title>[[ $title ]]</title>
</header>
<body>
[[ $body ]]
</body>
</html>

then the perl script is like this:

use strict;
use warnings;
use SQL::Textify;
use File::Slurp;

my $t = Sql::Textify->new(
    conn => "dbi:SQLite:dbname=samples.db",
    username => "username",
    password => "password",
    format => 'html'
);

# read the template file
my $html = read_file( 'main.html' );

# set the title
my $title = "Report Indicizzazione";

# set the content of the main component
my $body = <<'BODY';
<h1>Daily report</h1>

<h2>Public Web Service</h2>

<% $t->textify("select * from view_web_service_status where eventdate>=currentdate"); %>

<h2>Pentaho Integrations Log</h2>

<% "select * from view_pentaho_log" | $t->textify %>
BODY

# first syntax, evaluates code between <% and %>
$body =~ s /\<\%\s+(.*?);\s+\%\>/$1/eeg;

# second syntax, apply $t->textify to the query (works as a filter)
$body =~ s /\<\%\s+(\".*?\")\s*\|(.*?)\s+\%\>/"$2\($1\)"/eeg;

# convert all [[ $variable ]] to the actual value
$html =~ s/\[\[ (\$\w*) \]\]/$1/eeg;

print $html;

This is not a perfect solution, but I just needed a quick tool to export my data and I wanted it to look good. The syntax is inspired somehow to the Mason2 syntax. A real Mason2/(or anything else) component of course is much more flexible but at the same time is little slower to write and more difficult to mantain.

Dump Markdown plugin for Adminer

Few months ago I started working on a Markdown dump plugin for Adminer, now I can finally write a post about it.

I am using Markdown almost every day, I use it for writing blog posts (like this one), I am using it for writing documents, wikis, report, notes, and also for composing e-mails - thanks to the powerful plugin Markdown Here - so I needed a tool to quickly export SQL tables and queries from Adminer.php to Markdown tables.

Another tool I've been working on is my perl module SQL::Textify, it is essentially an improved version of my old SQL Markdown Builder tool (now considered obsolete), and I find it great to run SQL query from the command line, and get the result in Markdown, HTML, JSON format.

Install the plugin

I am using a directory /var/www/tools where I put all of my tools which I want to be accessible through a web interface. This directory will be published at https://localhost/tools, make sure that php files put here will be handled correctly.

There I my adminer.php file which will load the original adminer.php along with my own plugin plus other plugins I've downloaded:

adminer.php

<?php
function adminer_object() {
    // required to run any plugin
    include_once "./plugins/plugin.php";

    // autoloader
    foreach (glob("plugins/*.php") as $filename) {
        include_once "./$filename";
    }

    $plugins = array(
        // specify enabled plugins here
        new AdminerDumpMarkdown,
    );

    /* It is possible to combine customization and plugins:
    class AdminerCustomization extends AdminerPlugin {
    }
    return new AdminerCustomization($plugins);
    */

    return new AdminerPlugin($plugins);
}

// include original Adminer or Adminer Editor
include "./adminer-4.3.1-en.php";
?>

on the same directory I have put the original adminer adminer-4.3.1-en.php, and in the directory plugins I have put the plugin.php file, which is required to run any plugin, and my own plugin dump-markdown.php:

you can make sure that only the adminer.php will be accessible to the outside world while any other file will be blocked.

dumpFormat() function

This function will return the Markdown output options:

return array('markdown' => 'Markdown');

this output will be added to the existing list.

dumpTable() function

This function will add slashes to the table name before some special characters:

  • \r
  • \n
  • \"
  • \

and will return it as a header h2 (preceded by two ##). The return true; will make sure that the default output of Adminer won't be processed.

The header of the table won't be printed here, as at this point we still don't know the width of every single column.

dumpData() function

Here is where Markdown table will be printed to the output. This function will sample the first 100 rows to calculate the width of every column, then it will start to output all sampled rows and then every single additional row.

Here we also use return true; so the default output of Adminer won't be processed.

To Do

There are few things I want to implement:

  • add slashes before every special character of every single field. There's no a "standard" way to quote Markdown strings and there's not a standard list of special character, I am trying to get best results with stackedit.io, Markdown Here, other php and perl libraries;
  • add some parameters in the query, to control its output e.g. maximum width, record or table format, etc. same as SQL::Textify
  • add groups and levels in the output query. This will be implemented in SQL::Textify also.