Tuesday, 22 July 2014

Dettare l’agenda: come la comunità open data è entrata nel radar delle politiche europee



Non scrivo mai post "proxy", ma questa volta faccio un'eccezione per raccontarvi un'aneddotto che mi inorgoglisce un po.
Per prima cosa leggetivi l'articolo dello spaghettaro Alberto Cottica:

http://www.cottica.net/2014/07/22/dettare-lagenda-come-la-comunita-open-data-e-entrata-nel-radar-delle-politiche-europee/

Vorrei aggiugnere un punto 0 gogliardico alle puntate precedenti:
Raduno Spaghetti Open Data, cena del Venerdi 28 Marzo.
Dopo aver fatto i caciaroni alla Cantina Bentivoglio manco fossimo una gita delle medie stiamo in giro per bologna a bere e chiacchierare.
Ad un certo punto sono con Alberto Cottica e Matteo Brunati e si ragiona sul fatto che non ci siano esponenti del progetto ePSI tranne Matteo, italiano, già super conosciuto tra noi spaghettari.
Questa è la politica del progetto sulla partecipazione agli eventi ci spiega. Forse per motivi economici penso io, comunque sbagliata.
Alberto non ci sta, è abituato a "fare rete" in Europa, e subito partiamo con il capire cosa darebbe uno scambio culturale a tema Open Data.
Pensiamo a quanto potrebbe essere utile far venire un ragazzino che ne so, portoghese, a capire che facciamo qui in Italia, a quanto potrebbe diffondere idee in giro per l'Europa.
Un viaggio pagato? Una borsa di studio?
Ecco che Alberto la spara: ci vorrebbe un Erasmus degli Open Data.
Sorrido, penso "magari".. :)

Thursday, 13 March 2014

My Cinema Knowledge: "my movies 2" aka kiss open refine

In this episode of " My Cinema Knowledge" I will try to link video files I have on my pc with Freebase. I tried to keep the process as simple as possible and this time the requiremt is only a normal laptop with open refine and its extensions.
Use of this data in the next post!

Step By Step How-to

  • Build a csv with all the film names. I used folders names in my disk extracted using a linux command ( find . -type d > myMovies.csv )
  • Import in open refine (I used ver2.6)
  • Extracted movie name from folder name taking only the last part of the location. For me it was something like this in GREL:
    value.split("/")[value.split("/").length()-1]
    Here a guide:
    https://github.com/OpenRefine/OpenRefine/wiki/Understanding-Expressions
  • Reconciliation: There is now a cloud based reconciliation service for freebase, now working also with italian language. It should be included in open refine but it does not work for this bug https://github.com/OpenRefine/OpenRefine/issues/805. You can make it work creating a new standard one using this address:   http://reconcile.freebaseapps.com/reconcile  

  • Run reconciliation selecting the "film" type
  • Select the match with a "high" match using the facet on best match score
  • match all cell to the highest candidate (from reconciliation->action)
  • Manually find a match for other items
  • Add new column based on the reconciliated one with expression: cell.recon.match.id
  • In order to have the freebase rdf id add new column based on the last one with expression: "http://rdf.freebase.com/ns/" +  "m." +
    value.split("/")[value.split("/").length()-1]
  • Download and compile rdf-extension: https://github.com/fadmaa/grefine-rdf-extension
  • Edit rdf skeleton like this: (preview does not work because of https://github.com/fadmaa/grefine-rdf-extension/issues/89)


In order to make a uri out of row index I just used a custom vale: "http://www.mycinemaknowledge/video/" + value
  • Export in rdf

    Wednesday, 20 November 2013

    My Cinema Knowledge: "my movies" aka Multi-language reconciliation using Freebase


    In this first episode of " My Cinema Knowledge" I will try to describe my film catalog mixing private information (my disk folders) with public ones (Freebase)
    I will use Ubuntu, Open Refine, a little python script and a RDF store, Virtuoso.




    Step By Step How-to

    • Build a csv with all the folders names in my disk using a linux command ( find . -type d > myMovies.csv )
    • import in open refine (I used lod refine, a package including open refine and the rdf extension)
    • Extracted movie name from folder name taking only the last part of the location
    • Added a reconciliation service based on the freebase dump created previously (the making of is described in This post) imported in a Virtuoso triple store
      For this I used the SPARQL based reconciliation service feature of the RDF extension
      Using a custom reconciliation service over freebase I will not be limited to the english languages provided by the Freebase reconciliation service
    • After 10 minutes on a 8gb Ram machine, this the results (out of about 310):
      • 138 movies automatically recognized 
      • 66 movies with multiple choices (semi automatic)
      • 109 without a match
    • Reason for the missing matches are:
      • Missing in Freebase (mostly italina movies)
      • Missing italian title in Freebase
      • Missing in my Freebase copy
      • Some intermediate folder (about 15)
    • I also got a severe BUG in selecting new matches: https://github.com/fadmaa/grefine-rdf-extension/issues/82 (grrrr)
    • UPDATE!!!
      There is also a cloud based reconciliation service for freebase, now working also with italian language. It should be included in open refine but it does not work for this bug https://github.com/OpenRefine/OpenRefine/issues/805. You can make it work creating a new standard one using this address:

      http://reconcile.freebaseapps.com/reconcile 
    • Copy reconciled data in a new column 
    • Exported csv. on raw for example is:
      ./doppiati/1984 , 1984, http://rdf.freebase.com/ns/m.03kp2l
    • Transformed the csv to rdf using a Python script as simple as this  using python-rdflibsudo apt-get install python-pip (ubuntu)
      sudo pip install rdflib
      To use in this way:
      python myMoviesToRDF.py myMovies-csv.csv myMovies.ttl
    • Upload data into my RDF store
    • Enjoy data analysis NOW!
      In the first attemp i used SPARQL queries in order to get the genre ranking, the directors ranking and the director nationality ranking. A first attemp now, some more will come soon!


      An interactive version here



    Monday, 11 November 2013

    Fact checking : Fassina vs wallstreetitalia


    who : viceministro dell'Economia, Stefano Fassina:

     Il taglio delle pensioni d'oro, anche nell'ipotesi di considerare 'd'oro' le pensioni superiori a 3500 euro netti mensili, implica risparmi di alcune centinaia di milioni di euro all'anno".

    source:
    http://www.repubblica.it/politica/2013/11/08/news/reddito_di_cittadinanza_la_proposta_di_grillo_copertura_con_imu_su_immobili_della_chiesa_e_taglio_pensioni_d_oro-70523273/

    who: wallstreetitalia

    Nel 2011, il 5,2% dei pensionati (861mila persone in tutto), che percepisce un assegno mensile superiore ai tremila euro, ha assorbito in tutto 45 miliardi, vale a dire il 17% della spesa previdenziale. Poco meno di quanto sborsato per i 7,3 milioni di italiani, il 44% del totale, il cui reddito non supera i mille euro al mese. In cifre 51 miliardi in tutto, pari al 19,2% della spesa complessiva.

    source: 
    http://www.wallstreetitalia.com/article/1641597/le-pensioni-d-oro-costano-45-miliardi.aspx

    Io:
    Mi sono perso qsa?

    Wednesday, 30 October 2013

    The best open dataset about Cinema: Freebase. aka "how I fight with Freebase in order to push it in my database"


    For my work about Linked Data in the Digital Library domain (and for personal passion in Cinema too) i investigated on how to have the most complete open dataset about movies in my local database.

    After some quality evaluation I identified Freebase as my target.
    These are the steps I used:

    0) Try out freebase export by google. It does not work, their rdf has some syntactic problems.

    1) Discovered :baseKb . Downloaded the most recent version (some month ago) :A copy of :BaseKB Lime derived from the 2012-02-10 Freebase RDF dump obtained using  this tool:
    https://github.com/paulhoule/infovore/wiki/:BaseKB%20Lime
    Thanks paul Houle!
    This tool cleans the original dump and make it loadable in a database.

    2) Realize that it's very big.
    With my test machine (4 cpu and 16 gb RAM) and one of the best RDF triple stores, Virtuoso 7 I could not end the data loading.

    3) Built a Virtuoso script (and some manual iteratio):
    • Load the dataset in parts (4)
    • Created the list of object types related to the Cinema domain, manually from the Freebase website
    • Get the IDs of all the resources related to the cinema domain.
    • Load the dataset in parts 
    • Export all the triples related to the IDs selected before.

    LessonS learned:
    • The list of types is not complete (awards are not there for example)
    • Use a database for this kind of bulk processing is not ideal but it lets me use a tool I am familiar. Alternatives will be the topic of another source.
    • A lot of data are not usefull for me, the latest version of BaseKb is split in parts (see the news) and I could choose for which part to download. It's not free to download (someone has to pay for the big transfert!) but it will be the next step
    Where I use it:
    • Used in this hackathon i organized. Very nice.
    • I will enrich the information of a video library
    • Just started to play with queries
    • Dreaming, some movies reccomendation.

    Monday, 14 October 2013

    And the winner is.. Team Fungo! #hackindustry #hackathon #h-farm

    Abbiamo vintoooooooo!
    H-industry è il nuovo form di hackathon lanciato lo scorso weekend da H-farm Ventures, incubatore aziendale che opera a livello internazionale in ambito Web, Digital e New Media, favorendo lo sviluppo di startup basate su innovativi modelli di business.


    Ho partecipato alla 24 ore questo weekend dedicandomi a un progetto nel settore Automotive che sfrutta componenti di Texa, azienda leader nella costruzione di strumenti per diagnosi e autodiagnosi.

    La nostra idea è stata quella di sfruttare in un'ottica nuova il device fornitoci, considerando i benefici per il guidatore.
    Semplificando, l'idea è quella di legare la musica ascoltata in macchina alla situazione in cui ci troviamo con la nostra quattro ruote.

    Il mio ruolo nel team è stato variegato ma provo a riassumerlo qui:
    • Product design, legato allo sviluppo dell'idea iniziale
    • Software Architect, quindi studio di fattibilità tecnologica e design dell'architettura del sistema
    • Data Scientist: studio dei dati a disposizione e del miglior modo di utilizzarli
    • Presentatore: ho presentato l'idea davanti a un pubblico di un centinaio di persone, tra cui uno dei fondatori di Texa e molti dirigenti di H-Farm
    Il nostro progetto è stato molto apprezzato e.. abbiamo vinto!

    In particolare i punti che mi è parso siano stati più graditi:
    • L'idea originale, laterale rispetto all'usuale settore dell'azienda
    • La solida idea di business elaborata
    • Il design del brand e la conseguente coerenza grafica (Grazie Designer MAI DORMIENTE!)
    • La presentazione che ha mescoltato visioni aziendali a un video introduttivo accattivante e simpatico per il pubblico
    • La completa analisi tecnica
    Più di tutto mi sono divertito e ho conosciuto persone interessanti partendo dal Team Fungo per arrivare ai ragazzi di H-Farm e alle persone di Texa.
    Grazie a tutti!
    Ps. Qui di seguito la presentazione che abbiamo portato

    Monday, 30 September 2013

    The first Hackathon

    During an internal activity in the place where I work i got the opportunity to organize a little hackathon about Open Data and Linked Data.

    The activity was a use case in the "Cinema" domain where we used the Linked Data technologies in order to mix public and private Data with the aim to discover new insight in the Data and imagine new uses of these Graph of data.
    I split the activity in many sub activity. In this way everyone could work on what preferred using the technologies one could.

    Data Sources
    We analyzed the domain and we discovered interesting source of data released as Open Data such as Freebase in RDF and the list of the "Cinema Theatre" of Milan.
    Someone imported this data in the graph and created a sample link between a movie and the place where a person saw it.

    We also tried to think about the private information the graph could contains and we realized that a lot of them are already in digital form, but are in the Social Network silos.
    Someone else discovered this really cool open data release about cultural events in rome. We couldn't find the data at the end, what a pity!

    Data Process
    I described some methodology of simple data import, cleaning and reconciliation before the hackathon.
    Someone in the group used Open Refine and the RDF extension in order to reconciliate a list of movies and places in data from an existing source.
    In this way some movies from a local libray were linked to the freebase dataset.

    Data Store
    I presented a lab instance of the Virtuoso Store and the people used that in order to store and query the graph of data.

    Data Queries
    I presented some basic SPARQL queries useful in order to discover what's inside a triple store.
    Someone edited that in order to get an idea of the Freebase graph and all the data linked to it.
    Some interesting queries about "Genre of movies i am more interested in" were produced.
    People with query skills found SPARQL very similar to SQL

    Data View
    I presented some tools useful in exploring a distribuited graph of data and for make a graph view from the result of a query.
    Someone used lodlive for getting ispired and discover new relationship in the data.
    Someone else used VisualBox in order to build a tag cloud of the genre linked with the movies he had in his library: take a look!




    This basic hackathon has been appreciated by the public. They found the "Hands-on" session helpful in order to understand better the use case of the semantic web technologies and a good annex to the more theoretical courses we organized