Showing posts with label york. Show all posts
Showing posts with label york. Show all posts

Thursday, September 17, 2015

Apps Used in York's Archaeology Data Service

Following a short presentation about online apps we're looking at at York, Michael Charno got in touch and said..


The following are apps that we use at the Archaeology Data Service:

* Asana [https://asana.com/]: Its a really simple task management app
that enables task allocation, commenting, prioritising, creating
deadlines, etc. Its free for use amongst 10 colleagues, so we've
been fine with it so far.
* New Relic [http://newrelic.com/]: Systems analytics software for
understanding where problems exist in servers/web
applications/interfaces. Obviously more useful for people managing
servers or web applications, so might not be widely useful. However
if the university was going to get a license we'd happily join in!
* Slack [https://slack.com/]: We used the free version but quit after
we found ourselves moving to the 10,000 message limit quickly and
didn't want to purchase a license. We haven't replaced it, but would
certainly start using it again if the university was going to get it.



It's not the first time someone at York has mentioned Asana to me. I went to the tool, logged in with my York account and it tells me that 277 York members are already there (including Dan and Paul from the Web Office). After a quick look, I do like the simplicity of Asana.

Slack is like a twitter for your team application. Anyone else tried it or like it?

Tuesday, January 20, 2015

The Solution: Rendering video onto the inside walls of a 3D room

So after a lot of experimentation, I decided that WebGL was a good way to go ( see an earlier post  about automatically showing videos on a 3D models walls).

I took the video example and simply hacked around, watching where objects move to when I changed values, and then added extra objects, in this case walls.



And it worked! Which is pretty impressive ( I think ) for someone who knows nothing about 3D programming. Here is a live version showing music I loved from the 70s.

Thursday, February 6, 2014

360 Degrees of Tom Smith ( What More Could You Want? )

Yesterday was fun. Sara Perry is planning to use the amazing 3 sixty space in The Hub on Heslington East for a module on museum exhibition design. The 3 sixty is a room in which you can display images on all four walls and play audio. It's quite a big space as you can see. If Sara is 5'10" how big do you reckon that wall is? Anyone?


So, before heading off there, with only a lunch break to spare I decided to get a better idea of what it could do. I downloaded the PowerPoint template file that you can use to create the content that you might display. 

I decided, like a megalomanic to see if I could make a room that was filled with my head in a really, nightmareish and ominous way. So I used the built in camera on my laptop to video myself, slowed it down, added effects and some audio from Sunn O))) and put the videos I made into the Powerpoint.

The video was like this... ( play them both at the same time )... as is meant to be displayed on opposite walls.


... and this ...



We had a few glitches along the way to do with syncing the video playback, but I'm sure they will be easily fixable. 









This was really fun because making the media to drive this room was really, really easy. I'd added textual slides ( with transitions ) and just by being shown at this size felt different somehow. The strangest part is how emotive large images and sound can be. The thing I displayed was put together in a few minutes but it felt great being *in* it, by massive nose and eyes and teeth slowly bearing down on you. It was really OPPRESSIVE, which is what I'd wanted.

I didn't have time to think long and hard about how I might use the four walls. The size would be great for all sorts of shifts in presentation style. I'd like to have a go at creating a David Hockney style video, maybe with 4 iphones, all videoing as they walk along but pointing in different directions.


What if the area of interest in your presentation, slowly swept around like a lighthouse? What if data being represented using units of life-sized human bodies, something that when you are in the space you are already more aware of, perhaps, than if you were watching a regular Powerpoint presentation.

It's given me lots of ideas of what I might do, or what could be done. Makes me want to be a student again, if only for this module.







Wednesday, May 29, 2013

The Next Generation of Aggregation and RSS Readers

I've been thinking ( again ) about RSS readers, what they are, what they are for, where ( if anywhere ) they're going next and lastly, how I might maybe make my own ( on the cheap ) .

What is an RSS reader anyway?

Avoiding the technical specification of what RSS is... it is just a way of collecting explicit subscriptions to news from various sites. It is markedly NOT email.

Along the way, RSS forgot, or were told to forget, they were also aggregators. Aggregators were cousins of RSS readers in that they collected lots of news together and re-published it, normally as a web page. In the olden days ( 2005ish ) there were lots of aggregators. There were tools to make your own aggregators and aggregators were important, in my opinion because they did the hard curatorial work or selecting related news sources and making them available, normally in a format you could subscribe to.

As more and more people came online, producing news, consuming news, many aggregators - who were normally making no money whatsoever - went to the wall. Most desktop-based RSS software tools simply failed under the demands of people who had thousands and thousands of subscriptions. Google Reader, one of the few tools that seemed to handle this scale, slowly wiped out both other RSS readers ( online or desktop based ) AND aggregator tools and sites.


The Next Generation of RSS Readers

The next generation of RSS readers, weren't really RSS readers at all. Tools like
Flipboard and Feedly looked to broaden your news reading reach to bring you both the good stuff and also include your personal connections. One of the big problems with RSS was it's subscribing mechanism, which to this day is way too geeky really. Just explaining to someone what to look for if they wanted to subscribe to a site was a usability nightmare. "It will be called RSS or Atom or Latest Entries" or maybe it's in the source code, oh it might have an orange icon or be called XML".

Removing the "subscription" aspect from news reading kills it dead. The loathsome Summly, is a good example, bringing you the sort of news worse than you find in a free newspaper scattered all over the bus floor.

Feedly tries to extend your news reading range, but in my opinion, does it poorly. One of the BIG problems with new subscriptions is that after a while you get bored of them. Your interests change and migrate and a good news reader tool should allow this to happen naturally and not try to do it, with AI magic, for you.



Where They're All Going Wrong


In an article called Produce Organise Consume ( Jan 2011) I point out the strange division between "bookmarking" or tagging, blogging and reading and that for me they were essentially different aspects of the same thing.

These similar activities start to really come into their own when they start feeding the others. When what you are writing about, or bookmarking starts affecting what you are brought to read, for example.

The hidden glue in these three blobs is Connection. Those connections might be explicit, based on another connection or deduced or part of some fantastic artificial back end joining mysterious pieces together ( although probably not ).




The then 2009ish "holy trinity" of Delicious (for organising), Google Reader(for reading) and Blogger(for writing) is looking like a dead threesome in the water with Delicious foundering and not know what it is anymore, Reader on death row and Blogger slowly dying from neglect in an unknown address.


Production and Consumption have moved to Twitter and Facebook ( Google+ maybe ) and Organisation lurks in Twitter hashtags and trending topics, loosely joining things together. Bringing these three activities together shouldn't be too hard. I don't even think it needs one uber-tool to dominate how they happen ( people are very pernickety about how they work ) but it does need to be thought about.

And it's clear that Google aren't thinking on these simple basic activity lines. Where Google Keep fits into my model is obvious, it's an Organising thing - but it doesn't fit well with the Production or Consumption thing, and it really could. The same is true of Google+ which fails awfully as a tool of Production or Organisation ( it'd be nice to see a page of the hashtags I'd used for example ).



Can I Make My Own Aggregator?

And so, given that the world isn't dancing to my tune, or even in the same beat, I wonder if, on a small scale I could make my own aggregator cum reader cum writing platform cum organiser using parts that already exist. 

There are remarkably few open source tools to run aggregators that I liked, Wp-o-matic and Planet Python being exceptions in that they were easily adaptable into something else, but still far from ideal. To be fair it's years since I looked but I recall trying HUNDREDS of them.

My aggregator would need to be:

  • Small - Not dealing with a huge amount of data
  • Shareable - browsable - subscribeable ( offering search term feeds is great )
  • Accessible via more than the obvious "what's new" interface, including tags and connections etc
  • Run in the cloud
  • Integrate with a Production and Collation process somehow.

The parts that already exist


You've probably guessed that the parts I want to work with are Google Spreadsheets, Google Docs and Apps Script. I don't think I can use a ScriptDB or Fusion table because there would be just too many read/writes to their databases even though they deal with large amounts of data very well.

It would be possible to create an aggregator that collected, say less than a thousand feeds and saved the articles in a Google Spreadsheet. But it's worth paying attention to Google Spreadsheet limits and quotas.

Spreadsheets: 400,000 cells, with a maximum of 256 columns per sheet. 
Number of Tabs: 200 sheets per workbook
Note: The limit on the number of ImportHtml functions per spreadsheet is 50. (from @mhawksey)


So, if I used Google Spreadsheets I might have to have monthly rollover, creating a new spreadsheet for each month. And so that connections between items ( via tags ) worked across months, I might need to render them as HTML and use Google Drive hosting to serve them.

I'm just mulling over if making a very simple aggregator with Google Parts is sensible or not. I would still like to be able to show a TagCloud of news concepts from all the various social media corners of the University of York like the one shown below ( from  of my PPPeoplePPPowered project in 2010 ).



And most interestingly ( to me ) was the ability to connect people and concepts in a network. This was achieved by processing each news article using Open Calais for the concepts contained.



I think what I'm trying to say, that in order to convince people that there is real value in working with blogs ( or whatever means of online connected Production ) you need to "hook them" and show them an explicit example of how everything fits together, but that example can't be one you have chosen, it has to fit their world model.

An real life example I've worn out a little from overuse is when, showing a lecturer how blogging works I suggested they added a tag to their post. They added the tag "ahrc" to a post about working on a funding bid. They clicked the tag and found another lecturer in another department working on the same bid. He said he'd immediately go for a chat and see if there was an opportunity for collaboration.

This stunningly simple example can be made to happen over and over if only we can connect the ephemera of writing and reading and organising from wherever people are doing. I think an aggregator of tweets and wiki edits and Google Site updates and blog posts from York staff would be a start in letting a lot more people into this area of what can seem at times, just too geeky for some.








Thursday, May 9, 2013

Information Freedom Fighting


My eye caught the City of York Council announcing that they publish all the "Freedom of Information" requests as PDFs ( here ).

The sharp-eyed amongst you will spot that the requests are organised by weeks. Each week's requests are stored in a PDF for that week. Each PDF would need clicking through to that week, then clicking through to that page and then downloading separately (using the handy "Download Now" link ) and then reading. The search engine is pretty hopeless and can't just return FOI requests and so gives you hundreds of results for any query.

Organising FOI requests by week is completely ridiculous, almost as ridiculous as ordering them by the number of words used or alphabetically. Now of course it probably makes sense from the point of view of compiling the requests - it sounds like a "once a week" job for somebody, but to then publish them once a week seems madness.

One of my pet hates is information that is made available but totally impossible to use. It's like saying, "Yes, of course you can have all the data we keep on you, we have written it on mist on these eggshells - would you like us to post it to you?". It's exactly like that.

So, one evening, I wondered if I could retrospectively do something more useful with their data. It is open, so why not.

First I made a crawler with ScraperWiki ( what an excellent tool this is ) that follows all those links and grabs the text from the PDFs. I based this on Martin's Scraper ( thanks Martin ).

I then downloaded the data collected as CSV and tried using the free statistical tool Sci2 for stemming and combining the words. Whilst this approach worked, I didn't like the resulting stemmed words, like glaz, because they just look so unfriendly.

Next, I wrote a python script to strip out stopwords like "the" and "where" etc and count the word popularity of each word in the downloaded file,  keeping track of the URL that it came from, and saved it back to a new .csv file.

from string import *
import re, HTMLParser
def get_urls(word):
f = '/Users/tomsmith/Downloads/pdfextractor_1.csv'
lines = open(f).readlines()
urls =[]
for line in lines:
url = line.split(",")[0].strip()
text = lower(line.split(",")[1].strip())
words = text.split(" ")
if word in words:
urls.append( url )
return urls

class MLStripper(HTMLParser.HTMLParser):
def __init__(self):
self.reset()
self.fed = []
def handle_data(self, d):
self.fed.append(d)
def get_fed_data(self):
return ''.join(self.fed)
def strip_tags(html):
#Warning this does all including script and javascript
x = MLStripper()
x.feed(html)
return x.get_fed_data()
def match(s, reg):
p = re.compile(reg, re.IGNORECASE| re.DOTALL)
results = p.findall(s)
return results

f = '/Users/tomsmith/Downloads/pdfextractor_1.csv'
lines = open(f).readlines()
stopwords = open( 'stopwords-en.txt').read().split()
d = {}
words = []
for line in lines:
url = line.split(",")[0].strip()
text = line.split(",")[1].strip()
words = text.split(" ")
for word in words:

word = lower( word )
word = match(word, "[a-z]*")[0]
print word
try:
float(word)
word = ''
int( word )
word = ''
except:
pass

if word != '' and word != '"\r\n' and word !='"' and len(word) > 1:
if word not in stopwords:
print word
try:
d[word] += 1
except:
d[word] = 1



finalFreq = sorted(d.iteritems(), key=lambda t: t[1], reverse=True)
out = open("tagcloud.csv", 'w')
out.write("word,frequency\r")
for item in finalFreq:
urls = get_urls( item[0] )
urls_str = "|".join(urls)

out.write(item[0] + "," + str(item[1]) + "," + urls_str + "\r" )
print item[0], item[1], urls_str

out.close()

I then uploaded it as a Google Spreadsheet and turned it into a Web Application.

Problems along the way


I found that displaying 1,000 words a bit of struggle for jQuery ( maybe I made it wrong ) so ended just showing 250 at a time.

I found, of course, that the tag clouding the word by popularity only revealed that the most popular words were fairly meaningless in this context like"foi" and "february" and "council". To do this properly you also need a list of contextual stopwords... some human intervention.

Another glitch was that occassionally, a FOI request would comprise of massive tables of suppliers ( in a PDF remember ) which would skew the words to be "B&Q" and "Wickes" or "building materials" etc. I had to avoid those.


The result


And here it is. A list of words that you can browse and find direct links to the PDFs from which it came. In the end, it seems that showing 250ish words is easier on the technology and the eye. The result is something that you could maybe browse and find a link to something relevant to you.



https://script.google.com/macros/s/AKfycbxAFOZjPBNnpg60MHzFQQjb2TkTsFUSDP_oPrRdNJDg83e3eCc/exec




Tuesday, April 24, 2012

Scraping The Festival of Ideas, June 2012

I noticed something on Twitter about the University's Festival of Ideas and thought I'd take a look at the events listing. Not long ago, the Web Office used to put microformat information in web pages so that I could easily add events to my calendar... Either they've stopped doing that, or it's stopped working, so I thought how easy would it be to grab the events listed and add them to my (or a separate calendar).

In order to do this, I'd need to...

  1. Scrape the HTML from the web page and find the event data
  2. Connect to Google Calendar and add the events found


Because I like programming in python, the first thing I did was to go get the latest copy of BeautifulSoup, which is a library that is unbelievably handy for scraping data out of HTML and also Google GData which lets me talk to Google Calendar.

I so I began...


import urllib, urlparse, gdata, time, datetime
from bs4 import BeautifulSoup
import atom
import gdata.calendar
import gdata.calendar.service

... and loaded the libraries.  Then I connected to Google Calendar, like this...


print "Connecting to Google Calendar"
calendar_service = gdata.calendar.service.CalendarService()
calendar_service.email = '*********@york.ac.uk'
calendar_service.password = '**********'
calendar_service.source = 'Google-Calendar_Python_Sample-1.0'
calendar_service.ProgrammaticLogin()



 .... then got the web page with the Festival of Ideas events on it like this...


url = 'http://yorkfestivalofideas.com/talks/'
print "reading ", url
u = urllib.urlopen( url )
html = u.read()

... At this point, I knew I wanted to create a separate calendar, so I made one in Google Calendar ( IMPORTANT! Set the timezone of your newly created calendar!!! ). Once I'd done this, I could then find what's called the calendar link which you use to specify which calendar you want events to go into...


def get_my_calendars_url(cal_name):
feed = calendar_service.GetOwnCalendarsFeed()
for i, a_calendar in enumerate(feed.entry):
name = a_calendar.title.text
print i, a_calendar.title.text, a_calendar.link[0].href
if name == cal_name:
return a_calendar.link[0].href


calendar_link = get_my_calendars_url("Festival of Ideas")




So, now I have some HTML with useful information in it and a way of connecting to my chosen calendar... I need to use Beautiful soup to fish out the data I need.  I begin like this...


soup = BeautifulSoup( html )
events = soup.find_all("div", {'class':'event'})


... Now the HTML has been turned into a "soup" which means I can do fancy things with it... like the 2nd line above where I grab any DIV that is of class "event" from code that looks like this..


<div class="event">
<div class="eventdate">
<div class="day">
Thu
</div>
<div class="date">
14
</div>
<div class="month">
Jun
</div>
</div>
<div class="eventdetails">
<p class="eventtitle">
<a href="/talks/2012/frenck/">
Where it all began: The Big Bang
</a>
</p>
<p class="eventteaser">
Professor Carlos Frenk will open this year's York Festival of Ideas with a talk on the biggest metamorphosis of all - that of the universe as a whole, from the simplicity of the Big Bang to the complexity of the universe of galaxies, stars, and the planet on which we live.
</p>
</div>
<div class="clear"></div>


...Once I've got a list of events I can then do this... which finds the title, and the text and the dates and times of the events....



for event in events:
try:
title = event.find('p', {'class':'eventtitle'}).find('a').contents[0].strip()
href = event.find('p', {'class':'eventtitle'}).find('a')['href']
href = urlparse.urljoin(url, href)

#Get the actual page in the href!
u = urllib.urlopen( href )
event_html = u.read()
small_soup = BeautifulSoup(event_html)
start_time = small_soup.find('abbr', {'class':'dtstart'})['title']
st = time.strptime(start_time, "%Y-%m-%dT%H:%M")
end_dt = datetime.datetime(2012, st.tm_mon, st.tm_mday, st.tm_hour+2, 0, 0)
end_time = end_dt.strftime("%Y-%m-%dT%H:%M:%S")
start_time = start_time + ":00" #HACK UG!


teaser = event.find('p', {'class':'eventteaser'}).contents[0].strip()
teaser =  teaser + "\n\n" + href

print "creating event:", title
print create_event(title, teaser, "York, UK", start_time, end_time) 

print "_" * 80
except Exception, err:
print err


.... and the create_event code, which uses that calendar_link mentioned earlier, is...


def create_event( title='A lovely event', 
    content='Some text about it', 
    where='York, UK', start_time=None, end_time=None):


    event = gdata.calendar.CalendarEventEntry()
    event.title = atom.Title(text=title)
    event.content = atom.Content(text=content)
    
    #time_zone = 'Europe/London'
    #event.timezone = gdata.calendar.data.TimeZoneProperty(value=time_zone)
    event.where.append(gdata.calendar.Where(value_string=where))


    if start_time is None:
      # Use current time for the start_time and have the event last 1 hour
      start_time = time.strftime('%Y-%m-%dT%H:%M:%S.000Z', time.gmtime())
      end_time = time.strftime('%Y-%m-%dT%H:%M:%S.000Z', time.gmtime(time.time() + 3600))
    event.when.append(gdata.calendar.When(start_time=start_time, end_time=end_time))


    new_event = calendar_service.InsertEvent(event, calendar_link)


    return new_event



... Putting it all together I got a events that can be displayed in a fairly rubbishy widget ( go to June 2012 to see the events! ) or a calendar that anyone can browse here.

https://www.google.com/calendar/embed?src=york.ac.uk_9d9et5aruukobiaqpgke4n63rk@group.calendar.google.com&ctz=Europe/London&gsessionid=OK









The End Result?


To be honest, presentation isn't Google Calendar's strongpoint is it? It's fugly. It's all about the utility though... and I suppose making sure you get to those events.

I guess my point was, and is, that more of this sort of data should be ending up in places that I can use it, i.e in Google Calendar rather than hiding on a web page somewhere. Maybe this little bit of code will help someone to get their events in a more usable form.



 

© 2013 Klick Dev. All rights resevered.

Back To Top