If it won't be simple, it simply won't be. [Hire me, source code] by Miki Tebeka, CEO, 353Solutions
Thursday, May 29, 2008
mikistools
I've placed some of the utilities I use daily at http://code.google.com/p/mikistools/.
Wednesday, May 21, 2008
next_n
Suppose you want to find the next n elements of a stream the matches a predicate.
(I just used it in web scraping with BeautifulSoup to get the next 5 sibling "tr" for a table).
(I just used it in web scraping with BeautifulSoup to get the next 5 sibling "tr" for a table).
#!/usr/bin/env python
from itertools import ifilter, islice
def next_n(items, pred, count):
return islice(ifilter(pred, items), count)
if __name__ == "__main__":
from gmpy import is_prime
from itertools import count
for prime in next_n(count(1), is_prime, 10):
print prime
Will print2(Using gmpy for is_prime)
3
5
7
11
13
17
19
23
29
Thursday, May 15, 2008
Tagcloud

Generating tagcloud (using Mako here, but you can use any other templating system)
tagcloud.cgi
tagcloud.mako
Fading Div
Add a new fading (background) div to your document:
<html>
<body>
</body>
<script>
function set_color(elem) {
var colorstr = elem.fade_color.toString(16).toUpperCase();
/* Pad to 6 digits */
while (colorstr.length < 6) {
colorstr = "0" + colorstr;
}
elem.style.background = "#" + colorstr;
}
function fade(elem, color) {
if (typeof(color) != "undefined") {
elem.fade_color = color;
}
else {
elem.fade_color += 0x001111;
}
set_color(elem);
if (elem.fade_color < 0xFFFFFF) {
setTimeout(function() { fade(elem); }, 200);
}
}
function initialize()
{
var div = document.createElement("div");
div.innerHTML = "I'm Fading";
document.body.appendChild(div);
fade(div, 0xFF0000); /* Red */
}
window.onload = initialize;
</script>
</html>
Wednesday, April 30, 2008
XML RPC File Server
#!/usr/bin/env python
'''Simple file client/server using XML RPC'''
from SimpleXMLRPCServer import SimpleXMLRPCServer
from xmlrpclib import ServerProxy, Error as XMLRPCError
import socket
def get_file(filename):
fo = open(filename, "rb")
try: # When will "with" be here?
return fo.read()
finally:
fo.close()
def main(argv=None):
if argv is None:
import sys
argv = sys.argv
default_port = "3030"
from optparse import OptionParser
parser = OptionParser("usage: %prog [options] [[HOST:]PORT]")
parser.add_option("--get", help="get file", dest="filename",
action="store", default="")
opts, args = parser.parse_args(argv[1:])
if len(args) not in (0, 1):
parser.error("wrong number of arguments") # Will exit
if args:
port = args[0]
else:
port = default_port
if ":" in port:
host, port = port.split(":")
else:
host = "localhost"
try:
port = int(port)
except ValueError:
raise SystemExit("error: bad port - %s" % port)
if opts.filename:
try:
proxy = ServerProxy("http://%s:%s" % (host, port))
print proxy.get_file(opts.filename)
raise SystemExit
except XMLRPCError, e:
error = "error: can't get %s (%s)" % (opts.filename, e.faultString)
raise SystemExit(error)
except socket.error, e:
raise SystemExit("error: can't connect (%s)" % e)
server = SimpleXMLRPCServer(("localhost", port))
server.register_function(get_file)
print "Serving files on port %d" % port
server.serve_forever()
if __name__ == "__main__":
main()
This is a huge security hole, use at your own risk.
Friday, April 18, 2008
web-install
#!/bin/bash
# Do the `./configure && make && sudo make install` dance, given a download URL
if [ $# -ne 1 ]; then
echo "usage: `basename $0` URL"
exit 1
fi
set -e # Fail on errors
url=$1
wget --no-check-certificate $url
archive=`basename $url`
if echo $archive | grep -q .tar.bz2; then
tar -xjf $archive
else
tar -xzf $archive
fi
cd ${archive/.tar*}
if [ -f setup.py ]; then
sudo python setup.py install
else
./configure && make && sudo make install
fi
cd ..
Tuesday, April 15, 2008
Registering URL clicks
Some sites (such as Google), gives you a "trampoline" URL so they can register what you have clicked on. I find it highly annoying since you can't tell where you are going just by hovering above the URL and you can't "copy link location" to a document.
The problem is that these people are just lazy:
Notes:
The problem is that these people are just lazy:
<html>
<body>
<a href="http://pythonwise.blogspot.com"
onclick="jump(this, 1);">Pythonwise</a> knows.
</body>
<script src="jquery.js"></script>
<script>
function jump(url, value)
{
$.post("jump.cgi", {
url: url,
value: value
});
return true;
}
</script>
</html>
Notes:
Wednesday, April 09, 2008
num_checkins
#!/bin/bash
# How many checking I did today?
# Without arguments will default to current directory
svn log -r"{`date +%Y%m%d`}:HEAD" $1 | grep "| $USER |" | wc -l
Thursday, April 03, 2008
FeedMe - A simple web-based RSS reader

A simple web-based RSS reader in less than 100 lines of code.
Using feedparser, jQuery and plain old CGI.
index.html
<html>
<head>
<title>FeedMe - A Minimal Web Based RSS Reader</title>
<link rel="stylesheet" type="text/css" href="feedme.css" />
<link rel="shortcut icon" href="feedme.ico" />
<style>
a {
text-decoration: none;
}
a:hover {
background-color: silver;
}
div.summary {
display: none;
position: absolute;
background: gray;
width: 70%;
font 18px monospace;
border: 1px solid black;
}
</style>
</head>
<body>
<h2>FeedMe - A Minimal Web Based RSS Reader</h2>
<div>
Feed URL: <input type="text" size="80" id="feed_url"/>
<button onclick="refresh_feed();">Load</button>
</div>
<hr />
<div id="items">
<div>
</body>
<script src="jquery.js"></script>
<script>
function refresh_feed() {
var url = $.trim($("#feed_url").val());
if ("" == url) {
return;
}
$("#items").load("feed.cgi", {"url" : url});
/* Update every minute */
setTimeout("refresh_feed();", 1000 * 60);
}
</script>
</html>
feed.cgi
#!/usr/bin/env python
import feedparser
from cgi import FieldStorage, escape
from time import ctime
ENTRY_TEMPLATE = '''
<a href="%(link)s"
onmouseover="$('#%(eid)s').show();"
onmouseout="$('#%(eid)s').hide();"
target="_new"
>
%(title)s
</a> <br />
<div class="summary" id="%(eid)s">
%(summary)s
</div>
'''
def main():
print "Content-type: text/html\n"
form = FieldStorage()
url = form.getvalue("url", "")
if not url:
raise SystemExit("error: not url given")
feed = feedparser.parse(url)
for enum, entry in enumerate(feed.entries):
entry.eid = "entry%d" % enum
try:
html = ENTRY_TEMPLATE % entry
print html
except Exception, e:
# FIXME: Log errors
pass
print "<br />%s" % ctime()
if __name__ == "__main__":
main()
How it works:
- The JavaScript script call loads the output of feed.cgi to the items div
- feed.cgi reads the RSS feed from the given URL and output an HTML fragment
- Hovering over a title will show the entry summary
- setTimeout makes sure we refresh the view every minute
Wednesday, March 26, 2008
httpserve
#!/bin/bash
# Quickly serve files over HTTP
# Miki Tebeka <miki.tebeka@gmail.com>
usage="usage: `basename $0` PATH [PORT]"
if [ $# -ne 1 ] && [ $# -ne 2 ]; then
echo $usage >&2
exit 1
fi
case $1 in
"-h" | "-H" | "--help" ) echo $usage; exit;;
* ) path=$1; port=$2;;
esac
if [ ! -d $path ]; then
echo "error: $path is not a directory" >&2
exit 1
fi
cd $path
python -m SimpleHTTPServer $port
Tuesday, March 18, 2008
unique
def unique(items):
'''Remove duplicate items from a sequence, preserving order
>>> unique([1, 2, 3, 2, 1, 4, 2])
[1, 2, 3, 4]
>>> unique([2, 2, 2, 1, 1, 1])
[2, 1]
>>> unique([1, 2, 3, 4])
[1, 2, 3, 4]
>>> unique([])
[]
'''
seen = set()
def is_new(obj, seen=seen, add=seen.add):
if obj in seen:
return 0
add(obj)
return 1
return filter(is_new, items)
Tuesday, March 04, 2008
Thursday, February 21, 2008
extract-audio
OK, not Python - but sometime bash is a better tool.
#!/bin/bash
# Extract audio from video files
# Uses ffmpeg and lame
# Miki Tebeka <miki.tebeka@gmail.com>
if [ $# -ne 2 ]; then
echo "usage: `basename $0` INPUT_VIDEO OUTPUT_MP3"
exit 1
fi
infile=$1
outfile=$2
if [ ! -f $infile ]; then
echo "error: can't find $infile"
exit 1
fi
if [ -f $outfile ]; then
echo "error: $outfile exists"
exit 1
fi
fifoname=/tmp/encode.$$
mkfifo $fifoname
mplayer -vc null -vo null -ao pcm:fast -ao pcm:file=$fifoname $1&
lame $fifoname $outfile
rm $fifoname
Wednesday, February 20, 2008
pfilter
#!/usr/bin/env python
'''Path filter, to be used in pipes to filter out paths.
* Unix test commands (such as -f can be specified as well)
* {} replaces file name
Examples:
# List only files in current directory
ls -a | pfilter -f
# Find files not versioned in svn
# (why, oh why, does svn *always* return 0?)
find . | pfilter 'test -n "`svn info {} 2>&1 | grep Not`"'
'''
__author__ = "Miki Tebeka <miki.tebeka@gmail.com>"
from os import system
def pfilter(path, command):
'''Filter path according to command'''
if "{}" in command:
command = command.replace("{}", path)
else:
command = "%s %s" % (command, path)
if command.startswith("-"):
command = "test %s" % command
# FIXME: win32 support
command += " 2>&1 > /dev/null"
return system(command) == 0
def main(argv=None):
if argv is None:
import sys
argv = sys.argv
from sys import stdin
from itertools import imap, ifilter
from string import strip
from functools import partial
if len(argv) != 2:
from os.path import basename
from sys import stderr
print >> stderr, "usage: %s COMMAND" % basename(argv[0])
print >> stderr
print >> stderr, __doc__
raise SystemExit(1)
command = argv[1]
# Don't you love functional programming?
for path in ifilter(partial(pfilter, command=command), imap(strip, stdin)):
print path
if __name__ == "__main__":
main()
Tuesday, February 12, 2008
Opening File according to mime type
Most of the modern desktops already have a command line utility to open file according to their mime type (GNOME/gnome-open, OSX/open, Windows/start, XFCE/exo-open, KDE/kfmclient ...)
However, most (all?) of them rely on the file extension, where I needed something to view attachments from mutt. Which passes the file data in stdin.
So, here we go (I call this attview):
However, most (all?) of them rely on the file extension, where I needed something to view attachments from mutt. Which passes the file data in stdin.
So, here we go (I call this attview):
#!/usr/bin/env python
'''View attachment with right application'''
__author__ = "Miki Tebeka <miki.tebeka@gmail.com>"
from os import popen, system
from os.path import isfile
import re
class ViewError(Exception):
pass
def view_attachment(data):
# In the .destop file, the file name is %u or %U
u_sub = re.compile("%u", re.I).sub
FILENAME = "/tmp/attview"
fo = open(FILENAME, "wb")
fo.write(data)
fo.close()
mime_type = popen("file -ib %s" % FILENAME).read().strip()
if ";" in mime_type:
mime_type = mime_type[:mime_type.find(";")]
if mime_type == "application/x-not-regular-file":
raise ViewError("can't guess mime type")
APPS_DIR = "/usr/share/applications"
for line in open("%s/defaults.list" % APPS_DIR):
if line.startswith(mime_type):
mime, appfile = line.strip().split("=")
break
else:
raise ViewError("can't find how to open %s" % mime_type)
appfile = "%s/%s" % (APPS_DIR, appfile)
if not isfile(appfile):
raise ViewError("can't find %s" % appfile)
for line in open(appfile):
line = line.strip()
if line.startswith("Exec"):
key, cmd = line.split("=")
fullcmd = u_sub(FILENAME, cmd)
if fullcmd == cmd:
fullcmd += " %s" % FILENAME
system(fullcmd + "&")
break
else:
raise ViewError("can't find Exec in %s" % appfile)
def main(argv=None):
from sys import stdin
if argv is None:
import sys
argv = sys.argv
from optparse import OptionParser
parser = OptionParser("usage: %prog [FILENAME]")
opts, args = parser.parse_args(argv[1:])
if len(args) not in (0, 1):
parser.error("wrong number of arguments") # Will exit
filename = args[0] if args else "-"
if filename == "-":
data = stdin.read()
else:
try:
data = open(filename, "rb").read()
except IOError, e:
raise SystemExit("error: %s" % e.strerror)
try:
view_attachment(data)
except ViewError, e:
raise SystemExit("error: %s" % e)
if __name__ == "__main__":
main()
Thursday, February 07, 2008
Playing with bits
def mask(size):
'''Mask for `size' bits
>>> mask(3)
7
'''
return (1L << size) - 1
def num2bits(num, width=32):
'''String represntation (in bits) of a number
>>> num2bits(3, 5)
'00011'
'''
s = ""
for bit in range(width - 1, -1, -1):
if num & (1L << bit):
s += "1"
else:
s += "0"
return s
def get_bit(value, bit):
'''Get value of bit
>>> num2bits(5, 5)
'00101'
>>> get_bit(5, 0)
1
>>> get_bit(5, 1)
0
'''
return (value >> bit) & 1
def get_range(value, start, end):
'''Get range of bits
>>> num2bits(5, 5)
'00101'
>>> get_range(5, 0, 1)
1
>>> get_range(5, 1, 2)
2
'''
return (value >> start) & mask(end - start + 1)
def set_bit(num, bit, value):
'''Set bit `bit' in num to `value'
>>> i = 5
>>> set_bit(i, 1, 1)
7
>>> set_bit(i, 0, 0)
4
'''
if value:
return num | (1L << bit)
else:
return num & (~(1L << bit))
def sign_extend(num, size):
'''Sign exten number who is `size' bits wide
>>> sign_extend(5, 2)
1
>>> sign_extend(5, 3)
-3
'''
m = mask(size - 1)
res = num & m
# Positive
if (num & (1L << (size - 1))) == 0:
return res
# Negative, 2's complement
res = ~res
res &= m
res += 1
return -res
Wednesday, February 06, 2008
rotate and stretch
from operator import itemgetter
from itertools import imap, chain, repeat
def rotate(matrix):
'''Rotate matrix 90 degrees'''
def row(row_num):
return map(itemgetter(row_num), matrix)
return map(row, range(len(matrix[0])))
def stretch(items, times):
'''stretch([1,2], 3) -> [1,1,1,2,2,2]'''
return reduce(add, map(lambda item: [item] * times, items), [])
def istretch(items, count):
'''istretch([1,2], 3) -> [1,1,1,2,2,2] (generator)'''
return chain(*imap(lambda i: repeat(i, count), items))
Friday, February 01, 2008
svnfind
#!/usr/bin/env python
# Find paths matching directories in subversion repository
__author__ = "Miki Tebeka <miki.tebeka@gmail.com>"
# TODO:
# * Limit search depth
# * Add option to case [in]sensitive
# * Handling of svn errors
# * Support more of "find" predicates (-type, -and, -mtime ...)
# * Another porject: Pre index (using swish-e ...) and update only from
# changelog
from os import popen
def join(path1, path2):
if not path1.endswith("/"):
path1 += "/"
return "%s%s" % (path1, path2)
def svn_walk(root):
command = "svn ls '%s'" % root
for path in popen(command):
path = join(root, path.strip())
yield path
if path.endswith("/"): # A directory
for subpath in svn_walk(path):
yield subpath
def main(argv=None):
if argv is None:
import sys
argv = sys.argv
import re
from itertools import ifilter
from optparse import OptionParser
parser = OptionParser("usage: %prog PATH EXPR")
opts, args = parser.parse_args(argv[1:])
if len(args) != 2:
parser.error("wrong number of arguments") # Will exit
path, expr = args
try:
pred = re.compile(expr, re.I).search
except re.error:
raise SystemExit("error: bad search expression: %s" % expr)
found = 0
for path in ifilter(pred, svn_walk(path)):
found = 1
print path
if not found:
raise SystemError("error: nothing matched %s" % expr)
if __name__ == "__main__":
main()
Friday, January 18, 2008
Simple Text Summarizer
Comments:
- About 50 lines of code
- Gives reasonable results (try it out)
- tokenize need to be improved much more (better detection, stop words ...)
- split_to_sentences need to be improved much more (handle 3.2, Mr. Smith ...)
- In real life you'll need to "clean" the text (Ads, credits, ...)
Subscribe to:
Posts (Atom)