Python rich text XSS filter

Foreword: That day I was developing the most critical part of the website-XSS filter, and the goddess suddenly called and said: "That thing is so difficult, stop developing it, come and play at my house!". I hung up the phone with a "pop", trying to make my website leak XSS vulnerability, no way~

Python has gradually become one of the mainstream for web development, but some related third-party modules and libraries are not as many as PHP and Node.js.

For example, the XSS filter component, PHP has the famous "HTML Purifier" (http://htmlpurifier.org/), and the non-famous filter component "XssHtml" (http://phith0n.github.io/XssHtml). Of course, the latter is developed by myself.

A library named "html-purifier" can also be installed under python's pip, but this purifier is quite different from the one under php. This library is responsible for filtering out tags and attributes outside the whitelist in html.

Note that it does not filter XSS, only tags and attributes that are not in the whitelist. In other words, similar JavaScript will not be filtered.

So I had to develop a python xss filter myself and use it in my future python project.

Talk about the specific implementation principle.

1. Parse HTML

To parse HTML, use the HTMLParser class that comes with python. In python2, the name is HTMLParser, in python3 it is called html.parser.

To use HTMLParser, you need your own class to inherit HTMLParser and implement the handle_starttag, handle_startendtag, handle_endtag, handle_data and other methods.

For example, the handle_starttag method is called when entering a tag. When we implement this method, we can get the tag and all attributes attrs being processed at this time.

We can check whether tag and attrs are in the whitelist, and do special processing on some of the special tags and attributes, as follows:

2. Link special handling

Some attributes can be executed with javascript pseudo-protocol, such as href of a and src of embed, so special treatment is needed: judge whether it starts with http|https|ftp://, if not, force it in front Add http://

In this way, against potential XSS injection.

Three, embed special treatment

Embed is a tag for embedding media files such as swf. In theory, sometimes our rich text editor allows inserting flash. But we need to ensure that arbitrary javascript code cannot be executed in the flash, nor can it be allowed to make some HTTP requests (which easily causes CSRF attacks).

So it is mandatory to set allowscriptaccess=never and allownetworking=none of the embed tag:

4. When splicing tags and attributes, prevent double quotation marks from going out and becoming new tags

I once found an XSS vulnerability (CVE-2015-1433) in Roundcube Webmail. The reason was that the double quotes were not filtered when the html tags and attributes were spliced after the whitelist check was completed, which caused the attribute value to go out and become a new one. The attribute name causes XSS.

So I use self.__htmlspecialchars to process the attribute value to prevent it from going out:

Finally, this module is also more convenient to use. The simplest demo is as follows:

import pxfilter
parser = pxfilter.XssHtml()
parser.feed('<html code>')
parser.close()
html = parser.getHtml()
print html

Then modify it according to the instructions in the source code. github project address: https://github.com/phith0n/python-xss-filter

I have built a demo with web.py, welcome to test and submit issues: http://python-xss-filter.leavesongs.com/, you still need some suggestions on functions and safety.

Recommended Posts

Python rich text XSS filter
Python implements text version minesweeper
How to filter numbers in python