Page 1 of 1
Is there a way to detect UTF off of a URL?
Posted: 01 Nov 2012 22:06
by SkyFrontier
Title says it all.
I'm wondering on a way XY scripting can detect encoding of a url content without having to readurl() it, then writing a file beforehand so filetype() can do the job.
Thanks.
Re: Is there a way to detect UTF off of a URL?
Posted: 02 Nov 2012 00:09
by Marco
Mmh, let's see if I get this correctly.
The HTML code of
www.xyplorer.com page starts as follows:
Code: Select all
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN" "http://www.w3.org/TR/html4/loose.dtd">
<html><head><title>XYplorer - A Windows File Manager and Explorer Replacement</title>
<meta http-equiv="Content-Type" content="text/html; charset=iso-8859-1">
<meta http-equiv="Content-Language" content="en">
You are interested in getting the info on the third line? If so then no, the only way is reading the HTML code, since the HTTP response (which you can inspect by going here
http://web-sniffer.net/ ) doesn't contain info about that.
Re: Is there a way to detect UTF off of a URL?
Posted: 02 Nov 2012 13:43
by SkyFrontier
Thanks, Marco.
As XY can via writefile()
t: [default] text; auto-detects whether text can be written as ASCII or needs to be written as UNICODE
I thought it should be a way to produce sample of (online or offline) documents written in original sources' encoding as the clipped section may not contain a special character thus not triggering writefile auto-detection.
In case of online sources, writing a local file then using readfile([numbytes]), then writefile, then deleting the source document is so CPU/HDD intensive that I scrapped the whole method.
Detection of ~charset=iso-8859-1~ can be done, but that's not accurate on the scenario I describe, see?
Re: Is there a way to detect UTF off of a URL?
Posted: 02 Nov 2012 13:48
by Marco
Maybe isunicode() can help you? [pag. 348 of the help]
Re: Is there a way to detect UTF off of a URL?
Posted: 02 Nov 2012 13:52
by SkyFrontier
Marco wrote:Maybe isunicode() can help you? [pag. 348 of the help]
It could, but it doesnt support URLs parsing...

Re: Is there a way to detect UTF off of a URL?
Posted: 02 Nov 2012 13:57
by Marco
You could do something like
Code: Select all
$code=readurl("desiredurl.com");
isunicode($code);
and check what isunicode returns. Keep in mind that the would html code must flow somewhere in the code. Even if isunicode could handle urls it still would have to make http requests and then check the whole code. As of now it simply is visible to the user.
Re: Is there a way to detect UTF off of a URL?
Posted: 02 Nov 2012 13:57
by SkyFrontier
:!:A combination of redurl+isunicode may solve the problem, Marco. Ill give this one another try. Thanks!
(dumb me...)
Re: Is there a way to detect UTF off of a URL?
Posted: 02 Nov 2012 13:58
by SkyFrontier
Marco wrote:You could do something like
Code: Select all
$code=readurl("desiredurl.com");
isunicode($code);
and check what isunicode returns. Keep in mind that the would html code must flow somewhere in the code. Even if isunicode could handle urls it still would have to make http requests and then check the whole code. As of now it simply is visible to the user.
Yes, just had the same idea... Will give it a go.