Модераторы: Sardar, Aliance
  

Поиск:

Ответ в темуСоздание новой темы Создание опроса
> Редактор Word, Неправильный код HTML 
:(
    Опции темы
t77
  Дата 21.4.2009, 11:28 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



Доброе время суток.
На странице(Update.html) имеется редактор(TEXTAREA), куда может пользователь вносить текст, картинки и так далее.Затем при нажатии на кнопку сохранить вызывается событие(функция), которое заносит все содержимое в CDATA и сохраняет в одной из таблиц базы данных.Так же если необходимо более расширенные возможности редактора, имеется кнопка над ним, при нажатии на которую, открывается редактор WORD, куда уже можно вносить, все то, с чем знаком WORD - картинки, таблицы, текст с разнообразными фонтами и так далее...В редакторе WORD, естественно имеются все существующие инструменты и в добавок к ним приделана кнопка, при нажатии на которую, происходит то же самое сохранение. Тесть все содержимое обертывается в CDATA и затем сохраняется в одной из таблиц базы данных.
Существует также страница(View.html), отображение которой, непосредственно зависит от содержимого редактора на странице(Update.html)...
Так вот если в редакторе на странице(Update.html), пользователь вставляет содержимое посредством COPY-PASTE. Тоесть копирует откуда-то текст и вставляет в редактор WORD. То возникает проблема следующего характера;
Взглянув на код HTML, что содержит редактор, видно открытые и не закрытые теги HTML или тег TR без тега TABLE перед ним...Короче болный балаган в коде HTML, что приводит к неправильному отображению страниц View.html и Update.html.
Проблема в неправильном коде HTML...
Как можно решить эту проблему?
PM MAIL   Вверх
Aliance
Дата 21.4.2009, 11:37 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


I ♥ <script>
****


Профиль
Группа: Модератор
Сообщений: 6418
Регистрация: 2.8.2004
Где: spb

Репутация: 17
Всего: 137



Если я правильно понял проблему, то нужно при выводе использовать ф-цию htmlentities
PM MAIL WWW ICQ Skype   Вверх
t77
Дата 21.4.2009, 14:27 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



Сылка указывает на примеры использования данной функции в PHP.
Я так понял, что данная функция не сущетвует в JavaScript.
Неужели придется изучать PHP...? В мои планы сейчас это не входило...
Можно ли решить проблему средствами JavaScript, XSL, XHTML ?
PM MAIL   Вверх
Aliance
Дата 21.4.2009, 14:45 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


I ♥ <script>
****


Профиль
Группа: Модератор
Сообщений: 6418
Регистрация: 2.8.2004
Где: spb

Репутация: 17
Всего: 137



Ты же написал, что ты данные хранишь в БД. А значит они от туда читаются с помощью серверного языка. Вот на этапе чтения (либо, еще записи) их и нужно преобразовывать этой функцией.
PM MAIL WWW ICQ Skype   Вверх
t77
Дата 21.4.2009, 17:56 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



Я хочу решить данную проблему на клиенте!
Скажите мне пожалуйста - это возможно ?
вот аналог PHP функции htmlspecialchars на JavaScript, который нашел в сети
Код

function htmlspecialchars(string, quote_style) 
{    
    // *     example 1: htmlspecialchars("<a href='test'>Test</a>", 'ENT_QUOTES');
    // *     returns 1: '<a href='test'>Test</a>'    
    string = string.toString();
    
    // Always encode
    string = string.replace(/&/g, '&');
    string = string.replace(/</g, '<');
    string = string.replace(/>/g, '>');
    
    // Encode depending on quote_style
    if (quote_style == 'ENT_QUOTES') 
    {
        string = string.replace(/"/g, '"');
        string = string.replace(/'/g, ''');
    } 
    else if (quote_style != 'ENT_NOQUOTES') 
    {
        // All other cases (ENT_COMPAT, default, but not ENT_NOQUOTES)
        string = string.replace(/"/g, '"');
    }
    
    return string;
}

У меня она не работает! Браузер ругается на строки, с replace...
Помогите пожалуйста разобратся с функцией, прокоментировав строки.
Я хотел бы понять, что именно она делает ?

PM MAIL   Вверх
Soah
Дата 21.4.2009, 18:25 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 512
Регистрация: 18.2.2009

Репутация: 13
Всего: 54



PM MAIL   Вверх
ksnk
Дата 21.4.2009, 18:47 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


прохожий
****


Профиль
Группа: Комодератор
Сообщений: 6855
Регистрация: 13.4.2007
Где: СПб

Репутация: 48
Всего: 386



Вероятно, правильным ответом будет HTMLTidy Это "чистильщик" html кода. Правда версия для PHP там чегой-то отсутствует (ссылка гнилая), но в принципе, не очень сложно бывает упросить хостера поставить его на сервер и заюзать online-ajax-корректировку.

Можно посмотреть на продвинутые ВИЗИВИГ редакторы. fck editor и mce. У них те-же проблемы с Вордом, что и у всех, они как-то приспособились код чистить smile





--------------------
Человеку свойственно ошибаться, программисту свойственно ошибаться профессионально ! user posted image
PM MAIL WWW Skype   Вверх
Aliance
Дата 22.4.2009, 08:22 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


I ♥ <script>
****


Профиль
Группа: Модератор
Сообщений: 6418
Регистрация: 2.8.2004
Где: spb

Репутация: 17
Всего: 137



На клиенте такие задачи не должны решаться, иначе это может сыграть с автором злую шутку smile
Советую, во-первых, использовать упомянутый выше FCKEditor, а во-вторых, если проблема сохранится, перед записью данных в БД, предварительно их преобразовывать.
PM MAIL WWW ICQ Skype   Вверх
t77
Дата 22.4.2009, 10:50 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



Следовав вышеприведенным советам, пришел к выводу, что FCKeditor лучшее решение! smile 
Большое спасибо всем!
PM MAIL   Вверх
t77
Дата 26.4.2009, 14:50 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



Следуя вышеприведенным советам, скачал и испробовал FCKEditor.
Он действительно справляется с данной проблемой и правильно отображает HTML код. smile 
Но, по некоторым причинам, нет возможности заменить существующий редактор Word на FCKEditor...
Приходится искать альтернативние решения данной проблемы...
Например, как было предложено выше, пробую чистить код перед сохранением в базе данных. 
Но ничего не выходит.  smile 
На данном этапе решил найти и воспользоваться той самой функцией, что чистит HTML код в FCKEditor-e...
Вроде нашел:
Код

// This function will be called from the PasteFromWord dialog (fck_paste.html)
// Input: oNode a DOM node that contains the raw paste from the clipboard
// bIgnoreFont, bRemoveStyles booleans according to the values set in the dialog
// Output: the cleaned string
function CleanWord( oNode, bIgnoreFont, bRemoveStyles )
{
    var html = oNode.innerHTML ;

    html = html.replace(/<o:p>\s*<\/o:p>/g, '') ;
    html = html.replace(/<o:p>[\s\S]*?<\/o:p>/g, '&nbsp;') ;

    // Remove mso-xxx styles.
    html = html.replace( /\s*mso-[^:]+:[^;"]+;?/gi, '' ) ;

    // Remove margin styles.
    html = html.replace( /\s*MARGIN: 0cm 0cm 0pt\s*;/gi, '' ) ;
    html = html.replace( /\s*MARGIN: 0cm 0cm 0pt\s*"/gi, "\"" ) ;

    html = html.replace( /\s*TEXT-INDENT: 0cm\s*;/gi, '' ) ;
    html = html.replace( /\s*TEXT-INDENT: 0cm\s*"/gi, "\"" ) ;

    html = html.replace( /\s*TEXT-ALIGN: [^\s;]+;?"/gi, "\"" ) ;

    html = html.replace( /\s*PAGE-BREAK-BEFORE: [^\s;]+;?"/gi, "\"" ) ;

    html = html.replace( /\s*FONT-VARIANT: [^\s;]+;?"/gi, "\"" ) ;

    html = html.replace( /\s*tab-stops:[^;"]*;?/gi, '' ) ;
    html = html.replace( /\s*tab-stops:[^"]*/gi, '' ) ;

    // Remove FONT face attributes.
    if ( bIgnoreFont )
    {
        html = html.replace( /\s*face="[^"]*"/gi, '' ) ;
        html = html.replace( /\s*face=[^ >]*/gi, '' ) ;

        html = html.replace( /\s*FONT-FAMILY:[^;"]*;?/gi, '' ) ;
    }

    // Remove Class attributes
    html = html.replace(/<(\w[^>]*) class=([^ |>]*)([^>]*)/gi, "<$1$3") ;

    // Remove styles.
    if ( bRemoveStyles )
        html = html.replace( /<(\w[^>]*) style="([^\"]*)"([^>]*)/gi, "<$1$3" ) ;

    // Remove style, meta and link tags
    html = html.replace( /<STYLE[^>]*>[\s\S]*?<\/STYLE[^>]*>/gi, '' ) ;
    html = html.replace( /<(?:META|LINK)[^>]*>\s*/gi, '' ) ;

    // Remove empty styles.
    html =  html.replace( /\s*style="\s*"/gi, '' ) ;

    html = html.replace( /<SPAN\s*[^>]*>\s*&nbsp;\s*<\/SPAN>/gi, '&nbsp;' ) ;

    html = html.replace( /<SPAN\s*[^>]*><\/SPAN>/gi, '' ) ;

    // Remove Lang attributes
    html = html.replace(/<(\w[^>]*) lang=([^ |>]*)([^>]*)/gi, "<$1$3") ;

    html = html.replace( /<SPAN\s*>([\s\S]*?)<\/SPAN>/gi, '$1' ) ;

    html = html.replace( /<FONT\s*>([\s\S]*?)<\/FONT>/gi, '$1' ) ;

    // Remove XML elements and declarations
    html = html.replace(/<\\?\?xml[^>]*>/gi, '' ) ;

    // Remove w: tags with contents.
    html = html.replace( /<w:[^>]*>[\s\S]*?<\/w:[^>]*>/gi, '' ) ;

    // Remove Tags with XML namespace declarations: <o:p><\/o:p>
    html = html.replace(/<\/?\w+:[^>]*>/gi, '' ) ;

    // Remove comments [SF BUG-1481861].
    html = html.replace(/<\!--[\s\S]*?-->/g, '' ) ;

    html = html.replace( /<(U|I|STRIKE)>&nbsp;<\/\1>/g, '&nbsp;' ) ;

    html = html.replace( /<H\d>\s*<\/H\d>/gi, '' ) ;

    // Remove "display:none" tags.
    html = html.replace( /<(\w+)[^>]*\sstyle="[^"]*DISPLAY\s?:\s?none[\s\S]*?<\/\1>/ig, '' ) ;

    // Remove language tags
    html = html.replace( /<(\w[^>]*) language=([^ |>]*)([^>]*)/gi, "<$1$3") ;

    // Remove onmouseover and onmouseout events (from MS Word comments effect)
    html = html.replace( /<(\w[^>]*) onmouseover="([^\"]*)"([^>]*)/gi, "<$1$3") ;
    html = html.replace( /<(\w[^>]*) onmouseout="([^\"]*)"([^>]*)/gi, "<$1$3") ;

    if ( FCKConfig.CleanWordKeepsStructure )
    {
        // The original <Hn> tag send from Word is something like this: <Hn style="margin-top:0px;margin-bottom:0px">
        html = html.replace( /<H(\d)([^>]*)>/gi, '<h$1>' ) ;

        // Word likes to insert extra <font> tags, when using MSIE. (Wierd).
        html = html.replace( /<(H\d)><FONT[^>]*>([\s\S]*?)<\/FONT><\/\1>/gi, '<$1>$2<\/$1>' );
        html = html.replace( /<(H\d)><EM>([\s\S]*?)<\/EM><\/\1>/gi, '<$1>$2<\/$1>' );
    }
    else
    {
        html = html.replace( /<H1([^>]*)>/gi, '<div$1><b><font size="6">' ) ;
        html = html.replace( /<H2([^>]*)>/gi, '<div$1><b><font size="5">' ) ;
        html = html.replace( /<H3([^>]*)>/gi, '<div$1><b><font size="4">' ) ;
        html = html.replace( /<H4([^>]*)>/gi, '<div$1><b><font size="3">' ) ;
        html = html.replace( /<H5([^>]*)>/gi, '<div$1><b><font size="2">' ) ;
        html = html.replace( /<H6([^>]*)>/gi, '<div$1><b><font size="1">' ) ;

        html = html.replace( /<\/H\d>/gi, '<\/font><\/b><\/div>' ) ;

        // Transform <P> to <DIV>
        var re = new RegExp( '(<P)([^>]*>[\\s\\S]*?)(<\/P>)', 'gi' ) ;    // Different because of a IE 5.0 error
        html = html.replace( re, '<div$2<\/div>' ) ;

        // Remove empty tags (three times, just to be sure).
        // This also removes any empty anchor
        html = html.replace( /<([^\s>]+)(\s[^>]*)?>\s*<\/\1>/g, '' ) ;
        html = html.replace( /<([^\s>]+)(\s[^>]*)?>\s*<\/\1>/g, '' ) ;
        html = html.replace( /<([^\s>]+)(\s[^>]*)?>\s*<\/\1>/g, '' ) ;
    }

    return html ;
}


Так как не знаток по регулярным выражениям, мне не совсем понятен данный код! И вообще я не совсем понимаю, что именно нужно очищать после редактора Ворд, чтобы ЧТМЛ код, отображался правильно ?
Помогите пожалуйста разобрать данный код и объясните мне, что теоретически мне необходимо сделать с содержимым редактора Word, перед его сохранением в базе данных.
Помогите почистить код от всякой ненужной гадости после редактора Word! smile 

PM MAIL   Вверх
t77
Дата 26.4.2009, 23:10 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



Покопавшись в исходниках ФСКЕдитора, приходит такая мысль, что то, что было приведено мной выше, не достаточно... 
Если иходить из того, что до сих пор все было сделанно правильно(теоритически), то первое, что необходимо сделать это:
1. Очистка содержимого, редактора Word, от 'мусора':
Код

// This function will be called from the PasteFromWord dialog (fck_paste.html)
// Input: oNode a DOM node that contains the raw paste from the clipboard
// bIgnoreFont, bRemoveStyles booleans according to the values set in the dialog
// Output: the cleaned string
function CleanWord( oNode, bIgnoreFont, bRemoveStyles )
{
    var html = oNode.innerHTML ;

    html = html.replace(/<o:p>\s*<\/o:p>/g, '') ;
    html = html.replace(/<o:p>[\s\S]*?<\/o:p>/g, '&nbsp;') ;

    // Remove mso-xxx styles.
    html = html.replace( /\s*mso-[^:]+:[^;"]+;?/gi, '' ) ;

    // Remove margin styles.
    html = html.replace( /\s*MARGIN: 0cm 0cm 0pt\s*;/gi, '' ) ;
    html = html.replace( /\s*MARGIN: 0cm 0cm 0pt\s*"/gi, "\"" ) ;

    html = html.replace( /\s*TEXT-INDENT: 0cm\s*;/gi, '' ) ;
    html = html.replace( /\s*TEXT-INDENT: 0cm\s*"/gi, "\"" ) ;

    html = html.replace( /\s*TEXT-ALIGN: [^\s;]+;?"/gi, "\"" ) ;

    html = html.replace( /\s*PAGE-BREAK-BEFORE: [^\s;]+;?"/gi, "\"" ) ;

    html = html.replace( /\s*FONT-VARIANT: [^\s;]+;?"/gi, "\"" ) ;

    html = html.replace( /\s*tab-stops:[^;"]*;?/gi, '' ) ;
    html = html.replace( /\s*tab-stops:[^"]*/gi, '' ) ;

    // Remove FONT face attributes.
    if ( bIgnoreFont )
    {
        html = html.replace( /\s*face="[^"]*"/gi, '' ) ;
        html = html.replace( /\s*face=[^ >]*/gi, '' ) ;

        html = html.replace( /\s*FONT-FAMILY:[^;"]*;?/gi, '' ) ;
    }

    // Remove Class attributes
    html = html.replace(/<(\w[^>]*) class=([^ |>]*)([^>]*)/gi, "<$1$3") ;

    // Remove styles.
    if ( bRemoveStyles )
        html = html.replace( /<(\w[^>]*) style="([^\"]*)"([^>]*)/gi, "<$1$3" ) ;

    // Remove style, meta and link tags
    html = html.replace( /<STYLE[^>]*>[\s\S]*?<\/STYLE[^>]*>/gi, '' ) ;
    html = html.replace( /<(?:META|LINK)[^>]*>\s*/gi, '' ) ;

    // Remove empty styles.
    html =  html.replace( /\s*style="\s*"/gi, '' ) ;

    html = html.replace( /<SPAN\s*[^>]*>\s*&nbsp;\s*<\/SPAN>/gi, '&nbsp;' ) ;

    html = html.replace( /<SPAN\s*[^>]*><\/SPAN>/gi, '' ) ;

    // Remove Lang attributes
    html = html.replace(/<(\w[^>]*) lang=([^ |>]*)([^>]*)/gi, "<$1$3") ;

    html = html.replace( /<SPAN\s*>([\s\S]*?)<\/SPAN>/gi, '$1' ) ;

    html = html.replace( /<FONT\s*>([\s\S]*?)<\/FONT>/gi, '$1' ) ;

    // Remove XML elements and declarations
    html = html.replace(/<\\?\?xml[^>]*>/gi, '' ) ;

    // Remove w: tags with contents.
    html = html.replace( /<w:[^>]*>[\s\S]*?<\/w:[^>]*>/gi, '' ) ;

    // Remove Tags with XML namespace declarations: <o:p><\/o:p>
    html = html.replace(/<\/?\w+:[^>]*>/gi, '' ) ;

    // Remove comments [SF BUG-1481861].
    html = html.replace(/<\!--[\s\S]*?-->/g, '' ) ;

    html = html.replace( /<(U|I|STRIKE)>&nbsp;<\/\1>/g, '&nbsp;' ) ;

    html = html.replace( /<H\d>\s*<\/H\d>/gi, '' ) ;

    // Remove "display:none" tags.
    html = html.replace( /<(\w+)[^>]*\sstyle="[^"]*DISPLAY\s?:\s?none[\s\S]*?<\/\1>/ig, '' ) ;

    // Remove language tags
    html = html.replace( /<(\w[^>]*) language=([^ |>]*)([^>]*)/gi, "<$1$3") ;

    // Remove onmouseover and onmouseout events (from MS Word comments effect)
    html = html.replace( /<(\w[^>]*) onmouseover="([^\"]*)"([^>]*)/gi, "<$1$3") ;
    html = html.replace( /<(\w[^>]*) onmouseout="([^\"]*)"([^>]*)/gi, "<$1$3") ;

    if ( FCKConfig.CleanWordKeepsStructure )
    {
        // The original <Hn> tag send from Word is something like this: <Hn style="margin-top:0px;margin-bottom:0px">
        html = html.replace( /<H(\d)([^>]*)>/gi, '<h$1>' ) ;

        // Word likes to insert extra <font> tags, when using MSIE. (Wierd).
        html = html.replace( /<(H\d)><FONT[^>]*>([\s\S]*?)<\/FONT><\/\1>/gi, '<$1>$2<\/$1>' );
        html = html.replace( /<(H\d)><EM>([\s\S]*?)<\/EM><\/\1>/gi, '<$1>$2<\/$1>' );
    }
    else
    {
        html = html.replace( /<H1([^>]*)>/gi, '<div$1><b><font size="6">' ) ;
        html = html.replace( /<H2([^>]*)>/gi, '<div$1><b><font size="5">' ) ;
        html = html.replace( /<H3([^>]*)>/gi, '<div$1><b><font size="4">' ) ;
        html = html.replace( /<H4([^>]*)>/gi, '<div$1><b><font size="3">' ) ;
        html = html.replace( /<H5([^>]*)>/gi, '<div$1><b><font size="2">' ) ;
        html = html.replace( /<H6([^>]*)>/gi, '<div$1><b><font size="1">' ) ;

        html = html.replace( /<\/H\d>/gi, '<\/font><\/b><\/div>' ) ;

        // Transform <P> to <DIV>
        var re = new RegExp( '(<P)([^>]*>[\\s\\S]*?)(<\/P>)', 'gi' ) ;    // Different because of a IE 5.0 error
        html = html.replace( re, '<div$2<\/div>' ) ;

        // Remove empty tags (three times, just to be sure).
        // This also removes any empty anchor
        html = html.replace( /<([^\s>]+)(\s[^>]*)?>\s*<\/\1>/g, '' ) ;
        html = html.replace( /<([^\s>]+)(\s[^>]*)?>\s*<\/\1>/g, '' ) ;
        html = html.replace( /<([^\s>]+)(\s[^>]*)?>\s*<\/\1>/g, '' ) ;
    }

    return html ;
}


Второй шаг будет:
2. Перевод, полученого результата, в правильный xhtml код, по всем стандартам.
(закрыть, если это необходимо открытый тег и все такое.)
Код

var FCKXHtml = new Object() ;

FCKXHtml.CurrentJobNum = 0 ;

FCKXHtml.GetXHTML = function( node, includeNode, format )
{
    FCKDomTools.CheckAndRemovePaddingNode( FCKTools.GetElementDocument( node ), FCKConfig.EnterMode ) ;
    FCKXHtmlEntities.Initialize() ;

    // Set the correct entity to use for empty blocks.
    this._NbspEntity = ( FCKConfig.ProcessHTMLEntities? 'nbsp' : '#160' ) ;

    // Save the current IsDirty state. The XHTML processor may change the
    // original HTML, dirtying it.
    var bIsDirty = FCK.IsDirty() ;

    // Special blocks are blocks of content that remain untouched during the
    // process. It is used for SCRIPTs and STYLEs.
    FCKXHtml.SpecialBlocks = new Array() ;

    // Create the XML DOMDocument object.
    this.XML = FCKTools.CreateXmlObject( 'DOMDocument' ) ;

    // Add a root element that holds all child nodes.
    this.MainNode = this.XML.appendChild( this.XML.createElement( 'xhtml' ) ) ;

    FCKXHtml.CurrentJobNum++ ;

//    var dTimer = new Date() ;

    if ( includeNode )
        this._AppendNode( this.MainNode, node ) ;
    else
        this._AppendChildNodes( this.MainNode, node, false ) ;

    // Get the resulting XHTML as a string.
    var sXHTML = this._GetMainXmlString() ;

//    alert( 'Time: ' + ( ( ( new Date() ) - dTimer ) ) + ' ms' ) ;

    this.XML = null ;

    // Safari adds xmlns="http://www.w3.org/1999/xhtml" to the root node (#963)
    if ( FCKBrowserInfo.IsSafari )
        sXHTML = sXHTML.replace( /^<xhtml.*?>/, '<xhtml>' ) ;

    // Strip the "XHTML" root node.
    sXHTML = sXHTML.substr( 7, sXHTML.length - 15 ).Trim() ;

    // According to the doctype set the proper end for self-closing tags
    // HTML: <br>
    // XHTML: Add a space, like <br/> -> <br />
    if (FCKConfig.DocType.length > 0 && FCKRegexLib.HtmlDocType.test( FCKConfig.DocType ) )
        sXHTML = sXHTML.replace( FCKRegexLib.SpaceNoClose, '>');
    else
        sXHTML = sXHTML.replace( FCKRegexLib.SpaceNoClose, ' />');

    if ( FCKConfig.ForceSimpleAmpersand )
        sXHTML = sXHTML.replace( FCKRegexLib.ForceSimpleAmpersand, '&' ) ;

    if ( format )
        sXHTML = FCKCodeFormatter.Format( sXHTML ) ;

    // Now we put back the SpecialBlocks contents.
    for ( var i = 0 ; i < FCKXHtml.SpecialBlocks.length ; i++ )
    {
        var oRegex = new RegExp( '___FCKsi___' + i ) ;
        sXHTML = sXHTML.replace( oRegex, FCKXHtml.SpecialBlocks[i] ) ;
    }

    // Replace entities marker with the ampersand.
    sXHTML = sXHTML.replace( FCKRegexLib.GeckoEntitiesMarker, '&' ) ;

    // Restore the IsDirty state if it was not dirty.
    if ( !bIsDirty )
        FCK.ResetIsDirty() ;

    FCKDomTools.EnforcePaddingNode( FCKTools.GetElementDocument( node ), FCKConfig.EnterMode ) ;
    return sXHTML ;
}

FCKXHtml._AppendAttribute = function( xmlNode, attributeName, attributeValue )
{
    try
    {
        if ( attributeValue == undefined || attributeValue == null )
            attributeValue = '' ;
        else if ( attributeValue.replace )
        {
            if ( FCKConfig.ForceSimpleAmpersand )
                attributeValue = attributeValue.replace( /&/g, '___FCKAmp___' ) ;

            // Entities must be replaced in the attribute values.
            attributeValue = attributeValue.replace( FCKXHtmlEntities.EntitiesRegex, FCKXHtml_GetEntity ) ;
        }

        // Create the attribute.
        var oXmlAtt = this.XML.createAttribute( attributeName ) ;
        oXmlAtt.value = attributeValue ;

        // Set the attribute in the node.
        xmlNode.attributes.setNamedItem( oXmlAtt ) ;
    }
    catch (e)
    {}
}

FCKXHtml._AppendChildNodes = function( xmlNode, htmlNode, isBlockElement )
{
    var oNode = htmlNode.firstChild ;

    while ( oNode )
    {
        this._AppendNode( xmlNode, oNode ) ;
        oNode = oNode.nextSibling ;
    }

    // Trim block elements. This is also needed to avoid Firefox leaving extra
    // BRs at the end of them.
    if ( isBlockElement && htmlNode.tagName && htmlNode.tagName.toLowerCase() != 'pre' )
    {
        FCKDomTools.TrimNode( xmlNode ) ;

        if ( FCKConfig.FillEmptyBlocks )
        {
            var lastChild = xmlNode.lastChild ;
            if ( lastChild && lastChild.nodeType == 1 && lastChild.nodeName == 'br' )
                this._AppendEntity( xmlNode, this._NbspEntity ) ;
        }
    }

    // If the resulting node is empty.
    if ( xmlNode.childNodes.length == 0 )
    {
        if ( isBlockElement && FCKConfig.FillEmptyBlocks )
        {
            this._AppendEntity( xmlNode, this._NbspEntity ) ;
            return xmlNode ;
        }

        var sNodeName = xmlNode.nodeName ;

        // Some inline elements are required to have something inside (span, strong, etc...).
        if ( FCKListsLib.InlineChildReqElements[ sNodeName ] )
            return null ;

        // We can't use short representation of empty elements that are not marked
        // as empty in th XHTML DTD.
        if ( !FCKListsLib.EmptyElements[ sNodeName ] )
            xmlNode.appendChild( this.XML.createTextNode('') ) ;
    }

    return xmlNode ;
}

FCKXHtml._AppendNode = function( xmlNode, htmlNode )
{
    if ( !htmlNode )
        return false ;

    switch ( htmlNode.nodeType )
    {
        // Element Node.
        case 1 :
            // If we detect a <br> inside a <pre> in Gecko, turn it into a line break instead.
            // This is a workaround for the Gecko bug here: https://bugzilla.mozilla.org/show_bug.cgi?id=92921
            if ( FCKBrowserInfo.IsGecko
                    && htmlNode.tagName.toLowerCase() == 'br'
                    && htmlNode.parentNode.tagName.toLowerCase() == 'pre' )
            {
                var val = '\r' ;
                if ( htmlNode == htmlNode.parentNode.firstChild )
                    val += '\r' ;
                return FCKXHtml._AppendNode( xmlNode, this.XML.createTextNode( val ) ) ;
            }

            // Here we found an element that is not the real element, but a
            // fake one (like the Flash placeholder image), so we must get the real one.
            if ( htmlNode.getAttribute('_fckfakelement') )
                return FCKXHtml._AppendNode( xmlNode, FCK.GetRealElement( htmlNode ) ) ;

            // Ignore bogus BR nodes in the DOM.
            if ( FCKBrowserInfo.IsGecko &&
                    ( htmlNode.hasAttribute('_moz_editor_bogus_node') || htmlNode.getAttribute( 'type' ) == '_moz' ) )
            {
                if ( htmlNode.nextSibling )
                    return false ;
                else
                {
                    htmlNode.removeAttribute( '_moz_editor_bogus_node' ) ;
                    htmlNode.removeAttribute( 'type' ) ;
                }
            }

            // This is for elements that are instrumental to FCKeditor and
            // must be removed from the final HTML.
            if ( htmlNode.getAttribute('_fcktemp') )
                return false ;

            // Get the element name.
            var sNodeName = htmlNode.tagName.toLowerCase()  ;

            if ( FCKBrowserInfo.IsIE )
            {
                // IE doens't include the scope name in the nodeName. So, add the namespace.
                if ( htmlNode.scopeName && htmlNode.scopeName != 'HTML' && htmlNode.scopeName != 'FCK' )
                    sNodeName = htmlNode.scopeName.toLowerCase() + ':' + sNodeName ;
            }
            else
            {
                if ( sNodeName.StartsWith( 'fck:' ) )
                    sNodeName = sNodeName.Remove( 0,4 ) ;
            }

            // Check if the node name is valid, otherwise ignore this tag.
            // If the nodeName starts with a slash, it is a orphan closing tag.
            // On some strange cases, the nodeName is empty, even if the node exists.
            if ( !FCKRegexLib.ElementName.test( sNodeName ) )
                return false ;

            // The already processed nodes must be marked to avoid then to be duplicated (bad formatted HTML).
            // So here, the "mark" is checked... if the element is Ok, then mark it.
            if ( htmlNode._fckxhtmljob && htmlNode._fckxhtmljob == FCKXHtml.CurrentJobNum )
                return false ;

            var oNode = this.XML.createElement( sNodeName ) ;

            // Add all attributes.
            FCKXHtml._AppendAttributes( xmlNode, htmlNode, oNode, sNodeName ) ;

            htmlNode._fckxhtmljob = FCKXHtml.CurrentJobNum ;

            // Tag specific processing.
            var oTagProcessor = FCKXHtml.TagProcessors[ sNodeName ] ;

            if ( oTagProcessor )
                oNode = oTagProcessor( oNode, htmlNode, xmlNode ) ;
            else
                oNode = this._AppendChildNodes( oNode, htmlNode, Boolean( FCKListsLib.NonEmptyBlockElements[ sNodeName ] ) ) ;

            if ( !oNode )
                return false ;

            xmlNode.appendChild( oNode ) ;

            break ;

        // Text Node.
        case 3 :
            if ( htmlNode.parentNode && htmlNode.parentNode.nodeName.IEquals( 'pre' ) )
                return this._AppendTextNode( xmlNode, htmlNode.nodeValue ) ;
            return this._AppendTextNode( xmlNode, htmlNode.nodeValue.ReplaceNewLineChars(' ') ) ;

        // Comment
        case 8 :
            // IE catches the <!DOTYPE ... > as a comment, but it has no
            // innerHTML, so we can catch it, and ignore it.
            if ( FCKBrowserInfo.IsIE && !htmlNode.innerHTML )
                break ;

            try { xmlNode.appendChild( this.XML.createComment( htmlNode.nodeValue ) ) ; }
            catch (e) { /* Do nothing... probably this is a wrong format comment. */ }
            break ;

        // Unknown Node type.
        default :
            xmlNode.appendChild( this.XML.createComment( "Element not supported - Type: " + htmlNode.nodeType + " Name: " + htmlNode.nodeName ) ) ;
            break ;
    }
    return true ;
}

// Append an item to the SpecialBlocks array and returns the tag to be used.
FCKXHtml._AppendSpecialItem = function( item )
{
    return '___FCKsi___' + ( FCKXHtml.SpecialBlocks.push( item ) - 1 ) ;
}

FCKXHtml._AppendEntity = function( xmlNode, entity )
{
    xmlNode.appendChild( this.XML.createTextNode( '#?-:' + entity + ';' ) ) ;
}

FCKXHtml._AppendTextNode = function( targetNode, textValue )
{
    var bHadText = textValue.length > 0 ;
    if ( bHadText )
        targetNode.appendChild( this.XML.createTextNode( textValue.replace( FCKXHtmlEntities.EntitiesRegex, FCKXHtml_GetEntity ) ) ) ;
    return bHadText ;
}

// Retrieves a entity (internal format) for a given character.
function FCKXHtml_GetEntity( character )
{
    // We cannot simply place the entities in the text, because the XML parser
    // will translate & to &amp;. So we use a temporary marker which is replaced
    // in the end of the processing.
    var sEntity = FCKXHtmlEntities.Entities[ character ] || ( '#' + character.charCodeAt(0) ) ;
    return '#?-:' + sEntity + ';' ;
}

// An object that hold tag specific operations.
FCKXHtml.TagProcessors =
{
    a : function( node, htmlNode )
    {
        // Firefox may create empty tags when deleting the selection in some special cases (SF-BUG 1556878).
        if ( htmlNode.innerHTML.Trim().length == 0 && !htmlNode.name )
            return false ;

        var sSavedUrl = htmlNode.getAttribute( '_fcksavedurl' ) ;
        if ( sSavedUrl != null )
            FCKXHtml._AppendAttribute( node, 'href', sSavedUrl ) ;


        // Anchors with content has been marked with an additional class, now we must remove it.
        if ( FCKBrowserInfo.IsIE )
        {
            // Buggy IE, doesn't copy the name of changed anchors.
            if ( htmlNode.name )
                FCKXHtml._AppendAttribute( node, 'name', htmlNode.name ) ;
        }

        node = FCKXHtml._AppendChildNodes( node, htmlNode, false ) ;

        return node ;
    },

    area : function( node, htmlNode )
    {
        var sSavedUrl = htmlNode.getAttribute( '_fcksavedurl' ) ;
        if ( sSavedUrl != null )
            FCKXHtml._AppendAttribute( node, 'href', sSavedUrl ) ;

        // IE ignores the "COORDS" and "SHAPE" attribute so we must add it manually.
        if ( FCKBrowserInfo.IsIE )
        {
            if ( ! node.attributes.getNamedItem( 'coords' ) )
            {
                var sCoords = htmlNode.getAttribute( 'coords', 2 ) ;
                if ( sCoords && sCoords != '0,0,0' )
                    FCKXHtml._AppendAttribute( node, 'coords', sCoords ) ;
            }

            if ( ! node.attributes.getNamedItem( 'shape' ) )
            {
                var sShape = htmlNode.getAttribute( 'shape', 2 ) ;
                if ( sShape && sShape.length > 0 )
                    FCKXHtml._AppendAttribute( node, 'shape', sShape.toLowerCase() ) ;
            }
        }

        return node ;
    },

    body : function( node, htmlNode )
    {
        node = FCKXHtml._AppendChildNodes( node, htmlNode, false ) ;
        // Remove spellchecker attributes added for Firefox when converting to HTML code (Bug #1351).
        node.removeAttribute( 'spellcheck' ) ;
        return node ;
    },

    // IE loses contents of iframes, and Gecko does give it back HtmlEncoded
    // Note: Opera does lose the content and doesn't provide it in the innerHTML string
    iframe : function( node, htmlNode )
    {
        var sHtml = htmlNode.innerHTML ;

        // Gecko does give back the encoded html
        if ( FCKBrowserInfo.IsGecko )
            sHtml = FCKTools.HTMLDecode( sHtml );

        // Remove the saved urls here as the data won't be processed as nodes
        sHtml = sHtml.replace( /\s_fcksavedurl="[^"]*"/g, '' ) ;

        node.appendChild( FCKXHtml.XML.createTextNode( FCKXHtml._AppendSpecialItem( sHtml ) ) ) ;

        return node ;
    },

    img : function( node, htmlNode )
    {
        // The "ALT" attribute is required in XHTML.
        if ( ! node.attributes.getNamedItem( 'alt' ) )
            FCKXHtml._AppendAttribute( node, 'alt', '' ) ;

        var sSavedUrl = htmlNode.getAttribute( '_fcksavedurl' ) ;
        if ( sSavedUrl != null )
            FCKXHtml._AppendAttribute( node, 'src', sSavedUrl ) ;

        // Bug #768 : If the width and height are defined inline CSS,
        // don't define it again in the HTML attributes.
        if ( htmlNode.style.width )
            node.removeAttribute( 'width' ) ;
        if ( htmlNode.style.height )
            node.removeAttribute( 'height' ) ;

        return node ;
    },

    // Fix orphaned <li> nodes (Bug #503).
    li : function( node, htmlNode, targetNode )
    {
        // If the XML parent node is already a <ul> or <ol>, then add the <li> as usual.
        if ( targetNode.nodeName.IEquals( ['ul', 'ol'] ) )
            return FCKXHtml._AppendChildNodes( node, htmlNode, true ) ;

        var newTarget = FCKXHtml.XML.createElement( 'ul' ) ;

        // Reset the _fckxhtmljob so the HTML node is processed again.
        htmlNode._fckxhtmljob = null ;

        // Loop through all sibling LIs, adding them to the <ul>.
        do
        {
            FCKXHtml._AppendNode( newTarget, htmlNode ) ;

            // Look for the next element following this <li>.
            do
            {
                htmlNode = FCKDomTools.GetNextSibling( htmlNode ) ;

            } while ( htmlNode && htmlNode.nodeType == 3 && htmlNode.nodeValue.Trim().length == 0 )

        }    while ( htmlNode && htmlNode.nodeName.toLowerCase() == 'li' )

        return newTarget ;
    },

    // Fix nested <ul> and <ol>.
    ol : function( node, htmlNode, targetNode )
    {
        if ( htmlNode.innerHTML.Trim().length == 0 )
            return false ;

        var ePSibling = targetNode.lastChild ;

        if ( ePSibling && ePSibling.nodeType == 3 )
            ePSibling = ePSibling.previousSibling ;

        if ( ePSibling && ePSibling.nodeName.toUpperCase() == 'LI' )
        {
            htmlNode._fckxhtmljob = null ;
            FCKXHtml._AppendNode( ePSibling, htmlNode ) ;
            return false ;
        }

        node = FCKXHtml._AppendChildNodes( node, htmlNode ) ;

        return node ;
    },

    pre : function ( node, htmlNode )
    {
        var firstChild = htmlNode.firstChild ;

        if ( firstChild && firstChild.nodeType == 3 )
            node.appendChild( FCKXHtml.XML.createTextNode( FCKXHtml._AppendSpecialItem( '\r\n' ) ) ) ;

        FCKXHtml._AppendChildNodes( node, htmlNode, true ) ;

        return node ;
    },

    script : function( node, htmlNode )
    {
        // The "TYPE" attribute is required in XHTML.
        if ( ! node.attributes.getNamedItem( 'type' ) )
            FCKXHtml._AppendAttribute( node, 'type', 'text/javascript' ) ;

        node.appendChild( FCKXHtml.XML.createTextNode( FCKXHtml._AppendSpecialItem( htmlNode.text ) ) ) ;

        return node ;
    },

    span : function( node, htmlNode )
    {
        // Firefox may create empty tags when deleting the selection in some special cases (SF-BUG 1084404).
        if ( htmlNode.innerHTML.length == 0 )
            return false ;

        node = FCKXHtml._AppendChildNodes( node, htmlNode, false ) ;

        return node ;
    },

    style : function( node, htmlNode )
    {
        // The "TYPE" attribute is required in XHTML.
        if ( ! node.attributes.getNamedItem( 'type' ) )
            FCKXHtml._AppendAttribute( node, 'type', 'text/css' ) ;

        var cssText = htmlNode.innerHTML ;
        if ( FCKBrowserInfo.IsIE )    // Bug #403 : IE always appends a \r\n to the beginning of StyleNode.innerHTML
            cssText = cssText.replace( /^(\r\n|\n|\r)/, '' ) ;

        node.appendChild( FCKXHtml.XML.createTextNode( FCKXHtml._AppendSpecialItem( cssText ) ) ) ;

        return node ;
    },

    title : function( node, htmlNode )
    {
        node.appendChild( FCKXHtml.XML.createTextNode( FCK.EditorDocument.title ) ) ;

        return node ;
    }
} ;

FCKXHtml.TagProcessors.ul = FCKXHtml.TagProcessors.ol ;


Если я правильно мыслю, то после реализации этих 2-ух шагов, должен получиться правильный xhtml-код, который будет нормально отображаться на странице...
Я понимаю, что привел очень много кода smile  и вполне вероятно, много лишнего и ненужного...
P.S. 
Этот код взят из исходников FCKEditor-а и для реализации моей задачи, придеться потратить не мало времени разбираясь в нем.Поэтому хотелось бы узнать мнение гуру - правильно ли рассуждаю ? 
Если, что не так, то прошу поправить. Приветствуются любые замечания по этому вопросу.
Мне очень необходимо ваше мнение и помощь!
Спасибо.
PM MAIL   Вверх
Aliance
Дата 27.4.2009, 09:52 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


I ♥ <script>
****


Профиль
Группа: Модератор
Сообщений: 6418
Регистрация: 2.8.2004
Где: spb

Репутация: 17
Всего: 137



Все правильно. Возможно - что-то еще, не могу так сказать. Но то, что ты сказал - точно. Так что дерзай! ;)
А по поводу регулярок, прочти литературу - станет полегче.
PM MAIL WWW ICQ Skype   Вверх
t77
Дата 27.4.2009, 14:36 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



Содержимое Word находится в Стринге- wordStr. Mне необходимо передать это содержимое в виде первого параметра функции - 
Код

GetXHTML = function(node, includeNode, format )


Никак не могу сообразить, как мне перевести Стринг в тип Node smile , для того, чтобы передать ввиде параметра.

Я пробую делать так:
Код

var xmlDocum = new ActiveXObject("Microsoft.XMLDOM");
xmlDocum.async=false;
xmlDocum.loadXML(wordStr);                                            
var newStr = get_xhtml(xmlDocum);

Но при дебагинге водно? что length, переданного нода, равен нулю!
Что я делаю не так?
Помогите пожалуйста...Как это сделать ?

Добавлено через 2 минуты и 39 секунд
Прошу прощения за опечатку... 
Функция, которой я хочу передать нод:
Код

// Parameters:
// node - dom node to convert
// lang - document lang (need it if whole page converted)
// encoding - document charset (need it if whole page converted)
// need_nl - if true, add \n before a tag if it is in list need_nl_before
// inside_pre - if true, do not change content, as it is inside a <pre>
function get_xhtml(node, lang, encoding, need_nl, inside_pre) 
{
    var i;
    var text = '';
    var children = node.childNodes;
    var child_length = children.length;
    var tag_name;
    var do_nl = need_nl ? true : false;
    var page_mode = true;
    
    for (i = 0; i < child_length; i++) 
    {
        var child = children[i];
        
        switch (child.nodeType) 
        {
            case 1: { //ELEMENT_NODE
                var tag_name = String(child.tagName).toLowerCase();
                
                if (tag_name == '') break;
                
                if (tag_name == 'meta') 
                {
                    var meta_name = String(child.name).toLowerCase();
                    if (meta_name == 'generator') break;
                }
                
                if (!need_nl && tag_name == 'body')  //html fragment mode
                {
                    page_mode = false;
                }
                
                if (tag_name == '!')  //COMMENT_NODE in IE 5.0/5.5
                {
                    //get comment inner text
                    var parts = re_comment.exec(child.text);
                    
                    if (parts) 
                    {
                        //the last char of the comment text must not be a hyphen
                        var inner_text = parts[1];
                        text += fix_comment(inner_text);
                    }
                } 
                else 
                    {
                    if (tag_name == 'html') 
                    {
                        text = '<?xml version="1.0" encoding="'+encoding+'"?>\n<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">\n';
                    }
                    
                    //inset \n to make code more neat
                    if (need_nl_before.indexOf('|'+tag_name+'|') != -1) 
                    {
                        if ((do_nl || text != '') && !inside_pre) text += '\n';
                    } 
                    else 
                    {
                        do_nl = true;
                    }
                    
                    text += '<'+tag_name;
                    
                    //add attributes
                    var attr = child.attributes;
                    var attr_length = attr.length;
                    var attr_value;
                    
                    var attr_lang = false;
                    var attr_xml_lang = false;
                    var attr_xmlns = false;
                    
                    var is_alt_attr = false;
                    
                    for (j = 0; j < attr_length; j++) 
                    {
                        var attr_name = attr[j].nodeName.toLowerCase();
                        
                        if (!attr[j].specified && 
                            (attr_name != 'selected' || !child.selected) && 
                            (attr_name != 'style' || child.style.cssText == '') && 
                            attr_name != 'value') continue; //IE 5.0
                        
                        if (attr_name == '_moz_dirty' || 
                            attr_name == '_moz_resizing' || 
                            tag_name == 'br' && 
                            attr_name == 'type' && 
                            child.getAttribute('type') == '_moz') continue;
                        
                        var valid_attr = true;
                        
                        switch (attr_name) 
                        {
                            case "style":
                                attr_value = child.style.cssText;
                                break;
                            case "class":
                                attr_value = child.className;
                                break;
                            case "http-equiv":
                                attr_value = child.httpEquiv;
                                break;
                            case "noshade": break; //this set of choices will extend
                            case "checked": break;
                            case "selected": break;
                            case "multiple": break;
                            case "nowrap": break;
                            case "disabled": break;
                                attr_value = attr_name;
                                break;
                            default:
                                try 
                                {
                                    attr_value = child.getAttribute(attr_name, 2);
                                } 
                                catch (e) 
                                {
                                    valid_attr = false;
                                }
                                break;
                        }
                        
                        //html tag attribs
                        if (attr_name == 'lang') 
                        {
                            attr_lang = true;
                            attr_value = lang;
                        }
                        if (attr_name == 'xml:lang') 
                        {
                            attr_xml_lang = true;
                            attr_value = lang;
                        }
                        if (attr_name == 'xmlns') attr_xmlns = true;
                        if (valid_attr) 
                        {
                            //value attribute set to "0" is not handled correctly in Mozilla
                            if (!(tag_name == 'li' && attr_name == 'value')) 
                            {
                                text += ' '+attr_name+'="'+fix_attribute(attr_value)+'"';
                            }
                        }
                        
                        if (attr_name == 'alt') is_alt_attr = true;
                    }
                    
                    if (tag_name == 'img' && !is_alt_attr) 
                    {
                        text += ' alt=""';
                    }
                    
                    if (tag_name == 'html') 
                    {
                        if (!attr_lang) text += ' lang="'+lang+'"';
                        if (!attr_xml_lang) text += ' xml:lang="'+lang+'"';
                        if (!attr_xmlns) text += ' xmlns="http://www.w3.org/1999/xhtml"';
                    }
                    
                    if (child.canHaveChildren || child.hasChildNodes())
                    {
                        text += '>';

                        text += get_xhtml(child, lang, encoding, true, inside_pre || tag_name == 'pre' ? true : false);
                        text += '</'+tag_name+'>';
                    } 
                    else 
                        {
                        if (tag_name == 'style' || tag_name == 'title' || tag_name == 'script') 
                        {
                            text += '>';
                            var inner_text;
                            if (tag_name == 'script') 
                            {
                                inner_text = child.text;
                            } 
                            else 
                            {
                                inner_text = child.innerHTML;
                            }
                            
                            if (tag_name == 'style') 
                            {
                                inner_text = String(inner_text).replace(/[\n]+/g,'\n');
                            }
                            
                            text += inner_text+'</'+tag_name+'>';
                        } 
                        else 
                        {
                            text += ' />';
                        }
                    }
                }
                break;
            }
            case 3: { //TEXT_NODE
                if (!inside_pre) //do not change text inside <pre> tag
                { 
                    if (child.nodeValue != '\n') 
                    {
                        text += fix_text(child.nodeValue);
                    }
                } 
                else 
                {
                    text += child.nodeValue;
                }
                break;
            }
            case 8:  //COMMENT_NODE
            {
                text += fix_comment(child.nodeValue);
                break;
            }
            default:
                break;
        }
    }
    
    if (!need_nl && !page_mode)  //delete head and body tags from html fragment
    {
        text = text.replace(/<\/?head>[\n]*/gi, "");
        text = text.replace(/<head \/>[\n]*/gi, "");
        text = text.replace(/<\/?body>[\n]*/gi, "");
    }
    
    return text;
}

//fix inner text of a comment
function fix_comment(text) 
{
    //delete double hyphens from the comment text
    text = text.replace(/--/g, "__");
    
    if(re_hyphen.exec(text))  //last char must not be a hyphen
    {
        text += " ";
    }
    
    return "<!--"+text+"-->";
}

//fix content of a text node
function fix_text(text) 
{
    //convert <,> and & to the corresponding entities
    return String(text).replace(/\n{2,}/g, "\n").replace(/\&/g, "&amp;").replace(/</g, "&lt;").replace(/>/g, "&gt;").replace(/\u00A0/g, "&nbsp;");
}

//fix content of attributes href, src or background
function fix_attribute(text) 
{
    //convert <,>, & and " to the corresponding entities
    return String(text).replace(/\&/g, "&amp;").replace(/</g, "&lt;").replace(/>/g, "&gt;").replace(/\"/g, "&quot;");
}


PM MAIL   Вверх
t77
Дата 27.4.2009, 17:11 (ссылка) | (нет голосов) Загрузка ... Загрузка ... Быстрая цитата Цитата


Опытный
**


Профиль
Группа: Участник
Сообщений: 459
Регистрация: 27.7.2008

Репутация: нет
Всего: нет



На последний свой вопрос нашел следующее решение:
Код

var contentNode = document.createElement('div');
contentNode.innerHTML = wordStr;
                                
//convert html to xhtml...                                
var newStr = get_xhtml(contentNode);


Функция вроде получает нод, как положено.
Теперь остается открытым основной вопрос сконвертировать ЧТМЛ код в ХЧТМЛ. Если кто может помочь буду признателен!
PM MAIL   Вверх
  
Ответ в темуСоздание новой темы Создание опроса
0 Пользователей читают эту тему (0 Гостей и 0 Скрытых Пользователей)
0 Пользователей:
« Предыдущая тема | JavaScript: для новичков | Следующая тема »


 




[ Время генерации скрипта: 0.0731 ]   [ Использовано запросов: 22 ]   [ GZIP включён ]


Реклама на сайте     Информационное спонсорство

 
По вопросам размещения рекламы пишите на vladimir(sobaka)vingrad.ru
Отказ от ответственности     Powered by Invision Power Board(R) 1.3 © 2003  IPS, Inc.