从HTML标签中删除样式属性

我用正则expression式不太好，但是用PHP，我想从TinyMCE返回的string中的HTML标签中删除style属性。

所以把<p style="...">Text</p>改成简单的<p>Test</p> 。

我怎样才能实现这个像preg_replace()函数？

实用的正则expression式(<[^>]+) style=".*?" 将在一切合理的情况下解决这个问题。应该删除不是第一个被捕获组的匹配部分，如下所示：

 $output = preg_replace('/(<[^>]+) style=".*?"/i', '$1', $input);

匹配一个<后跟一个或多个“不是> ”，直到我们来到space和style="..."部分。 /i甚至使用STYLE="..." 。将此匹配replace为$1 ，这是捕获的组。如果标签不包含style="..." ，它将保持原样。

像这样的东西应该工作（未经testing的代码警告）：

 <?php $html = '<p style="asd">qwe</p><br /><p class="qwe">qweqweqwe</p>'; $domd = new DOMDocument(); libxml_use_internal_errors(true); $domd->loadHTML($html); libxml_use_internal_errors(false); $domx = new DOMXPath($domd); $items = $domx->query("//p[@style]"); foreach($items as $item) { $item->removeAttribute("style"); } echo $domd->saveHTML();

我评论了@Mayerln的function。它确实工作，但DOMDocument真的塞满了编码。这是我的simplehtmldom版本

 function stripAttributes($html,$attribs) { $dom = new simple_html_dom(); $dom->load($html); foreach($attribs as $attrib) foreach($dom->find("*[$attrib]") as $e) $e->$attrib = null; $dom->load($dom->save()); return $dom->save(); }

干得好：

 <?php $html = '<p style="border: 1px solid red;">Test</p>'; echo preg_replace('/<p style="(.+?)">(.+?)<\/p>/i', "<p>$2</p>", $html); ?>

顺便说一下，正如其他人所指出的那样，正则expression式并不是为此而build议的。

我使用这个：

 function strip_word_html($text, $allowed_tags = '<a><ul><li><b><i><sup><sub><em><strong><u><br><br/><br /><p><h2><h3><h4><h5><h6>') { mb_regex_encoding('UTF-8'); //replace MS special characters first $search = array('/&lsquo;/u', '/&rsquo;/u', '/&ldquo;/u', '/&rdquo;/u', '/&mdash;/u'); $replace = array('\'', '\'', '"', '"', '-'); $text = preg_replace($search, $replace, $text); //make sure _all_ html entities are converted to the plain ascii equivalents - it appears //in some MS headers, some html entities are encoded and some aren't //$text = html_entity_decode($text, ENT_QUOTES, 'UTF-8'); //try to strip out any C style comments first, since these, embedded in html comments, seem to //prevent strip_tags from removing html comments (MS Word introduced combination) if(mb_stripos($text, '/*') !== FALSE){ $text = mb_eregi_replace('#/\*.*?\*/#s', '', $text, 'm'); } //introduce a space into any arithmetic expressions that could be caught by strip_tags so that they won't be //'<1' becomes '< 1'(note: somewhat application specific) $text = preg_replace(array('/<([0-9]+)/'), array('< $1'), $text); $text = strip_tags($text, $allowed_tags); //eliminate extraneous whitespace from start and end of line, or anywhere there are two or more spaces, convert it to one $text = preg_replace(array('/^\s\s+/', '/\s\s+$/', '/\s\s+/u'), array('', '', ' '), $text); //strip out inline css and simplify style tags $search = array('#<(strong|b)[^>]*>(.*?)</(strong|b)>#isu', '#<(em|i)[^>]*>(.*?)</(em|i)>#isu', '#<u[^>]*>(.*?)</u>#isu'); $replace = array('<b>$2</b>', '<i>$2</i>', '<u>$1</u>'); $text = preg_replace($search, $replace, $text); //on some of the ?newer MS Word exports, where you get conditionals of the form 'if gte mso 9', etc., it appears //that whatever is in one of the html comments prevents strip_tags from eradicating the html comment that contains //some MS Style Definitions - this last bit gets rid of any leftover comments */ $num_matches = preg_match_all("/\<!--/u", $text, $matches); if($num_matches){ $text = preg_replace('/\<!--(.)*--\>/isu', '', $text); } $text = preg_replace('/(<[^>]+) style=".*?"/i', '$1', $text); return $text; }

除了Lorenzo Marcon的回答：

使用preg_replace来select除了样式属性之外的所有内容：

 $html = preg_replace('/(<p.+?)style=".+?"(>.+?)/i', "$1$2", $html);

你可以处理它的客户端，最简单的将与jQuery。就像是：

 $("#tinyMce p").removeAttr("style");

从HTML标签中删除样式属性

如何从java中的string中删除非数字字符？

使用string格式显示小数点最多2个位置或简单的整数

如何从HTML提取img src，title和alt使用php？

如何使用正则expression式去除尾随空格？

正则expression式的字母数字，至less有1个数字和1个字符

有没有办法将恶意代码放入正则expression式？

正则expression式的名字

哪个正则expression式运算符意味着“不要”匹配这个字符？

Python中lambdaexpression式的赋值

如何在java中实现像'LIKE'运算符的SQL？