Communities

Writing
Writing
Codidact Meta
Codidact Meta
The Great Outdoors
The Great Outdoors
Photography & Video
Photography & Video
Scientific Speculation
Scientific Speculation
Cooking
Cooking
Electrical Engineering
Electrical Engineering
Judaism
Judaism
Languages & Linguistics
Languages & Linguistics
Software Development
Software Development
Mathematics
Mathematics
Christianity
Christianity
Code Golf
Code Golf
Music
Music
Physics
Physics
Linux Systems
Linux Systems
Power Users
Power Users
Tabletop RPGs
Tabletop RPGs
Community Proposals
Community Proposals
tag:snake search within a tag
answers:0 unanswered questions
user:xxxx search by author id
score:0.5 posts with 0.5+ score
"snake oil" exact phrase
votes:4 posts with 4+ votes
created:<1w created < 1 week ago
post_type:xxxx type of post
Search help
Notifications
Mark all as read See all your notifications »
Q&A

Welcome to Software Development on Codidact!

Will you help us build our independent community of developers helping developers? We're small and trying to grow. We welcome questions about all aspects of software development, from design to code to QA and more. Got questions? Got answers? Got code you'd like someone to review? Please join us.

Post History

66%
+2 −0
Q&A Python-compatible regex for Markdown tables

For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this: import re markdown = """| S...

posted 8mo ago by hkotsubo‭  ·  edited 8mo ago by hkotsubo‭

Answer
#5: Post edited by user avatar hkotsubo‭ · 2026-01-15T11:18:45Z (8 months ago)
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents of a cell:
  • - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
  • - `+` means "one or more occurrences"
  • Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
  • Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • The only case where a single substitution will work is when you know the exact number of columns and rows:
  • ```python
  • html = re.sub(
  • r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
  • r'''<table>
  • <tr>
  • <th>\1</th>
  • <th>\2</th>
  • </tr>
  • <tr>
  • <td>\3</td>
  • <td>\4</td>
  • </tr>
  • <tr>
  • <td>\5</td>
  • <td>\6</td>
  • </tr>
  • </table>''',
  • markdown
  • )
  • ```
  • In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
  • ---
  • # Without regex
  • IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
  • One example is the [Python-Markdown](https://pypi.org/project/Markdown/) lib:
  • ```python
  • import markdown
  • table = r"""| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • print(markdown.markdown(table, extensions=['tables']))
  • ```
  • Not only it's more straighforward, it'll also handle corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
  • ---
  • Or, if you want to keep it simple, just break the lines and cells using `split`:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • Which is easier to read and maintain, IMO.
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents of a cell:
  • - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
  • - `+` means "one or more occurrences"
  • Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
  • Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • The only case where a single substitution will work is when you know the exact number of columns and rows:
  • ```python
  • html = re.sub(
  • r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
  • r'''<table>
  • <tr>
  • <th>\1</th>
  • <th>\2</th>
  • </tr>
  • <tr>
  • <td>\3</td>
  • <td>\4</td>
  • </tr>
  • <tr>
  • <td>\5</td>
  • <td>\6</td>
  • </tr>
  • </table>''',
  • markdown
  • )
  • ```
  • In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
  • ---
  • # Without regex
  • IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
  • One example is the [Python-Markdown](https://pypi.org/project/Markdown/) lib:
  • ```python
  • import markdown
  • table = r"""| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • print(markdown.markdown(table, extensions=['tables']))
  • ```
  • Not only it's more straighforward, it'll also handle corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
  • ---
  • Or, if you don't want to use a lib for that, just break the lines and cells using `split`:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • Which is easier than a regex to read and maintain, IMO.
#4: Post edited by user avatar hkotsubo‭ · 2026-01-15T11:16:19Z (8 months ago)
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents of a cell:
  • - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
  • - `+` means "one or more occurrences"
  • Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
  • Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • The only case where a single substitution will work is when you know the exact number of columns and rows:
  • ```python
  • html = re.sub(
  • r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
  • r'''<table>
  • <tr>
  • <th>\1</th>
  • <th>\2</th>
  • </tr>
  • <tr>
  • <td>\3</td>
  • <td>\4</td>
  • </tr>
  • <tr>
  • <td>\5</td>
  • <td>\6</td>
  • </tr>
  • </table>''',
  • markdown
  • )
  • ```
  • In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
  • ---
  • # Without regex
  • IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
  • Or, if you want to keep it simple, just break the lines and cells using `split`:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • Which is easier to read and maintain, IMO.
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents of a cell:
  • - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
  • - `+` means "one or more occurrences"
  • Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
  • Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • The only case where a single substitution will work is when you know the exact number of columns and rows:
  • ```python
  • html = re.sub(
  • r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
  • r'''<table>
  • <tr>
  • <th>\1</th>
  • <th>\2</th>
  • </tr>
  • <tr>
  • <td>\3</td>
  • <td>\4</td>
  • </tr>
  • <tr>
  • <td>\5</td>
  • <td>\6</td>
  • </tr>
  • </table>''',
  • markdown
  • )
  • ```
  • In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
  • ---
  • # Without regex
  • IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
  • One example is the [Python-Markdown](https://pypi.org/project/Markdown/) lib:
  • ```python
  • import markdown
  • table = r"""| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • print(markdown.markdown(table, extensions=['tables']))
  • ```
  • Not only it's more straighforward, it'll also handle corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
  • ---
  • Or, if you want to keep it simple, just break the lines and cells using `split`:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • Which is easier to read and maintain, IMO.
#3: Post edited by user avatar hkotsubo‭ · 2026-01-13T11:57:26Z (8 months ago)
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents of a cell:
  • - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
  • - `+` means "one or more occurrences"
  • Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
  • Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • # Without regex
  • IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
  • Or, if you want to keep it simple, just break the lines and cells using `split`:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents of a cell:
  • - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
  • - `+` means "one or more occurrences"
  • Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
  • Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • The only case where a single substitution will work is when you know the exact number of columns and rows:
  • ```python
  • html = re.sub(
  • r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
  • r'''<table>
  • <tr>
  • <th>\1</th>
  • <th>\2</th>
  • </tr>
  • <tr>
  • <td>\3</td>
  • <td>\4</td>
  • </tr>
  • <tr>
  • <td>\5</td>
  • <td>\6</td>
  • </tr>
  • </table>''',
  • markdown
  • )
  • ```
  • In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
  • ---
  • # Without regex
  • IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
  • Or, if you want to keep it simple, just break the lines and cells using `split`:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • Which is easier to read and maintain, IMO.
#2: Post edited by user avatar hkotsubo‭ · 2026-01-13T11:46:05Z (8 months ago)
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents, which is "one or more characters that are not `|`". Then I remove the spaces, and check if it is a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • But honestly, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
  • Or, if you want to keep it simple, just split the lines and cells:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text.strip()}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
  • ```python
  • import re
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • regex_cell = re.compile(r'\|([^|]+)')
  • cell = 'th'
  • previous = ' '
  • html = '<table>\n<tr>'
  • for match in regex_cell.finditer(markdown):
  • text = match[1].strip(' ')
  • if text == '\n':
  • # new row, unless it's the all-hyphens line or the end of string
  • if previous[0] != '-' and match.span()[1] != len(markdown):
  • html += '</tr>\n<tr>'
  • elif text[0] == '-': # header already rendered, change to td
  • cell = 'td'
  • else:
  • html += f'<{cell}>{text}</{cell}>'
  • previous = text
  • html += '</tr>\n</table>'
  • print(html)
  • ```
  • I used `[^|]+` to get the contents of a cell:
  • - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
  • - `+` means "one or more occurrences"
  • Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
  • Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
  • ---
  • # Without regex
  • IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
  • Or, if you want to keep it simple, just break the lines and cells using `split`:
  • ```python
  • markdown = """| Syntax | Description |
  • | ----------- | ----------- |
  • | Header | Title |
  • | Paragraph | Text |
  • """
  • cell = 'th'
  • html = '<table>\n<tr>'
  • lines = markdown.split('\n')
  • for i, line in enumerate(lines):
  • # ignore line with hyphens, and change cell type from th to td
  • if '----' in line:
  • cell = 'td'
  • continue
  • for text in line.split('|'):
  • text = text.strip()
  • if text:
  • html += f'<{cell}>{text}</{cell}>'
  • if i < len(lines) - 2:
  • html += '</tr>\n<tr>'
  • html += '</tr>\n</table>'
  • print(html)
  • ```
#1: Initial revision by user avatar hkotsubo‭ · 2026-01-12T21:26:20Z (8 months ago)
For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:

```python
import re

markdown = """| Syntax      | Description |
| ----------- | ----------- |
| Header      | Title       |
| Paragraph   | Text        |
"""

regex_cell = re.compile(r'\|([^|]+)')
cell = 'th'
previous = ' '
html = '<table>\n<tr>'
for match in regex_cell.finditer(markdown):
    text = match[1].strip(' ')
    if text == '\n':
        # new row, unless it's the all-hyphens line or the end of string
        if previous[0] != '-' and match.span()[1] != len(markdown):
            html += '</tr>\n<tr>'
    elif text[0] == '-': # header already rendered, change to td
        cell = 'td'
    else:
        html += f'<{cell}>{text}</{cell}>'
    previous = text

html += '</tr>\n</table>'
print(html)
```

I used `[^|]+` to get the contents, which is "one or more characters that are not `|`". Then I remove the spaces, and check if it is a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).

---

But honestly, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.

Or, if you want to keep it simple, just split the lines and cells:

```python
markdown = """| Syntax      | Description |
| ----------- | ----------- |
| Header      | Title       |
| Paragraph   | Text        |
"""

cell = 'th'
html = '<table>\n<tr>'
lines = markdown.split('\n')
for i, line in enumerate(lines):
    # ignore line with hyphens, and change cell type from th to td
    if '----' in line:
        cell = 'td'
        continue

    for text in line.split('|'):
        text = text.strip()
        if text:
            html += f'<{cell}>{text.strip()}</{cell}>'

    if i < len(lines) - 2:
        html += '</tr>\n<tr>'

html += '</tr>\n</table>'
print(html)
```