Welcome to Software Development on Codidact!
Will you help us build our independent community of developers helping developers? We're small and trying to grow. We welcome questions about all aspects of software development, from design to code to QA and more. Got questions? Got answers? Got code you'd like someone to review? Please join us.
Post History
For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this: import re markdown = """| S...
#5: Post edited
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
- I used `[^|]+` to get the contents of a cell:
- - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
- - `+` means "one or more occurrences"
- Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
- Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
- ---
- The only case where a single substitution will work is when you know the exact number of columns and rows:
- ```python
- html = re.sub(
- r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
- r'''<table>
- <tr>
- <th>\1</th>
- <th>\2</th>
- </tr>
- <tr>
- <td>\3</td>
- <td>\4</td>
- </tr>
- <tr>
- <td>\5</td>
- <td>\6</td>
- </tr>
- </table>''',
- markdown
- )
- ```
- In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
- ---
- # Without regex
- IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
- One example is the [Python-Markdown](https://pypi.org/project/Markdown/) lib:
- ```python
- import markdown
- table = r"""| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- print(markdown.markdown(table, extensions=['tables']))
- ```
- Not only it's more straighforward, it'll also handle corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
- ---
Or, if you want to keep it simple, just break the lines and cells using `split`:- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
- html += f'<{cell}>{text}</{cell}>'
- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
- ```
Which is easier to read and maintain, IMO.
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
- I used `[^|]+` to get the contents of a cell:
- - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
- - `+` means "one or more occurrences"
- Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
- Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
- ---
- The only case where a single substitution will work is when you know the exact number of columns and rows:
- ```python
- html = re.sub(
- r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
- r'''<table>
- <tr>
- <th>\1</th>
- <th>\2</th>
- </tr>
- <tr>
- <td>\3</td>
- <td>\4</td>
- </tr>
- <tr>
- <td>\5</td>
- <td>\6</td>
- </tr>
- </table>''',
- markdown
- )
- ```
- In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
- ---
- # Without regex
- IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
- One example is the [Python-Markdown](https://pypi.org/project/Markdown/) lib:
- ```python
- import markdown
- table = r"""| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- print(markdown.markdown(table, extensions=['tables']))
- ```
- Not only it's more straighforward, it'll also handle corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
- ---
- Or, if you don't want to use a lib for that, just break the lines and cells using `split`:
- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
- html += f'<{cell}>{text}</{cell}>'
- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
- ```
- Which is easier than a regex to read and maintain, IMO.
#4: Post edited
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
- I used `[^|]+` to get the contents of a cell:
- - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
- - `+` means "one or more occurrences"
- Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
- Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
- ---
- The only case where a single substitution will work is when you know the exact number of columns and rows:
- ```python
- html = re.sub(
- r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
- r'''<table>
- <tr>
- <th>\1</th>
- <th>\2</th>
- </tr>
- <tr>
- <td>\3</td>
- <td>\4</td>
- </tr>
- <tr>
- <td>\5</td>
- <td>\6</td>
- </tr>
- </table>''',
- markdown
- )
- ```
- In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
- ---
- # Without regex
IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).- Or, if you want to keep it simple, just break the lines and cells using `split`:
- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
- html += f'<{cell}>{text}</{cell}>'
- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
- ```
- Which is easier to read and maintain, IMO.
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
- I used `[^|]+` to get the contents of a cell:
- - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
- - `+` means "one or more occurrences"
- Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
- Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
- ---
- The only case where a single substitution will work is when you know the exact number of columns and rows:
- ```python
- html = re.sub(
- r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
- r'''<table>
- <tr>
- <th>\1</th>
- <th>\2</th>
- </tr>
- <tr>
- <td>\3</td>
- <td>\4</td>
- </tr>
- <tr>
- <td>\5</td>
- <td>\6</td>
- </tr>
- </table>''',
- markdown
- )
- ```
- In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
- ---
- # Without regex
- IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
- One example is the [Python-Markdown](https://pypi.org/project/Markdown/) lib:
- ```python
- import markdown
- table = r"""| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- print(markdown.markdown(table, extensions=['tables']))
- ```
- Not only it's more straighforward, it'll also handle corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
- ---
- Or, if you want to keep it simple, just break the lines and cells using `split`:
- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
- html += f'<{cell}>{text}</{cell}>'
- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
- ```
- Which is easier to read and maintain, IMO.
#3: Post edited
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
- I used `[^|]+` to get the contents of a cell:
- - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
- - `+` means "one or more occurrences"
- Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
- Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
- ---
- # Without regex
- IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
- Or, if you want to keep it simple, just break the lines and cells using `split`:
- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
- html += f'<{cell}>{text}</{cell}>'
- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
```
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
- I used `[^|]+` to get the contents of a cell:
- - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
- - `+` means "one or more occurrences"
- Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
- Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
- ---
- The only case where a single substitution will work is when you know the exact number of columns and rows:
- ```python
- html = re.sub(
- r'\|([^|]+)\|([^|]+)\|\n\|[- ]+\|[- ]+\|\n\|([^|]+)\|([^|]+)\|\n\|([^|]+)\|([^|]+)\|',
- r'''<table>
- <tr>
- <th>\1</th>
- <th>\2</th>
- </tr>
- <tr>
- <td>\3</td>
- <td>\4</td>
- </tr>
- <tr>
- <td>\5</td>
- <td>\6</td>
- </tr>
- </table>''',
- markdown
- )
- ```
- In this case, each cell is a capturing group, but look how complicated and - IMO - unreadable the regex became.
- ---
- # Without regex
- IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
- Or, if you want to keep it simple, just break the lines and cells using `split`:
- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
- html += f'<{cell}>{text}</{cell}>'
- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
- ```
- Which is easier to read and maintain, IMO.
#2: Post edited
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
I used `[^|]+` to get the contents, which is "one or more characters that are not `|`". Then I remove the spaces, and check if it is a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).- ---
But honestly, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.Or, if you want to keep it simple, just split the lines and cells:- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
html += f'<{cell}>{text.strip()}</{cell}>'- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
- ```
- For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
- ```python
- import re
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- regex_cell = re.compile(r'\|([^|]+)')
- cell = 'th'
- previous = ' '
- html = '<table>\n<tr>'
- for match in regex_cell.finditer(markdown):
- text = match[1].strip(' ')
- if text == '\n':
- # new row, unless it's the all-hyphens line or the end of string
- if previous[0] != '-' and match.span()[1] != len(markdown):
- html += '</tr>\n<tr>'
- elif text[0] == '-': # header already rendered, change to td
- cell = 'td'
- else:
- html += f'<{cell}>{text}</{cell}>'
- previous = text
- html += '</tr>\n</table>'
- print(html)
- ```
- I used `[^|]+` to get the contents of a cell:
- - `[^|]` is a [negated character class](https://www.regular-expressions.info/charclass.html#negated), which means "anything that's not inside `[^` and `]`". In this case, it's any character that's not a `|`
- - `+` means "one or more occurrences"
- Therefore, `[^|]+` means "one or more characters that are not `|`". You were using `.*`, which means "zero or more characters" (any characters, including `|`), but I think it's better to restrict the characters to not include `|` itself, and to have at least one (not sure if it'll have empty cells, but anyway).
- Then I remove the spaces using `strip`, and check if it's a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
- ---
- # Without regex
- IMO, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML. It'll handle more corner cases, such as escaped characters (if the text is `Text \| more text`, the `|` is part of the text, not the separator).
- Or, if you want to keep it simple, just break the lines and cells using `split`:
- ```python
- markdown = """| Syntax | Description |
- | ----------- | ----------- |
- | Header | Title |
- | Paragraph | Text |
- """
- cell = 'th'
- html = '<table>\n<tr>'
- lines = markdown.split('\n')
- for i, line in enumerate(lines):
- # ignore line with hyphens, and change cell type from th to td
- if '----' in line:
- cell = 'td'
- continue
- for text in line.split('|'):
- text = text.strip()
- if text:
- html += f'<{cell}>{text}</{cell}>'
- if i < len(lines) - 2:
- html += '</tr>\n<tr>'
- html += '</tr>\n</table>'
- print(html)
- ```
#1: Initial revision
For an arbitrary number of rows and columns, a single regex won't work. You'll need to loop through all cells, replacing them as they are found. Something like this:
```python
import re
markdown = """| Syntax | Description |
| ----------- | ----------- |
| Header | Title |
| Paragraph | Text |
"""
regex_cell = re.compile(r'\|([^|]+)')
cell = 'th'
previous = ' '
html = '<table>\n<tr>'
for match in regex_cell.finditer(markdown):
text = match[1].strip(' ')
if text == '\n':
# new row, unless it's the all-hyphens line or the end of string
if previous[0] != '-' and match.span()[1] != len(markdown):
html += '</tr>\n<tr>'
elif text[0] == '-': # header already rendered, change to td
cell = 'td'
else:
html += f'<{cell}>{text}</{cell}>'
previous = text
html += '</tr>\n</table>'
print(html)
```
I used `[^|]+` to get the contents, which is "one or more characters that are not `|`". Then I remove the spaces, and check if it is a new line (which indicates a new `tr`), a hyphen (so I ignore the whole line) or a text (which becomes a `td` or `th` cell).
---
But honestly, regex is not the best tool for this job. You could search for specialized libraries that convert Markdown to HTML.
Or, if you want to keep it simple, just split the lines and cells:
```python
markdown = """| Syntax | Description |
| ----------- | ----------- |
| Header | Title |
| Paragraph | Text |
"""
cell = 'th'
html = '<table>\n<tr>'
lines = markdown.split('\n')
for i, line in enumerate(lines):
# ignore line with hyphens, and change cell type from th to td
if '----' in line:
cell = 'td'
continue
for text in line.split('|'):
text = text.strip()
if text:
html += f'<{cell}>{text.strip()}</{cell}>'
if i < len(lines) - 2:
html += '</tr>\n<tr>'
html += '</tr>\n</table>'
print(html)
```
