Merge pull request #2467 from opendatalab/release-1.3.11

Release 1.3.11
Merge pull request #2466 from opendatalab/dev
2026-03-27 11:08:32 +07:00 · 2025-05-14 10:33:00 +08:00 · 2025-05-14 10:32:31 +08:00 · 2025-05-14 10:31:55 +08:00 · 2025-05-14 10:30:55 +08:00 · 2025-05-14 10:22:18 +08:00
22 changed files with 435 additions and 69 deletions
--- a/README.md
+++ b/README.md
@@ -48,6 +48,10 @@ Easier to use: Just grab MinerU Desktop. No coding, no login, just a simple inte
 </div>

 # Changelog
+- 2025/04/29 1.3.10 Released
+  - Support for custom formula delimiters can be achieved by modifying the `latex-delimiter-config` item in the `magic-pdf.json` file under the user directory.
+- 2025/04/27 1.3.9 Released  
+  - Optimized the formula parsing function to improve the success rate of formula rendering
 - 2025/04/23 1.3.8 Released
  - The default `ocr` model (`ch`) has been updated to `PP-OCRv4_server_rec_doc` (model update required)
    - `PP-OCRv4_server_rec_doc` is trained on a mix of more Chinese document data and PP-OCR training data, enhancing recognition capabilities for some traditional Chinese characters, Japanese, and special characters. It supports over 15,000 recognizable characters, improving text recognition in documents while also boosting general text recognition.
@@ -349,7 +353,7 @@ There are three different ways to experience MinerU:
    </tr>
    <tr>
        <td colspan="3">Python Version</td>
-        <td colspan="3">>=3.10</td>
+        <td colspan="3">3.10~3.13</td>
    </tr>
    <tr>
        <td colspan="3">Nvidia Driver Version</td>
@@ -359,8 +363,7 @@ There are three different ways to experience MinerU:
    </tr>
    <tr>
        <td colspan="3">CUDA Environment</td>
-        <td>11.8/12.4/12.6/12.8</td>
-        <td>11.8/12.4/12.6/12.8</td>
+        <td colspan="2"><a href="https://pytorch.org/get-started/locally/">Refer to the PyTorch official website</a></td>
        <td>None</td>
    </tr>
    <tr>
@@ -374,7 +377,7 @@ There are three different ways to experience MinerU:
        <td colspan="2">GPU VRAM 6GB or more</td>
        <td colspan="2">All GPUs with Tensor Cores produced from Volta(2017) onwards.<br>
        More than 6GB VRAM </td>
-        <td rowspan="2">apple slicon</td>
+        <td rowspan="2">Apple silicon</td>
    </tr>
 </table>

@@ -391,7 +394,7 @@ Synced with dev branch updates:
 #### 1. Install magic-pdf

 ```bash
-conda create -n mineru 'python>=3.10' -y
+conda create -n mineru 'python=3.12' -y
 conda activate mineru
 pip install -U "magic-pdf[full]"
 ```
--- a/README_zh-CN.md
+++ b/README_zh-CN.md
@@ -47,6 +47,10 @@
 </div>

 # 更新记录
+- 2025/04/29 1.3.10 发布
+  - 支持使用自定义公式标识符，可通过修改用户目录下的`magic-pdf.json`文件中的`latex-delimiter-config`项实现。
+- 2025/04/27 1.3.9 发布
+  - 优化公式解析功能，提升公式渲染的成功率
 - 2025/04/23 1.3.8 发布
  - `ocr`默认模型(`ch`)更新为`PP-OCRv4_server_rec_doc`（需更新模型）
    - `PP-OCRv4_server_rec_doc`是在`PP-OCRv4_server_rec`的基础上，在更多中文文档数据和PP-OCR训练数据的混合数据训练而成，增加了部分繁体字、日文、特殊字符的识别能力，可支持识别的字符为1.5万+，除文档相关的文字识别能力提升外，也同时提升了通用文字的识别能力。
@@ -338,7 +342,7 @@ https://github.com/user-attachments/assets/4bea02c9-6d54-4cd6-97ed-dff14340982c
    </tr>
    <tr>
        <td colspan="3">python版本</td>
-        <td colspan="3">>=3.10</td>
+        <td colspan="3">3.10~3.13</td>
    </tr>
    <tr>
        <td colspan="3">Nvidia Driver 版本</td>
@@ -348,8 +352,7 @@ https://github.com/user-attachments/assets/4bea02c9-6d54-4cd6-97ed-dff14340982c
    </tr>
    <tr>
        <td colspan="3">CUDA环境</td>
-        <td>11.8/12.4/12.6/12.8</td>
-        <td>11.8/12.4/12.6/12.8</td>
+        <td colspan="2"><a href="https://pytorch.org/get-started/locally/">Refer to the PyTorch official website</a></td>
        <td>None</td>
    </tr>
    <tr>
@@ -364,7 +367,7 @@ https://github.com/user-attachments/assets/4bea02c9-6d54-4cd6-97ed-dff14340982c
        <td colspan="2">
        Volta(2017)及之后生产的全部带Tensor Core的GPU <br>
        6G显存及以上</td>
-        <td rowspan="2">apple slicon</td>
+        <td rowspan="2">Apple silicon</td>
    </tr>
 </table>

@@ -384,7 +387,7 @@ https://github.com/user-attachments/assets/4bea02c9-6d54-4cd6-97ed-dff14340982c
 > 最新版本国内镜像源同步可能会有延迟，请耐心等待

 ```bash
-conda create -n mineru 'python>=3.10' -y
+conda create -n mineru 'python=3.12' -y
 conda activate mineru
 pip install -U "magic-pdf[full]" -i https://mirrors.aliyun.com/pypi/simple
 ```
--- a/docker/china/Dockerfile
+++ b/docker/china/Dockerfile
@@ -45,7 +45,7 @@ RUN /bin/bash -c "wget https://gcore.jsdelivr.net/gh/opendatalab/MinerU@master/m
    pip3 install -U magic-pdf[full] -i https://mirrors.aliyun.com/pypi/simple"

 # Download models and update the configuration file
-RUN /bin/bash -c "pip3 install modelscope && \
+RUN /bin/bash -c "pip3 install modelscope -i https://mirrors.aliyun.com/pypi/simple && \
    wget https://gcore.jsdelivr.net/gh/opendatalab/MinerU@master/scripts/download_models.py -O download_models.py && \
    python3 download_models.py && \
    sed -i 's|cpu|cuda|g' /root/magic-pdf.json"
--- a/docs/README_Ubuntu_CUDA_Acceleration_en_US.md
+++ b/docs/README_Ubuntu_CUDA_Acceleration_en_US.md
@@ -54,7 +54,7 @@ In the final step, enter `yes`, close the terminal, and reopen it.
 ### 4. Create an Environment Using Conda

 ```bash
-conda create -n mineru 'python>=3.10' -y
+conda create -n mineru 'python=3.12' -y
 conda activate mineru
 ```

--- a/docs/README_Ubuntu_CUDA_Acceleration_zh_CN.md
+++ b/docs/README_Ubuntu_CUDA_Acceleration_zh_CN.md
@@ -54,7 +54,7 @@ bash Anaconda3-2024.06-1-Linux-x86_64.sh
 ## 4. 使用conda 创建环境

 ```bash
-conda create -n mineru 'python>=3.10' -y
+conda create -n mineru 'python=3.12' -y
 conda activate mineru
 ```

--- a/docs/README_Windows_CUDA_Acceleration_en_US.md
+++ b/docs/README_Windows_CUDA_Acceleration_en_US.md
@@ -2,11 +2,12 @@

 ### 1. Install CUDA and cuDNN

-You need to install a CUDA version that is compatible with torch's requirements. Currently, torch supports CUDA 11.8/12.4/12.6.
+You need to install a CUDA version that is compatible with torch's requirements. For details, please refer to the [official PyTorch website](https://pytorch.org/get-started/locally/).

 - CUDA 11.8 https://developer.nvidia.com/cuda-11-8-0-download-archive
 - CUDA 12.4 https://developer.nvidia.com/cuda-12-4-0-download-archive
 - CUDA 12.6 https://developer.nvidia.com/cuda-12-6-0-download-archive
+- CUDA 12.8 https://developer.nvidia.com/cuda-12-8-0-download-archive

 ### 2. Install Anaconda

@@ -17,7 +18,7 @@ Download link: https://repo.anaconda.com/archive/Anaconda3-2024.06-1-Windows-x86
 ### 3. Create an Environment Using Conda

 ```bash
-conda create -n mineru 'python>=3.10' -y
+conda create -n mineru 'python=3.12' -y
 conda activate mineru
 ```

@@ -63,7 +64,7 @@ If your graphics card has at least 6GB of VRAM, follow these steps to test CUDA-
 1. **Overwrite the installation of torch and torchvision** supporting CUDA.(Please select the appropriate index-url based on your CUDA version. For more details, refer to the [PyTorch official website](https://pytorch.org/get-started/locally/).)

   ```
-   pip install --force-reinstall torch torchvision "numpy<=2.1.1" --index-url https://download.pytorch.org/whl/cu124
+   pip install --force-reinstall torch torchvision --index-url https://download.pytorch.org/whl/cu124
   ```

 2. **Modify the value of `"device-mode"`** in the `magic-pdf.json` configuration file located in your user directory.
--- a/docs/README_Windows_CUDA_Acceleration_zh_CN.md
+++ b/docs/README_Windows_CUDA_Acceleration_zh_CN.md
@@ -1,12 +1,13 @@
 # Windows10/11

-## 1. 安装cuda和cuDNN
+## 1. 安装cuda环境

-需要安装符合torch要求的cuda版本，torch目前支持11.8/12.4/12.6
+需要安装符合torch要求的cuda版本，具体可参考[torch官网](https://pytorch.org/get-started/locally/)

 - CUDA 11.8 https://developer.nvidia.com/cuda-11-8-0-download-archive
 - CUDA 12.4 https://developer.nvidia.com/cuda-12-4-0-download-archive
 - CUDA 12.6 https://developer.nvidia.com/cuda-12-6-0-download-archive
+- CUDA 12.8 https://developer.nvidia.com/cuda-12-8-0-download-archive

 ## 2. 安装anaconda

@@ -18,7 +19,7 @@ https://mirrors.tuna.tsinghua.edu.cn/anaconda/archive/Anaconda3-2024.06-1-Window
 ## 3. 使用conda 创建环境

 ```bash
-conda create -n mineru 'python>=3.10' -y
+conda create -n mineru 'python=3.12' -y
 conda activate mineru
 ```

@@ -64,7 +65,7 @@ pip install -U magic-pdf[full] -i https://mirrors.aliyun.com/pypi/simple
 **1.覆盖安装支持cuda的torch和torchvision**(请根据cuda版本选择合适的index-url，具体可参考[torch官网](https://pytorch.org/get-started/locally/))

 ```bash
-pip install --force-reinstall torch torchvision "numpy<=2.1.1" --index-url https://download.pytorch.org/whl/cu124
+pip install --force-reinstall torch torchvision --index-url https://download.pytorch.org/whl/cu124
 ```

 **2.修改【用户目录】中配置文件magic-pdf.json中"device-mode"的值**
--- a/magic-pdf.template.json
+++ b/magic-pdf.template.json
@@ -20,6 +20,16 @@
        "enable": true,
        "max_time": 400
    },
+    "latex-delimiter-config": {
+        "display": {
+            "left": "$$",
+            "right": "$$"
+        },
+        "inline": {
+            "left": "$",
+            "right": "$"
+        }
+    },
    "llm-aided-config": {
        "formula_aided": {
            "api_key": "your_api_key",
@@ -40,5 +50,5 @@
            "enable": false
        }
    },
-    "config_version": "1.2.0"
+    "config_version": "1.2.1"
 }
--- a/magic_pdf/dict2md/ocr_mkcontent.py
+++ b/magic_pdf/dict2md/ocr_mkcontent.py
@@ -5,6 +5,7 @@ from loguru import logger
 from magic_pdf.config.make_content_config import DropMode, MakeMode
 from magic_pdf.config.ocr_content_type import BlockType, ContentType
 from magic_pdf.libs.commons import join_path
+from magic_pdf.libs.config_reader import get_latex_delimiter_config
 from magic_pdf.libs.language import detect_lang
 from magic_pdf.libs.markdown_utils import ocr_escape_special_markdown_char
 from magic_pdf.post_proc.para_split_v3 import ListLineTag
@@ -145,6 +146,19 @@ def full_to_half(text: str) -> str:
            result.append(char)
    return ''.join(result)

+latex_delimiters_config = get_latex_delimiter_config()
+
+default_delimiters = {
+    'display': {'left': '$$', 'right': '$$'},
+    'inline': {'left': '$', 'right': '$'}
+}
+
+delimiters = latex_delimiters_config if latex_delimiters_config else default_delimiters
+
+display_left_delimiter = delimiters['display']['left']
+display_right_delimiter = delimiters['display']['right']
+inline_left_delimiter = delimiters['inline']['left']
+inline_right_delimiter = delimiters['inline']['right']

 def merge_para_with_text(para_block):
    block_text = ''
@@ -168,9 +182,9 @@ def merge_para_with_text(para_block):
            if span_type == ContentType.Text:
                content = ocr_escape_special_markdown_char(span['content'])
            elif span_type == ContentType.InlineEquation:
-                content = f"${span['content']}$"
+                content = f"{inline_left_delimiter}{span['content']}{inline_right_delimiter}"
            elif span_type == ContentType.InterlineEquation:
-                content = f"\n$$\n{span['content']}\n$$\n"
+                content = f"\n{display_left_delimiter}\n{span['content']}\n{display_right_delimiter}\n"

            content = content.strip()

--- a/magic_pdf/libs/config_reader.py
+++ b/magic_pdf/libs/config_reader.py
@@ -125,6 +125,15 @@ def get_llm_aided_config():
    else:
        return llm_aided_config

+def get_latex_delimiter_config():
+    config = read_config()
+    latex_delimiter_config = config.get('latex-delimiter-config')
+    if latex_delimiter_config is None:
+        logger.warning(f"'latex-delimiter-config' not found in {CONFIG_FILE_NAME}, use 'None' as default")
+        return None
+    else:
+        return latex_delimiter_config
+

 if __name__ == '__main__':
    ak, sk, endpoint = get_s3_config('llm-raw')
--- a/magic_pdf/libs/version.py
+++ b/magic_pdf/libs/version.py
@@ -1 +1 @@
-__version__ = "1.3.8"
+__version__ = "1.3.10"
--- a/magic_pdf/model/doc_analyze_by_custom_model.py
+++ b/magic_pdf/model/doc_analyze_by_custom_model.py
@@ -156,7 +156,10 @@ def doc_analyze(
        batch_images = [images_with_extra_info]

    results = []
-    for batch_image in batch_images:
+    processed_images_count = 0
+    for index, batch_image in enumerate(batch_images):
+        processed_images_count += len(batch_image)
+        logger.info(f'Batch {index + 1}/{len(batch_images)}: {processed_images_count} pages/{len(images_with_extra_info)} pages')
        result = may_batch_image_analyze(batch_image, ocr, show_log,layout_model, formula_enable, table_enable)
        results.extend(result)

--- a/magic_pdf/model/sub_modules/mfr/unimernet/unimernet_hf/modeling_unimernet.py
+++ b/magic_pdf/model/sub_modules/mfr/unimernet/unimernet_hf/modeling_unimernet.py
@@ -5,6 +5,7 @@ from typing import Optional

 import torch
 from ftfy import fix_text
+from loguru import logger

 from transformers import AutoConfig, AutoModel, AutoModelForCausalLM, AutoTokenizer, PretrainedConfig, PreTrainedModel
 from transformers import VisionEncoderDecoderConfig, VisionEncoderDecoderModel
@@ -57,22 +58,322 @@ class TokenizerWrapper:
        return toks


-def latex_rm_whitespace(s: str):
-    """Remove unnecessary whitespace from LaTeX code.
+LEFT_PATTERN = re.compile(r'(\\left)(\S*)')
+RIGHT_PATTERN = re.compile(r'(\\right)(\S*)')
+LEFT_COUNT_PATTERN = re.compile(r'\\left(?![a-zA-Z])')
+RIGHT_COUNT_PATTERN = re.compile(r'\\right(?![a-zA-Z])')
+LEFT_RIGHT_REMOVE_PATTERN = re.compile(r'\\left\.?|\\right\.?')
+
+def fix_latex_left_right(s):
    """
-    text_reg = r'(\\(operatorname|mathrm|text|mathbf)\s?\*? {.*?})'
-    letter = r'[a-zA-Z]'
-    noletter = r'[\W_^\d]'
-    names = [x[0].replace(' ', '') for x in re.findall(text_reg, s)]
-    s = re.sub(text_reg, lambda _: str(names.pop(0)), s)
-    news = s
-    while True:
-        s = news
-        news = re.sub(r'(?!\\ )(%s)\s+?(%s)' % (noletter, noletter), r'\1\2', s)
-        news = re.sub(r'(?!\\ )(%s)\s+?(%s)' % (noletter, letter), r'\1\2', news)
-        news = re.sub(r'(%s)\s+?(%s)' % (letter, noletter), r'\1\2', news)
-        if news == s:
-            break
+    修复LaTeX中的\\left和\\right命令
+    1. 确保它们后面跟有效分隔符
+    2. 平衡\\left和\\right的数量
+    """
+    # 白名单分隔符
+    valid_delims_list = [r'(', r')', r'[', r']', r'{', r'}', r'/', r'|',
+                         r'\{', r'\}', r'\lceil', r'\rceil', r'\lfloor',
+                         r'\rfloor', r'\backslash', r'\uparrow', r'\downarrow',
+                         r'\Uparrow', r'\Downarrow', r'\|', r'\.']
+
+    # 为\left后缺失有效分隔符的情况添加点
+    def fix_delim(match, is_left=True):
+        cmd = match.group(1)  # \left 或 \right
+        rest = match.group(2) if len(match.groups()) > 1 else ""
+        if not rest or rest not in valid_delims_list:
+            return cmd + "."
+        return match.group(0)
+
+    # 使用更精确的模式匹配\left和\right命令
+    # 确保它们是独立的命令，不是其他命令的一部分
+    # 使用预编译正则和统一回调函数
+    s = LEFT_PATTERN.sub(lambda m: fix_delim(m, True), s)
+    s = RIGHT_PATTERN.sub(lambda m: fix_delim(m, False), s)
+
+    # 更精确地计算\left和\right的数量
+    left_count = len(LEFT_COUNT_PATTERN.findall(s))  # 不匹配\lefteqn等
+    right_count = len(RIGHT_COUNT_PATTERN.findall(s))  # 不匹配\rightarrow等
+
+    if left_count == right_count:
+        # 如果数量相等，检查是否在同一组
+        return fix_left_right_pairs(s)
+    else:
+        # 如果数量不等，移除所有\left和\right
+        # logger.debug(f"latex:{s}")
+        # logger.warning(f"left_count: {left_count}, right_count: {right_count}")
+        return LEFT_RIGHT_REMOVE_PATTERN.sub('', s)
+
+
+def fix_left_right_pairs(latex_formula):
+    """
+    检测并修复LaTeX公式中\\left和\\right不在同一组的情况
+
+    Args:
+        latex_formula (str): 输入的LaTeX公式
+
+    Returns:
+        str: 修复后的LaTeX公式
+    """
+    # 用于跟踪花括号嵌套层级
+    brace_stack = []
+    # 用于存储\left信息: (位置, 深度, 分隔符)
+    left_stack = []
+    # 存储需要调整的\right信息: (开始位置, 结束位置, 目标位置)
+    adjustments = []
+
+    i = 0
+    while i < len(latex_formula):
+        # 检查是否是转义字符
+        if i > 0 and latex_formula[i - 1] == '\\':
+            backslash_count = 0
+            j = i - 1
+            while j >= 0 and latex_formula[j] == '\\':
+                backslash_count += 1
+                j -= 1
+
+            if backslash_count % 2 == 1:
+                i += 1
+                continue
+
+        # 检测\left命令
+        if i + 5 < len(latex_formula) and latex_formula[i:i + 5] == "\\left" and i + 5 < len(latex_formula):
+            delimiter = latex_formula[i + 5]
+            left_stack.append((i, len(brace_stack), delimiter))
+            i += 6  # 跳过\left和分隔符
+            continue
+
+        # 检测\right命令
+        elif i + 6 < len(latex_formula) and latex_formula[i:i + 6] == "\\right" and i + 6 < len(latex_formula):
+            delimiter = latex_formula[i + 6]
+
+            if left_stack:
+                left_pos, left_depth, left_delim = left_stack.pop()
+
+                # 如果\left和\right不在同一花括号深度
+                if left_depth != len(brace_stack):
+                    # 找到\left所在花括号组的结束位置
+                    target_pos = find_group_end(latex_formula, left_pos, left_depth)
+                    if target_pos != -1:
+                        # 记录需要移动的\right
+                        adjustments.append((i, i + 7, target_pos))
+
+            i += 7  # 跳过\right和分隔符
+            continue
+
+        # 处理花括号
+        if latex_formula[i] == '{':
+            brace_stack.append(i)
+        elif latex_formula[i] == '}':
+            if brace_stack:
+                brace_stack.pop()
+
+        i += 1
+
+    # 应用调整，从后向前处理以避免索引变化
+    if not adjustments:
+        return latex_formula
+
+    result = list(latex_formula)
+    adjustments.sort(reverse=True, key=lambda x: x[0])
+
+    for start, end, target in adjustments:
+        # 提取\right部分
+        right_part = result[start:end]
+        # 从原位置删除
+        del result[start:end]
+        # 在目标位置插入
+        result.insert(target, ''.join(right_part))
+
+    return ''.join(result)
+
+
+def find_group_end(text, pos, depth):
+    """查找特定深度的花括号组的结束位置"""
+    current_depth = depth
+    i = pos
+
+    while i < len(text):
+        if text[i] == '{' and (i == 0 or not is_escaped(text, i)):
+            current_depth += 1
+        elif text[i] == '}' and (i == 0 or not is_escaped(text, i)):
+            current_depth -= 1
+            if current_depth < depth:
+                return i
+        i += 1
+
+    return -1  # 未找到对应结束位置
+
+
+def is_escaped(text, pos):
+    """检查字符是否被转义"""
+    backslash_count = 0
+    j = pos - 1
+    while j >= 0 and text[j] == '\\':
+        backslash_count += 1
+        j -= 1
+
+    return backslash_count % 2 == 1
+
+
+def fix_unbalanced_braces(latex_formula):
+    """
+    检测LaTeX公式中的花括号是否闭合，并删除无法配对的花括号
+
+    Args:
+        latex_formula (str): 输入的LaTeX公式
+
+    Returns:
+        str: 删除无法配对的花括号后的LaTeX公式
+    """
+    stack = []  # 存储左括号的索引
+    unmatched = set()  # 存储不匹配括号的索引
+    i = 0
+
+    while i < len(latex_formula):
+        # 检查是否是转义的花括号
+        if latex_formula[i] in ['{', '}']:
+            # 计算前面连续的反斜杠数量
+            backslash_count = 0
+            j = i - 1
+            while j >= 0 and latex_formula[j] == '\\':
+                backslash_count += 1
+                j -= 1
+
+            # 如果前面有奇数个反斜杠，则该花括号是转义的，不参与匹配
+            if backslash_count % 2 == 1:
+                i += 1
+                continue
+
+            # 否则，该花括号参与匹配
+            if latex_formula[i] == '{':
+                stack.append(i)
+            else:  # latex_formula[i] == '}'
+                if stack:  # 有对应的左括号
+                    stack.pop()
+                else:  # 没有对应的左括号
+                    unmatched.add(i)
+
+        i += 1
+
+    # 所有未匹配的左括号
+    unmatched.update(stack)
+
+    # 构建新字符串，删除不匹配的括号
+    return ''.join(char for i, char in enumerate(latex_formula) if i not in unmatched)
+
+
+def process_latex(input_string):
+    """
+        处理LaTeX公式中的反斜杠：
+        1. 如果\后跟特殊字符(#$%&~_^\\{})或空格，保持不变
+        2. 如果\后跟两个小写字母，保持不变
+        3. 其他情况，在\后添加空格
+
+        Args:
+            input_string (str): 输入的LaTeX公式
+
+        Returns:
+            str: 处理后的LaTeX公式
+        """
+
+    def replace_func(match):
+        # 获取\后面的字符
+        next_char = match.group(1)
+
+        # 如果是特殊字符或空格，保持不变
+        if next_char in "#$%&~_^|\\{} \t\n\r\v\f":
+            return match.group(0)
+
+        # 如果是字母，检查下一个字符
+        if 'a' <= next_char <= 'z' or 'A' <= next_char <= 'Z':
+            pos = match.start() + 2  # \x后的位置
+            if pos < len(input_string) and ('a' <= input_string[pos] <= 'z' or 'A' <= input_string[pos] <= 'Z'):
+                # 下一个字符也是字母，保持不变
+                return match.group(0)
+
+        # 其他情况，在\后添加空格
+        return '\\' + ' ' + next_char
+
+    # 匹配\后面跟一个字符的情况
+    pattern = r'\\(.)'
+
+    return re.sub(pattern, replace_func, input_string)
+
+# 常见的在KaTeX/MathJax中可用的数学环境
+ENV_TYPES = ['array', 'matrix', 'pmatrix', 'bmatrix', 'vmatrix',
+             'Bmatrix', 'Vmatrix', 'cases', 'aligned', 'gathered']
+ENV_BEGIN_PATTERNS = {env: re.compile(r'\\begin\{' + env + r'\}') for env in ENV_TYPES}
+ENV_END_PATTERNS = {env: re.compile(r'\\end\{' + env + r'\}') for env in ENV_TYPES}
+ENV_FORMAT_PATTERNS = {env: re.compile(r'\\begin\{' + env + r'\}\{([^}]*)\}') for env in ENV_TYPES}
+
+def fix_latex_environments(s):
+    """
+    检测LaTeX中环境（如array）的\\begin和\\end是否匹配
+    1. 如果缺少\\begin标签则在开头添加
+    2. 如果缺少\\end标签则在末尾添加
+    """
+    for env in ENV_TYPES:
+        begin_count = len(ENV_BEGIN_PATTERNS[env].findall(s))
+        end_count = len(ENV_END_PATTERNS[env].findall(s))
+
+        if begin_count != end_count:
+            if end_count > begin_count:
+                format_match = ENV_FORMAT_PATTERNS[env].search(s)
+                default_format = '{c}' if env == 'array' else ''
+                format_str = '{' + format_match.group(1) + '}' if format_match else default_format
+
+                missing_count = end_count - begin_count
+                begin_command = '\\begin{' + env + '}' + format_str + ' '
+                s = begin_command * missing_count + s
+            else:
+                missing_count = begin_count - end_count
+                s = s + (' \\end{' + env + '}') * missing_count
+
+    return s
+
+
+UP_PATTERN = re.compile(r'\\up([a-zA-Z]+)')
+COMMANDS_TO_REMOVE_PATTERN = re.compile(
+    r'\\(?:lefteqn|boldmath|ensuremath|centering|textsubscript|sides|textsl|textcent|emph|protect|null)')
+REPLACEMENTS_PATTERNS = {
+    re.compile(r'\\underbar'): r'\\underline',
+    re.compile(r'\\Bar'): r'\\hat',
+    re.compile(r'\\Hat'): r'\\hat',
+    re.compile(r'\\Tilde'): r'\\tilde',
+    re.compile(r'\\slash'): r'/',
+    re.compile(r'\\textperthousand'): r'‰',
+    re.compile(r'\\sun'): r'☉',
+    re.compile(r'\\textunderscore'): r'\\_',
+    re.compile(r'\\fint'): r'⨏',
+    re.compile(r'\\up '): r'\\ ',
+    re.compile(r'\\vline = '): r'\\models ',
+    re.compile(r'\\vDash '): r'\\models ',
+    re.compile(r'\\sq \\sqcup '): r'\\square ',
+}
+QQUAD_PATTERN = re.compile(r'\\qquad(?!\s)')
+
+def latex_rm_whitespace(s: str):
+    """Remove unnecessary whitespace from LaTeX code."""
+    s = fix_unbalanced_braces(s)
+    s = fix_latex_left_right(s)
+    s = fix_latex_environments(s)
+
+    # 使用预编译的正则表达式
+    s = UP_PATTERN.sub(
+        lambda m: m.group(0) if m.group(1) in ["arrow", "downarrow", "lus", "silon"] else f"\\{m.group(1)}", s
+    )
+    s = COMMANDS_TO_REMOVE_PATTERN.sub('', s)
+
+    # 应用所有替换
+    for pattern, replacement in REPLACEMENTS_PATTERNS.items():
+        s = pattern.sub(replacement, s)
+
+    # 处理LaTeX中的反斜杠和空格
+    s = process_latex(s)
+
+    # \qquad后补空格
+    s = QQUAD_PATTERN.sub(r'\\qquad ', s)
+
    return s


--- a/magic_pdf/model/sub_modules/model_utils.py
+++ b/magic_pdf/model/sub_modules/model_utils.py
@@ -172,8 +172,8 @@ def filter_nested_tables(table_res_list, overlap_threshold=0.8, area_threshold=0
        tables_inside = [j for j in range(len(table_res_list))
                         if i != j and is_inside(table_info[j], table_info[i], overlap_threshold)]

-        # Continue if there are at least 2 tables inside
-        if len(tables_inside) >= 2:
+        # Continue if there are at least 3 tables inside
+        if len(tables_inside) >= 3:
            # Check if inside tables overlap with each other
            tables_overlap = any(do_overlap(table_info[tables_inside[idx1]], table_info[tables_inside[idx2]])
                                 for idx1 in range(len(tables_inside))
--- a/next_docs/en/user_guide/install/boost_with_cuda.rst
+++ b/next_docs/en/user_guide/install/boost_with_cuda.rst
@@ -76,11 +76,11 @@ In the final step, enter ``yes``, close the terminal, and reopen it.
 4. Create an Environment Using Conda
 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

-Specify Python version 3.10.
+Specify Python version 3.10～3.13.

 .. code:: sh

-    conda create -n mineru 'python>=3.10' -y
+    conda create -n mineru 'python=3.12' -y
    conda activate mineru

 5. Install Applications
@@ -155,14 +155,15 @@ to test CUDA acceleration:
 Windows 10/11
 --------------

-1. Install CUDA and cuDNN
+1. Install CUDA
 ~~~~~~~~~~~~~~~~~~~~~~~~~

-You need to install a CUDA version that is compatible with torch's requirements. Currently, torch supports CUDA 11.8/12.4/12.6.
+You need to install a CUDA version that is compatible with torch's requirements. For details, please refer to the [official PyTorch website](https://pytorch.org/get-started/locally/).

 - CUDA 11.8 https://developer.nvidia.com/cuda-11-8-0-download-archive
 - CUDA 12.4 https://developer.nvidia.com/cuda-12-4-0-download-archive
 - CUDA 12.6 https://developer.nvidia.com/cuda-12-6-0-download-archive
+- CUDA 12.8 https://developer.nvidia.com/cuda-12-8-0-download-archive


 2. Install Anaconda
@@ -177,7 +178,7 @@ Download link: https://repo.anaconda.com/archive/Anaconda3-2024.06-1-Windows-x86

 ::

-    conda create -n mineru 'python>=3.10' -y
+    conda create -n mineru 'python=3.12' -y
    conda activate mineru

 4. Install Applications
--- a/next_docs/en/user_guide/install/install.rst
+++ b/next_docs/en/user_guide/install/install.rst
@@ -61,7 +61,7 @@ Also you can try `online demo <https://www.modelscope.cn/studios/OpenDataLab/Min
    </tr>
    <tr>
        <td colspan="3">Python Version</td>
-        <td colspan="3">3.10~3.12</td>
+        <td colspan="3">3.10~3.13</td>
    </tr>
    <tr>
        <td colspan="3">Nvidia Driver Version</td>
@@ -71,8 +71,7 @@ Also you can try `online demo <https://www.modelscope.cn/studios/OpenDataLab/Min
    </tr>
    <tr>
        <td colspan="3">CUDA Environment</td>
-        <td>11.8/12.4/12.6/12.8</td>
-        <td>11.8/12.4/12.6/12.8</td>
+        <td colspan="2"><a href="https://pytorch.org/get-started/locally/">Refer to the PyTorch official website</a></td>
        <td>None</td>
    </tr>
    <tr>
@@ -86,7 +85,7 @@ Also you can try `online demo <https://www.modelscope.cn/studios/OpenDataLab/Min
        <td colspan="2">GPU VRAM 6GB or more</td>
        <td colspan="2">All GPUs with Tensor Cores produced from Volta(2017) onwards.<br>
        More than 6GB VRAM </td>
-        <td rowspan="2">apple slicon</td>
+        <td rowspan="2">Apple silicon</td>
    </tr>
    </table>

@@ -97,7 +96,7 @@ Create an environment

 .. code-block:: shell

-    conda create -n mineru 'python>=3.10' -y
+    conda create -n mineru 'python=3.12' -y
    conda activate mineru
    pip install -U "magic-pdf[full]"

--- a/projects/gradio_app/app.py
+++ b/projects/gradio_app/app.py
@@ -117,8 +117,12 @@ def to_markdown(file_path, end_pages, is_ocr, layout_mode, formula_enable, table
    return md_content, txt_content, archive_zip_path, new_pdf_path


-latex_delimiters = [{'left': '$$', 'right': '$$', 'display': True},
-                    {'left': '$', 'right': '$', 'display': False}]
+latex_delimiters = [
+    {'left': '$$', 'right': '$$', 'display': True},
+    {'left': '$', 'right': '$', 'display': False},
+    {'left': '\\(', 'right': '\\)', 'display': False},
+    {'left': '\\[', 'right': '\\]', 'display': True},
+]


 def init_model():
@@ -218,7 +222,8 @@ if __name__ == '__main__':
                with gr.Tabs():
                    with gr.Tab('Markdown rendering'):
                        md = gr.Markdown(label='Markdown rendering', height=1100, show_copy_button=True,
-                                         latex_delimiters=latex_delimiters, line_breaks=True)
+                                         latex_delimiters=latex_delimiters,
+                                         line_breaks=True)
                    with gr.Tab('Markdown text'):
                        md_text = gr.TextArea(lines=45, show_copy_button=True)
        file.change(fn=to_pdf, inputs=file, outputs=pdf_show)
--- a/projects/multi_gpu/README.md
+++ b/projects/multi_gpu/README.md
@@ -4,9 +4,7 @@
 ## 环境配置
 请使用以下命令配置所需的环境：
 ```bash
-pip install -U litserve python-multipart filetype
-pip install -U magic-pdf[full] --extra-index-url https://wheels.myhloli.com
-pip install paddlepaddle-gpu==3.0.0b1 -i https://www.paddlepaddle.org.cn/packages/stable/cu118
+pip install -U magic-pdf[full] litserve python-multipart filetype
 ```

 ## 快速使用
--- a/projects/web_api/app.py
+++ b/projects/web_api/app.py
@@ -21,6 +21,7 @@ from magic_pdf.libs.config_reader import get_bucket_name, get_s3_config
 from magic_pdf.model.doc_analyze_by_custom_model import doc_analyze
 from magic_pdf.operators.models import InferenceResult
 from magic_pdf.operators.pipes import PipeResult
+from fastapi import Form

 model_config.__use_inside_model__ = True

@@ -102,6 +103,7 @@ def init_writers(
        # 处理上传的文件
        file_bytes = file.file.read()
        file_extension = os.path.splitext(file.filename)[1]
+
        writer = FileBasedDataWriter(output_path)
        image_writer = FileBasedDataWriter(output_image_path)
        os.makedirs(output_image_path, exist_ok=True)
@@ -176,14 +178,14 @@ def encode_image(image_path: str) -> str:
 )
 async def file_parse(
    file: UploadFile = None,
-    file_path: str = None,
-    parse_method: str = "auto",
-    is_json_md_dump: bool = False,
-    output_dir: str = "output",
-    return_layout: bool = False,
-    return_info: bool = False,
-    return_content_list: bool = False,
-    return_images: bool = False,
+    file_path: str = Form(None),
+    parse_method: str = Form("auto"),
+    is_json_md_dump: bool = Form(False),
+    output_dir: str = Form("output"),
+    return_layout: bool = Form(False),
+    return_info: bool = Form(False),
+    return_content_list: bool = Form(False),
+    return_images: bool = Form(False),
 ):
    """
    Execute the process of converting PDF to JSON and MD, outputting MD and JSON files
--- a/requirements.txt
+++ b/requirements.txt
@@ -7,9 +7,9 @@ numpy>=1.21.6
 pydantic>=2.7.2,<2.11
 PyMuPDF>=1.24.9,<1.25.0
 scikit-learn>=1.0.2
-torch>=2.2.2,!=2.5.0,!=2.5.1
+torch>=2.2.2,!=2.5.0,!=2.5.1,<3
 torchvision
 transformers>=4.49.0,!=4.51.0,<5.0.0
-pdfminer.six==20231228
+pdfminer.six==20250506
 tqdm>=4.67.1
 # The requirements.txt must ensure that only necessary external dependencies are introduced. If there are new dependencies to add, please contact the project administrator.
--- a/setup.py
+++ b/setup.py
@@ -81,7 +81,7 @@ if __name__ == '__main__':
            "Programming Language :: Python :: 3.12",
            "Programming Language :: Python :: 3.13",
        ],
-        python_requires=">=3.10,<4",  # 项目依赖的 Python 版本
+        python_requires=">=3.10,<3.14",  # 项目依赖的 Python 版本
        entry_points={
            "console_scripts": [
                "magic-pdf = magic_pdf.tools.cli:cli",
--- a/signatures/version1/cla.json
+++ b/signatures/version1/cla.json
@@ -247,6 +247,22 @@
      "created_at": "2025-04-17T03:54:59Z",
      "repoId": 765083837,
      "pullRequestNo": 2267
+    },
+    {
+      "name": "kowyo",
+      "id": 110339237,
+      "comment_id": 2829263082,
+      "created_at": "2025-04-25T02:54:20Z",
+      "repoId": 765083837,
+      "pullRequestNo": 2367
+    },
+    {
+      "name": "CharlesKeeling65",
+      "id": 94165417,
+      "comment_id": 2841356871,
+      "created_at": "2025-04-30T09:25:31Z",
+      "repoId": 765083837,
+      "pullRequestNo": 2411
    }
  ]
 }
Author	SHA1	Message	Date
Xiaomeng Zhao	ea619281ef	Merge pull request #2467 from opendatalab/release-1.3.11 Release 1.3.11	2025-05-14 10:33:00 +08:00
Xiaomeng Zhao	212cfcf24a	Merge pull request #2466 from opendatalab/dev docs(changelog): remove pdfminer.six version pinning from release notes	2025-05-14 10:32:31 +08:00
Xiaomeng Zhao	cda85d6262	Merge pull request #2465 from myhloli/dev docs(changelog): remove pdfminer.six version pinning from release notes	2025-05-14 10:31:55 +08:00
myhloli	51ceb48014	docs(changelog): remove pdfminer.six version pinning from release notes	2025-05-14 10:30:55 +08:00
Xiaomeng Zhao	0b8c614280	Merge pull request #2464 from opendatalab/release-1.3.11 Release 1.3.11	2025-05-14 10:22:18 +08:00
Xiaomeng Zhao	c1b387abe6	Merge pull request #2451 from myhloli/dev fix(modeling): escape backslashes in LaTeX command descriptions	2025-05-10 00:37:50 +08:00
myhloli	1ab54ac2e3	fix(modeling): escape backslashes in LaTeX command descriptions	2025-05-10 00:34:11 +08:00
myhloli	78a0208425	docs(installation): remove numpy version restriction from PyTorch installation instructions	2025-05-10 00:28:55 +08:00
Xiaomeng Zhao	cd785f6af8	Merge pull request #2450 from myhloli/dev fix(requirements): update pdfminer.six version and restrict torch version upper limit	2025-05-09 23:58:42 +08:00
myhloli	a8f752f753	fix(requirements): update pdfminer.six version and restrict torch version upper limit	2025-05-09 23:57:22 +08:00
Xiaomeng Zhao	65f332ffae	Merge pull request #2449 from myhloli/dev fix(setup): update python_requires to support Python 3.10 to 3.13	2025-05-09 23:45:16 +08:00
myhloli	c4b04ae642	Merge remote-tracking branch 'origin/dev' into dev	2025-05-09 23:38:50 +08:00
myhloli	3858d918dd	fix(setup): update python_requires to support Python 3.10 to 3.13	2025-05-09 23:38:37 +08:00
Xiaomeng Zhao	70696165c7	Merge pull request #2446 from myhloli/dev fix(Dockerfile): update modelscope installation command to use mirror	2025-05-09 18:23:08 +08:00
myhloli	b799d302c2	Merge remote-tracking branch 'origin/dev' into dev	2025-05-09 17:35:01 +08:00
myhloli	9351d64a41	fix(Dockerfile): update modelscope installation command to use mirror	2025-05-09 17:33:47 +08:00
Xiaomeng Zhao	3230793b55	Merge pull request #2440 from myhloli/dev docs(installation): update Python version and CUDA installation instructions	2025-05-09 11:10:09 +08:00
myhloli	9f0d45bb58	docs(installation): update Python version and CUDA installation instructions	2025-05-09 10:48:14 +08:00
Xiaomeng Zhao	6c9645aa0c	Merge pull request #2437 from myhloli/dev docs(README): reorder installation commands for clarity	2025-05-08 18:56:34 +08:00
myhloli	96fb646a86	Merge remote-tracking branch 'origin/dev' into dev	2025-05-08 18:55:49 +08:00
myhloli	71a429a32e	docs(README): reorder installation commands for clarity	2025-05-08 18:54:39 +08:00
Xiaomeng Zhao	201e338b3a	Merge pull request #2429 from myhloli/dev feat(modeling): add regex patterns for LaTeX symbol replacements	2025-05-08 11:27:57 +08:00
myhloli	2a28f604c6	feat(modeling): add regex patterns for LaTeX symbol replacements	2025-05-08 11:26:42 +08:00
Xiaomeng Zhao	38d0a622d9	Merge pull request #2423 from myhloli/dev feat(modeling): add 'protect' command to removal patterns	2025-05-06 18:22:18 +08:00
myhloli	a8ca183094	feat(modeling): add 'protect' command to removal patterns	2025-05-06 18:21:03 +08:00
Xiaomeng Zhao	11bf98d0aa	Merge pull request #2411 from CharlesKeeling65/patch-1 Update app.py: Fix parameter parsing in /file_parse endpoint	2025-04-30 17:51:08 +08:00
github-actions[bot]	50700646e4	@CharlesKeeling65 has signed the CLA in opendatalab/MinerU#2411	2025-04-30 09:25:44 +00:00
Wang Yubo	862891e294	Update app.py: Fix parameter parsing in /file_parse endpoint I have updated the `/file_parse` endpoint in `app.py` to correctly handle boolean and string parameters when they are sent via `multipart/form-data` requests (commonly used for file uploads). Previously, these parameters were not being properly parsed because FastAPI expects them to be passed as query or JSON body parameters by default. ### Changes Made: - Added `Form(...)` to all non-file parameters (`parse_method`, `is_json_md_dump`, `output_dir`, and return flags like `return_layout`, etc.). - This ensures that FastAPI correctly reads these fields from form-data, allowing clients to send both files and structured configuration options in the same request. ### Why This Change Was Needed: - When using `requests.post(..., data=data, files=files)`, the `data` dictionary is sent as form-encoded data. - Without explicitly declaring these fields with `Form(...)`, FastAPI does not bind them correctly, leading to default values always being used (e.g., `False` for boolean flags). - This change allows the API to accurately reflect the client's intent and enables features like `return_layout`, `return_images`, etc., to work as expected. This update improves compatibility with HTTP clients that rely on standard form-based file upload mechanisms while preserving the existing behavior of the API.	2025-04-30 17:15:54 +08:00
Xiaomeng Zhao	f0b66d3aab	Merge pull request #2410 from myhloli/dev feat(model): add logging for batch image processing	2025-04-30 17:09:49 +08:00
myhloli	b29b73af21	feat(model): add logging for batch image processing - Add logger info for each batch processed - Include batch number and page count in log message	2025-04-30 17:08:20 +08:00
Xiaomeng Zhao	5e8656c74f	Merge pull request #2406 from opendatalab/master update version	2025-04-29 16:09:37 +08:00
myhloli	2aaf2310f2	Update version.py with new version	2025-04-29 08:06:04 +00:00
Xiaomeng Zhao	8802687934	Merge pull request #2404 from opendatalab/release-1.3.10 Release 1.3.10	2025-04-29 15:48:55 +08:00
Xiaomeng Zhao	2c2fcbe832	Merge pull request #2403 from myhloli/dev feat(model_utils): adjust table detection threshold and add features	2025-04-29 15:27:44 +08:00
myhloli	9c37d65fab	docs(README_zh-CN): update doc	2025-04-29 15:26:08 +08:00
myhloli	49a8f8be0a	feat(model_utils): adjust table detection threshold and add features - Adjust the threshold for considering tables inside other tables from2 to 3 - Add support for custom formula delimiters through user configuration - Pin pdfminer.six to version 20250324 to prevent parsing failures	2025-04-29 15:24:28 +08:00
Xiaomeng Zhao	5e15d9b664	Merge pull request #2402 from myhloli/dev build(deps): pin pdfminer.six version to 20250324	2025-04-29 14:56:21 +08:00
myhloli	81daf298b5	build(deps): pin pdfminer.six version to 20250324 - Update pdfminer.six dependency from >=20250416 to ==20250324 - This change ensures compatibility with specific project requirements	2025-04-29 14:55:07 +08:00
myhloli	2d4e9e544e	Merge remote-tracking branch 'origin/dev' into dev	2025-04-29 10:54:34 +08:00
myhloli	dfd13fa2ab	fix(mfr): add LaTeX symbol replacements for fint and up - Add regex patterns for replacing LaTeX symbols \fint and \up with their Unicode equivalents	2025-04-29 10:53:40 +08:00
Xiaomeng Zhao	2cf55ce1d1	Merge pull request #2395 from myhloli/dev feat(latex): enhance LaTeX delimiter support and configurability	2025-04-28 14:37:33 +08:00
myhloli	100e9c17a5	feat(latex): enhance LaTeX delimiter support and configurability - Add support for \(\) and \[\] delimiters in addition to $$ and $$- Make LaTeX delimiter configuration more flexible and user-defined - Update configuration file to include LaTeX delimiter settings - Modify OCR content generation to use configurable delimiters	2025-04-28 14:35:39 +08:00
Xiaomeng Zhao	cf33cb882d	Merge pull request #2389 from myhloli/dev fix(mfr): add underscore symbol to unimernet	2025-04-28 01:56:17 +08:00
myhloli	98dd179053	Merge remote-tracking branch 'origin/dev' into dev	2025-04-28 01:55:20 +08:00
myhloli	7d77d614ec	fix(mfr): add underscore symbol to unimernet - Add \textunderscore to the list of LaTeX patterns - This allows the model to properly render underscore characters	2025-04-28 01:54:29 +08:00
Xiaomeng Zhao	c060413b19	Merge pull request #2388 from opendatalab/master update version	2025-04-27 18:30:05 +08:00
myhloli	1e715d026d	Update version.py with new version	2025-04-27 10:23:03 +00:00
Xiaomeng Zhao	0d5762e57a	Merge pull request #2381 from opendatalab/release-1.3.9 Release 1.3.9	2025-04-27 18:18:46 +08:00
Xiaomeng Zhao	d68fe15bde	Merge pull request #2386 from opendatalab/dev Dev	2025-04-27 18:18:28 +08:00
Xiaomeng Zhao	9bdc254456	Merge pull request #2385 from myhloli/dev docs: correct typo for Apple Silicon in install guide and README	2025-04-27 18:18:01 +08:00
myhloli	ebb7df984e	docs: correct typo for Apple Silicon in install guide and README - Fix typo in install.rst and README_zh-CN.md - Change 'apple slicon' to 'Apple silicon'	2025-04-27 18:16:46 +08:00
Xiaomeng Zhao	e54f8fd31e	Merge pull request #2384 from opendatalab/dev docs(README): fix typo	2025-04-27 18:14:46 +08:00
Xiaomeng Zhao	9f892a5e9d	Merge pull request #2367 from kowyo/patch-1 docs(README): fix typo	2025-04-27 18:14:08 +08:00
Xiaomeng Zhao	623537dd9c	Merge pull request #2383 from opendatalab/dev update readme	2025-04-27 18:12:26 +08:00
Xiaomeng Zhao	c1fbf01c43	Merge pull request #2382 from myhloli/dev feat(pdf): optimize formula parsing and update pdfminer.six	2025-04-27 18:11:47 +08:00
myhloli	0807e971fe	feat(pdf): optimize formula parsing and update pdfminer.six - Improve formula parsing success rate for better formula rendering - Upgrade pdfminer.six to the latest version to fix PDF parsing issues- Update changelog in both English and Chinese README files	2025-04-27 18:10:02 +08:00
Xiaomeng Zhao	ef854b23aa	Merge pull request #2380 from myhloli/dev build(deps): update pdfminer.six to latest version	2025-04-27 17:38:42 +08:00
myhloli	2d1a0f2ca6	fix(mfr): optimize LaTeX formula repair functionality - Improve \left and \right command handling in LaTeX formulas - Enhance environment type matching for array, matrix, and other structures - Refactor code for better readability and maintainability	2025-04-27 17:35:36 +08:00
myhloli	c8747cffb4	fix(magic_pdf): improve LaTeX formula processing and environment handling - Refactor LaTeX left/right pair fixing logic for better balance - Add environment detection and correction for common math environments - Implement more robust whitespace handling and command substitution - Optimize regex patterns for improved performance and readability	2025-04-27 17:10:15 +08:00
myhloli	0299dea199	build(deps): update pdfminer.six to latest version - Change pdfminer.six dependency from ==20231228 to >=20250416 - This update ensures compatibility with the latest version of pdfminer.six	2025-04-27 16:38:34 +08:00
myhloli	2e91fb3f52	fix(mfr): improve LaTeX formula processing and repair - Add functions to fix LaTeX left and right commands - Implement brace matching and repair in LaTeX formulas - Remove unnecessary whitespace and repair LaTeX code - Replace specific LaTeX commands with appropriate alternatives - Add logging for debugging purposes	2025-04-25 20:43:39 +08:00
myhloli	6c1511517a	fix(mfr): improve LaTeX formula processing and repair - Add functions to fix LaTeX left and right commands - Implement brace matching and repair in LaTeX formulas - Remove unnecessary whitespace and repair LaTeX code - Replace specific LaTeX commands with appropriate alternatives - Add logging for debugging purposes	2025-04-25 20:12:50 +08:00
github-actions[bot]	b864062a4f	@kowyo has signed the CLA in opendatalab/MinerU#2367	2025-04-25 02:54:31 +00:00
小林在忙毕业设计	c1558af3ef	docs(README): fix typo	2025-04-24 23:08:19 +08:00
Xiaomeng Zhao	2a9ac8939f	Merge pull request #2365 from myhloli/dev fix(mfr): improve LaTeX whitespace handling in unimernet model	2025-04-24 19:34:45 +08:00
myhloli	bfb80cb2e5	fix(mfr): improve LaTeX whitespace handling in unimernet model - Preserve "\ " sequences during whitespace removal - Add temporary substitution to prevent incorrect processing of "\ " sequences - Restore "\ " sequences after removing unnecessary whitespace	2025-04-24 19:33:03 +08:00
Xiaomeng Zhao	80a80482f3	Merge pull request #2356 from opendatalab/master master->dev	2025-04-23 18:50:42 +08:00