Multi-mode PDF intelligent analysis system and method based on food safety standard

The multimodal PDF intelligent parsing system solves the problem of missing multimodal information caused by the limitation of existing PDF parsing technology to a single modality, and realizes accurate identification and management decision support for image and table information in food safety standard documents.

CN120996022AInactive Publication Date: 2025-11-21JIANGXI INSTITUTE OF QUALITY & STANDARDIZATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511111905.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing PDF parsing technologies are limited to a single modality, resulting in the omission of multimodal information such as images and tables in food safety standard documents, which affects management decisions.

Method used

The system employs a multimodal PDF intelligent parsing system, which includes a PDF access module, a multimodal traffic splitting module, a text extraction module, an image parsing module, a table parsing module, a cross-modal fusion hub, a self-optimizing output engine, and an encrypted transmission unit. Through intelligent algorithms, OCR technology, deep learning, and encryption algorithms, it achieves accurate identification, extraction, and fusion of multimodal data.

Benefits of technology

It enables comprehensive and accurate parsing of multimodal information in food safety standard documents, ensuring the integrity of management decisions and secure data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996022A_ABST
    Figure CN120996022A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal PDF intelligent analysis system and method based on food safety standards, and the system comprises a PDF access module, a multi-modal distribution module, a text extraction module, an image analysis module, a table analysis module, a cross-modal fusion center, a self-optimization output engine, an encryption transmission unit and a client. The multi-modal distribution module is connected with the PDF access module, the text extraction module, the image analysis module and the table analysis module are connected with the multi-modal distribution module, the cross-modal fusion center is connected with the text extraction module, the image analysis module and the table analysis module, and the self-optimization output engine is connected with the cross-modal fusion center. In this way, the technical problem that in the prior art, a PDF analysis technology bureau is limited to a single mode, so that multi-mode information in a food safety standard document is omitted, and management decisions are affected is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a multi-modal PDF intelligent analysis system and method based on food safety standards. BACKGROUND

[0002] In today's digital age, the management of food safety standards documents faces many challenges. With the rapid development of information technology, a large number of food safety standards exist and spread in the form of PDF format files. PDF format has become a common format for storing and sharing various files due to its good cross-platform compatibility and document fidelity. In the field of food safety, many standard files, test reports, etc. are presented in PDF form. However, PDF files have multi-modal characteristics, and their content not only contains text information, but also may contain images, tables and other forms of data. Different modal data carries key information in food safety standards.

[0003] However, existing PDF analysis technology is limited to a single modality (such as pure text), resulting in the omission of multi-modal information (such as images, tables) in food safety standard documents, affecting management decisions. SUMMARY

[0004] The purpose of the present application is to provide a multi-modal PDF intelligent analysis system and method based on food safety standards, which aims to solve the technical problem that the existing PDF analysis technology is limited to a single modality (such as pure text), resulting in the omission of multi-modal information (such as images, tables) in food safety standard documents, affecting management decisions.

[0005] To achieve the above-mentioned purpose, a multi-modal PDF intelligent analysis system based on food safety standards is adopted, which comprises a PDF access module, a multi-modal shunt module, a text extraction module, an image analysis module, a table analysis module, a cross-modal fusion hub, a self-optimizing output engine, an encrypted transmission unit and a client. The multi-modal shunt module is connected to the PDF access module, the text extraction module, the image analysis module and the table analysis module are connected to the multi-modal shunt module, the cross-modal fusion hub is connected to the text extraction module, the image analysis module and the table analysis module, the self-optimizing output engine is connected to the cross-modal fusion hub, and the encrypted transmission unit is connected to the self-optimizing output engine and the client.

[0006] The PDF access module is used to receive PDF files, supporting the uploading and preprocessing of PDF files in multiple formats;

[0007] The multi-modal shunt module adopts an intelligent algorithm to perform multi-modal analysis on the preprocessed PDF file, and automatically identifies and divides the text region, image region and table region according to the document content;

[0008] The text extraction module uses advanced OCR technology and natural language processing technology to perform high-precision extraction, correction and semantic understanding of the text content, and especially focuses on accurate identification of professional terms and key indicator information in food safety standards;

[0009] The image analysis module uses image recognition and understanding technology, combined with an image feature library in the food safety field, to identify key content in the image and convert the analysis result into structured data;

[0010] The table analysis module uses a table structure recognition algorithm based on deep learning (such as a CNN-RNN hybrid model) to identify information in the table, accurately extract food safety standard data in the table, and organize it into a standardized structured format;

[0011] The cross-modal fusion hub uses a data fusion engine to associate the structured data of text, images and tables;

[0012] The self-optimizing output engine is used to adaptively convert and optimize the output of the fused structured data;

[0013] The encryption transmission unit is used to encrypt the output result and transmit it to the user.

[0014] The encryption transmission unit includes an encryption module, a decryption module, a key management module, a key generation module and a transmission module. The transmission module is arranged between the self-optimizing output engine and the client. The encryption module and the decryption module are connected with the self-optimizing output engine and the client respectively. The key management module is connected with the encryption module, the decryption module and the key generation module.

[0015] The self-optimizing output engine includes an adaptive template generation module, a template database and a template update module. The adaptive template generation module is connected with the cross-modal fusion hub and the encryption transmission unit. The template database is connected with the adaptive template generation module. The template update module is connected with the template database.

[0016] The multi-modal shunt module includes a streaming content detector and a dynamic allocator. The streaming content detector identifies text streams, images and table regions based on PDF object tree topology features;

[0017] The dynamic allocator automatically allocates computing resources according to the content.

[0018] The multi-modal PDF intelligent analysis system based on food safety standards further comprises a term reinforcement identification module, which is connected with the text extraction module.

[0019] The multi-modal PDF intelligent analysis system based on food safety standards further comprises a running monitoring module and an alarm module, wherein the running monitoring module is connected with the cross-modal fusion hub, and the alarm module is connected with the running monitoring module.

[0020] The multi-modal PDF intelligent analysis system based on food safety standards further comprises a space alignment engine, which is connected with the cross-modal fusion hub.

[0021] The multi-modal PDF intelligent analysis system based on food safety standards further comprises a space alignment engine, which is connected with the cross-modal fusion hub.

[0022] The multi-modal PDF intelligent analysis method based on food safety standards comprises the following steps:

[0023] Firstly, the PDF access module is responsible for receiving tasks, supports uploading of PDF files in multiple formats, and performs preprocessing work on the uploaded files.

[0024] Then, the multi-modal shunt module uses intelligent algorithms to perform multi-modal analysis on the preprocessed PDF files, accurately identifies the document content, and automatically divides the text area, image area and table area.

[0025] Subsequently, the text extraction module uses advanced OCR technology and natural language processing technology to perform high-precision extraction, correction and semantic understanding on the text area content; the image analysis module uses image recognition and understanding technology, combined with the image feature library in the food safety field, to identify the key content of the image and convert it into structured data; the table analysis module uses a table structure recognition algorithm based on deep learning to automatically identify table information, accurately extract food safety standard data and organize it into a standard structured format.

[0026] After that, the cross-modal fusion hub performs deep fusion and correlation analysis on the extracted structured data, mines the potential relationship between the data, and forms a comprehensive and accurate food safety standard information set.

[0027] Then, the self-optimizing output engine performs adaptive conversion and optimization processing on the fused structured data according to user requirements or preset formats.

[0028] Finally, the encryption transmission unit uses an encryption algorithm to encrypt the output result and transmits it to the client specified by the user, completing the entire analysis process.

[0029] This invention discloses a multimodal PDF intelligent parsing system and method based on food safety standards. In practical use, the PDF access module first undertakes the task of receiving PDF files, supporting uploads of various formats and performing preprocessing on the uploaded files. The multimodal splitting module uses intelligent algorithms to perform multimodal analysis on the preprocessed PDF files, accurately identifying document content and automatically dividing them into text regions, image regions, and table regions. Subsequently, the text extraction module uses advanced OCR and natural language processing technologies to perform high-precision extraction, correction, and semantic understanding of the text region content. The image parsing module uses image recognition and understanding technologies, combined with an image feature library in the food safety field, to identify key image content and convert it into structured data. The table parsing module uses deep learning-based table structure... The algorithm automatically identifies table information, accurately extracts food safety standard data, and organizes it into a standardized structured format. Then, the cross-modal fusion hub performs deep fusion and correlation analysis on the extracted structured data, uncovering potential relationships between data points to form a comprehensive and accurate set of food safety standard information. Next, the self-optimizing output engine adaptively transforms and optimizes the fused structured data according to user needs or preset formats. Finally, the encrypted transmission unit encrypts the output results using an encryption algorithm and transmits them to the user-specified client, completing the entire parsing process. This approach solves the technical problem of existing PDF parsing technologies being limited to a single modality (such as plain text), leading to the omission of multimodal information (such as images and tables) in food safety standard documents, thus affecting management decisions. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a schematic diagram of the first embodiment of the present invention.

[0032] Figure 2 This is a schematic diagram of the second embodiment of the present invention.

[0033] Figure 3 This is a schematic diagram of the third embodiment of the present invention.

[0034] 101-PDF access module, 102-multimodal shunt module, 103-text extraction module, 104-image analysis module, 105-table analysis module, 106-cross-modal fusion hub, 107-self-optimizing output engine, 108-encrypted transmission unit, 109-client, 110-encryption module, 111-decryption module, 112-key management module, 113-key generation module, 114-transmission module, 115-adaptive template generation module, 116-template database, 117-template update module, 118-streaming content detector, 119-dynamic allocator, 201-terminology reinforcement identification module, 202-operation monitoring module, 203-alarm module, 204-space alignment engine, 301-file verification module, 302-log recording module, 303-log audit module. DETAILED DESCRIPTION

[0035] Embodiments of the present application are described in detail below with reference to examples shown in the accompanying drawings, which are exemplary and intended to explain the present application, and cannot be understood as a limitation of the present application.

[0036] First embodiment, please refer to Figure 1 , Figure 1 is the principle block diagram of the first embodiment of the present application.

[0037] The present application provides a multimodal PDF intelligent analysis system based on food safety standards, comprising a PDF access module 101, a multimodal shunt module 102, a text extraction module 103, an image analysis module 104, a table analysis module 105, a cross-modal fusion hub 106, a self-optimizing output engine 107, an encrypted transmission unit 108 and a client 109; the encrypted transmission unit 108 comprises an encryption module 110, a decryption module 111, a key management module 112, a key generation module 113 and a transmission module 114, the self-optimizing output engine 107 comprises an adaptive template generation module 115, a template database 116 and a template update module 117, the multimodal shunt module 102 comprises a streaming content detector 118 and a dynamic allocator 119, the foregoing technical solution solves the technical problem that the PDF analysis technology in the prior art is limited to a single mode (such as pure text), resulting in the omission of multimodal information (such as images, tables) in food safety standard documents, affecting management decisions.

[0038] For this specific embodiment, the PDF access module 101 is used to receive PDF files, supporting multiple formats of PDF file upload and preprocessing;

[0039] The multi-modal shunt module 102 uses intelligent algorithms to perform multi-modal analysis on the pre-processed PDF files, automatically identifies and divides the text area, image area and table area according to the document content;

[0040] The text extraction module 103 uses advanced OCR technology and natural language processing technology to perform high-precision extraction, correction and semantic understanding of the text content, and especially focuses on accurate identification of professional terms and key indicator information in food safety standards;

[0041] The image analysis module 104 uses image recognition and understanding technology, combined with the image feature library in the field of food safety, to identify key content in the image and convert the analysis results into structured data;

[0042] The image analysis module 104 includes the following cascade processing mode:

[0043] First level: Document element detection network based on YOLOv7, positioning formula / icon / signature area;

[0044] Second level: Enhanced OCR preprocessing using CycleGAN to eliminate perspective distortion;

[0045] Third level: Fusing CLIP model for image-text matching degree evaluation, outputting description text with confidence;

[0046] The table analysis module 105 uses a deep learning-based table structure recognition algorithm (such as a CNN-RNN hybrid model) to identify the information of the table, accurately extract the food safety standard data in the table, and organize it into a standardized structured format;

[0047] The cross-modal fusion hub 106 uses a data fusion engine to associate the structured data of text, images and tables;

[0048] The self-optimizing output engine 107 is used for adaptive conversion and optimization output of the fused structured data;

[0049] The encryption transmission unit 108 is used for encrypted transmission of the output results to the user;

[0050] The multi-modal distribution module 102 is connected with the PDF access module 101, the text extraction module 103, the image analysis module 104 and the table analysis module 105 are all connected with the multi-modal distribution module 102, the cross-modal fusion hub 106 is connected with the text extraction module 103, the image analysis module 104 and the table analysis module 105, the self-optimizing output engine 107 is connected with the cross-modal fusion hub 106, and the encrypted transmission unit 108 is connected with the self-optimizing output engine 107 and the client 109. In specific use, the PDF access module 101 first undertakes the receiving task, supports multiple formats of PDF file uploading, and carries out pretreatment work on the uploaded file; the multi-modal distribution module 102 uses intelligent algorithm to analyze the pretreated PDF file, accurately identifies the document content, and automatically divides the text area, image area and table area; then, the text extraction module 103 uses advanced OCR technology and natural language processing technology to extract, correct and understand the semantic of the text area content with high precision; the image analysis module 104 uses image recognition and understanding technology, combined with the image feature library of the food safety field, to identify the key content of the image and convert it into structured data; the table analysis module 105 uses a table structure recognition algorithm based on deep learning to automatically identify table information, accurately extract food safety standard data and arrange it into a standard structured format; then, the cross-modal fusion hub 106 performs deep fusion and correlation analysis on the extracted structured data, mines the potential relationship between the data, and forms a comprehensive and accurate food safety standard information set; then, the self-optimizing output engine 107 performs adaptive conversion and optimization processing on the fused structured data according to user demand or preset format; finally, the encrypted transmission unit 108 uses encryption algorithm to encrypt the output result and transmits it to the user-specified client 109, completes the entire analysis process, and solves the technical problem that the PDF analysis technology in the prior art is limited to a single mode (such as pure text), resulting in the omission of multi-modal information (such as images and tables) in the food safety standard document, affecting management decisions.

[0051] Secondly, the transmission module 114 is arranged between the self-optimizing output engine 107 and the client 109, the encryption module 110 and the decryption module 111 are connected with the self-optimizing output engine 107 and the client 109 respectively, the key management module 112 is connected with the encryption module 110, the decryption module 111 and the key generation module 113, when the transmission module 114 detects that the self-optimizing output engine 107 has data to be transmitted to the client 109, the key generation module 113 immediately generates a key by using an elliptic curve cryptography algorithm, and delivers the generated key to the key management module 112; the key management module 112 securely stores the key and accurately allocates the key to the encryption module 110 and the decryption module 111 according to the requirement; the encryption module 110 acquires data from the self-optimizing output engine 107 and acquires the key from the key management module 112, and then encrypts the data according to a preset strategy by using an AES or RSA algorithm, and then the transmission module 114 transmits the encrypted ciphertext to the client 109; the decryption module 111 of the client 109 acquires the key from the key management module 112 to perform decryption operation, and also can verify the integrity and authenticity of the data by means of a digital signature algorithm, so as to ensure safe transmission of the data.

[0052] And, the adaptive template generation module 115 is connected with the cross-modal fusion hub 106 and the encrypted transmission unit 108, the template database 116 is connected with the adaptive template generation module 115, and the template updating module 117 is connected with the template database 116; when the cross-modal fusion hub 106 completes deep fusion of multi-source data, the adaptive template generation module 115 first loads a basic template set from the template database 116, calculates a semantic matching degree of the fused data and the template by means of an improved TF-IDF algorithm combined with a food safety field vocabulary; if the matching degree is lower than a threshold value, a dynamic template generation process is activated: a BERT-based sequence labeling model is used to identify key data entities, a GAT (graph attention network) is used to construct a data relationship graph, and then a Transformer decoder is used to generate a customized template structure; the generated template is compared with an existing template by means of an incremental learning algorithm (based on online gradient descent) of the template updating module 117 to find feature differences, only the significant differences are reserved and updated to the database, so that the template library always contains the latest data expression paradigm of the food safety standard.

[0053] Thirdly, the streaming content detector 118 identifies text streams, images and table regions based on PDF object tree topology features.

[0054] The dynamic allocator 119 automatically allocates computing resources according to the content.

[0055] Using a food safety standard-based multi-modal PDF intelligent analysis system of the embodiment, in specific use, first, the PDF access module 101 undertakes the receiving task, supports the uploading of PDF files in multiple formats, and carries out preprocessing work on the uploaded files; the multi-modal shunt module 102 uses intelligent algorithms to analyze the preprocessed PDF files, accurately identifies the document content, and automatically divides the text area, image area and table area; then, the text extraction module 103 uses advanced OCR technology and natural language processing technology to extract, correct and understand the semantic of the text area content with high precision; the image analysis module 104 uses image recognition and understanding technology, combined with the image feature library in the field of food safety, to identify the key content of the image and convert it into structured data; the table analysis module 105 uses a table structure recognition algorithm based on deep learning to automatically identify table information, accurately extract food safety standard data and organize it into a standard structured format; then, the cross-modal fusion hub 106 performs deep fusion and correlation analysis on the extracted structured data, mines the potential relationship between the data, and forms a comprehensive and accurate food safety standard information set; then, the self-optimizing output engine 107 performs adaptive conversion and optimization processing on the fused structured data according to user requirements or preset formats; finally, the encryption transmission unit 108 uses encryption algorithms to encrypt the output results and transmit them to the client 109 specified by the user, completing the entire analysis process, thereby solving the technical problem that the PDF analysis technology in the prior art is limited to a single mode (such as pure text), resulting in the omission of multi-modal information (such as images and tables) in food safety standard documents, affecting management decisions.

[0056] The second embodiment is a principle block diagram of the second embodiment of the present application. Figure 2 , Figure 2 The second embodiment is a principle block diagram of the second embodiment of the present application.

[0057] Based on the technology of the first embodiment, the present application provides a food safety standard-based multi-modal PDF intelligent analysis system, which further comprises a term reinforcement identification module 201, a running monitoring module 202, an alarm module 203 and a space alignment engine 204.

[0058] For this specific embodiment, the term reinforcement identification module 201 is connected with the text extraction module 103, and the term reinforcement identification module 201 uses a BiLSTM-CRF sequence labeling model to identify the extracted text through deep cooperation with the text extraction module 103; at the same time, the attention mechanism is introduced to strengthen the context association, ensuring the accuracy of term analysis and effectively solving the problem of incorrect identification of professional symbols by traditional OCR.

[0059] The operation monitoring module 202 is connected with the cross-modal fusion hub 106, the alarm module 203 is connected with the operation monitoring module 202, the operation monitoring module 202 collects the processing state of the cross-modal fusion hub 106 in real time, calculates the baseline threshold value through the sliding window statistical method, when detecting the abnormality, triggers the alarm module 203 to generate the hierarchical early warning: the first level alarm (system level failure) is pushed to the administrator mailbox through the SMTP protocol, the second level alarm (module performance decline) is prompted in the system interface pop-up window, and is recorded to the blockchain log to ensure traceability; at the same time, the self-healing mechanism (such as restarting the abnormal microservice) is started, and the system continuity is ensured.

[0060] Secondly, the space alignment engine 204 is connected with the cross-modal fusion hub 106, the space alignment engine 204 locates the multi-modal region coordinate through YOLOv8, constructs the modal correlation graph in combination with GNN, and adopts projection transformation to unify the data space coordinate, so that the semantic level three-dimensional linkage mapping of text, table and image is realized.

[0061] Using a multi-modal PDF intelligent analysis system based on food safety standards according to the embodiment, the term strengthening recognition module 201 cooperates with the text extraction module 103 in depth, adopts a BiLSTM-CRF sequence labeling model to perform entity recognition on the extracted text, and simultaneously introduces an attention mechanism to strengthen the context association, so that the term analysis accuracy is ensured, and the problem of recognition error of professional symbols in traditional OCR is effectively solved.

[0062] The space alignment engine 204 locates the multi-modal region coordinate through YOLOv8, constructs the modal correlation graph in combination with GNN, and adopts projection transformation to unify the data space coordinate, so that the semantic level three-dimensional linkage mapping of text, table and image is realized.

[0063] Third embodiment, please refer to Figure 3 , Figure 3 is the principle block diagram of the third embodiment of the application.

[0064] On the basis of the technology of the second embodiment, the application provides a multi-modal PDF intelligent analysis system based on food safety standards, further comprising a file verification module 301, a log recording module 302 and a log audit module 303.

[0065] For this specific embodiment, the file verification module 301 is used for integrity verification (such as hash value comparison) and format compliance detection of the uploaded PDF file.

[0066] The log recording module 302 is connected with the cross-modal fusion hub 106, the log audit module 303 is connected with the log recording module 302, the log recording module 302 records the whole process operation log (such as template update time, user query behavior), and adopts a blockchain structure to store and prevent tampering; the log audit module 303 obtains the whole process operation log (including template update, user query, etc.) in the log recording module 302 in real time through a secure interface, hashes the key data to the alliance chain for storage, and performs real-time risk analysis based on preset rules (such as high-frequency abnormal access).

[0067] Using a multi-modal PDF intelligent analysis system based on food safety standards, the log audit module 303 is connected with the log recording module 302, the log recording module 302 records the whole process operation log (such as template update time, user query behavior), and adopts a blockchain structure to store and prevent tampering; the log audit module 303 obtains the whole process operation log (including template update, user query, etc.) in the log recording module 302 in real time through a secure interface, hashes the key data to the alliance chain for storage, and performs real-time risk analysis based on preset rules (such as high-frequency abnormal access).

[0068] The application also provides a multi-modal PDF intelligent analysis method based on food safety standards, which is applied to the multi-modal PDF intelligent analysis system based on food safety standards,

[0069] Comprising the following steps:

[0070] Firstly, the PDF access module 101 undertakes the receiving task, supports the uploading of PDF files in multiple formats, and carries out preprocessing work on the uploaded files;

[0071] Then, the multi-modal shunt module 102 uses intelligent algorithms to analyze the preprocessed PDF files, accurately identifies the document content, and automatically divides the text area, image area and table area;

[0072] Subsequently, the text extraction module 103 uses advanced OCR technology and natural language processing technology to accurately extract, correct and understand the semantic understanding of the text area content; the image analysis module 104 uses image recognition and understanding technology, combined with the image feature library in the field of food safety, to identify the key content of the image and convert it into structured data; the table analysis module 105 uses a table structure recognition algorithm based on deep learning to automatically identify table information, accurately extract food safety standard data and organize it into a standard structured format;

[0073] After that, the cross-modal fusion center 106 performs deep fusion and correlation analysis on the extracted structured data, mines the potential relationship between the data, and forms a comprehensive and accurate food safety standard information set;

[0074] Then, the self-optimization output engine 107 performs adaptive conversion and optimization processing on the fused structured data according to user requirements or preset formats.

[0075] Finally, the encryption transmission unit 108 encrypts the output result using an encryption algorithm and transmits it to the client 109 specified by the user, completing the entire analysis process.

[0076] Among them, the OCR technology is an optical character recognition technology, which is used to convert the text in the image into editable text format.

[0077] YOLOv7 is a high-efficiency target detection algorithm that can quickly locate specific areas in an image (such as formulas, icons, signature areas).

[0078] YOLOv8 is the latest target detection model in the YOLO series, which is used to accurately locate multi-modal region coordinates to achieve spatial alignment.

[0079] CycleGAN is an image conversion model that eliminates image perspective distortion through a cycle-consistent generative adversarial network to improve the OCR preprocessing effect.

[0080] CLIP is a multi-modal model that is used to evaluate the matching degree of text and images and output description text with confidence, realizing the semantic association between text and images.

[0081] The above disclosure is only a preferred embodiment of the present application, and of course cannot limit the scope of the present application. Those skilled in the art can understand that the implementation of all or part of the above-mentioned embodiments still belongs to the scope covered by the present application.

Claims

1. A multi-modal PDF intelligent analysis system based on food safety standards, characterized in that, comprising a PDF access module, a multi-modal shunt module, a text extraction module, an image analysis module, a table analysis module, a cross-modal fusion hub, a self-optimizing output engine, an encrypted transmission unit and a client; the multi-modal shunt module is connected with the PDF access module, the text extraction module, the image analysis module and the table analysis module are connected with the multi-modal shunt module, the cross-modal fusion hub is connected with the text extraction module, the image analysis module and the table analysis module, the self-optimizing output engine is connected with the cross-modal fusion hub, and the encrypted transmission unit is connected with the self-optimizing output engine and the client; the PDF access module is used for receiving PDF files, supporting multiple formats of PDF file uploading and preprocessing; the multi-modal shunt module uses intelligent algorithms to analyze the preprocessed PDF files in multiple modes, and automatically identifies and divides the text area, image area and table area according to the document content; the text extraction module uses advanced OCR technology and natural language processing technology to extract, correct and understand the text content with high precision, especially focusing on the accurate identification of professional terms and key index information in food safety standards; the image analysis module uses image recognition and understanding technology, combined with the image feature library in the field of food safety, to identify the key content in the image and convert the analysis results into structured data; the table analysis module uses a table structure recognition algorithm based on deep learning to identify the information of the table, accurately extract the food safety standard data in the table, and arrange it into a standard structured format; the cross-modal fusion hub uses a data fusion engine to associate the structured data of text, image and table; the self-optimizing output engine is used for adaptive conversion and optimization output of the fused structured data; the encrypted transmission unit is used for encrypted transmission of the output results to the user.

2. The multi-modal PDF intelligent analysis system based on food safety standards according to claim 1, characterized in that, the encrypted transmission unit includes an encryption module, a decryption module, a key management module, a key generation module and a transmission module, the transmission module is arranged between the self-optimizing output engine and the client, the encryption module and the decryption module are connected with the self-optimizing output engine and the client respectively, and the key management module is connected with the encryption module, the decryption module and the key generation module.

3. The multi-modal PDF intelligent analysis system based on food safety standards according to claim 2, characterized in that, the self-optimizing output engine includes an adaptive template generation module, a template database and a template update module, the adaptive template generation module is connected with the cross-modal fusion hub and the encrypted transmission unit, the template database is connected with the adaptive template generation module, and the template update module is connected with the template database.

4. The multi-modal PDF intelligent analysis system based on food safety standards according to claim 3, wherein the multi-modal shunting module comprises a streaming content detector and a dynamic allocator, the streaming content detector identifies text, image and table regions based on PDF object tree topology features; the dynamic allocator automatically allocates computing resources according to content.

5. The multi-modal PDF intelligent analysis system based on food safety standards according to claim 4, wherein the multi-modal PDF intelligent analysis system based on food safety standards further comprises a term reinforcement identification module connected with the text extraction module.

6. The multi-modal PDF intelligent analysis system based on food safety standards according to claim 5, wherein the multi-modal PDF intelligent analysis system based on food safety standards further comprises a running monitoring module connected with the cross-modal fusion hub and an alarm module connected with the running monitoring module.

7. The multi-modal PDF intelligent analysis system based on food safety standards according to claim 6, wherein the multi-modal PDF intelligent analysis system based on food safety standards further comprises a spatial alignment engine connected with the cross-modal fusion hub.

8. A multi-modal PDF intelligent analysis method based on food safety standards, applied to the multi-modal PDF intelligent analysis system based on food safety standards according to claim 7, comprising the following steps: firstly, the PDF access module is responsible for receiving tasks, supports uploading of PDF files in multiple formats, and carries out pretreatment work on the uploaded files; then, the multi-modal shunting module uses intelligent algorithms to perform multi-modal analysis on the pretreated PDF files, accurately identifies document content, and automatically divides text, image and table regions; subsequently, the text extraction module uses advanced OCR technology and natural language processing technology to extract, correct and understand the semantics of the text region content with high precision; the image analysis module uses image recognition and understanding technology, combined with the image feature library of the food safety field, to identify key image content and convert it into structured data; the table analysis module uses a deep learning-based table structure recognition algorithm to automatically identify table information, accurately extract food safety standard data and organize it into a standard structured format; after that, the cross-modal fusion hub performs deep fusion and correlation analysis on the extracted structured data, mines potential relationships between data, and forms a comprehensive and accurate food safety standard information set; then, the self-optimizing output engine performs adaptive conversion and optimization processing on the fused structured data according to user requirements or preset formats; finally, the encryption transmission unit uses encryption algorithms to encrypt the output results and transmits them to the client specified by the user, completing the entire analysis process. ​ ​ ​ ​ ​