Textile detection report structured analysis method based on multi-modal deep learning

Through a multimodal deep learning method, the text and images in PDF documents are structured and multimodal fusion are solved, and the problems of weak image processing capabilities and inaccurate information extraction in the prior art are achieved, and high-accurate information extraction and deep understanding are achieved.

CN119940292APending Publication Date: 2025-05-06SHENZHEN PURCOTTON TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411787846.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art cannot effectively process images embedded in PDF documents, and lacks comprehensive processing and correlation analysis, resulting in low accuracy of information extraction.

Method used

The structured analysis method of textile inspection report based on multimodal deep learning is adopted, and the text and image information are divided into structured data through document analysis tools, image processing and feature extraction are performed, image areas and text blocks are matched, and semantic analysis and multimodal data fusion are used using the depth bidirectional representation model BERT and natural language processing algorithms.

Benefits of technology

The precise extraction and multimodal fusion of text and image content in PDF documents are realized, which improves the accuracy and depth of information extraction, and solves the problems of weak image processing capabilities and inaccurate information extraction in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940292A_ABST
    Figure CN119940292A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a textile detection report structured analysis method based on multi-modal deep learning, and relates to the technical field of document content extraction and processing.The method comprises the steps that a document analysis tool is used for dividing text content and image information in a document to obtain text blocks and image areas, and the text blocks and the image areas are used for analyzing the text blocks and the image areas; performing marking to obtain structured data, performing image processing and feature extraction on an image region to obtain image features, matching the image region and a text block to obtain a multi-modal content block, performing semantic analysis, semantic enhancement and context association on the multi-modal content block by using BERT and NLP to obtain multi-modal data, and formatting the multi-modal data to obtain multi-modal data; and obtaining a document content file by combining the relevance between the text content and the image data in the multi-modal data through a deep bidirectional representation model BERT. The problems that in the prior art, images embedded in documents cannot be effectively processed, comprehensive processing and correlation analysis are lacked, and the accuracy of information extraction is low are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document content extraction and processing, and in particular to a structured analysis method for textile test reports based on multimodal deep learning. Background Art

[0002] PDF document content extraction technology has received extensive attention and development in recent years, and its application scenarios cover multiple fields such as information retrieval, automated processing, content analysis, etc. With the continuous advancement of information technology, PDF documents, as a commonly used electronic file format, are widely used in academic exchanges, business reports, government documents and other fields.

[0003] However, most existing technologies only focus on the extraction of text content, and have weak processing capabilities for images embedded in documents. Especially when the image contains important information (such as charts, scanned documents, illustrations), it is often impossible to effectively extract or recognize the image content.

[0004] In addition, the extraction of text and images is usually carried out separately, lacking comprehensive processing and association analysis, which makes it difficult to fuse text and images at the semantic level, affecting the accuracy of information extraction and the depth of understanding of document content.

[0005] At the same time, the existing technology lacks further semantic analysis and structured processing capabilities after extracting content, which limits the application value of the extraction results in information retrieval, content analysis, etc.

[0006] Therefore, a structured analysis method for textile test reports is needed, which can accurately extract text and image content from PDF documents, perform multimodal data fusion and analysis, and generate structured data output. Summary of the invention

[0007] The embodiments of the present invention provide a structured analysis method for textile test reports based on multimodal deep learning to solve the problems that the prior art cannot effectively process images embedded in documents, lacks comprehensive processing and association analysis, and has low accuracy in information extraction. The technical solution is as follows:

[0008] According to one aspect of the present invention, a structured analysis method for textile test reports based on multimodal deep learning comprises: using a document parsing tool to divide the text content and image information in a document into text blocks and image areas, respectively, marking the text blocks and image areas to obtain structured data; performing image processing and feature extraction on the image area according to the structured data to obtain image features, matching the image area and the text block according to the image features to obtain a multimodal content block; using a deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the multimodal content block to obtain multimodal data; formatting the multimodal data, and obtaining a document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT.

[0009] In one of the embodiments, a document parsing tool is used to divide the text content and image information in a document into text blocks and image areas respectively, and the text blocks and image areas are marked to obtain structured data through the following steps: using a document parsing tool to extract text content from each page of the document and divide it into blocks to obtain multiple text blocks, and recording the position, size and content of each text block; using a document parsing tool to obtain the position, size, and format of the image from each page of the document to obtain multiple image areas; filtering the text blocks and image areas, marking and integrating the information of the filtered text blocks and image areas to obtain structured data.

[0010] In one of the embodiments, filtering the text blocks and image areas is achieved by the following steps: removing text blocks located at the edge of the page, text blocks whose sizes do not meet set values, and text blocks whose contents do not meet set amounts according to the positions, sizes, and contents of the text blocks; performing semantic relevance analysis on the text blocks using a natural language processing algorithm to remove text blocks whose relevance is lower than a set value; removing image areas located at the edge of the page, deviating from the text, and having sizes that do not meet set values ​​according to the positions and sizes of the image areas, and filtering out image areas that are not related to the text content by an image processing algorithm.

[0011] In one of the embodiments, performing image processing and feature extraction on the image area according to the structured data to obtain image features is achieved by the following steps: cropping and black-and-white binarization of the image area according to the structured data, denoising, adjusting contrast and brightness of the image area by using the image processing tool OpenCV; performing text extraction on the adjusted image area by using the character recognition tool Tesseract to obtain image text, and associating the image text with the location information of the image area; classifying and identifying the image area by using the deep learning models ResNet and YOLO-V5 to obtain the image type, and performing feature extraction on the chart image according to the image type to obtain image features; the image features include data, coordinate axis labels and legends.

[0012] In one of the embodiments, matching the image area and the text block according to the image features to obtain the multimodal content block is achieved by the following steps: organizing and linearizing the text block, and matching the text block and the image area on the same page according to the position information in the structured data to obtain the multimodal content block.

[0013] In one of the embodiments, after matching the image area and the text block according to the image features to obtain the multimodal content block, the following steps are also included: using a multimodal deep learning model CLIP to perform data fusion on the multimodal content block according to the text block, image text and image features, and fusing the text and image content in the multimodal content block into the same semantic space.

[0014] In one of the embodiments, the use of the deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the comprehensive data to obtain multimodal data is achieved through the following steps: using the deep bidirectional representation model BERT to perform semantic analysis on the text in the comprehensive data to obtain key information of the text; combining the image features in the comprehensive data and the key information to semantically enhance the text, and integrating the comprehensive data at the semantic level through a multimodal fusion algorithm; using the natural language processing algorithm HuggingFace function to analyze the semantic relationship between the text and the image in the comprehensive data to obtain the context association between the text and the image to generate multimodal data.

[0015] According to one aspect of the present invention, a textile test report structured analysis device based on multimodal deep learning comprises: a document parsing module, used to use a document parsing tool to divide the text content and image information in a document into text blocks and image areas, respectively, and mark the text blocks and image areas to obtain structured data; a content extraction module, used to perform image processing and feature extraction on the image area according to the structured data to obtain image features, and match the image area and the text block according to the image features to obtain a multimodal content block; a multimodal fusion module, using a deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the multimodal content block to obtain multimodal data; a document content extraction module, used to format the multimodal data, and obtain a document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT.

[0016] According to one aspect of the present invention, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the structured analysis method of textile testing reports based on multimodal deep learning as described above.

[0017] According to one aspect of the present invention, a storage medium stores computer-readable instructions thereon, and the computer-readable instructions are executed by one or more processors to implement the above-mentioned method for structuring analysis of textile test reports based on multimodal deep learning.

[0018] The beneficial effects brought by the technical solution provided by the present invention are:

[0019] In the above technical scheme, the present invention first uses a document parsing tool to divide the text content and image information in the document to obtain text blocks and image areas respectively, marks the text blocks and image areas to obtain structured data, performs image processing and feature extraction on the image area according to the structured data to obtain image features, matches the image area and the text block according to the image features to obtain a multimodal content block, uses a deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the multimodal content block to obtain multimodal data, formats the multimodal data, and obtains a document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT, accurately extracts the image area by combining an image processing tool and a deep learning model, and further analyzes the chart-like image to extract the key features to ensure that the key information in the image can be fully utilized and accurately parsed, and performs multimodal fusion of the text and image content to achieve deep association between the text and the image at the semantic level, which has significant advantages in comprehensive content analysis and information accuracy, thereby effectively solving the problems that the prior art cannot effectively process images embedded in documents, lacks comprehensive processing and association analysis, and has low accuracy in information extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in describing the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative work.

[0021] Figure 1 is a flow chart of a method for structured analysis of textile test reports based on multimodal deep learning according to an exemplary embodiment;

[0022] Figure 2 It is a flowchart of text block processing in a structured analysis method of a textile test report based on multimodal deep learning, shown in an exemplary embodiment;

[0023] Figure 3 yes Figure 2 Flow chart of image content processing in the corresponding embodiment;

[0024] Figure 4 is a block diagram of a textile test report structured analysis device based on multimodal deep learning according to an exemplary embodiment;

[0025] Figure 5 is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0026] Figure 6It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0027] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.

[0028] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0029] To this end, the present invention provides a textile test report structured analysis method based on multimodal deep learning, which accurately extracts image areas by combining image processing tools and deep learning models, and multimodally fuses text and image content to achieve deep association between text and image at the semantic level, which can effectively solve the problems that the prior art cannot effectively process images embedded in documents, lacks comprehensive processing and association analysis, and has low accuracy in information extraction. The textile test report structured analysis method based on multimodal deep learning is suitable for a textile test report structured analysis device, and the document content extraction device can be an electronic device. The textile test report structured analysis method based on multimodal deep learning in the embodiment of the present invention can be applied to a variety of scenarios, such as textile test report structured analysis.

[0030] See also Figure 1 , an embodiment of the present invention provides a structured analysis method for textile test reports based on multimodal deep learning, which is applicable to electronic devices.

[0031] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.

[0032] like Figure 1 As shown, the method may include the following steps:

[0033] Step 110: Use a document parsing tool to divide the text content and image information in the document into text blocks and image areas, respectively, and mark the text blocks and image areas to obtain structured data.

[0034] In one possible implementation, a document parsing tool is used to extract text content from each page of a document and divide the text content into blocks to obtain multiple text blocks, and the position, size, and content of each text block are recorded. The document parsing tool is used to obtain the position, size, and format of images from each page of the document to obtain multiple image areas. The text blocks and image areas are filtered, and the filtered text blocks and image areas are marked and integrated to obtain structured data.

[0035] In one possible implementation, text blocks located at the edge of the page, text blocks whose sizes do not meet set values, and text blocks whose contents do not meet set amounts are removed based on the position, size, and content of the text blocks, a natural language processing algorithm is used to perform semantic relevance analysis on the text blocks, and text blocks whose relevance is lower than a set value are removed, and image areas located at the edge of the page, deviating from the text, and having a size that does not meet the set value are removed based on the position and size of the image area, and image areas that are not related to the text content are filtered out using an image processing algorithm.

[0036] Step 130: performing image processing and feature extraction on the image region according to the structured data to obtain image features, and matching the image region and the text block according to the image features to obtain a multimodal content block.

[0037] In one possible implementation, the image area is cropped and binarized according to the structured data, the image area is denoised, the contrast and brightness are adjusted by the image processing tool OpenCV, the text is extracted from the adjusted image area by the character recognition tool Tesseract to obtain the image text, and the image text is associated with the position information of the image area, the image area is classified and identified by the deep learning models ResNet and YOLO-V5 to obtain the image type, and the image features are extracted for the chart image according to the image type.

[0038] Among them, image features include data, axis labels and legends.

[0039] In one possible implementation, the text blocks are sorted and linearized, and the text blocks and image areas on the same page are matched according to the position information in the structured data to obtain multimodal content blocks.

[0040] In one possible implementation, a multimodal deep learning model CLIP is used to perform data fusion on multimodal content blocks according to text blocks, image text, and image features, so as to fuse the text and image content in the multimodal content blocks into the same semantic space.

[0041] Step 150: Use the deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement, and context association on the multimodal content block to obtain multimodal data.

[0042] In one possible implementation, the deep bidirectional representation model BERT is used to perform semantic analysis on the text in the comprehensive data to obtain the key information of the text, and the image features and key information in the comprehensive data are combined to perform semantic enhancement on the text. The comprehensive data is integrated at the semantic level through a multimodal fusion algorithm, and the natural language processing algorithm HuggingFace function is used to analyze the semantic relationship between the text and the image in the comprehensive data to obtain the contextual association between the text and the image to generate multimodal data.

[0043] Step 170, format the multimodal data, and obtain a document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT.

[0044] Through the above process, the embodiment of the present invention first uses a document parsing tool to divide the text content and image information in the document into text blocks and image areas respectively, marks the text blocks and image areas to obtain structured data, performs image processing and feature extraction on the image area according to the structured data to obtain image features, matches the image area and the text block according to the image features to obtain a multimodal content block, uses the deep bidirectional representation model BERT and the natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the multimodal content block to obtain multimodal data, formats the multimodal data, and obtains a document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT, accurately extracts the image area by combining the image processing tool and the deep learning model, and further analyzes the chart image to extract the key features to ensure that the key information in the image can be fully utilized and accurately parsed, and performs multimodal fusion of the text and image content to achieve deep association between the text and the image at the semantic level, which has significant advantages in comprehensive content analysis and information accuracy, thereby effectively solving the problem that the prior art cannot effectively process the images embedded in the document, lacks comprehensive processing and association analysis, and has low accuracy in information extraction.

[0045] In an exemplary embodiment, taking a PDF document as an example, using the document content extraction method provided by the present invention to extract document content may specifically include the following steps:

[0046] In the first step, use the tool PyMuPDF to load the PDF file, parse the PDF file into page objects, and separate the text blocks and image areas from each page.

[0047] Specifically, use PyMuPDF to extract the text content in the page, divide it into text blocks and record its position information, use PyMuPDF to obtain the image information in the page, extract the image's position, size, format and other information, and associate it with the page content, mark the extracted text blocks and image areas, and record the relative position, size, type (text or image) and other information of the elements in each page to form a structured page representation.

[0048] In the second step, the text blocks and image areas in the PDF document are further processed and analyzed for use in subsequent multimodal data fusion.

[0049] like Figure 2 As shown, text block recognition and extraction includes the following steps:

[0050] Step S1, extracting text blocks.

[0051] Specifically, use a PDF parsing tool such as the text extraction function in PyMuPDF to obtain text blocks within the page. Each text block is described by a bounding box (bbox) that indicates the position of the text block on the page. The content of the text block is extracted and stored together with the position information.

[0052] Step S2, filtering and cleaning.

[0053] Specifically, the text content includes all characters, words and sentences on the page, retaining the original text format information, and performing preliminary filtering on the extracted text blocks to remove insignificant small text blocks (such as headers, footers, page numbers) to ensure that the retained text blocks have practical significance.

[0054] Step S3: structured storage.

[0055] Specifically, the content, location information and format information of each text block are stored in a structured text data format.

[0056] like Figure 3 As shown, image content extraction includes the following steps:

[0057] Step S1, extracting the image area.

[0058] Specifically, according to the image area annotation in the page, the image content in the page is extracted using the image processing tool OpenCV, the page is cropped using the bbox information, the precise image area is extracted, the image data is saved in binary format, and its page number and location information are recorded.

[0059] Step S2: image preprocessing.

[0060] Specifically, OpenCV is used to perform basic preprocessing on the image, including denoising, adjusting contrast and brightness to improve image quality. Tesseract is used to extract text content from images containing text, and the text content extracted by OCR is associated with the location information and page number of the image.

[0061] Step S3: image classification and feature extraction.

[0062] Specifically, the extracted images are classified using a pre-trained ResNet model to identify the type of image (chart, photo, illustration). For chart images, the target detection algorithm YOLO-V5 is used to extract key features, including data points, axis labels, legends, etc. The extracted features will be used for multimodal data fusion and semantic analysis. Photos only need to be labeled.

[0063] The third step is to conduct a fusion analysis on all data based on the structured text and image data output in the second step. The fusion analysis includes data fusion and comprehensive content analysis.

[0064] Specifically, data fusion includes: matching text blocks on the same page with image areas based on page numbers and location information to form multimodal content blocks; integrating OCR text in images with adjacent or associated text blocks to ensure semantic consistency between text and image content; using multimodal deep learning models such as CLIP to fuse text and image content in the same semantic space to ensure that they can be understood and analyzed simultaneously; and finally generating a comprehensive content description that integrates text interpretation, image content, and extracted features into a holistic view.

[0065] Specifically, comprehensive content analysis includes: receiving multimodal content blocks, using the pre-trained BERT model to perform semantic analysis on the fused text, extracting key information, and then combining image features to semantically enhance the overall content to generate a more complete understanding result, using the HuggingFace library function to analyze the semantic relationship between text and images, identifying the contextual association between them, and finally generating a comprehensive analysis of structured data.

[0066] The fourth step is to output structured data and obtain the document content file.

[0067] Specifically, the structured data of the comprehensive analysis outputted in the third step is received, and according to user needs, the analyzed multimodal data is formatted into a commonly used data format to ensure that the correlation between the text content and the image data is retained during the formatting process, and then the data is passed through the pre-trained Bert model to obtain the final PDF text output.

[0068] In the above process, the present invention can accurately extract the image area, and further analyze the chart image, extract key features (such as data points, coordinate axes, legends, etc.), can perform multimodal fusion of text and image content, and realize the deep association between text and image at the semantic level, and has significant advantages in comprehensive content analysis and information accuracy. After extracting the PDF content, the present invention further performs semantic analysis and structural processing on the extracted text and image content, extracts key information in the document through natural language processing technology (such as named entity recognition, content summary generation, etc.), and organizes it into structured data that is easy to understand and use. Compared with the prior art, it not only provides text and image extraction, but also enhances the utilization value of information, and is suitable for various scenarios that require in-depth content analysis.

[0069] Through the above process, the embodiment of the present invention accurately extracts and processes text and image content from PDF documents, and performs multimodal data fusion and analysis. First, the PDF document is parsed by a PDF parsing tool, and the text blocks and image areas of each page are separated and the position information is marked. Then, the extracted text blocks are sorted and linearized to ensure that the logical order and format information of the text content are complete. At the same time, the image area is extracted and classified, and the OCR technology is applied to extract the text information in the image when necessary. Then, the text and image content are fused and analyzed using a multimodal deep learning model to achieve semantic association and comprehensive understanding of text and image. Finally, the processed data is formatted and output to generate a structured data file or report. The embodiment of the present invention has significant advantages in processing PDF documents with complex layouts, image content recognition, multimodal data fusion, OCR text extraction and semantic analysis, and is suitable for a variety of document analysis and information extraction application scenarios.

[0070] The following is an embodiment of the device of the present invention, which can be used to execute the method for structuring analysis of textile test reports based on multimodal deep learning involved in the present invention. For details not disclosed in the embodiment of the device of the present invention, please refer to the embodiment of the method for structuring analysis of textile test reports based on multimodal deep learning involved in the present invention.

[0071] See also Figure 4 In an embodiment of the present invention, a textile test report structured analysis device 800 based on multimodal deep learning is provided.

[0072] The device 800 includes but is not limited to: a document parsing module 810 , a content extraction module 830 , a multimodal fusion module 850 and a document content extraction module 870 .

[0073] The document parsing module 810 is used to use a document parsing tool to divide the text content and image information in the document into text blocks and image areas, and mark the text blocks and image areas to obtain structured data.

[0074] The content extraction module 830 is used to perform image processing and feature extraction on the image area according to the structured data to obtain image features, and match the image area and the text block according to the image features to obtain a multimodal content block.

[0075] The multimodal fusion module 850 uses the deep bidirectional representation model BERT and the natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the multimodal content blocks to obtain multimodal data.

[0076] The document content extraction module 870 is used to format the multimodal data and obtain the document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT.

[0077] It should be noted that, in the document content extraction provided in the above embodiment, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the document content extraction device will be divided into different functional modules to complete all or part of the functions described above.

[0078] In addition, the above-mentioned embodiments provide a structured analysis device for textile test reports based on multimodal deep learning and an embodiment of a structured analysis method for textile test reports based on multimodal deep learning, and the specific manner in which each module performs operations has been described in detail in the method embodiments and will not be repeated here.

[0079] Figure 5 The present invention schematically illustrates the structure of an electronic device according to an exemplary embodiment.

[0080] It should be noted that the electronic device is only an example adapted to the present invention and cannot be considered to provide any limitation on the scope of use of the present invention. The electronic device cannot be interpreted as requiring dependence on or having Figure 5 One or more components of exemplary electronic device 2000 are shown.

[0081] The hardware structure of the electronic device 2000 may vary greatly due to differences in configuration or performance, such as Figure 5As shown, the electronic device 2000 includes: a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .

[0082] Specifically, the power supply 210 is used to provide operating voltage for each hardware device on the electronic device 2000 .

[0083] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. Of course, in other examples adapted by the present invention, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input-output interface 235, and at least one USB interface 237, etc. Figure 5 As shown, this is not intended to be a specific limitation.

[0084] The memory 250 is a carrier for storing resources, which may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 251, an application 253 and data 255, etc. The storage method may be temporary storage or permanent storage.

[0085] Among them, the operating system 251 is used to manage and control various hardware devices and application programs 253 on the electronic device 2000 to enable the central processor 270 to calculate and process the massive data 255 in the memory 250. It can be WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0086] The application program 253 is a computer-readable instruction that performs at least one specific task based on the operating system 251, and may include at least one module ( Figure 5 (not shown), each module may respectively include computer-readable instructions for the electronic device 2000. For example, the document content extraction device may be regarded as an application 253 deployed on the electronic device 2000.

[0087] The data 255 may be signal information, etc., stored in the memory 250 .

[0088] The central processor 270 may include one or more processors, and is configured to communicate with the memory 250 through at least one communication bus to read the computer-readable instructions stored in the memory 250, thereby realizing the operation and processing of the mass data 255 in the memory 250. For example, the structured analysis method of textile test reports based on multimodal deep learning is completed by the central processor 270 reading a series of computer-readable instructions stored in the memory 250.

[0089] In addition, the present invention can also be implemented by hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present invention is not limited to any specific hardware circuits, software, or a combination of the two.

[0090] See also Figure 6 In an embodiment of the present invention, an electronic device 4000 is provided. The electronic device 400 may include: a desktop computer, a laptop computer, a server, etc. with sensor recognition capability.

[0091] exist Figure 6 In the embodiment, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0092] The data interaction between the processor 4001 and the memory 4003 can be realized through at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0093] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception, etc. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0094] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0095] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program instructions or codes in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to this.

[0096] Computer-readable instructions are stored in the memory 4003 , and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .

[0097] The computer-readable instructions are executed by one or more processors 4001 to implement the structured analysis method of textile testing reports based on multimodal deep learning in the above-mentioned embodiments.

[0098] In addition, a storage medium is provided in an embodiment of the present invention, on which computer-readable instructions are stored, and the computer-readable instructions are executed by one or more processors to implement the structured analysis method of textile testing reports based on multimodal deep learning as described above.

[0099] A computer program product is provided in an embodiment of the present invention. The computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the structured analysis method of textile inspection reports based on multimodal deep learning as described above.

[0100] Compared with the related art, the present invention has the following beneficial effects:

[0101] 1. The present invention first uses a document parsing tool to divide the text content and image information in the document into text blocks and image areas, marks the text blocks and image areas to obtain structured data, performs image processing and feature extraction on the image area according to the structured data to obtain image features, matches the image area and the text block according to the image features to obtain a multimodal content block, uses a deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the multimodal content block to obtain multimodal data, formats the multimodal data, obtains a document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT, accurately extracts the image area by combining an image processing tool and a deep learning model, further analyzes the chart-like image, extracts the key features, ensures that the key information in the image can be fully utilized and accurately parsed, performs multimodal fusion of the text and image content to achieve deep association between the text and the image at the semantic level, has significant advantages in comprehensive content analysis and information accuracy, and can effectively solve the problems that the prior art cannot effectively process images embedded in documents, lacks comprehensive processing and association analysis, and has low accuracy in information extraction.

[0102] 2. The present invention can not only extract text from PDF documents, but also provide comprehensive processing capabilities for embedded images. By combining image processing tools and deep learning models, it can accurately extract image areas, analyze chart images, and extract key features to ensure that the key information in the image can be fully utilized and accurately parsed.

[0103] 3. The present invention can perform multimodal fusion of text and image content, realize deep association between text and image at the semantic level, integrate text blocks and image areas on the same page through a multimodal deep learning model, and analyze the semantic relationship between them, thereby providing an overall understanding of the document content.

[0104] 4. The present invention can perform semantic analysis and structural processing on the extracted text and image content, extract key information from the document through natural language processing technology, and organize it into structured data that is easy to understand and use, thereby enhancing the utilization value of the information.

[0105] 5. The present invention can improve the accuracy and depth of information extraction, has significant advantages in image content recognition, multimodal data fusion and semantic analysis, and can significantly improve the accuracy of information extraction and the depth of understanding of document content.

[0106] 6. The present invention is not only suitable for simple PDF document content extraction, but also for scenarios that require in-depth analysis and understanding of document content, such as legal analysis, business report interpretation, etc., thereby broadening the application scenarios of PDF document content extraction methods and providing structured data output to make the extracted information easier to understand and use, thereby improving the user experience.

[0107] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0108] The above descriptions are only some embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A structured analysis method for textile test reports based on multimodal deep learning, characterized in that: The method comprises: Using a document parsing tool to divide the text content and image information in the document into text blocks and image areas respectively, and marking the text blocks and image areas to obtain structured data; Performing image processing and feature extraction on the image region according to the structured data to obtain image features, and matching the image region and the text block according to the image features to obtain a multimodal content block; Using a deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement, and context association on the multimodal content block to obtain multimodal data; The multimodal data is formatted, and a document content file is obtained by combining the relevance between text content and image data in the multimodal data through a deep bidirectional representation model BERT.

2. The method for structural analysis of textile test reports based on multimodal deep learning according to claim 1, characterized in that: The document parsing tool is used to divide the text content and image information in the document into text blocks and image areas respectively, and the text blocks and image areas are marked to obtain structured data, including: Using a document parsing tool to extract text content from each page of the document and divide it into blocks to obtain a plurality of text blocks, and recording the position, size and content of each of the text blocks; Using a document parsing tool to obtain the position, size, and format of the image from each page of the document to obtain multiple image areas; The text blocks and image areas are filtered, and information is marked and integrated on the filtered text blocks and image areas to obtain structured data.

3. The method for structural analysis of textile test reports based on multimodal deep learning according to claim 2, characterized in that: The filtering of the text block and the image area comprises: According to the position, size and content of the text block, remove the text block located at the edge of the page, the text block whose size does not meet the set value and the text block whose content does not meet the set amount; Using a natural language processing algorithm to perform semantic relevance analysis on the text blocks, and removing text blocks whose relevance is lower than a set value; Image regions located at the edge of the page, deviating from the text and having a size that does not meet the set value are removed according to the position and size of the image region, and image regions irrelevant to the text content are filtered out through an image processing algorithm.

4. The method for structural analysis of textile test reports based on multimodal deep learning according to claim 1, characterized in that: The performing image processing and feature extraction on the image region according to the structured data to obtain image features includes: The image region is cropped and converted to black and white binary according to the structured data, and denoised, contrast and brightness are adjusted on the image region using the image processing tool OpenCV; Extracting text from the adjusted image area using a character recognition tool, Tesseract, to obtain image text, and associating the image text with position information of the image area; The image area is classified and identified by deep learning models ResNet and YOLO-V5 to obtain the image type, and the feature extraction of the chart image is performed according to the image type to obtain the image features; the image features include data, coordinate axis labels and legends.

5. The method for structural analysis of textile test reports based on multimodal deep learning according to claim 4, characterized in that: The step of matching the image region and the text block according to the image feature to obtain a multimodal content block includes: The text blocks are sorted and linearized, and the text blocks and image areas on the same page are matched according to the position information in the structured data to obtain multimodal content blocks.

6. The method for structural analysis of textile test reports based on multimodal deep learning according to claim 4, characterized in that: After obtaining the multimodal content block by matching the image region and the text block according to the image features, the method further includes: The multimodal deep learning model CLIP is used to perform data fusion on the multimodal content block according to the text block, image text and image features, so as to fuse the text and image content in the multimodal content block into the same semantic space.

7. The method for structural analysis of textile test reports based on multimodal deep learning according to claim 1, characterized in that: The method of using the deep bidirectional representation model BERT and the natural language processing algorithm to perform semantic analysis, semantic enhancement and context association on the comprehensive data to obtain multimodal data includes: Using the deep bidirectional representation model BERT to perform semantic analysis on the text in the comprehensive data to obtain key information of the text; Combining the image features in the comprehensive data with the key information, the text is semantically enhanced, and the comprehensive data is semantically integrated through a multimodal fusion algorithm; The natural language processing algorithm HuggingFace function is used to analyze the semantic relationship between the text and the image in the comprehensive data to obtain the contextual association between the text and the image to generate multimodal data.

8. A structured analysis device for textile test reports based on multimodal deep learning, characterized in that: The device comprises: A document parsing module, used to use a document parsing tool to divide the text content and image information in the document into text blocks and image areas, and mark the text blocks and image areas to obtain structured data; A content extraction module, configured to perform image processing and feature extraction on the image region according to the structured data to obtain image features, and match the image region and the text block according to the image features to obtain a multimodal content block; A multimodal fusion module uses a deep bidirectional representation model BERT and a natural language processing algorithm to perform semantic analysis, semantic enhancement, and context association on the multimodal content block to obtain multimodal data; The document content extraction module is used to format the multimodal data and obtain a document content file by combining the correlation between the text content and the image data in the multimodal data through the deep bidirectional representation model BERT.

9. An electronic device, characterized in that: include: at least one processor and at least one memory, wherein: The memory has computer-readable instructions stored thereon; The computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the textile test report structured analysis method based on multimodal deep learning as described in any one of claims 1 to 7.

10. A storage medium having computer-readable instructions stored thereon, characterized in that: The computer-readable instructions are executed by one or more processors to implement the structured analysis method for textile test reports based on multimodal deep learning as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Information marking method and device based on large model and information marking method and device

    CN121033881A

  • Multi-modal document analysis method and device, electronic equipment and storage medium

    CN121236783A

  • Multi-format text analysis method and system based on large language model

    CN122113896A