Data processing method and device, equipment and storage medium

Through a multimodal parsing model, mixed-format documents are jointly parsed with tables, tags, and text to generate structured output and calculate comprehensive confidence. This solves the problems of error accumulation and lack of confidence assessment in existing technologies and improves the accuracy and reliability of automated processing.

CN120635926APending Publication Date: 2025-09-12CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510728175.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When processing mixed-format documents, existing technologies have problems such as error accumulation caused by module cascade processing, poor correlation between multimodal data, and lack of result confidence assessment. Especially in scenarios with high reliability requirements such as medical diagnosis and insurance underwriting, the accuracy and reliability of automated processing results are limited.

Method used

By acquiring the target image and loading it into the preset multimodal parsing model, the table, marked area and unstructured text area are parsed separately to generate their respective documents and confidence levels. The comprehensive confidence calculation module is then used to fuse multiple rounds of parsing results to generate a target information list in a unified format, thus achieving dynamic confidence assessment.

Benefits of technology

It effectively eliminates erroneous transmission in intermediate links, improves the accuracy and reliability of data processing, and generates structured output that is directly connected to downstream business systems without the need for additional data cleaning or format conversion, thereby improving processing efficiency and result credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635926A_ABST
    Figure CN120635926A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer software, is applied to the field of medical health and the field of financial science and technology, and discloses a data processing method, device and equipment and a storage medium. Responding to the first instruction, calling a preset model to analyze a first detection area in the target image, and generating a first document and a first confidence coefficient; responding to the second instruction, calling a preset model to analyze a second detection area in the target image, and generating a second document and a second confidence coefficient; responding to the third instruction, calling a preset model to analyze a third detection area in the target image, and generating a third document and a third confidence coefficient; summarizing the target information list; and calculating a comprehensive confidence coefficient. The invention provides a data processing method and device, equipment and a storage medium, and solves the problems of error accumulation, poor multi-modal data relevance and lack of result confidence evaluation caused by module cascade processing in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer software technology and is applied to the fields of medical health and financial technology, and in particular relates to a data processing method, device, equipment and storage medium. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, image-based multimodal data processing has been widely used in automated document analysis in fields such as finance, healthcare, and industry. However, existing technologies still face significant technical bottlenecks when processing mixed-format documents containing structured data (such as tables), visually significant markers (such as symbols and arrows), and unstructured text (such as paragraph descriptions). These bottlenecks are reflected in the following aspects:

[0003] Traditional methods typically employ a phased process, first converting document images into text using optical character recognition (OCR), then segmenting table areas and text paragraphs through layout analysis, and finally parsing the tables and performing named entity recognition (NER) on the text. However, such approaches suffer from severe cascading error propagation. For example, character recognition errors in the OCR stage can lead to field misalignment in subsequent table parsing; while mis-segmentation of complex layouts by the layout analysis module can lead to loss of the relationship between tables and text.

[0004] In scenarios requiring high reliability, such as medical diagnosis and insurance underwriting, automated processing results must be accompanied by quantifiable confidence levels to support manual review or decision-making. However, in traditional staged processing workflows, each module (such as OCR and NER) typically outputs only discrete results, lacking a joint assessment of the overall confidence level. Although some studies have attempted to calculate confidence levels through rule weighting or posterior probability, these results often lack calibration with the true error rate, making them ineffective in guiding automated decision-making.

[0005] In summary, when processing mixed-format documents, existing technologies are limited in accuracy, reliability and business practicality of automated processing results due to problems such as module fragmentation, insufficient multimodal association capabilities, lack of confidence assessment and poor adaptability in vertical fields. Summary of the Invention

[0006] The present invention provides a data processing method, apparatus, device and storage medium to solve the problems of error accumulation, poor correlation of multimodal data and lack of result confidence assessment caused by module cascade processing in the prior art.

[0007] In a first aspect, the present invention provides a data processing method, comprising:

[0008] Acquire a target image, and load the target image into a preset model;

[0009] In response to the first instruction, calling the preset model to analyze the first detection area in the target image, and generating a first document and a first confidence level;

[0010] In response to the second instruction, calling the preset model to analyze the second detection area in the target image and generate a second document and a second confidence level;

[0011] In response to a third instruction, calling the preset model to analyze a third detection area in the target image, and generating a third document and a third confidence level;

[0012] compiling a target information list according to the first document, the second document, and the third document;

[0013] A comprehensive confidence level is calculated based on the first confidence level, the second confidence level, and the third confidence level.

[0014] In a second aspect, the present invention provides a data processing device, comprising:

[0015] An acquisition module is used to acquire a target image and load the target image into a preset model;

[0016] A first generating module is configured to call the preset model to analyze the first detection area in the target image in response to the first instruction, and generate a first document and a first confidence level;

[0017] A second generating module is configured to call the preset model to analyze the second detection area in the target image in response to the second instruction, and generate a second document and a second confidence level;

[0018] A third generating module is configured to call the preset model to analyze the third detection area in the target image in response to a third instruction, and generate a third document and a third confidence level;

[0019] a summarizing module, configured to summarize a target information list according to the first document, the second document, and the third document;

[0020] A calculation module is used to calculate a comprehensive confidence level based on the first confidence level, the second confidence level, and the third confidence level.

[0021] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned data processing method when executing the computer program.

[0022] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned data processing method are implemented.

[0023] In the above-mentioned data processing methods, devices, equipment and storage media, multiple detection areas in the image are jointly analyzed by responding to different instructions, overcoming the limitations of the single analysis module of the traditional method, ensuring that the key information of tabular data, symbolic marks and text descriptions are captured synchronously. Structured output is generated directly based on image input, eliminating the error transmission in the intermediate links. The comprehensive confidence is calculated based on the first confidence, second confidence and third confidence of multiple rounds of analysis results, and the credibility of the results is dynamically evaluated using a multi-task cross-validation mechanism to achieve dynamic confidence fusion to enhance the reliability of the results. A unified format of target information list and confidence annotation is generated, which can be directly connected to the downstream business system without the need for additional data cleaning or format conversion, thereby improving processing efficiency.

[0024] In summary, this solution can solve the problems of error accumulation, poor correlation of multimodal data, and lack of result confidence assessment caused by module cascade processing in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0026] Figure 1 is a flow chart of a data processing method according to an embodiment of the present invention;

[0027] Figure 2 yes Figure 1 A flow chart of step S130;

[0028] Figure 3 yes Figure 1 A flow chart of step S140;

[0029] Figure 4 yes Figure 1 A flow chart of step S150;

[0030] Figure 5 yes Figure 1 A flow chart of step S160;

[0031] Figure 6 yes Figure 5 A flow chart of step S170;

[0032] Figure 7 is another flow chart of a data processing method according to an embodiment of the present invention;

[0033] Figure 8 is a structural diagram of a data processing device in one embodiment of the present invention;

[0034] Figure 9 is a structural diagram of a computer device in one embodiment of the present invention;

[0035] Figure 10 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0037] See also Figure 1 As shown, an embodiment of the present invention provides a flowchart of a data processing method, which includes the following steps.

[0038] Step S120: Acquire a target image and load the target image into a preset model.

[0039] It should be noted that the target image in step S120 is the visual data that needs to be parsed, and its type is strongly related to the specific business scenario. For example, in the field of medical health, the target image is a visual data carrier related to disease diagnosis, treatment or health management, including but not limited to scans of structured tables (such as blood routine, biochemical indicators) and unstructured text (such as doctors' diagnostic opinions). Extract the values ​​of indicators such as blood sugar and cholesterol from the scanned copy of the physical examination report, and associate the reference range to generate structured data. In the field of financial technology, the target image is a document or image related to financial transactions, identity authentication or compliance audits, including: financial documents (such as checks, contracts, invoices, insurance policies); handwritten records (such as customer signatures, handwritten notes). Parse key contract terms (such as amount, term) and the validity of signatures.

[0040] It should be noted that the preset model in step S120 is a multimodal parsing model pre-trained based on a deep learning framework. The multimodal model here supports joint recognition and association analysis of different areas such as tables, marked areas, and unstructured text, converts image content into standardized documents (such as JSON / XML), adapts to business system import, and outputs confidence for each parsing result (such as numerical values, text fields) for result credibility verification. For example, in the field of medical health, medical examination reports are parsed, and scanned copies of blood routine reports (including tables, abnormal marks, and doctor's handwritten suggestions) are parsed. In the field of financial technology, insurance policies are parsed, and scanned copies of auto insurance policies (including insurance information, terms, and signatures) are parsed.

[0041] Step S130 , in response to the first instruction, calling the preset model to parse the first detection area in the target image, and generating a first document and a first confidence level.

[0042] It should be noted that the first instruction triggers the preset model to parse the table area in the target image. This instruction can be initiated by the user (such as clicking the "Table Parse" button) or triggered by the system automatically identifying table features in the image. The first detection area refers to the table area in the image containing structured data, such as the test index table in a medical report or the expense details table in a financial insurance policy.

[0043] In some embodiments of the present invention, Figure 2 As shown, step S130 includes the following steps:

[0044] Step S131, locating the table area in the target image and extracting the data fields in the table;

[0045] Step S132, identifying the value and reference range corresponding to the data field;

[0046] Step S133: Generate a first document, which includes data fields, values, and reference ranges.

[0047] Specifically, in step S131, the table area in the target image is located, and the method of extracting the data fields in the table can be any feasible method. For example, a table detection method can be used to identify the table outline through the target detection model using a pre-trained model such as YOLO, MaskR-CNN, and output the coordinate area of ​​the table (such as [x1, y1, x2, y2] represents the coordinates of the upper left corner and the lower right corner). Or traditional image processing can be used to identify table lines through edge detection (such as Canny operator) and Hough transform, or to determine the table boundary by analyzing the pixel distribution using horizontal / vertical projection. Or deep learning can be combined with traditional methods, such as first locating the approximate area of ​​the table through the model, and then refining the cell boundary using projection analysis. When processing complex tables, cells can be merged, and the range of the merged cells can be identified by analyzing the connectivity and size ratio of adjacent cells, and then processed as a whole in subsequent extraction. The table structure of the borderless table can be inferred by text alignment (such as vertically aligned column headers) and spacing analysis. In the field of medical health, the table of medical reports is usually a test item table, which may contain irregular merged cells (such as the "Test Item" column merging multiple rows). In the field of financial technology, financial document forms may contain complex header hierarchies (such as multi-level account classifications), and the field hierarchical relationship needs to be determined through text semantic association.

[0048] Specifically, in step S132, the method of identifying the numerical value and reference range corresponding to the data field can be any feasible method. For example, the numerical value and reference range corresponding to the data field can be identified by combining optical character recognition (OCR) technology with natural language processing (NLP) algorithm. Compared with the method of simply extracting text by relying on OCR, the text semantics can be understood more accurately, and the precise association between the numerical value and the reference range can be achieved. For example, OCR tools such as Tesseract and PaddleOCR can be used to recognize the text in the table cells and convert the text in the image into editable character information. However, the simple OCR recognition results may have errors or lack of semantic understanding, so it is necessary to use an NLP model, such as the BERT model based on the Transformer architecture, to perform semantic analysis on the recognized text and distinguish information such as field names, numerical values ​​and reference ranges. In actual operation, for the extraction of numerical values, regular expressions can be used to match common numerical patterns, such as integer and decimal formats, to quickly locate the numerical content in the text. For the reference range, regular expressions can also be used to define specific patterns (such as (\d+\.\d+)-(\d+\.\d+)) to match reference range content in a format similar to (3.9-6.1). When regular expressions cannot accurately extract, the semantic analysis model is combined with the domain knowledge base (such as the medical test index reference range library, financial industry standard data, etc.) to perform data matching and inference to obtain accurate reference range information. In the medical and health field, the numerical values ​​and reference ranges in the test reports are expressed in various ways, and there may be problems such as unit conversion. The above method can combine medical semantic knowledge to accurately identify and associate relevant information. In the field of financial technology, the numerical values ​​in financial documents involve various types such as amounts, ratios, and interest rates. The reference range is also closely related to business rules. Through semantic analysis and rule matching, these key information can be effectively identified and extracted.

[0049] Specifically, in step S133, a first document is generated by constructing a structured data model. Compared to simply listing textual information, this makes the document more standardized, readable, and practical, facilitating subsequent data processing and analysis. When generating the document, common structured data formats such as JSON or XML can be used. Taking JSON format as an example, data fields, values, reference ranges, and related confidence information are organized into key-value pairs to form clearly structured data. For example: data field: blood glucose, value: 5.6, unit: mmol / L, reference range: 3.9-6.1, confidence level: 0.95. During the generation process, considering multilingual and different unit requirements, for multilingual text, data fields and related content can be converted into the target language using a translation model or language mapping table. For different units, unified standardization is performed according to pre-set unit conversion rules to ensure data consistency and comparability. Furthermore, confidence information from each processing step is integrated into the document to provide a basis for subsequent data reliability assessment. In healthcare, the generated primary document can be directly connected to electronic medical record systems or medical data analysis platforms, providing structured data support for clinical diagnosis, medical quality assessment, and more. In fintech, structured primary documents facilitate data interaction with financial systems and risk assessment models, helping to automate and intelligentize financial business processes.

[0050] It's understandable that locating table regions by integrating deep learning object detection with a multimodal model algorithm can more accurately identify structured data regions within complex documents than data extraction methods that rely on manual annotation or fixed templates. This technology is particularly critical in medical report parsing scenarios, as accurately capturing biochemical index tables in laboratory reports or measurement parameter tables in diagnostic imaging reports helps clinical research institutions efficiently acquire standardized data, providing automated support for digital archiving of medical records and scientific research data analysis. This feature fusion-based table recognition framework demonstrates strong robustness across a wide range of medical document types. Whether it's handwritten test reports, printed physical examination reports, or mixed-format outpatient records, as long as the document contains a table structure that conforms to medical data standards, the region localization algorithm can extract valid data boundaries, enabling structured parsing. Unlike traditional OCR technologies that focus solely on character recognition accuracy, this method's semantic associations significantly enhance the contextual matching capabilities of medical data fields. Accurately extracting table data forms a crucial foundation for the subsequent establishment of standardized medical databases and the generation of structured reports. For example, when generating electronic health records, key indicators can be automatically filled in according to the field specifications of different departments, and data interfaces that comply with HL7 / FHIR standards can be generated to improve the interoperability of medical information systems and effectively meet the medical institutions' needs for intelligent processing and cross-platform exchange of massive medical record data.

[0051] Step S140 , in response to the second instruction, calling the preset model to parse the second detection area in the target image, and generating a second document and a second confidence level.

[0052] It should be noted that the second instruction triggers the preset model to analyze the marked area in the target image. This instruction can be initiated by the user (e.g., clicking the "Marker Identification" button) or by the system automatically identifying the marked features in the image. The second detection area refers to the area in the image containing the test mark. For example, in the healthcare field, in response to the second instruction, the preset model is invoked to analyze the abnormal marked area in the medical image to generate a structured diagnostic report and confidence level. The second instruction can be triggered by a doctor clicking the "Smart Analysis" button or by the system automatically identifying the abnormal marked features in the test report. The second detection area can be the area in the medical report containing the test result mark (e.g., the abnormality mark next to the blood biochemical index table). In the financial technology field, in response to the second instruction, the preset model is invoked to analyze the risk mark area in the financial document to generate a risk assessment report and confidence level. The second instruction can be triggered automatically by the risk control system (e.g., detecting a suspicious transaction) or manually by the auditor clicking the "Risk Analysis" button. The second detection area refers to the area in the financial document containing the risk mark (e.g., the high-risk mark on the credit card application form).

[0053] In some embodiments of the present invention, Figure 3 As shown, step S140 includes the following steps:

[0054] Step S141, detecting a second detection area in the target image, and identifying a marker type and position coordinates;

[0055] Step S142, associating the inspection items in the table according to the marked positions;

[0056] Step S143, generating an exception description text based on the tag type;

[0057] Step S144: Generate a second document, which includes a marking type, inspection items, and anomaly description.

[0058] Specifically, in step S141, the second detection area in the target image is detected, and the method of identifying the marker type and position coordinates can be any feasible method. For example, a target detection model can be used to identify the marker contour in the second detection area through pre-trained YOLO, FasterR-CNN and other models, and output the marker type (such as "↑", "↓", "*", etc.) and position coordinates (such as [x, y, w, h] represents the center point coordinates and width and height of the marker). Or through traditional image processing methods, first perform pre-processing such as binarization and noise reduction on the image, and then use contour detection, morphological analysis and other methods to identify the shape and position of the marker. Or combine deep learning with traditional methods, such as first locating the approximate area of ​​the marker through the model, and then use template matching or feature comparison to determine the specific marker type. When processing complex markers, the color, size, direction and other attributes of the marker can be considered, and different types of markers can be distinguished through methods such as color space segmentation and geometric feature calculation. In healthcare, medical reports often contain markers indicating abnormal test results. These markers may appear in different colors (e.g., a red "↑" indicates a significant increase, a green "↓" indicates a decrease) or in different shapes (e.g., a circle indicates a condition pending review). In fintech, risk markers may include a variety of special symbols (e.g., "!", "★"), requiring identification through multi-classification models.

[0059] Specifically, in step S142, the method of associating the inspection items in the table according to the mark position can be any feasible method, for example, the mark can be associated with the corresponding inspection item by establishing a spatial mapping relationship between the mark position and the table cell. In step S121, the table area has been located and the data field has been extracted. The row and column information of the table can be used to construct a coordinate index, and the cell where the mark is located can be determined according to the position coordinates of the mark. For example, the distance from the center point of the mark to the center of each cell is calculated, and the mark is associated with the inspection item corresponding to the nearest cell. Or, combined with text semantic analysis, the text content around the mark is understood through the NLP model to further confirm the associated inspection item. When dealing with merged cells, the cell range needs to be expanded according to the merge rules to ensure that the mark can be correctly associated with the corresponding inspection item. In the field of medical health, the mark in the test report may be located in a specific cell of the table and needs to be associated with the inspection item corresponding to the cell (such as blood sugar, blood pressure, etc.). In the field of financial technology, risk markers need to be associated with corresponding financial accounts (such as "cash", "accounts receivable", etc.).

[0060] Specifically, in step S143, the method of generating the abnormal description text based on the tag type can be any feasible method, for example, the corresponding abnormal description text can be generated by combining the preset mapping relationship between the tag type and the abnormal description with the associated examination item. For example, when the tag type is "×" and the associated examination item is "blood sugar", the generated abnormal description text can be "The blood sugar value is abnormal, further examination is recommended". Or, using the NLP model, the tag type and examination item are used as input to generate abnormal information described in natural language. During the generation process, it is possible to consider adding specific numerical information (such as the numerical value and reference range identified in step S122) to make the description more detailed and accurate. For example, when the blood sugar value is 7.8mmol / L and the reference range is 3.9-6.1mmol / L, the generated abnormal description text can be "The blood sugar value 7.8mmol / L exceeds the reference range (3.9-6.1mmol / L), further examination is recommended". In the field of medical health, different tag types may correspond to different preset rules: {"↑": "Above normal range", "↓": "Below normal range", "*": "Recheck required"}. In the field of FinTech, the preset risk level mapping is: {"!": "Medium risk", "★": "High risk", "?": "Needs verification"}.

[0061] Specifically, in step S144, the second document can be generated using any feasible method. For example, the second document can be generated by constructing a structured data model, organizing the marker type, examination items, abnormality description, and related confidence information into a standardized document format. Common document formats such as Word, PDF, and JSON can be used when generating the document. Taking JSON format as an example, each piece of information is organized into key-value pairs to form clearly structured data. For example: {Marker type: ×, Examination item: Blood glucose, Abnormality description: Blood glucose value 7.8mmol / L exceeds the reference range (3.9-6.1mmol / L), further examination recommended, Confidence level: 0.95}. During the generation process, taking into account multilingual and different industry standards, for multilingual text, relevant content can be converted into the target language using a translation model or language mapping table. For different industry standards, adaptation is performed based on preset templates and terminology libraries to ensure that the document complies with industry standards. Furthermore, confidence information from each processing step is integrated into the document to provide a basis for subsequent data reliability assessment. In healthcare, the generated secondary document can be directly connected to electronic medical record systems or medical decision support systems to provide doctors with auxiliary diagnostic information. In the field of financial technology, structured secondary documents can generate standard risk assessment documents to display the risk distribution of each subject.

[0062] It's understandable that the fusion of deep learning object detection and multimodal model algorithms to identify marker regions can more accurately identify marker information in complex documents than marker recognition methods that rely on manual annotation or fixed templates. This technology is particularly critical in medical report parsing scenarios, as accurately capturing abnormal markers in test reports helps clinicians quickly obtain key information, providing a basis for diagnosis and treatment. This feature fusion-based marker recognition framework demonstrates strong robustness across multiple medical document types. Whether it's handwritten test reports, printed physical examination reports, or mixed-format outpatient records, as long as the document contains a marker structure that conforms to medical data standards, the region localization algorithm can extract valid marker information, enabling structured parsing. Unlike traditional OCR technology, which focuses solely on character recognition accuracy, this method's semantic association characteristics significantly enhance the ability to understand the context of medical markers. Accurately extracting marker information forms a crucial foundation for the subsequent establishment of standardized medical databases and the generation of structured reports. For example, when generating electronic health records, key tag information can be automatically extracted according to the specifications of different departments, and standard data interfaces can be generated to improve the interoperability of medical information systems and effectively meet the medical institutions' needs for intelligent processing and cross-platform exchange of massive medical record data.

[0063] Step S150 , in response to the third instruction, calling the preset model to parse the third detection area in the target image, and generating a third document and a third confidence level.

[0064] It should be noted that the third instruction triggers the preset model to parse the unstructured text paragraphs in the target image. This instruction can be triggered manually by the user (such as clicking the "Text Analysis" button) or automatically by the system based on the presence of free text paragraphs in the image (such as continuous text blocks). The third detection area refers to free text paragraphs in the image that are not presented in a structured form such as a table, such as diagnostic opinions in medical reports and risk clauses in financial contracts.

[0065] In some embodiments of the present invention, Figure 4 As shown, step S150 includes the following steps:

[0066] Step S151, locating an unstructured text paragraph in the target image;

[0067] Step S152, extracting abnormality description keywords and related inspection items in the paragraph;

[0068] Step S153: Generate a third document, wherein the third document includes an exception category and description content.

[0069] Specifically, in step S151, the unstructured text paragraphs in the target image are located, which can be achieved in a variety of ways. For example, a text detection model based on deep learning, such as pre-trained models such as TextBoxes and EAST, can be used to identify the text area in the image and output its coordinate range (such as [x1, y1, x2, y2] represents the coordinates of the upper left and lower right corners of the text box). This type of model extracts image features through a convolutional neural network and combines regression and classification algorithms to predict the position and boundaries of the text area. Traditional image processing methods can also be used to first grayscale and reduce noise on the image, and then locate the text block through methods such as connected domain analysis and projection histogram. In addition, optical character recognition (OCR) technology can be combined to assist in positioning. First, OCR tools such as Tesseract and PaddleOCR are used to perform text recognition on the image, and then the position of the text paragraph is inferred based on the coordinate information of the recognition result. In the field of medical health, diagnostic opinion paragraphs may be mixed with test indicator tables, and pure text paragraphs need to be separated from the complex layout through methods such as text density analysis and line spacing judgment; in the field of financial technology, long paragraphs of explanatory text in contract terms may be interspersed with special symbols or formats, and visual features such as paragraph indentation and font style can be used to assist in positioning.

[0070] Specifically, in step S152, extracting abnormal description keywords and associated inspection items from the paragraph can be achieved with the help of natural language processing (NLP) technology. First, the text paragraph is converted into editable character information using an OCR tool, and then semantic analysis is performed using an NLP model. For example, models based on the Transformer architecture, such as BERT and RoBERTa, are fine-tuned with domain-specific corpora to identify abnormal keywords in the text (such as "lesion" and "infection" in the medical field, and "default" and "overdue" in the financial field). At the same time, named entity recognition (NER) technology is used to extract key entities in the text (such as the names of inspection items) and establish an association between keywords and inspection items. In actual operation, regular expressions can also be used to match keywords with specific patterns. For example, in medical reports, r'(suspected|confirmed).*? (tumor|inflammation)' is used to capture disease-related descriptions; in financial contracts, r'(not on time|delayed).*? (repayment|payment)' is used to identify default risk statements. When the text semantics are ambiguous, domain knowledge bases (such as medical diagnostic standard libraries and financial regulatory terms libraries) can be combined for reasoning to improve the accuracy and relevance of keyword extraction.

[0071] Specifically, in step S153, the generation of the third document can be achieved by constructing a structured data model. To facilitate subsequent data processing and analysis, the anomaly category and description content can be organized into a structured format such as JSON or XML. Taking JSON format as an example, the following structure can be constructed: {"anomaly category":"cardiovascular disease","description":"Examination revealed the presence of atherosclerosis in the coronary arteries, suspected of causing myocardial ischemia","confidence level"}. During the document generation process, issues such as multilingualism and unified terminology need to be considered. For multilingual text, machine translation models (such as Transformer-based translation models) or language mapping tables can be used to convert the content into the target language. For specialized terminology in different fields, pre-set terminology standardization rules are used to standardize the content. Furthermore, confidence level information for each processing step is integrated into the document to provide a basis for subsequent data reliability assessment. In the healthcare sector, the generated third document can be directly connected to the electronic medical record system to provide doctors with auxiliary diagnostic references. In the fintech sector, structured third documents help risk warning systems quickly identify potential risks in contract terms and support intelligent decision-making in financial services.

[0072] It is understandable that by fusing deep learning text detection with natural language processing technology to parse unstructured text, it is possible to more efficiently and accurately mine key information from the text than with traditional manual extraction or simple keyword matching methods. In medical report scenarios, this technology can automatically extract disease descriptions, treatment recommendations, and other information from doctors' handwritten or printed diagnostic opinions, accelerating the digitization of medical records and providing data support for clinical research and intelligent auxiliary diagnosis. In the processing of financial contracts, it can quickly identify risk clauses, breach of contract liability, and other content, reducing manual review costs and improving the risk control efficiency and compliance of financial services. This text analysis framework based on multimodal fusion shows good adaptability to texts with different layouts and different professional backgrounds, effectively meeting the needs of the medical and financial fields for intelligent processing of unstructured data.

[0073] Step S160: compile a target information list based on the first document, the second document, and the third document.

[0074] It should be noted that the third instruction is an instruction that triggers the system to integrate the three previous documents. This instruction can be triggered manually by the user (such as clicking the "Information Summary" button), or the system can be automatically triggered after the generation of the first document, the second document, and the third document. The target information list is intended to integrate structured, semi-structured, and unstructured information scattered in different documents into a unified, concise, and easy-to-use format, which can be applied to scenarios such as diagnostic decision-making in the medical and health field and risk assessment in the financial technology field.

[0075] In some embodiments of the present invention, Figure 5 As shown, step S160 includes the following steps:

[0076] Step S161: Merge the first document, the second document, and the third document;

[0077] Step S162: remove redundant items and generate a target information list.

[0078] Specifically, in step S161, the way of fusing the first document, the second document and the third document can be any feasible way. For example, data alignment and fusion can be performed based on public identifiers. By identifying key identifiers in the documents, such as the names of examination items in the medical field, financial subjects or clause names in the financial field, the relevant information in different documents is associated and integrated. In terms of technical implementation, for structured first documents (such as tabular data), semi-structured second documents (such as tag parsing results) and unstructured third documents (such as text summaries), knowledge graph technology can be used to construct an entity relationship network to achieve fusion. First, use natural language processing technology to extract entities (such as medical examination items, financial terms) and attributes (numerical values, exception descriptions) in the document, and then establish associations between entities through a graph database.

[0079] In the field of medical health, taking diabetes diagnosis as an example, the blood sugar test value in the first document, the abnormal mark of the blood sugar index in the second document (such as "↑"), and the symptom description of excessive drinking and eating in the medical record text in the third document can be integrated through the key identifier "blood sugar". By constructing a medical knowledge graph, the blood sugar value is compared with the normal range, and combined with the abnormal mark and symptom description, a comprehensive judgment of the patient's blood sugar status is formed. In the field of financial technology, for a loan business, the loan amount, repayment period and other data in the first document, the risk mark of the loan in the second document (such as "★" indicates high risk), and the default clause in the loan contract in the third document can be based on the public identifier of "loan business number" and integrated with the financial knowledge graph to clearly present the overall risk status of the loan business.

[0080] Furthermore, the fusion process requires addressing the confidence of multi-source data. A weighted average algorithm can be used to assign weights based on the reliability of different document data sources. For example, in the medical field, laboratory test data (the first document) would be given a higher weight, while doctor's handwritten annotations (the third document) would be given a lower weight due to the potential for typos. In the financial field, contract terms (the third document) would be given a higher weight than manually labeled risk levels (the second document).

[0081] Specifically, in step S162, redundant items are removed and a target information list is generated. This can be achieved in a variety of ways. For redundant item detection, text similarity calculation and duplicate key value detection can be utilized. The similarity between text descriptions is calculated using algorithms such as cosine similarity, and a threshold is set to filter out duplicate descriptions with excessive similarity. For data with the same key value (e.g., the same inspection item or the same financial account), the record with the highest confidence is retained. When generating the target information list, sorting and structuring can be performed according to business needs.

[0082] In healthcare, integrated information is sorted by disease risk level, with information related to high-risk diseases prioritized. This resulting list of targeted information can directly assist doctors in diagnosis and provide a comprehensive reference for developing treatment plans. For example, cardiovascular disease-related test data, abnormality markers, and medical records are integrated and sorted by risk, allowing doctors to quickly identify key issues. In fintech, loan business information is sorted based on risk scores, with high-risk businesses at the forefront. The resulting list allows risk control personnel to quickly identify risky businesses and take timely action. Furthermore, the output of lists in structured formats such as JSON and XML facilitates data exchange with other systems.

[0083] It's understandable that the aforementioned fusion and processing techniques, when used to compile a target information list, can more comprehensively and accurately reflect the actual situation than simply using each document's data separately. In healthcare, integrating multi-dimensional information helps doctors avoid missed or misdiagnosed cases, improving diagnostic efficiency and accuracy, and providing patients with higher-quality medical services. In fintech, it helps companies promptly identify potential risks, optimize risk control processes, and enhance the scientific and rational nature of financial business decisions, thereby strengthening their competitiveness.

[0084] Step S170: Calculate a comprehensive confidence level based on the first confidence level, the second confidence level, and the third confidence level.

[0085] It should be noted that the first, second, and third confidence levels correspond to the reliability of the first document (structured table data parsing results), the second document (markup parsing results), and the third document (unstructured text parsing results), respectively. This command is automatically triggered after the target information list is generated. It aims to provide a credibility reference for the final output information through quantitative assessment, assisting diagnostic decisions in the healthcare field and risk assessment in the financial technology field.

[0086] In some embodiments of the present invention, Figure 6 As shown, step S170 includes the following steps:

[0087] Step S171, determining the first confidence level, the second confidence level, and the third confidence level, and taking the minimum value as the comprehensive confidence level;

[0088] Step S172: If the target information is repeatedly parsed by the first document, the second document, and the third document, the comprehensive confidence is 1.

[0089] Specifically, in step S171, the first confidence level, the second confidence level, and the third confidence level are determined, and the minimum value is taken as the comprehensive confidence level. This is based on the robustness principle of the "barrel effect". In practical applications, the reliability of data processing depends on the most unreliable link. For example, in the field of medical health, the first document may be a test index table generated by high-precision equipment detection, and its first confidence level is 0.95; the second document is the analysis of abnormal marks in the test report, and due to fuzzy marks and other reasons, the second confidence level is 0.8; the third document is a text analysis of the doctor's handwritten diagnosis opinion, and due to handwriting clarity issues, the third confidence level is 0.85. At this time, taking the minimum value of 0.8 as the comprehensive confidence level means that the reliability of the overall result is limited by the most uncertain mark analysis link, avoiding the masking of overall risks due to local high confidence levels, and ensuring that doctors have a clear understanding of reliability when referring to information.

[0090] In the fintech sector, the first document might be a detailed expense statement from a financial statement, with a first confidence level of 0.9; the second document is an interpretation of the risk markers in the statement, with a second confidence level of 0.88; and the third document is an analysis of the risk description in the contract terms, with a third confidence level of 0.82. By taking the minimum value of 0.82 as the overall confidence level, risk control personnel can intuitively understand the reliability of the risk assessment results, making decisions more cautious and avoiding misjudgments caused by over-reliance on a few high-confidence data points.

[0091] Specifically, in step S172, if the target information is repeatedly parsed by the first document, the second document, and the third document, the comprehensive confidence is 1, which is based on the logic of cross-validation of multi-source data. When the same target information obtains consistent conclusions in structured data, tag parsing, and text description, it means that the accuracy of the information has been verified from multiple angles. For example, in the field of medical health, for the diagnosis of "hypertension", the blood pressure test value in the first document exceeds the standard, corresponding to a first confidence of 0.9; the second document marks "abnormal" next to the blood pressure indicator, and the second confidence is 0.85; the doctor's handwritten diagnosis in the third document explicitly mentions "hypertension", and the third confidence is 0.88. Since the three documents all point to the same conclusion, the comprehensive confidence is set to 1 at this time, indicating that the diagnosis result is highly reliable, and doctors can formulate treatment plans based on it with more confidence.

[0092] In the fintech sector, a high-risk assessment for a loan involves assessing the financial data in the first document indicating an excessive debt ratio, with a first confidence level of 0.92; a second document with a "★" sign indicating high risk, with a second confidence level of 0.89; and a third document with contractual clauses clearly listing multiple conditions that hinder repayment, with a third confidence level of 0.9. These three documents corroborate each other, resulting in a combined confidence level of 1. This helps risk management departments take decisive action, such as rejecting loan applications or increasing collateral requirements, to ensure financial security.

[0093] It's understandable that the aforementioned method for calculating comprehensive confidence can scientifically quantify information reliability. In healthcare, this helps doctors quickly determine the credibility of diagnostic evidence, reducing the likelihood of misdiagnosis and missed diagnoses, and improving the quality of medical services. In fintech, it provides a clear reliability reference for risk management decisions, optimizes risk assessment processes, reduces financial risks, and promotes efficient and stable business development.

[0094] In some embodiments of the present invention, step 110 is further included before step 120, such as Figure 7 As shown, the following steps are also included before step S120.

[0095] The following steps aim to provide reliable model support for subsequent image parsing tasks. By acquiring image data, performing data augmentation, and model training, the model can adapt to complex and diverse image scenarios in fields such as healthcare and finance.

[0096] Step S111: acquiring an image dataset, generating text annotations based on an OCR service, and screening high-confidence images;

[0097] Step S112, performing data enhancement processing such as noise injection, geometric deformation or resolution adjustment on the image data set;

[0098] Step S113: Using the target image as input and the OCR stitched text as output, the model is trained to generate text and confidence.

[0099] Specifically, in step S111, an image dataset is acquired, and text annotations are generated using an OCR service, followed by screening for high-confidence images. In the medical field, image datasets primarily come from hospital electronic medical record systems, image archiving, and communication systems, and include various types of test reports and medical imaging diagnostic reports. These reports may contain handwritten text or blurry printing, increasing recognition difficulty. OCR services such as PaddleOCR and Tesseract are used to recognize the text in the images and generate initial annotations. Due to the extremely high professional and accuracy requirements of medical data, a high confidence threshold, such as 0.9, is set to screen for images with reliable recognition results. For example, for key data such as "blood sugar: 5.6mmol / L" on a test report, an image is included in the dataset only if the OCR recognition confidence exceeds the threshold. In the financial field, image datasets cover a variety of documents, including bank statements, loan agreements, and insurance policies. These documents often contain complex headers, seals, handwritten signatures, and other elements. Information such as the amount, date, and terms in the documents is annotated by invoking services such as Google Cloud Vision and Smart Cloud OCR. Given the rigor of financial data, strict screening criteria are also applied to eliminate low-confidence images. For example, when identifying the amount clause in a loan contract, if the OCR recognition result is not confident enough, the contract image will be excluded.

[0100] Specifically, in step S112, the image dataset is subjected to data enhancement processing such as noise injection, geometric deformation, or resolution adjustment. In the medical field, various data enhancement operations are performed on images to simulate complex situations in actual applications. Noise injection is performed to add Gaussian noise to simulate signal interference during the scanning process, or salt and pepper noise is added to simulate spots caused by film aging; geometric deformation is performed to randomly rotate and tilt medical imaging reports to simulate shooting or scanning at different angles; resolution adjustment is performed to reduce image resolution to simulate the use of low-quality scanning equipment in primary medical institutions. These operations can expand the diversity of the dataset and improve the model's adaptability to various complex medical images. In the financial field, data enhancement operations are closely centered around the characteristics of financial documents. By adding occlusion simulations of seals and signatures, the model learns to accurately recognize text in the presence of interfering elements; perspective transformation is performed on contracts and insurance policies to simulate the scanning effect of folded and curled documents; and image resolution is adjusted to test the model's ability to extract key information such as amounts and terms from low-definition images. For example, bank statements can be rotated and stained to enable the model to cope with irregular and stained documents in actual business.

[0101] Specifically, in step S113, the model is trained to generate text and confidence scores using the target image as input and the OCR-joined text as output. In the medical field, a model suitable for medical text recognition, such as TrOCR based on the Transformer architecture, is selected. Professional medical corpus is introduced during training for fine-tuning. The filtered and enhanced medical image is used as input, and the accurate OCR-joined text is used as the output label. The model parameters are continuously optimized using a backpropagation algorithm. During training, the model focuses on the accuracy of recognition of key information such as medical terminology and laboratory parameters, and generates corresponding confidence scores for the text output by the model. For example, for the recognition of professional terms such as "coronary atherosclerosis," the model must not only output accurate text but also provide a reliable confidence score to provide a reference for subsequent medical diagnosis. In the financial field, models that excel at processing document layout information, such as LayoutLMv3, are used for training. Financial document images are input into the model, and the accurately annotated text content serves as a supervisory signal. The model is trained to learn the structural and semantic features of financial text. Recognizing the characteristics of financial data, the model is trained to strengthen its ability to recognize key information such as amounts, dates, and terms. At the same time, by improving the model structure or training algorithm, the model can output the confidence level of the text. For example, when identifying the repayment terms in a loan contract, the model can accurately extract the text and give the corresponding confidence level, making it easier for financial practitioners to assess the reliability of the information and assist in risk decision-making.

[0102] It can be understood that through the series of operations in step S110, a high-quality, diverse image dataset has been constructed, and a high-precision OCR model suitable for the medical and financial fields has been trained. In the medical field, this model can more accurately identify text information such as test reports and diagnostic records, helping doctors quickly obtain key data and improve diagnostic efficiency and accuracy. In the financial field, it can efficiently process various documents and accurately extract information such as amounts and terms, helping financial institutions to achieve automation and intelligent business processes, reduce labor costs, and enhance risk prevention and control capabilities.

[0103] Thus, the solution of the embodiment of the present invention achieves intelligent processing of complex documents in the medical and financial fields through multi-stage image analysis and cross-modal data fusion. First, in step S110, a high-quality dataset is constructed and the OCR model is optimized to solve the text recognition problem in scenarios such as blurred handwriting in medical reports and occluded seals on financial documents. For example, in the medical field, by injecting film noise and handwriting GAN enhancement, the model can accurately identify key indicators such as "blood sugar level" in test reports; in the financial field, by simulating seal coverage and perspective deformation, the recognition robustness of key terms such as "amount capitalization" in loan contracts is improved.

[0104] The solution uses multi-document parsing and information aggregation to cross-modally fuse structured table data (first document), semi-structured tags (second document), and unstructured text (third document). In medical scenarios, test indicator values ​​(such as "white blood cell count 12.3×10^9 / L"), abnormality marks ("↑"), and diagnostic recommendations ("indicating possible infection") can be integrated into a structured diagnostic summary to avoid missing information from a single modality. In financial scenarios, financial account data (such as "accounts receivable 5 million yuan"), risk marks ("★"), and contract terms ("daily penalty interest of 0.05%") can be linked to form a three-dimensional risk assessment report, addressing the one-sidedness of traditional single-modality analysis.

[0105] Through confidence quantification and reliability assessment, the "minimum value principle" and "multi-source cross-validation" mechanisms are adopted to scientifically quantify the reliability of information. For example, in medical scenarios, when there is a conflict between the test value (confidence 0.95), tag analysis (0.8), and text summary (0.85), the minimum value of 0.8 is used as the comprehensive confidence to avoid overestimating the reliability of the data; if the conclusions of the three are consistent (such as all pointing to "high blood pressure"), the confidence is set to 1 to strengthen the credibility of the diagnostic basis. In the financial field, this mechanism can assist risk control personnel in quickly identifying high-confidence risk points (such as "penalty interest rate exceeds the regulatory red line") and improve decision-making efficiency.

[0106] The multi-domain processing of the solution significantly improves the accuracy and practicality of complex document parsing. In the field of healthcare, the generated structured reports can be directly connected to the electronic medical record system of the HL7 / FHIR standard, providing standardized data support for clinical diagnosis and scientific research; in the field of financial technology, risk assessment reports that comply with the BaselIII standard can be seamlessly connected to the risk control system to assist in automated compliance review. Through multi-stage technology integration and domain knowledge injection, the present invention effectively solves the problems of diverse data modalities, complex formats, and high professional barriers in medical and financial documents, and provides an efficient data processing paradigm for smart healthcare and smart finance.

[0107] In summary, the solution implemented by the embodiment of the present invention can solve the problems of error accumulation, poor correlation of multimodal data, and lack of result confidence assessment caused by module cascade processing in the prior art.

[0108] It should be understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation of the embodiments of the present invention. The use of non-Company software tools or components in the embodiments of this application is for illustrative purposes only and does not represent actual use.

[0109] In one embodiment, a data processing device is provided, which corresponds one-to-one to the data processing method in the above embodiment. Figure 8 As shown, the data processing device includes an acquisition module 810, an analysis module 820, a segmentation module 830, an extraction module 840 and a prediction module 850. The functional modules are described in detail as follows:

[0110] An acquisition module 810 is used to acquire a target image and load the target image into a preset model;

[0111] A first generating module 820 is configured to call the preset model to analyze the first detection area in the target image in response to the first instruction, and generate a first document and a first confidence level;

[0112] A second generating module 830 is configured to call the preset model to analyze the second detection area in the target image in response to the second instruction, and generate a second document and a second confidence level;

[0113] A third generating module 840 is configured to call the preset model to analyze the third detection area in the target image in response to the third instruction, and generate a third document and a third confidence level;

[0114] A summarizing module 850 is configured to summarize a target information list based on the first document, the second document, and the third document;

[0115] The calculation module 860 is configured to calculate a comprehensive confidence level based on the first confidence level, the second confidence level, and the third confidence level.

[0116] In one embodiment, the first generating module 820 is specifically configured to:

[0117] Locating a table area in the target image and extracting data fields in the table;

[0118] Identify the numerical value and reference range corresponding to the data field;

[0119] A first document is generated, the first document including data fields, values, and reference ranges.

[0120] In one embodiment, the second generating module 830 is specifically configured to:

[0121] Detecting a second detection area in the target image and identifying a marker type and position coordinates;

[0122] Associate the inspection items in the table according to the marked positions;

[0123] Generate exception description text based on tag type;

[0124] A second document is generated, wherein the second document includes a marking type, inspection items, and an exception description.

[0125] In one embodiment, the third generating module 840 is specifically configured to:

[0126] locating an unstructured text paragraph in the target image;

[0127] Extract abnormal description keywords and related inspection items in the paragraph;

[0128] A third document is generated, wherein the third document includes an exception category and description content.

[0129] In one embodiment, the aggregation module 850 is specifically configured to:

[0130] fusing the first document, the second document, and the third document;

[0131] Remove redundant items and generate a target information list.

[0132] In one embodiment, the calculation module 860 is specifically configured to:

[0133] Determine the first confidence level, the second confidence level, and the third confidence level, and take the minimum value as the comprehensive confidence level;

[0134] If the target information is repeatedly parsed by the first document, the second document, and the third document, the comprehensive confidence is 1.

[0135] In one embodiment, the acquisition module 810 is specifically configured to:

[0136] Obtain image datasets, generate text annotations based on OCR services, and filter high-confidence images;

[0137] Performing data enhancement processing such as noise injection, geometric deformation or resolution adjustment on the image dataset;

[0138] The target image is used as input and the OCR stitched text is used as output, and the model is trained to generate text and confidence.

[0139] The present invention provides a solution provided by a data processing device, which jointly analyzes multiple detection areas in an image by responding to different instructions, overcomes the limitations of a single analysis module of traditional methods, and ensures that key information of tabular data, symbolic labels, and text descriptions are captured synchronously. Structured output is generated directly based on image input, eliminating erroneous transmission in intermediate links. The comprehensive confidence is calculated based on the first confidence, second confidence, and third confidence of multiple rounds of analysis results, and the credibility of the results is dynamically evaluated using a multi-task cross-validation mechanism, realizing dynamic confidence fusion to enhance the reliability of the results. A unified format of target information lists and confidence annotations is generated, which can be directly connected to downstream business systems without the need for additional data cleaning or format conversion, thereby improving processing efficiency.

[0140] In summary, the solution implemented by the above data processing method can solve the problems of error accumulation, poor correlation of multimodal data and lack of result confidence assessment caused by module cascade processing in the existing technology.

[0141] For the specific definition of the data processing device, please refer to the definition of the data processing method above and will not be repeated here. Each module in the above-mentioned data processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software so that the processor can call and execute the operations corresponding to each of the above modules.

[0142] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the service-side of a method for recommending an optimal strategy.

[0143] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the client side of a method for recommending an optimal strategy.

[0144] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0145] Acquire a target image, and load the target image into a preset model;

[0146] In response to the first instruction, calling the preset model to analyze the first detection area in the target image, and generating a first document and a first confidence level;

[0147] In response to the second instruction, calling the preset model to analyze the second detection area in the target image and generate a second document and a second confidence level;

[0148] In response to a third instruction, calling the preset model to analyze a third detection area in the target image, and generating a third document and a third confidence level;

[0149] compiling a target information list according to the first document, the second document, and the third document;

[0150] A comprehensive confidence level is calculated based on the first confidence level, the second confidence level, and the third confidence level.

[0151] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0152] Acquire a target image, and load the target image into a preset model;

[0153] In response to the first instruction, calling the preset model to analyze the first detection area in the target image, and generating a first document and a first confidence level;

[0154] In response to the second instruction, calling the preset model to analyze the second detection area in the target image and generate a second document and a second confidence level;

[0155] In response to a third instruction, calling the preset model to analyze a third detection area in the target image, and generating a third document and a third confidence level;

[0156] compiling a target information list according to the first document, the second document, and the third document;

[0157] A comprehensive confidence level is calculated based on the first confidence level, the second confidence level, and the third confidence level.

[0158] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0159] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0160] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0161] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A data processing method, characterized in that: include: Acquire a target image, and load the target image into a preset model; In response to the first instruction, calling the preset model to analyze the first detection area in the target image, and generating a first document and a first confidence level; In response to the second instruction, calling the preset model to analyze the second detection area in the target image and generate a second document and a second confidence level; In response to a third instruction, calling the preset model to analyze a third detection area in the target image, and generating a third document and a third confidence level; compiling a target information list according to the first document, the second document, and the third document; A comprehensive confidence level is calculated based on the first confidence level, the second confidence level, and the third confidence level.

2. The data processing method according to claim 1, wherein: The step of calling the preset model in response to the first instruction to analyze the first detection area in the target image and generating the first document and the first confidence level includes: Locating a table area in the target image and extracting data fields in the table; Identify the numerical value and reference range corresponding to the data field; A first document is generated, the first document including data fields, values, and reference ranges.

3. The data processing method according to claim 1, wherein: The step of calling the preset model in response to the second instruction to analyze the second detection area in the target image and generating a second document and a second confidence level includes: Detecting a second detection area in the target image and identifying a marker type and position coordinates; Associate the inspection items in the table according to the marked positions; Generate exception description text based on tag type; A second document is generated, wherein the second document includes a marking type, inspection items, and an exception description.

4. The data processing method according to claim 1, wherein: The step of calling the preset model in response to the third instruction to analyze the third detection area in the target image and generating a third document and a third confidence level includes: locating an unstructured text paragraph in the target image; Extract abnormal description keywords and related inspection items in the paragraph; A third document is generated, wherein the third document includes an exception category and description content.

5. The data processing method according to claim 1, wherein: The step of compiling a target information list according to the first document, the second document, and the third document includes: fusing the first document, the second document, and the third document; Remove redundant items and generate a target information list.

6. The data processing method according to claim 1, wherein: The step of calculating the comprehensive confidence level according to the first confidence level, the second confidence level, and the third confidence level comprises: Determine the first confidence level, the second confidence level, and the third confidence level, and take the minimum value as the comprehensive confidence level; If the target information is repeatedly parsed by the first document, the second document, and the third document, the comprehensive confidence is 1.

7. The data processing method according to claim 1, wherein: The step of acquiring a target image and loading the target image into a preset model includes: Obtain image datasets, generate text annotations based on OCR services, and filter high-confidence images; Performing data enhancement processing such as noise injection, geometric deformation or resolution adjustment on the image dataset; The target image is used as input and the OCR stitched text is used as output, and the model is trained to generate text and confidence.

8. A data processing device, characterized in that: include: An acquisition module is used to acquire a target image and load the target image into a preset model; A first generating module is configured to call the preset model to analyze the first detection area in the target image in response to the first instruction, and generate a first document and a first confidence level; A second generating module is configured to call the preset model to analyze the second detection area in the target image in response to the second instruction, and generate a second document and a second confidence level; A third generating module is configured to call the preset model to analyze the third detection area in the target image in response to a third instruction, and generate a third document and a third confidence level; a summarizing module, configured to summarize a target information list according to the first document, the second document, and the third document; A calculation module is used to calculate a comprehensive confidence level based on the first confidence level, the second confidence level, and the third confidence level.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the data processing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Hybrid PDF (Portable Document Format) document analysis and knowledge fragment construction method and device

    CN121527785A