File content identification method and file content identification device
By analyzing file contents and selecting and fusion, the problem of low accuracy of file content recognition is solved, efficient automatic recognition is achieved, and manual intervention is reduced.
Patent Information
- Application Number
- CN202510327034.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the file content recognition model has low accuracy in identifying various types of content, resulting in low overall recognition efficiency and requires more manual intervention.
By parsing the files uploaded by users, selecting target recognition models that are suitable for different content types for identification, and integrating and fusion of the recognition results to reduce manual intervention.
It improves the accuracy of automatic identification of file content, reduces manual participation, improves recognition efficiency, and avoids single model identification errors and resource utilization.
Smart Images

Figure CN120337924A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a file content recognition method and a file content recognition device. Background Art
[0002] With the widespread application of knowledge-based question-answering systems and large language models, it is becoming increasingly important to extract content from files and convert it into structured data. The key to content extraction is to identify various types of content in files (such as text content, table content, and image content). Currently, character recognition models (such as PaddleOCR's PP-Structure model) are mainly used to automatically identify the content in files.
[0003] However, in practice, it has been found that since there may be multiple types of content in a file at the same time, and the layout formats of different types of content in the file may be more complicated, the character recognition model may easily misrecognize certain types of content in the file, resulting in low recognition accuracy of the overall content of the file. Further processing requires less efficient manual recognition methods, which reduces the overall efficiency of the file content recognition process.
[0004] Therefore, how to improve the accuracy of automatic recognition of file content, thereby reducing the degree of manual participation and improving content recognition efficiency is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The present invention provides a file content recognition method and a file content recognition device, which can improve the accuracy of automatic file content recognition, thereby reducing the degree of manual participation to improve content recognition efficiency.
[0006] In order to solve the above technical problem, the first aspect of the present invention discloses a method for identifying file content, the method comprising:
[0007] Parsing the file to be detected uploaded by the user to obtain at least one set of content to be identified; wherein the set of content to be identified is a set of content formed by the same type of content in the file to be detected, and the file to be detected contains at least one of text content, table content and image content;
[0008] For each of the content sets to be identified, at least one target recognition model adapted to the content set to be identified is selected from the preset model library to perform content recognition on the content set to be identified, and the content recognition results of all the target recognition models for the content set to be identified are integrated to obtain an initial content recognition result corresponding to the content set to be identified; wherein the target recognition model includes a content recognition model in the preset model library that meets a preset screening condition;
[0009] Fuse the initial content recognition results corresponding to all the to-be-recognized content sets to obtain the target content recognition result corresponding to the to-be-detected file.
[0010] As an optional implementation manner, in the first aspect of the present invention, the preset screening condition includes a weight screening condition;
[0011] For each of the to-be-recognized content sets, select at least one target recognition model adapted to the to-be-recognized content set from the preset model library to perform content recognition on the to-be-recognized content set, and integrate the content recognition results of all the target recognition models for the to-be-recognized content set to obtain the initial content recognition result corresponding to the to-be-recognized content set, including:
[0012] For each of the to-be-recognized content sets, determine the selection weights of all the content recognition models in the preset model library according to the content type corresponding to the to-be-recognized content set, and determine at least one target recognition model adapted to the to-be-recognized content set from the preset model library according to the weight screening condition and the selection weights of all the content recognition models;
[0013] For each of the target recognition models, perform content recognition on the to-be-recognized content set corresponding to the target recognition model to obtain the content recognition result corresponding to the target recognition model, and send the content recognition result to the user;
[0014] Obtain the user judgment data feedback by the user; wherein, the user judgment data is the data generated by the user's judgment on the content recognition results corresponding to all the target recognition models;
[0015] Screen the content recognition results corresponding to all the target recognition models according to the user judgment data to obtain the initial content recognition result.
[0016] As an optional implementation manner, in the first aspect of the present invention, for each of the to-be-recognized content sets, determine the selection weights of all the content recognition models in the preset model library according to the content type corresponding to the to-be-recognized content set, and determine at least one target recognition model adapted to the to-be-recognized content set from the preset model library according to the weight screening condition and the selection weights of all the content recognition models, including:
[0017] For each of the to-be-recognized content sets, obtain the historical recognition data corresponding to each content recognition model in the preset model library, assign an initial weight to each content recognition model in the preset model library according to all the historical recognition data, obtain the target content characteristics of the to-be-recognized content set, and adjust the initial weights of all content recognition models according to the target content characteristics to obtain the selection weights corresponding to each content recognition model in the preset model library. Screen all content recognition models in the preset model library according to the weight screening conditions and the selection weights corresponding to each content recognition model to obtain at least one target recognition model adapted to the to-be-recognized content set;
[0018] Among them, the target content characteristics include one of the text character quantity, table row-column complexity, and image resolution. Meeting the weight screening condition specifically means that the selection weight corresponding to the target recognition model is greater than or equal to a preset weight threshold;
[0019] For each content recognition model in the preset model library, the historical recognition data corresponding to the content recognition model includes recognition performance data and model trustworthiness. The recognition performance data can characterize the high or low recognition performance of the content recognition model in recognizing content of the target type during the historical time period, and the target type is the same as the content type included in the to-be-detected file. The model trustworthiness can characterize the high or low reliability of the content recognition model when performing the content recognition task. Both the model trustworthiness and the recognition performance data are used to determine the size of the initial weight that can be assigned to the content recognition model.
[0020] As an optional implementation manner, in the first aspect of the present invention, the screening of the content recognition results corresponding to all the target recognition models according to the user evaluation data to obtain the initial content recognition results includes:
[0021] Perform score conversion on the user evaluation data to obtain the user score data corresponding to each target recognition model; wherein, the user score data corresponding to the target recognition model is used to characterize the accuracy and reliability of the user's determination of the content recognition result corresponding to the target recognition model, as well as the high or low degree of the recognition efficiency of the target recognition model;
[0022] For each target recognition model, obtain the historical score data of the target recognition model, and update the historical score data according to the user score data of the target recognition model to obtain the current score data of the target recognition model; wherein, the historical score data of the target recognition model is the score data updated after the target recognition model performs the historical recognition task;
[0023] Filter the current scoring data of all the target recognition models according to the preset scoring conditions to obtain target scoring data, and use the content recognition result corresponding to the target scoring data as the initial content recognition result.
[0024] As an optional implementation manner, in the first aspect of the present invention, for each of the target recognition models, obtain the historical scoring data of the target recognition model, and update the historical scoring data according to the user scoring data of the target recognition model to obtain the current scoring data of the target recognition model, including:
[0025] For each of the target recognition models, obtain the historical scoring data and the current contribution degree of the target recognition model, calculate the participation weight of the target recognition model according to the current contribution degree, calculate the difference between the user scoring data corresponding to the target recognition model and the historical scoring data to obtain scoring difference data, calculate the product of the scoring difference data, the participation weight, and a preset scoring adjustment parameter to obtain scoring update data, and calculate the sum of the historical scoring data of the target recognition model and the scoring update data to obtain the current scoring data of the target recognition model;
[0026] Wherein, the current contribution degree of the target recognition model is used to characterize the participation degree of the target recognition model in the current recognition task.
[0027] As an optional implementation manner, in the first aspect of the present invention, after filtering the content recognition results corresponding to all the target recognition models according to the user judgment data to obtain the initial content recognition result, the method further includes:
[0028] For each of the target recognition models, determine a classified adjustment coefficient according to the deployment type of the target recognition model and the classified level of the content to be recognized in the set of content to be recognized corresponding to the target recognition model, and update the model trust degree of the content recognition model according to the classified adjustment coefficient and the current scoring data of the content recognition model to obtain the updated model trust degree of the target recognition model; wherein, the updated model trust degree of the target recognition model is used to determine the size of the initial weight that can be assigned to the target recognition model in the next recognition task, and the deployment type of the target recognition model includes one of local deployment and cloud deployment.
[0029] As an optional implementation manner, in the first aspect of the present invention, parsing the file to be detected uploaded by the user to obtain at least one set of content to be recognized, including:
[0030] Perform format detection on the file to be detected uploaded by the user to obtain a format detection result;
[0031] When the format detection result indicates that the format of the file to be detected is legal, a preset file parsing tool is called to parse the file to be detected and generate at least one content block to be recognized with content type annotation, and all the content blocks to be recognized are merged to obtain at least one content set to be recognized;
[0032] Wherein, the content type annotation is used to indicate the content type of the content block to be recognized corresponding to the content type annotation.
[0033] As an optional implementation manner, in the first aspect of the present invention, the initial content recognition results corresponding to all the content sets to be recognized are fused to obtain the target content recognition result corresponding to the file to be detected, including:
[0034] Parse the sorting format of the file to be detected to obtain sorting format data; wherein, the sorting format data is used to characterize the content arrangement order and content arrangement format among multiple content blocks to be recognized;
[0035] Fuse the initial content recognition results corresponding to all the content sets to be recognized according to the sorting format data to obtain an initial fused recognition result;
[0036] Perform a verification and repair operation on the initial fused recognition result according to the content types of all the contents in the file to be detected to obtain the target content recognition result corresponding to the file to be detected;
[0037] Wherein, the verification and repair operation includes at least one of a text verification operation corresponding to text content, a table integrity repair operation corresponding to table content, and a graphic and text coherent combination operation corresponding to image content.
[0038] The second aspect of the present invention discloses a file content recognition device, the device includes:
[0039] A file parsing module, configured to parse a file to be detected uploaded by a user to obtain at least one content set to be recognized; wherein, the content set to be recognized is a content set formed by the same type of content in the file to be detected, and the file to be detected includes at least one of text content, table content, and image content;
[0040] A content recognition module, configured to, for each content set to be recognized, select at least one target recognition model adapted to the content set to be recognized from a preset model library to perform content recognition on the content set to be recognized, and integrate the content recognition results of all the target recognition models for the content set to be recognized to obtain an initial content recognition result corresponding to the content set to be recognized; wherein, the target recognition model includes a content recognition model in the preset model library that meets a preset screening condition;
[0041] A result fusion module, configured to fuse the initial content recognition results corresponding to all the to-be-recognized content sets to obtain the target content recognition result corresponding to the to-be-detected file.
[0042] As an alternative implementation manner, in the second aspect of the present invention, the preset screening condition includes a weight screening condition;
[0043] For each of the to-be-recognized content sets, the content recognition module selects at least one target recognition model adapted to the to-be-recognized content set from a preset model library to perform content recognition on the to-be-recognized content set, and integrates the content recognition results of all the target recognition models for the to-be-recognized content set to obtain the initial content recognition result corresponding to the to-be-recognized content set. The specific manner includes:
[0044] For each of the to-be-recognized content sets, determine the selection weights of all the content recognition models in the preset model library according to the content type corresponding to the to-be-recognized content set, and determine at least one target recognition model adapted to the to-be-recognized content set from the preset model library according to the weight screening condition and the selection weights of all the content recognition models;
[0045] For each of the target recognition models, perform content recognition on the to-be-recognized content set corresponding to the target recognition model through the target recognition model to obtain the content recognition result corresponding to the target recognition model, and send the content recognition result to the user;
[0046] Obtain user evaluation data fed back by the user; wherein, the user evaluation data is data generated by the user evaluating the content recognition results corresponding to all the target recognition models;
[0047] Screen the content recognition results corresponding to all the target recognition models according to the user evaluation data to obtain the initial content recognition result.
[0048] As an alternative implementation manner, in the second aspect of the present invention, for each of the to-be-recognized content sets, the content recognition module determines the selection weights of all the content recognition models in the preset model library according to the content type corresponding to the to-be-recognized content set, and determines at least one target recognition model adapted to the to-be-recognized content set from the preset model library according to the weight screening condition and the selection weights of all the content recognition models. The specific manner includes:
[0049] For each of the to-be-recognized content sets, obtain the historical recognition data corresponding to each content recognition model in the preset model library, assign an initial weight to each content recognition model in the preset model library according to all the historical recognition data, obtain the target content characteristics of the to-be-recognized content set, and adjust the initial weights of all content recognition models according to the target content characteristics to obtain the selection weights corresponding to each content recognition model in the preset model library. Screen all content recognition models in the preset model library according to the weight screening conditions and the selection weights corresponding to each content recognition model to obtain at least one target recognition model adapted to the to-be-recognized content set;
[0050] Among them, the target content characteristics include one of the text character quantity, the table row-column complexity, and the image resolution. Meeting the weight screening condition specifically means that the selection weight corresponding to the target recognition model is greater than or equal to a preset weight threshold;
[0051] For each content recognition model in the preset model library, the historical recognition data corresponding to the content recognition model includes recognition performance data and model trustworthiness. The recognition performance data can characterize the high or low recognition performance of the content recognition model in recognizing content of the target type during the historical time period, and the target type is the same as the content type included in the to-be-detected file. The model trustworthiness can characterize the high or low reliability of the content recognition model when performing the content recognition task. Both the model trustworthiness and the recognition performance data are used to determine the magnitude of the initial weight that can be assigned to the content recognition model.
[0052] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the content recognition module screens the content recognition results corresponding to all the target recognition models according to the user judgment data to obtain the initial content recognition results includes:
[0053] Perform score conversion on the user judgment data to obtain the user score data corresponding to each target recognition model; among them, the user score data corresponding to the target recognition model is used to characterize the accuracy and reliability of the user's determination of the content recognition result corresponding to the target recognition model, as well as the high or low degree of the recognition efficiency of the target recognition model;
[0054] For each target recognition model, obtain the historical score data of the target recognition model, and update the historical score data according to the user score data of the target recognition model to obtain the current score data of the target recognition model; among them, the historical score data of the target recognition model is the score data updated after the target recognition model performs the historical recognition task;
[0055] Filter the current scoring data of all the target recognition models according to the preset scoring conditions to obtain target scoring data, and use the content recognition result corresponding to the target scoring data as the initial content recognition result.
[0056] As an alternative implementation, in the second aspect of the present invention, for each of the target recognition models, the content recognition module obtains the historical scoring data of the target recognition model, and updates the historical scoring data according to the user scoring data of the target recognition model. The specific method for obtaining the current scoring data of the target recognition model includes:
[0057] For each of the target recognition models, obtain the historical scoring data and the current contribution degree of the target recognition model, calculate the participation weight of the target recognition model according to the current contribution degree, calculate the difference between the user scoring data corresponding to the target recognition model and the historical scoring data to obtain scoring difference data, calculate the product of the scoring difference data, the participation weight and a preset scoring adjustment parameter to obtain scoring update data, and calculate the sum of the historical scoring data of the target recognition model and the scoring update data to obtain the current scoring data of the target recognition model;
[0058] Among them, the current contribution degree of the target recognition model is used to represent the participation degree of the target recognition model in the current recognition task.
[0059] As an alternative implementation, in the second aspect of the present invention, the device further includes:
[0060] A trust degree update module. After the content recognition module screens the content recognition results corresponding to all the target recognition models according to the user judgment data to obtain the initial content recognition results, the trust degree update module is used to, for each of the target recognition models, determine a confidentiality adjustment coefficient according to the deployment type of the target recognition model and the content confidentiality level of the set of content to be recognized corresponding to the target recognition model, and update the model trust degree of the content recognition model according to the confidentiality adjustment coefficient and the current scoring data of the content recognition model to obtain the updated model trust degree of the target recognition model; among them, the updated model trust degree of the target recognition model is used to determine the size of the initial weight that can be assigned to the target recognition model in the next recognition task, and the deployment type of the target recognition model includes one of local deployment and cloud deployment.
[0061] As an alternative implementation, in the second aspect of the present invention, the specific method for the file parsing module to parse the file to be detected uploaded by the user to obtain at least one set of content to be recognized includes:
[0062] Perform format detection on the file to be detected uploaded by the user to obtain a format detection result;
[0063] When the format detection result indicates that the format of the file to be detected is legal, a preset file parsing tool is called to parse the file to be detected and generate at least one content block to be recognized with content type annotation, and all the content blocks to be recognized are merged to obtain at least one content set to be recognized;
[0064] Among them, the content type annotation is used to indicate the content type of the content block to be recognized corresponding to the content type annotation.
[0065] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the result fusion module fuses the initial content recognition results corresponding to all the content sets to be recognized to obtain the target content recognition result corresponding to the file to be detected includes:
[0066] Parse the sorting format of the file to be detected to obtain sorting format data; among them, the sorting format data is used to characterize the content arrangement order and content arrangement format among multiple content blocks to be recognized;
[0067] Fuse the initial content recognition results corresponding to all the content sets to be recognized according to the sorting format data to obtain an initial fusion recognition result;
[0068] Perform a verification and repair operation on the initial fusion recognition result according to the content types of all the contents in the file to be detected to obtain the target content recognition result corresponding to the file to be detected;
[0069] Among them, the verification and repair operation includes at least one of a text verification operation corresponding to text content, a table integrity repair operation corresponding to table content, and a graphic and text coherence combination operation corresponding to image content.
[0070] The third aspect of the present invention discloses another file content recognition device, and the device includes:
[0071] A memory storing executable program code;
[0072] A processor coupled to the memory;
[0073] The processor calls the executable program code stored in the memory to execute the file content recognition method disclosed in the first aspect of the present invention.
[0074] The fourth aspect of the present invention discloses a computer storage medium, and the computer storage medium stores computer instructions, and when the computer instructions are called by a processor, they are used to execute the file content recognition method disclosed in the first aspect of the present invention.
[0075] Compared with the prior art, the present invention has the following beneficial effects:
[0076] For a set of content to be recognized with different content types, one or more target recognition models adapted to the content type of the set of content to be recognized are selected from a preset model library according to preset screening conditions, that is, different recognition models are selected according to different content types for content recognition, avoiding the situation of content recognition errors due to the limited performance of the recognition model when the same recognition model recognizes different types of content, thereby improving the accuracy of automatic file content recognition and reducing the degree of manual participation to improve the content recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0078] Figure 1 is a flowchart of a method for recognizing file content disclosed in an embodiment of the present invention;
[0079] Figure 2 is a structural diagram of a device for recognizing file content disclosed in an embodiment of the present invention;
[0080] Figure 3 is a structural diagram of another device for recognizing file content disclosed in an embodiment of the present invention;
[0081] Figure 4 is a structural diagram of yet another device for recognizing file content disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0082] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0083] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or end including a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or ends.
[0084] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0085] With the widespread application of knowledge-based question-answering systems and large language models, it is becoming increasingly important to extract content from files and convert it into structured data. The key to content extraction is to identify various types of content in files (such as text content, table content, and image content). Currently, character recognition models (such as PaddleOCR's PP-Structure model) are mainly used to automatically identify the content in files.
[0086] However, in practice, it has been found that since there may be multiple types of content in a file at the same time, and the layout formats of different types of content in the file may be more complicated, the character recognition model may easily misrecognize certain types of content in the file, resulting in low recognition accuracy of the overall content of the file. Further processing requires less efficient manual recognition methods, which reduces the overall efficiency of the file content recognition process.
[0087] Therefore, how to improve the accuracy of automatic recognition of file content, thereby reducing the degree of manual participation and improving content recognition efficiency is a technical problem that needs to be solved urgently.
[0088] In order to solve the above technical problems, the present invention discloses a file content recognition method and a file content recognition device, which are intended to improve the accuracy of automatic file content recognition, thereby reducing the degree of manual participation and improving content recognition efficiency. Detailed descriptions are given below.
[0089] Embodiment 1
[0090] See also Figure 1 , Figure 1 : is a flowchart of a method for identifying file content disclosed in an embodiment of the present invention.Figure 1 The method shown can be applied to a document content recognition device that can recognize the document content of a document to be detected. Further, the document content recognition device can be integrated into a natural language processing system or exist independently of the natural language processing system, which is not limited in the embodiments of the present invention. As Figure 1 shown, a document content recognition method disclosed in an embodiment of the present invention includes but is not limited to the following operations:
[0091] 101. Parse the document to be detected uploaded by the user to obtain at least one content set to be recognized; wherein, the content set to be recognized is a content set formed by the same type of content in the document to be detected, and the document to be detected includes at least one of text content, table content, and image content.
[0092] When the document content recognition method of the embodiment of the present invention is applied to a system with natural language processing functions such as a document processing system and a document management system, the document to be detected uploaded by the user can be obtained through the system front-end interaction interface.
[0093] 102. For each content set to be recognized, select at least one target recognition model adapted to the content set to be recognized from a preset model library to perform content recognition on the content set to be recognized, and integrate the content recognition results of all target recognition models for the content set to be recognized to obtain an initial content recognition result corresponding to the content set to be recognized; wherein, the target recognition model includes a content recognition model in the preset model library that meets the preset screening conditions.
[0094] The content recognition models in the prior art usually only show excellent performance in the recognition tasks of some content types. For example, the PP-Structure model of PaddleOCR can quickly and accurately process text, but has low accuracy in the recognition of complex table and image structures. Therefore, in order to improve the recognition accuracy of different types of content in the document to be detected, a recognition model adapted to each type of content is selected for recognition.
[0095] 103. Fuse the initial content recognition results corresponding to all content sets to be recognized to obtain a target content recognition result corresponding to the document to be detected.
[0096] It can be seen that in the embodiments of the present invention, for the content sets to be recognized with different content types, one or more target recognition models adapted to the content types of the content sets to be recognized are selected from the preset model library according to the preset screening conditions for content recognition, that is, different recognition models are selected according to different content types for content recognition, avoiding the situation of content recognition errors due to the limitations of the performance of the recognition model when the same recognition model recognizes different types of content, thereby improving the accuracy of automatic file content recognition and reducing the degree of manual participation to improve the content recognition efficiency. In addition, by calling different recognition models as needed, it is possible to avoid the high resource occupancy of a single model and achieve better operation effects even under limited hardware conditions.
[0097] In an optional embodiment, the preset screening conditions include weight screening conditions;
[0098] For each content set to be recognized, at least one target recognition model adapted to the content set to be recognized is selected from the preset model library to perform content recognition on the content set to be recognized, and the content recognition results of all target recognition models for the content set to be recognized are integrated to obtain the initial content recognition result corresponding to the content set to be recognized, which may include:
[0099] For each content set to be recognized, determine the selection weights of all content recognition models in the preset model library according to the content type corresponding to the content set to be recognized, and determine at least one target recognition model adapted to the content set to be recognized from the preset model library according to the weight screening conditions and the selection weights of all content recognition models;
[0100] For each target recognition model, perform content recognition on the content set to be recognized corresponding to the target recognition model through the target recognition model to obtain the content recognition result corresponding to the target recognition model, and send the content recognition result to the user;
[0101] Obtain the user judgment data fed back by the user; wherein, the user judgment data is the data generated by the user judging the content recognition results corresponding to all target recognition models;
[0102] Screen the content recognition results corresponding to all target recognition models according to the user judgment data to obtain the initial content recognition result.
[0103] In the embodiments of the present invention, for different types of content, target recognition models adapted thereto are selected from the preset model library according to the weight screening conditions. For example, when processing image content, a model with excellent image content recognition effect is selected. The recognition results of each target recognition model will be sent to the user for judgment, and the recognition results of the target recognition models will be screened according to the judgment of the user to obtain the recognition results corresponding to each type of content set respectively.
[0104] In an optional embodiment, for each set of content to be recognized, the selection weights of all content recognition models in the preset model library are determined according to the content type corresponding to the set of content to be recognized, and at least one target recognition model adapted to the set of content to be recognized is determined from the preset model library according to the weight screening condition and the selection weights of all content recognition models, which may include:
[0105] For each set of content to be recognized, obtain the historical recognition data corresponding to each content recognition model in the preset model library, assign an initial weight to each content recognition model in the preset model library according to all the historical recognition data, obtain the target content characteristics of the set of content to be recognized, and adjust the initial weights of all content recognition models according to the target content characteristics to obtain the selection weights corresponding to each content recognition model in the preset model library. Screen all the content recognition models in the preset model library according to the weight screening condition and the selection weights corresponding to each content recognition model to obtain at least one target recognition model adapted to the set of content to be recognized;
[0106] Among them, the target content characteristics include one of the text character quantity, the table row-column complexity, and the image resolution. Meeting the weight screening condition specifically means that the selection weight corresponding to the target recognition model is greater than or equal to a preset weight threshold;
[0107] For each content recognition model in the preset model library, the historical recognition data corresponding to the content recognition model includes recognition performance data and model trustworthiness. The recognition performance data can characterize the high or low recognition performance of the content recognition model in recognizing the content of the target type during the historical time period, and the target type is the same as the content type included in the file to be detected. The model trustworthiness can characterize the high or low reliability of the model when performing the content recognition task. Both the model trustworthiness and the recognition performance data are used to determine the size of the initial weight that can be assigned to the content recognition model.
[0108] It can be seen that in the embodiment of the present invention, for each set of content to be recognized, the initial weight is assigned according to the historical performance and model trustworthiness of the content recognition model. For example, for the set of content to be recognized of the text content type, the PP-Structure model of PaddleOCR can obtain a relatively high initial weight. After assigning the initial weight, the initial weight is dynamically adjusted according to the corresponding target content characteristics, and the target recognition model adapted to the current content type is selected in combination with the weight screening condition, so as to improve the accuracy of automatic recognition of file content.
[0109] In an optional embodiment, screening the content recognition results corresponding to all target recognition models according to the user evaluation data to obtain the initial content recognition result may include:
[0110] Perform a scoring conversion on the user judgment data to obtain the user scoring data corresponding to each target recognition model; wherein, the user scoring data corresponding to the target recognition model is used to characterize the accuracy and reliability of the content recognition result determined by the user for the target recognition model, as well as the level of the recognition efficiency of the target recognition model.
[0111] For each target recognition model, obtain the historical scoring data of the target recognition model, and update the historical scoring data according to the user scoring data of the target recognition model to obtain the current scoring data of the target recognition model; wherein, the historical scoring data of the target recognition model is the scoring data updated after the target recognition model performs historical recognition tasks.
[0112] Filter the current scoring data of all target recognition models according to the preset scoring conditions to obtain the target scoring data, and use the content recognition result corresponding to the target scoring data as the initial content recognition result.
[0113] In the embodiments of the present invention, the user scoring data includes user accuracy scoring, user reliability scoring, and user performance scoring. The user accuracy scoring can reflect the recognition accuracy of the corresponding target recognition model, the user reliability scoring can characterize the consistency and reproducibility of the output results of the corresponding target recognition model, and the user performance scoring can reflect the performance of the model in terms of the calculation efficiency and resource consumption of the target recognition model.
[0114] In an alternative embodiment, for each target recognition model, obtaining the historical scoring data of the target recognition model and updating the historical scoring data according to the user scoring data of the target recognition model to obtain the current scoring data of the target recognition model may include:
[0115] For each target recognition model, obtain the historical scoring data and the current contribution degree of the target recognition model, calculate the participation weight of the target recognition model according to the current contribution degree, calculate the difference between the user scoring data corresponding to the target recognition model and the historical scoring data to obtain the scoring difference data, calculate the product of the scoring difference data, the participation weight, and the preset scoring adjustment parameter to obtain the scoring update data, and calculate the sum of the historical scoring data of the target recognition model and the scoring update data to obtain the current scoring data of the target recognition model;
[0116] Wherein, the current contribution degree of the target recognition model is used to characterize the participation degree of the target recognition model in the current recognition task.
[0117] Specifically, if the number of target recognition models is n, the calculation formula for the participation weight of the i-th target recognition model is:
[0118]
[0119] Among them, ω i represents the participation weight of the i-th target recognition model, and C i represents the current contribution degree of the i-th target recognition model, and represents the total contribution degree of all target recognition models.
[0120] Furthermore, the user rating data includes the user accuracy rating S, the user reliability rating R, and the user performance rating P. The historical rating data of the i-th target recognition model includes the historical accuracy rating S i old , the historical reliability rating R i old and the historical performance rating P i old , and the current rating data of the i-th target recognition model includes the current accuracy rating S i new , the current reliability rating R i new and the current performance rating P i new , then the calculation formula for the current rating data of the i-th target recognition model is as follows:
[0121]
[0122] Among them, ω i represents the participation weight of the i-th target recognition model, λ represents a preset first adjustment parameter, μ represents a preset second adjustment parameter, and ν represents a preset third adjustment parameter.
[0123] In some implementation scenarios, the document content recognition method according to the embodiments of the present invention can further introduce a dynamic model training mechanism to continuously optimize the performance of the content recognition models in the preset model library. Specifically, after performing the content recognition task, a model training data set is generated based on the user evaluation data and the comparison differences between the target content recognition results and the original file to be detected; the target recognition model participating in the current recognition task in the preset model library is incrementally trained or fine-tuned according to the model training data set to update its model parameters; at the same time, multi-type content samples are regularly extracted from public data sources or historical task data to construct a cross-domain training set, and all or part of the models in the preset model library are periodically retrained to improve the generalization ability of the models for complex layouts and emerging content types. Among them, incremental training uses an online learning algorithm to fuse user feedback data in real time and dynamically adjust the model weights; while periodic retraining is based on a distributed computing framework and uses transfer learning technology to adapt the general feature extraction ability of the basic model to specific task data. Through the dynamic training mechanism, each content recognition model in the preset model library can adapt to the changes in content recognition requirements, avoid the problem of decreased recognition accuracy caused by model aging, and thus further improve the long-term stability of automatic document content recognition.
[0124] In an alternative embodiment, after screening the content recognition results corresponding to all target recognition models according to the user evaluation data to obtain the initial content recognition results, the document content recognition method according to the embodiments of the present invention further includes:
[0125] For each target recognition model, a classified adjustment coefficient is determined according to the deployment type of the target recognition model and the classified level of the content to be recognized corresponding to the target recognition model, and the model trust degree of the content recognition model is updated according to the classified adjustment coefficient and the current scoring data of the content recognition model to obtain the updated model trust degree of the target recognition model; wherein, the updated model trust degree of the target recognition model is used to determine the size of the initial weight that can be assigned to the target recognition model in the next recognition task, and the deployment type of the target recognition model includes one of local deployment and cloud deployment.
[0126] Specifically, the calculation formula for the updated model trust degree of the i-th target recognition model is:
[0127]
[0128] Wherein, W i represents the updated model trust degree of the i-th target recognition model, α represents the first update parameter, β represents the second update parameter, γ represents the third update parameter, and P y represents the classified adjustment coefficient.
[0129] In the prior art, although the online content recognition service based on large models can handle a relatively large number of content recognition tasks simultaneously, it is difficult to meet the requirements of classified scenarios and there is a risk of data leakage.
[0130] In the embodiments of the present invention, when a certain target recognition model is of the cloud deployment type, if the classified level of the content in the content set to be recognized is higher, the value of the classified adjustment coefficient corresponding to the target recognition model is larger, thereby reducing the trust level of the target recognition model. Conversely, if the classified level of the content in the content set to be recognized is lower, the value of the classified adjustment coefficient corresponding to the target recognition model is smaller, thereby increasing the trust level of the target recognition model. When a certain target recognition model is of the local deployment type, regardless of whether the content set to be recognized is classified or not, the classified adjustment coefficient of the target recognition model takes a relatively small value, that is, maintaining a relatively high trust level of the target recognition model.
[0131] It can be seen that for classified content, in the embodiments of the present invention, by introducing a classified adjustment coefficient to limit the invocation of the content recognition model deployed in the cloud and preferentially invoking the content recognition model deployed locally, the confidentiality and security of classified content are improved.
[0132] In an optional embodiment, parsing the file to be detected uploaded by the user to obtain at least one content set to be recognized may include:
[0133] Performing format detection on the file to be detected uploaded by the user to obtain a format detection result;
[0134] When the format detection result indicates that the format of the file to be detected is legal, calling a preset file parsing tool to parse the file to be detected and generate at least one content block to be recognized with content type annotation, and merging all the content blocks to be recognized to obtain at least one content set to be recognized;
[0135] Wherein, the content type annotation is used to indicate the content type of the content block to be recognized corresponding to the content type annotation.
[0136] Specifically, the situation where the format of the file to be detected is legal includes that the file to be detected is one of the PDF, WORD, and EXCEL formats. After detecting that the file format is legal, the file is stored in a temporary directory and waits for further processing; then a file parsing tool (such as the PyMuPDF tool) is called to parse the document and generate content blocks to be recognized with content type annotation.
[0137] Examples of content type annotations are as follows:
[0138]
[0139] In an optional embodiment, fusing the initial content recognition results corresponding to all sets of content to be recognized to obtain the target content recognition results corresponding to the file to be detected may include:
[0140] Analyze the sorting format of the file to be detected to obtain sorting format data; wherein, the sorting format data is used to characterize the content arrangement order and content arrangement format among multiple content blocks to be recognized;
[0141] Fuse the initial content recognition results corresponding to all sets of content to be recognized according to the sorting format data to obtain an initial fused recognition result;
[0142] Perform a verification and repair operation on the initial fused recognition result according to the content types of all the content in the file to be detected to obtain the target content recognition results corresponding to the file to be detected;
[0143] Among them, the verification and repair operation includes at least one of a text verification operation corresponding to text content, a table integrity repair operation corresponding to table content, and a graphic and text coherence combination operation corresponding to image content.
[0144] Specifically, the text verification operation is to verify whether the recognition results are arranged in the order of the file pages and content positions. The table integrity repair operation is specifically to complete the missing cells in the table. The graphic and text coherence combination operation is specifically to verify whether the relationship between the chart and the descriptive text is complete to improve the coherence of the file content structure.
[0145] An example of the structure of the fused target content recognition result is as follows:
[0146]
[0147] Embodiment 2
[0148] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a file content recognition device disclosed in an embodiment of the present invention. Among them, Figure 2 the shown file content recognition device is used to implement the file content recognition method described in Embodiment 1. Further, the file content recognition device may be integrated in a natural language processing system or exist independently of the natural language processing system, which is not limited in the embodiments of the present invention. As Figure 2 shown, a file content recognition device disclosed in an embodiment of the present invention includes but is not limited to the following modules:
[0149] The file parsing module 201 is used to parse the file to be detected uploaded by the user to obtain at least one set of content to be recognized. Among them, the set of content to be recognized is a set of content formed by the same type of content in the file to be detected, and the file to be detected includes at least one of text content, table content, and image content.
[0150] The content recognition module 202 is used to, for each set of content to be recognized, select at least one target recognition model adapted to the set of content to be recognized from the preset model library to perform content recognition on the set of content to be recognized, and integrate the content recognition results of all target recognition models for the set of content to be recognized to obtain the initial content recognition result corresponding to the set of content to be recognized. Among them, the target recognition model includes the content recognition model in the preset model library that meets the preset screening conditions.
[0151] The result fusion module 203 is used to fuse the initial content recognition results corresponding to all sets of content to be recognized to obtain the target content recognition result corresponding to the file to be detected.
[0152] It can be seen that in the embodiment of the present invention, for sets of content to be recognized of different content types, one or more target recognition models adapted to the content type of the set of content to be recognized are selected from the preset model library according to the preset screening conditions for content recognition, that is, different recognition models are selected according to different content types for content recognition, avoiding the situation of content recognition errors due to the limitations of the performance of the recognition model when the same recognition model recognizes different types of content, thereby improving the accuracy of automatic file content recognition and reducing the degree of manual participation to improve the content recognition efficiency. In addition, by calling different recognition models as needed, the high resource occupancy of a single model can be avoided, and good operation effects can also be achieved under limited hardware conditions.
[0153] In an optional embodiment, the preset screening conditions include weight screening conditions;
[0154] The specific manner in which the content recognition module 202 selects at least one target recognition model adapted to the set of content to be recognized from the preset model library to perform content recognition on the set of content to be recognized and integrates the content recognition results of all target recognition models for the set of content to be recognized to obtain the initial content recognition result corresponding to the set of content to be recognized includes:
[0155] For each set of content to be recognized, determine the selection weights of all content recognition models in the preset model library according to the content type corresponding to the set of content to be recognized, and determine at least one target recognition model adapted to the set of content to be recognized from the preset model library according to the weight screening conditions and the selection weights of all content recognition models;
[0156] For each target recognition model, the target recognition model performs content recognition on the set of content to be recognized corresponding to the target recognition model, obtains the content recognition result corresponding to the target recognition model, and sends the content recognition result to the user;
[0157] Obtain the user judgment data feedback by the user; wherein, the user judgment data is data generated by the user's judgment on the content recognition results corresponding to all target recognition models;
[0158] Screen the content recognition results corresponding to all target recognition models according to the user judgment data to obtain the initial content recognition results.
[0159] In the embodiment of the present invention, for different types of content, an appropriate target recognition model is selected from the preset model library according to the weight screening condition for processing. For example, when processing image content, a model with excellent image content recognition effect is selected. The recognition results of each target recognition model will be sent to the user for judgment, and the recognition results of the target recognition models will be screened according to the user's judgment situation to obtain the recognition results corresponding to each type of content set respectively.
[0160] In an optional embodiment, for each set of content to be recognized, the content recognition module 202 determines the selection weights of all content recognition models in the preset model library according to the content type corresponding to the set of content to be recognized, and determines at least one target recognition model adapted to the set of content to be recognized from the preset model library according to the weight screening condition and the selection weights of all content recognition models. The specific method includes:
[0161] For each set of content to be recognized, obtain the historical recognition data corresponding to each content recognition model in the preset model library, assign an initial weight to each content recognition model in the preset model library according to all the historical recognition data, obtain the target content characteristics of the set of content to be recognized, and adjust the weight values of the initial weights of all content recognition models according to the target content characteristics to obtain the selection weights corresponding to each content recognition model in the preset model library. Screen all content recognition models in the preset model library according to the weight screening condition and the selection weights corresponding to each content recognition model to obtain at least one target recognition model adapted to the set of content to be recognized;
[0162] Among them, the target content characteristics include one of the text character quantity, the table row and column complexity, and the image resolution. Meeting the weight screening condition specifically means that the selection weight corresponding to the target recognition model is greater than or equal to the preset weight threshold;
[0163] For each content recognition model in the preset model library, the historical recognition data corresponding to the content recognition model includes recognition performance data and model trust level. The recognition performance data can characterize the high or low recognition performance of the content recognition model in recognizing content of the target type within a historical time period, and the target type is the same as the content type included in the file to be detected. The model trust level can characterize the high or low reliability of the content recognition model when performing the content recognition task. Both the model trust level and the recognition performance data are used to determine the magnitude of the initial weight that can be assigned to the content recognition model.
[0164] It can be seen that in the embodiments of the present invention, for each set of content to be recognized, an initial weight is assigned according to the historical performance and model trust level of the content recognition model. For example, for a set of content to be recognized of the text content type, the PP-Structure model of PaddleOCR can obtain a relatively high initial weight. After the initial weight is assigned, the initial weight is dynamically adjusted according to the corresponding target content characteristics, and the target recognition model suitable for the current content type is selected in combination with the weight screening conditions, thereby improving the accuracy of automatic recognition of file content.
[0165] In an optional embodiment, the specific manner in which the content recognition module 202 screens the content recognition results corresponding to all target recognition models according to the user evaluation data to obtain the initial content recognition results includes:
[0166] Perform score conversion on the user evaluation data to obtain the user score data corresponding to each target recognition model; wherein, the user score data corresponding to the target recognition model is used to characterize the accuracy and reliability of the user's determination of the content recognition result corresponding to the target recognition model, as well as the high or low degree of the recognition efficiency of the target recognition model;
[0167] For each target recognition model, obtain the historical score data of the target recognition model, and update the historical score data according to the user score data of the target recognition model to obtain the current score data of the target recognition model; wherein, the historical score data of the target recognition model is the score data updated after the target recognition model performs the historical recognition task;
[0168] Screen the current score data of all target recognition models according to the preset score conditions to obtain the target score data, and use the content recognition result corresponding to the target score data as the initial content recognition result.
[0169] In the embodiments of the present invention, the user rating data includes user accuracy rating, user reliability rating, and user performance rating. The user accuracy rating can reflect the recognition accuracy of the corresponding target recognition model. The user reliability rating can characterize the consistency and reproducibility of the output results of the corresponding target recognition model. The user performance rating can reflect the performance of the model in terms of the computational efficiency and resource consumption of the target recognition model.
[0170] In an alternative embodiment, for each target recognition model, the content recognition module 202 obtains the historical rating data of the target recognition model, and updates the historical rating data according to the user rating data of the target recognition model to obtain the current rating data of the target recognition model. The specific method includes:
[0171] For each target recognition model, obtain the historical rating data and the current contribution degree of the target recognition model, calculate the participation weight of the target recognition model according to the current contribution degree, calculate the difference between the user rating data and the historical rating data corresponding to the target recognition model to obtain the rating difference data, calculate the product of the rating difference data, the participation weight, and a preset rating adjustment parameter to obtain the rating update data, and calculate the sum of the historical rating data and the rating update data of the target recognition model to obtain the current rating data of the target recognition model;
[0172] Among them, the current contribution degree of the target recognition model is used to characterize the participation degree of the target recognition model in the current recognition task.
[0173] Specifically, if the number of target recognition models is n, the calculation formula for the participation weight of the i-th target recognition model is:
[0174]
[0175] where ω i represents the participation weight of the i-th target recognition model, C i represents the current contribution degree of the i-th target recognition model, represents the total contribution degree of all target recognition models.
[0176] Furthermore, the user rating data includes user accuracy rating S, user reliability rating R, and user performance rating P. The historical rating data of the i-th target recognition model includes historical accuracy rating Siold, historical reliability rating Riold, and historical performance rating Piold. The current rating data of the i-th target recognition model includes current accuracy rating Sinew, current reliability rating Rinew, and current performance rating Pinew. Then the calculation formula for the current rating data of the i-th target recognition model is as follows:
[0177]
[0178] where ω i represents the participation weight of the i-th target recognition model, λ represents a preset first adjustment parameter, μ represents a preset second adjustment parameter, and ν represents a preset third adjustment parameter.
[0179] In an alternative embodiment, refer to Figure 3 , Figure 3 which is a schematic structural diagram of another document content recognition device disclosed in an embodiment of the present invention. Another document content recognition device disclosed in an embodiment of the present invention further includes:
[0180] A confidence update module 204. After the content recognition module 202 screens the content recognition results corresponding to all target recognition models according to the user evaluation data to obtain the initial content recognition results, the confidence update module 204 is used to, for each target recognition model, determine a classified adjustment coefficient according to the deployment type of the target recognition model and the classified level of the content to be recognized in the set of content to be recognized corresponding to the target recognition model, and update the model confidence of the target recognition model according to the classified adjustment coefficient and the current scoring data of the content recognition model to obtain the updated model confidence of the target recognition model; wherein, the updated model confidence of the target recognition model is used to determine the magnitude of the initial weight that can be assigned to the target recognition model in the next recognition task, and the deployment type of the target recognition model includes one of local deployment and cloud deployment.
[0181] Specifically, the calculation formula for the updated model confidence of the i-th target recognition model is:
[0182]
[0183] where W i represents the updated model confidence of the i-th target recognition model, α represents a first update parameter, β represents a second update parameter, γ represents a third update parameter, and P y represents the classified adjustment coefficient.
[0184] In the prior art, although the online content recognition service based on large models can handle a relatively large number of content recognition tasks simultaneously, it is difficult to meet the requirements of classified scenarios and there is a risk of data leakage.
[0185] Further, the content recognition module 202 can deploy a federated learning framework to achieve co-evolution of models. Specifically, each target recognition model in the preset model library acts as a federated learning node and conducts distributed training under the coordination of the central server. After each content recognition task ends, the locally deployed model uploads the desensitized feature gradients to the server, while the cloud model provides the global model parameters. The central server aggregates the gradients of each node using differential privacy technology, updates the global model, and then distributes it to the local nodes. In addition, for classified content, a gradient masking mechanism is designed to ensure that the feature representation of sensitive information will not be leaked during the federated learning process. Through federated learning, each recognition model can share cross-domain knowledge, especially improving the recognition ability for small-sample content types (such as handwritten chemical formulas and rare language texts), while strictly complying with data privacy protection requirements.
[0186] In the embodiment of the present invention, when a certain target recognition model is of the cloud deployment type, if the classified level of the content in the content set to be recognized is higher, the value of the classified adjustment coefficient corresponding to the target recognition model is larger, thereby reducing the trust level of the target recognition model. On the contrary, if the classified level of the content in the content set to be recognized is lower, the value of the classified adjustment coefficient corresponding to the target recognition model is smaller, thereby increasing the trust level of the target recognition model. When a certain target recognition model is of the local deployment type, regardless of whether the content set to be recognized is classified, the classified adjustment coefficient of the target recognition model takes a relatively small value, that is, maintaining a relatively high trust level of the target recognition model.
[0187] It can be seen that for classified content, in the embodiment of the present invention, by introducing a classified adjustment coefficient to limit the invocation of the content recognition model deployed in the cloud and preferentially invoking the content recognition model deployed locally, the confidentiality and security of classified content are improved.
[0188] In an alternative embodiment, the specific manner in which the file parsing module 201 parses the file to be detected uploaded by the user to obtain at least one content set to be recognized includes:
[0189] Perform format detection on the file to be detected uploaded by the user to obtain a format detection result;
[0190] When the format detection result indicates that the format of the file to be detected is legal, call a preset file parsing tool to parse the file to be detected and generate at least one content block to be recognized with content type annotation, and merge all the content blocks to be recognized to obtain at least one content set to be recognized;
[0191] Among them, the content type annotation is used to indicate the content type of the content block to be recognized corresponding to the content type annotation.
[0192] Specifically, the cases where the format of the file to be detected is legal include that the file to be detected is one of the PDF, WORD, and EXCEL formats. When it is detected that the file format is legal, the file is stored in a temporary directory and waits for further processing; then a file parsing tool (such as the PyMuPDF tool) is called to parse the document and generate a content block to be recognized that includes content type annotations.
[0193] Examples of content type annotations are as follows:
[0194]
[0195] Furthermore, the file parsing module 201 can support multi-version format compatibility and adaptive preprocessing. Specifically, in the format detection stage, in addition to verifying the legality of the file type, the specific version is identified through file header feature analysis (such as PDF 1.4 and PDF 2.0), and a version-adapted parser is called for processing. For old-version files, a format conversion middleware is automatically enabled to convert them to a standard format before parsing. In addition, when generating the content block to be recognized, super-resolution reconstruction technology is used to enhance the low-quality image content, and an OCR error correction model is used to perform anti-aliasing processing on the blurred text. Such preprocessing operations can significantly improve the parsing robustness of complex files and lay a high-quality data foundation for subsequent content recognition.
[0196] In an optional embodiment, the specific manner in which the result fusion module 203 fuses the initial content recognition results corresponding to all content sets to be recognized to obtain the target content recognition result corresponding to the file to be detected includes:
[0197] Parse the sorting format of the file to be detected to obtain sorting format data; among them, the sorting format data is used to represent the content arrangement order and content arrangement format among multiple content blocks to be recognized;
[0198] Fuse the initial content recognition results corresponding to all content sets to be recognized according to the sorting format data to obtain an initial fusion recognition result;
[0199] Perform a verification and repair operation on the initial fusion recognition result according to the content types of all contents in the file to be detected to obtain the target content recognition result corresponding to the file to be detected;
[0200] Among them, the verification and repair operation includes at least one of a text verification operation corresponding to the text content, a table integrity repair operation corresponding to the table content, and a graphic and text coherence combination operation corresponding to the image content.
[0201] The text verification operation is specifically to verify whether the recognition results are arranged in the order of file pages and content positions. The table integrity repair operation is specifically to complete the missing cells in the table. The graphic and text coherence operation is specifically to verify whether the relationship between the chart and the descriptive text is complete to improve the coherence of the file content structure.
[0202] The structure of the fused target content recognition result is as follows:
[0203] {
[0204] "text":"The contents of this document are as follows...",
[0205] "table":[
[0206] ["Title","Data 1","Data 2"],
[0207] ["content","100","200"]
[0208] ],
[0209] "image":["image1.jpg","image2.jpg"]
[0210] }
[0211] Further, the verification and repair operation of the result fusion module 203 can integrate an artificial intelligence-assisted verification engine. Specifically, after generating the initial fusion recognition result, the system calls the pre-trained natural language processing model to perform semantic coherence analysis on the text content to detect logical faults or semantic contradictions; at the same time, the table structure reconstruction algorithm is used to automatically complete the missing cells based on the row and column topological relationship, and the image description generation model (such as CLIP) is used to verify the correlation between the identified image content and the adjacent text description. If the image and text are inconsistent, the re-identification process is triggered: the associated text and image content blocks are extracted from the set of content to be identified, and their semantic similarity is calculated through the multimodal alignment model. If the similarity is lower than the preset threshold, the locally deployed high-precision image recognition model is preferentially selected to re-parse the image content, and the contextual expression of the text content is adjusted synchronously. In addition, for the repair of table integrity, a sequence generation model based on the attention mechanism is introduced to predict missing values based on the content of the identified cells, and the rationality of the prediction results is iteratively optimized through the Monte Carlo tree search algorithm. This type of intelligent verification mechanism significantly reduces the need for manual intervention, while improving the accuracy and efficiency of the repair operation.
[0212] In addition, in some other implementation scenarios, the result fusion module 203 can introduce a cognitive consistency verification mechanism. After completing the initial fusion of recognition results, the system constructs a knowledge graph of the file content, where "nodes" represent text entities, table data items, or image objects, and "edges" represent their semantic association relationships (such as "contains", "describes", "compares", etc.). Abnormal situations such as broken logical loops and data contradictions are detected through the graph traversal algorithm. For example, if the text mentions that "data A has increased by 20%", while the corresponding data item in the table shows negative growth, a contradiction alarm is triggered. The system automatically marks the conflicting content and starts a multi-model voting mechanism: all compatible models in the preset model library are recalled to independently identify the conflicting content, and the result supported by the majority of models is selected as the final solution. For persistent conflicts, a pending review report is generated and submitted to the user for decision-making. Such a mechanism combines symbolic logical reasoning with statistical learning, significantly improving the logical self-consistency of the content recognition results.
[0213] Embodiment III
[0214] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of another file content recognition device disclosed in the embodiments of the present invention. Among them, Figure 4 the shown file content recognition device is used to implement the file content recognition method described in Embodiment I. Further, the file content recognition device can be integrated into a natural language processing system or exist independently of the natural language processing system, which is not limited in the embodiments of the present invention. As Figure 4 shown, a file content recognition device disclosed in the embodiments of the present invention includes but is not limited to the following modules:
[0215] A memory 301 storing executable program code;
[0216] A processor 302 coupled to the memory 301;
[0217] The processor 302 calls the executable program code stored in the memory 301 and executes some or all of the steps in a file content recognition method described in Embodiment I of the present invention.
[0218] Embodiment IV
[0219] The embodiments of the present invention disclose a computer storage medium. The computer storage medium stores computer instructions, which are used to execute some or all of the steps in a file content recognition method described in Embodiment I of the present invention when called by a processor.
[0220] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0221] Through the above specific descriptions of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0222] Finally, it should be noted that: the file content recognition method and the file content recognition device disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying file content, characterized in that, The method includes: Parsing the file to be detected uploaded by the user to obtain at least one content set to be recognized; wherein, the content set to be recognized is a content set formed by the same type of content in the file to be detected, and the file to be detected includes at least one of text content, table content, and image content; For each content set to be recognized, select at least one target recognition model adapted to the content set to be recognized from the preset model library to perform content recognition on the content set to be recognized, and integrate the content recognition results of all the target recognition models for the content set to be recognized to obtain the initial content recognition result corresponding to the content set to be recognized; wherein, the target recognition model includes the content recognition model in the preset model library that meets the preset screening conditions; Fusing the initial content recognition results corresponding to all the content sets to be recognized to obtain the target content recognition result corresponding to the file to be detected.
2. The method for identifying document content according to claim 1, wherein The preset screening conditions include weight screening conditions; The step of, for each content set to be recognized, selecting at least one target recognition model adapted to the content set to be recognized from the preset model library to perform content recognition on the content set to be recognized, and integrating the content recognition results of all the target recognition models for the content set to be recognized to obtain the initial content recognition result corresponding to the content set to be recognized, includes: For each content set to be recognized, determine the selection weights of all content recognition models in the preset model library according to the content type corresponding to the content set to be recognized, and determine at least one target recognition model adapted to the content set to be recognized from the preset model library according to the weight screening conditions and the selection weights of all the content recognition models; For each target recognition model, perform content recognition on the content set to be recognized corresponding to the target recognition model through the target recognition model to obtain the content recognition result corresponding to the target recognition model, and send the content recognition result to the user; Obtain the user judgment data fed back by the user; wherein, the user judgment data is data generated by the user's judgment on the content recognition results corresponding to all the target recognition models; Screen the content recognition results corresponding to all the target recognition models according to the user judgment data to obtain the initial content recognition result.
3. The document content recognition method according to claim 2, characterized in that The step of, for each content set to be recognized, determining the selection weights of all content recognition models in the preset model library according to the content type corresponding to the content set to be recognized, and determining at least one target recognition model adapted to the content set to be recognized from the preset model library according to the weight screening conditions and the selection weights of all the content recognition models, includes: For each of the content sets to be recognized, obtain the historical recognition data corresponding to each content recognition model in the preset model library, assign an initial weight to each content recognition model in the preset model library according to all the historical recognition data, obtain the target content characteristics of the content set to be recognized, and adjust the initial weights of all content recognition models according to the target content characteristics to obtain the selection weights corresponding to each content recognition model in the preset model library. Screen all content recognition models in the preset model library according to the weight screening conditions and the selection weights corresponding to each content recognition model to obtain at least one target recognition model adapted to the content set to be recognized; Among them, the target content characteristics include one of the text character quantity, the table row-column complexity, and the image resolution. Meeting the weight screening conditions specifically means that the selection weight corresponding to the target recognition model is greater than or equal to a preset weight threshold; For each content recognition model in the preset model library, the historical recognition data corresponding to the content recognition model includes recognition performance data and model trustworthiness. The recognition performance data can characterize the high or low recognition performance of the content recognition model in recognizing target-type content during a historical time period, and the target type is the same as the content type included in the file to be detected. The model trustworthiness can characterize the high or low reliability of the content recognition model when performing content recognition tasks. Both the model trustworthiness and the recognition performance data are used to determine the size of the initial weight that can be assigned to the content recognition model.
4. The document content recognition method according to claim 3, wherein The screening of the content recognition results corresponding to all the target recognition models according to the user judgment data to obtain the initial content recognition results includes: Performing score conversion on the user judgment data to obtain the user score data corresponding to each target recognition model; among them, the user score data corresponding to the target recognition model is used to characterize the accuracy and reliability of the user's determination of the content recognition result corresponding to the target recognition model, as well as the high or low degree of the recognition efficiency of the target recognition model; For each target recognition model, obtain the historical score data of the target recognition model, and update the historical score data according to the user score data of the target recognition model to obtain the current score data of the target recognition model; among them, the historical score data of the target recognition model is the score data updated after the target recognition model executes a historical recognition task; Screen the current score data of all the target recognition models according to the preset score conditions to obtain the target score data, and use the content recognition result corresponding to the target score data as the initial content recognition result.
5. The method for identifying document content according to claim 4, wherein The step of, for each target recognition model, obtaining the historical score data of the target recognition model and updating the historical score data according to the user score data of the target recognition model to obtain the current score data of the target recognition model includes: For each of the target recognition models, obtain the historical scoring data and the current contribution degree of the target recognition model, calculate the participation weight of the target recognition model according to the current contribution degree, calculate the difference between the user scoring data corresponding to the target recognition model and the historical scoring data to obtain the scoring difference data, calculate the product of the scoring difference data, the participation weight, and a preset scoring adjustment parameter to obtain the scoring update data, and calculate the sum of the historical scoring data of the target recognition model and the scoring update data to obtain the current scoring data of the target recognition model; Among them, the current contribution degree of the target recognition model is used to represent the participation degree of the target recognition model in the current recognition task.
6. The method for identifying the content of a document according to claim 4, characterized in that After screening the content recognition results corresponding to all the target recognition models according to the user judgment data to obtain the initial content recognition results, the method further includes: For each of the target recognition models, determine a classified adjustment coefficient according to the deployment type of the target recognition model and the classified level of the content to be recognized in the set of content to be recognized corresponding to the target recognition model, and update the model trust degree of the content recognition model according to the classified adjustment coefficient and the current scoring data of the content recognition model to obtain the updated model trust degree of the target recognition model; among them, the updated model trust degree of the target recognition model is used to determine the size of the initial weight that can be assigned to the target recognition model in the next recognition task, and the deployment type of the target recognition model includes one of local deployment and cloud deployment.
7. The method for identifying document content according to any one of claims 1 to 6, characterized in that, The parsing of the file to be detected uploaded by the user to obtain at least one set of content to be recognized includes: Perform format detection on the file to be detected uploaded by the user to obtain a format detection result; When the format detection result indicates that the format of the file to be detected is legal, call a preset file parsing tool to parse the file to be detected and generate at least one content block to be recognized with content type annotation, and merge all the content blocks to be recognized to obtain at least one set of content to be recognized; Among them, the content type annotation is used to indicate the content type of the content block to be recognized corresponding to the content type annotation.
8. The method for identifying document content according to claim 7, wherein The fusion of the initial content recognition results corresponding to all the sets of content to be recognized to obtain the target content recognition result corresponding to the file to be detected includes: Parse the sorting format of the file to be detected to obtain sorting format data; among them, the sorting format data is used to represent the content arrangement order and content arrangement format among multiple content blocks to be recognized; Fuse the initial content recognition results corresponding to all the sets of content to be recognized according to the sorting format data to obtain an initial fusion recognition result; Perform a verification and repair operation on the initial fusion recognition result according to the content types of all the contents in the file to be detected to obtain the target content recognition result corresponding to the file to be detected; Among them, the verification and repair operation includes at least one of a text verification operation corresponding to text content, a table integrity repair operation corresponding to table content, and a graphic and text coherence combination operation corresponding to image content.
9. A file content recognition device, characterized in that The device includes: A file parsing module, configured to parse a file to be detected uploaded by a user to obtain at least one content set to be recognized; wherein, the content set to be recognized is a content set formed by the same type of content in the file to be detected, and the file to be detected includes at least one of text content, table content, and image content; A content recognition module, configured to, for each content set to be recognized, select at least one target recognition model adapted to the content set to be recognized from a preset model library to perform content recognition on the content set to be recognized, and integrate the content recognition results of all the target recognition models for the content set to be recognized to obtain an initial content recognition result corresponding to the content set to be recognized; wherein, the target recognition model includes a content recognition model in the preset model library that meets a preset screening condition; A result fusion module, configured to fuse the initial content recognition results corresponding to all the content sets to be recognized to obtain a target content recognition result corresponding to the file to be detected.
10. A file content recognition device, characterized in that, The apparatus includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the file content recognition method according to any one of claims 1 to 8.