Text detection method and system

By structuring rich text historical archives, identifying and processing multimedia converged content, the problem of insufficient information consistency and comparability in unstructured rich texts is solved, and rapid and accurate information detection and extraction is achieved, improving data analysis efficiency.

CN119149738BActive Publication Date: 2025-08-19CHENGDA CULTURE TECH (GUANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411179019.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2025-08-19
Estimated Expiration
2044-08-27

AI Technical Summary

Technical Problem

In the prior art, unstructured free rich text recording methods lead to insufficient information consistency and comparability, making it difficult to quickly and accurately detect and extract key information, affecting the efficiency of subsequent data analysis.

Method used

By obtaining text usage requirements, analyzing rich text historical archives, determining structured multi-type rich text, identifying multimedia fusion content, dividing text data and image data, calculating correlation and positional relationships, building a detection result information set, and performing structured text output.

Benefits of technology

It improves the accuracy and efficiency of information extraction, ensures the readability and comprehensibility of the detection results, facilitates subsequent data analysis and processing, and enhances the consistency and comparability of archival content records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119149738B_ABST
    Figure CN119149738B_ABST
Patent Text Reader

Abstract

The present application relates to the field of text detection technology, and in particular to a text detection method and system. The method comprises: obtaining text usage requirements; obtaining rich text historical archives based on the text usage requirements; parsing the rich text historical archives to determine structured multi-type rich texts; determining detection requirement-oriented information based on the text usage requirements and the structured multi-type rich texts; judging whether there is multimedia fusion content in the detection requirement-oriented information; if so, determining the detection result information set based on the multimedia fusion content; and outputting the detection result information set as a structured text. This ensures the accuracy and reliability of data and reduces the interference of irrelevant data; makes information more convenient for data analysis and processing, improves the practical value of information, makes the detection result information set readable and understandable, facilitates users to quickly obtain information, facilitates computer processing and analysis, and provides support for subsequent data mining and application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text detection technology, and in particular to a text detection method and system. Background Art

[0002] In information content recording systems, content providers must ensure that recorded information is accurate and reproducible for subsequent use and analysis. During this process, they must document a wide range of information, including event status, processing steps, key elements involved, and any problems encountered. This information is crucial for subsequent content management and in-depth analysis.

[0003] Traditional recording methods often rely on unstructured, free-form rich text. While flexible, this approach suffers from significant shortcomings in the consistency and comparability of historical archives recorded by content providers. Key information in free-form rich text is difficult to quickly and accurately detect, making it difficult to extract and integrate the required information, which in turn impacts the efficiency of subsequent data analysis.

[0004] In areas such as news content submission and the establishment of corporate knowledge bases, accurate and structured information records are equally important. They help improve the efficiency of content management and provide strong support for subsequent analysis and utilization. Summary of the Invention

[0005] This application provides a text detection method and system to solve the above problems.

[0006] In a first aspect, the present application provides a text detection method, the method comprising:

[0007] Obtain text usage requirements; obtain rich text historical archives based on the text usage requirements; parse the rich text historical archives to determine structured multi-type rich texts; determine detection requirement-oriented information based on the text usage requirements and the structured multi-type rich texts; determine whether multimedia fusion content exists in the detection requirement-oriented information; if so, determine a detection result information set based on the multimedia fusion content; and output the detection result information set in structured text.

[0008] Through the above technical solution, rich text historical archives are obtained in a targeted manner according to text usage requirements, ensuring the accuracy and reliability of the data and reducing the interference of irrelevant data; by parsing rich text historical archives, deeply understanding their content and structure, and determining structured multi-type rich texts, information can be extracted more specifically, improving the efficiency of information extraction, ensuring the accuracy and relevance of detection demand-oriented information, and making the extracted detection demand-oriented information more convenient for subsequent data analysis and processing; identifying whether there is multimedia fusion content in the detection demand-oriented information, and if so, marking the multimedia fusion content, and sorting out the detection result information set, which is convenient for subsequent utilization and analysis, and improves the practical value of the information; structured text output improves the readability and comprehensibility of the detection result information set, facilitates users to quickly obtain information, and facilitates computer processing and analysis, providing support for subsequent data mining and application.

[0009] Optionally, the method of determining a detection result information set based on multimedia fusion content includes: dividing the multimedia fusion content into text data and image data; parsing the text data to determine paragraph line information and semantic information of the text data; parsing the image data to determine image coordinates; determining a target image and target image coordinates referring to a corresponding text based on the semantic information, the image coordinates and the paragraph line information; determining a valid text component represented by the target image based on the target image and the semantic information; and determining a detection result information set based on the valid text component, the semantic information, the paragraph line information and the target image coordinates.

[0010] Through the above technical solution, multimedia fusion content is effectively divided into text data and image data, ensuring accurate extraction of text data and image data, avoiding information loss, and providing a basis for subsequent processing of these two types of data separately. Parsing text data to obtain its detailed structural information (paragraph line information) and semantic information helps understand the text layout and provides a basis for positional relationship judgment for the subsequent determination of the relationship between text and image. Parsing image data to obtain its coordinate information plays an important role in the subsequent extraction of valid text components and providing accurate target images and their coordinates corresponding to the text content. Extracting the valid text components represented by the target image supplements the deficiencies in the text data and provides supplementary data for the subsequent construction of the detection result information set. The formation of a detection result information set containing rich information such as text, images, and coordinates facilitates subsequent output and use, meets user needs, improves the consistency and comparability of archival content records, and quickly and accurately detects, extracts, and integrates the required information.

[0011] Optionally, determining the target image and target image coordinates referring to the corresponding text according to the semantic information, the image coordinates, and the paragraph line information includes: calculating the association degree between each text data and each image data based on the semantic information, the image coordinates, and the paragraph line information, and determining the target image according to the association degree, with reference to the following formula:

[0012] ;in, is the target image; For image data data points; A collection of image data; is the degree of association; is the correlation threshold;

[0013] According to the correlation degree, the target image coordinates are determined, referring to the following formula: in, is the target image coordinate; for The image in the image data pointed to by the index; is the index of the target image that refers to the corresponding text; is the index of any image in the image data set; is the degree of association; is the correlation threshold.

[0014] Through the above technical solution, the correlation between each text data and each image data is calculated, providing a calculation basis for the subsequent determination of the target image and the target image coordinates. According to the correlation, the target image that is most relevant to the data point of the text data at the location of the key information in the detected structured multi-type rich text is determined, providing a basis for the subsequent extraction of the target image coordinates. According to the correlation, the coordinates of the target image that is most relevant to the text data point at the location of the key information in the detected structured multi-type rich text are determined, providing accurate image position information for subsequent data analysis, information extraction and integration. Through the above calculation method, it is possible to accurately find images and texts with a strong correlation, thereby avoiding the image from being distorted and the image interpretation from failing to express the true meaning of the placement here.

[0015] Optionally, the calculating the association between each text data and each image data based on the semantic information, the image coordinates, and the paragraph line information includes: extracting a keyword set from the set of text data; extracting a keyword set from the set of image data; calculating the intersection size of the keyword set in the set of text data and the keyword set in the set of image data to obtain a keyword matching score, referring to the following formula: in, Score keyword matching; is a set of keywords in a set of text data; is a set of keywords in a set of image data; convert the set of text data into a text semantic vector; convert the set of image data into an image semantic vector; calculate the cosine similarity between the text semantic vector and the image semantic vector to obtain a semantic matching score, referring to the following formula; in; is the cosine similarity; is the dot product formula of text semantic vector and image semantic vector; is the length of the text semantic vector; is the length of the image semantic vector;

[0016] Calculate the proximity between the target image coordinates and the paragraph line information corresponding to the keyword set in any text data to obtain a position relationship score; and obtain the association degree between each text data and each image data based on the keyword matching score, the semantic matching score, and the position relationship score, referring to the following formula: in, is the degree of association; is the weight coefficient of the position relationship score; Score keyword matching; is the weight coefficient of semantic matching score; Score for semantic matching; is the weight coefficient of the position relationship score; Score the positional relationship.

[0017] Through the above technical solution, combined with the keyword sets of text and images, key information in the archival content records can be more accurately extracted, which is more comprehensive and accurate than methods that rely solely on text or images. The calculation of the correlation takes into account the keywords, semantics and positional relationships between text and images, which helps to identify consistent information in the archival content records, thereby improving the comparability between data. By calculating the correlation between text and images, the required information can be quickly and accurately located, avoiding the tedious process of manual search and sorting in unstructured free rich text, greatly improving the efficiency of data analysis; using the above formula to calculate the correlation makes the information in the archival content records easier to integrate and utilize; by adjusting the weight coefficients corresponding to the keyword matching score, semantic matching score and positional relationship score, it can flexibly adapt to different application scenarios and needs, which makes this method very scalable and adaptable.

[0018] Optionally, the calculation of the proximity between the target image coordinates and the paragraph line information corresponding to the keywords in the keyword set in any text data to obtain the positional relationship score includes: parsing the proximity between the target image coordinates and the paragraph line information corresponding to the keywords in the keyword set in any text data to obtain the difference between the row numbers of the text and the image, the difference between the column numbers of the text and the image, and the surrounding situation between the text and the image; obtaining the surrounding distance of the text on the X-axis of the image and the surrounding distance of the text on the Y-axis of the image based on the surrounding situation of the image and the text; and calculating the positional relationship score based on the difference in the row numbers, the difference in the column numbers, the surrounding distance on the X-axis, and the surrounding distance on the Y-axis, with reference to the following formula:

[0019] ;in, is the weight coefficient of the absolute value of the difference between row numbers; is the absolute value of the difference between row numbers; is the weight coefficient of the difference in column numbers; is the difference in column numbers; is the weight coefficient of the surrounding distance on the X axis; is the orbital distance on the X axis; is the weight coefficient of the surrounding distance on the Y axis; is the surrounding distance on the Y axis. Through the above technical solution, by carefully analyzing the relative positions of text and image in the document, including the difference in line number, column number and surrounding situation, it is possible to more accurately locate the text information closely related to the image; the calculation of the position relationship score takes into account the spatial layout relationship between text and image, which helps to identify and match the consistent information in the archival content record. Even in different archival content records, as long as the spatial relationship between text and image is similar, their position relationship scores will be similar, thereby improving the comparability between data. By calculating the position relationship score, we can quickly and accurately locate the text information closely related to the image, avoiding the tedious process of manual search and sorting in unstructured free rich text, and greatly improving the efficiency of data analysis. This method is not only applicable to text data, but also combines image data to support the fusion analysis of multimedia data, which is particularly important for the analysis of image information such as photos that contain a large amount of key information in archival content records.

[0020] Optionally, determining the valid text component represented by the target image based on the target image and the semantic information includes: parsing the target image to determine whether the target image has text content; if so, extracting the image text content of the target image; interpreting based on the image text content and the semantic information of the image screen connection context, and extracting the valid text component in the image through the image text content, image interpretation information and the semantic information; if not, interpreting based on the semantic information of the image screen connection context, and using the image interpretation information as the valid text component.

[0021] The above technical solution not only considers possible textual content within an image but also the information conveyed by the image itself, improving the comprehensiveness of information extraction. By extracting textual content from images and combining it with semantic information, unstructured free-form rich text can be transformed into a more structured and understandable form, facilitating better understanding and utilization of information within archival content records. Consistent key information is extracted from different types of archival content records (including text and images), improving the consistency and comparability of archival content records and facilitating comparison and analysis across multiple records to identify potential issues and areas for improvement. The process of extracting effective text components is highly applicable to image data and supports the integrated analysis of multimedia data, particularly for the analysis of image information such as photographs contained within archival content records. By extracting effective textual content from images, key information within archival content records can be more quickly and accurately located, eliminating the tedious manual search and sorting of unstructured free-form rich text and thus optimizing data analysis efficiency.

[0022] Optionally, the semantic information based on the image-screen connection context is interpreted, and the image interpretation information is used as a valid text component, including: interpreting based on the image-screen connection context to obtain interpretation information; calculating the similarity between the interpretation information and the semantic information; if the similarity is greater than a preset threshold, the interpretation information is used as a valid text component.

[0023] Through this technical solution, combining imagery with contextual semantic information for interpretation, we can more accurately understand the specific meaning of images within archival content records, thereby extracting more precise information. Furthermore, this approach considers not only the textual content within the image but also the information conveyed by the image itself, improving the comprehensiveness of information extraction. It also transforms unstructured free-form rich text and image information into structured interpretation information, facilitating subsequent data analysis and processing. This structured interpretation information improves the consistency and comparability of archival content records, facilitating comparison and analysis across multiple archival content records. By extracting valid textual components from images, we can more accurately locate key information within archival content records.

[0024] Optionally, the detection result information set is determined based on the valid text components, the semantic information, the paragraph line information and the target image coordinates, including: based on the valid text components, determining the distribution status of the valid text components according to the target image coordinates, the paragraph line information and the semantic information; performing multimodal information fusion on the valid text components and the semantic information according to the distribution status to obtain the detection result information set.

[0025] This technical solution combines valid text components, semantic information, paragraph and line information, and target image coordinates to more accurately determine the distribution of text within an image, improving the accuracy of information extraction. It converts unstructured free-form rich text and image information into a structured detection result information set, facilitating subsequent data analysis and processing. Through multimodal information fusion, valid text components are integrated with text semantic information to form a more unified, rich, and comprehensive detection result information set.

[0026] Optionally, the structured text output of the detection result information set includes: filtering the detection result information set according to the query information to obtain a filtered information set; obtaining a preset structured template; filling the structured template with information through the filtered information set to obtain a streamlined information set; and outputting the streamlined detection result information set as structured text.

[0027] Through the above technical solution, the test result information set is filtered according to the query information, and the required information is quickly and accurately located, avoiding the tedious process of manually searching in unstructured text and improving the efficiency of subsequent processing. The filtered information is filled with a preset structured template to ensure a unified format and standard for the output information, improve the consistency and comparability of the archival content records, make the information clearer and easier to understand, and facilitate subsequent data analysis. It can be directly used for report generation, data visualization, or further data processing and analysis. The structured text output allows users to more clearly understand the specific circumstances and background in the archival records, which helps to make more accurate judgments and decisions. Because the information has been filtered and structured, the efficiency and accuracy of data analysis are improved.

[0028] In a second aspect, the present application provides a text detection system, the system comprising:

[0029] A rich text history archive acquisition module is used to obtain text usage requirements; and obtain rich text history archives according to the text usage requirements;

[0030] A structured multi-type rich text parsing module, used to parse the rich text historical archives and determine the structured multi-type rich text;

[0031] A detection demand-oriented information determination module, configured to determine detection demand-oriented information based on the text usage requirements and the structured multi-type rich text;

[0032] A multimedia content determination module, configured to determine whether the detection demand-oriented information contains multimedia fusion content;

[0033] a detection result information set construction module, for determining the detection result information set based on the multimedia fusion content, if any;

[0034] The structured text output module is used to output the detection result information set in structured text. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0036] Figure 1 A schematic diagram of an application scenario provided in one embodiment of the present application;

[0037] Figure 2 A flowchart of a text detection method provided in one embodiment of the present application;

[0038] Figure 3 A structural diagram of a text detection system provided in one embodiment of the present application. DETAILED DESCRIPTION

[0039] To make the purpose, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0040] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document, unless otherwise specified, generally indicates an "or" relationship between the related objects.

[0041] The embodiments of the present application are described in further detail below with reference to the accompanying drawings.

[0042] Archival content is often recorded in unstructured, free-form rich text. While flexible, this method suffers from low consistency and comparability in the rich text historical archives recorded by users. Key information in free-form rich text cannot be quickly and accurately detected, making it difficult to extract and integrate the required information, making subsequent data analysis extremely difficult.

[0043] Based on this, the present application provides a text detection method and system to obtain text usage requirements; according to the text usage requirements, the rich text historical archives are obtained in a targeted manner, the accuracy and reliability of the data are guaranteed, and the interference of irrelevant data is reduced; the rich text historical archives are parsed to determine the structured multi-type rich text; by parsing the archives, the content and structure are deeply understood, and after determining the structured multi-type rich text, information can be extracted more targetedly to improve the efficiency of information extraction; according to the text usage requirements and the structured multi-type rich text, the detection demand-oriented information is determined; the accuracy and relevance of the detection demand-oriented information are ensured, and the extracted detection demand-oriented information is more convenient for subsequent data analysis and Processing; judging whether there is multimedia fusion content in the detection demand-oriented information; being able to identify the multimedia fusion content in the detection demand-oriented information, such as images, text and other information forms, and marking the multimedia fusion content to facilitate subsequent targeted processing and analysis; if it exists, then based on the multimedia fusion content, determining the detection result information set, the sorted detection result information set is more convenient for subsequent use and analysis, and improves the practical value of the information; outputting the detection result information set in structured text; the structured text output improves the readability and comprehensibility of the information, facilitates users to quickly obtain information, facilitates computer processing and analysis, and provides support for subsequent data mining and application.

[0044] Figure 1A schematic diagram of an application scenario provided by this application is provided. When quickly and accurately detecting key information in a free rich text, the method provided by this application is applied. The method provided by this application is applied to any server, and the server, the text query terminal system and the archive content recording system interact to obtain text usage requirements from the text query terminal system; according to the text usage requirements, the rich text historical archives are obtained from the archive content recording system in a targeted manner, ensuring the accuracy and reliability of the data and reducing the interference of irrelevant data; by parsing the rich text historical archives, deeply understanding its content and structure, and determining the structured multi-type rich text, information can be extracted more targetedly, improving the efficiency of information extraction, ensuring the accuracy and relevance of the detection demand-oriented information, and making the extracted detection demand-oriented information more convenient for subsequent data analysis and processing; identifying whether there is multimedia fusion content in the detection demand-oriented information, if so, marking the multimedia fusion content, and sorting out the detection result information set, which is convenient for subsequent use and analysis, and improving the practical value of the information; the structured text output improves the readability and comprehensibility of the detection result information set, facilitates users to quickly obtain information, facilitates computer processing and analysis, and provides support for subsequent data mining and application. The specific implementation method can refer to the following embodiments.

[0045] Figure 2 This is a flowchart of a text detection method provided in one embodiment of the present application. The method of this embodiment can be applied to the server in the above scenario. Figure 2 As shown, the method includes:

[0046] S201. Obtain text usage requirements; and obtain rich text historical archives based on the text usage requirements.

[0047] Text usage requirements can include regular inspection, retrieval, content extraction, or summarization of rich text historical archives. Rich text historical archives can be detailed records of an event, usually in the form of text or images.

[0048] Specifically, the text usage requirements of users are obtained from the text query terminal system, and relevant rich text historical archives are retrieved from the archive content recording system based on the key information in the text usage requirements.

[0049] S202: Parse the rich text historical archives to determine structured multi-type rich texts.

[0050] Structured, multi-type rich text can be structured with specific organizational methods and information hierarchies. It can include multiple sections, such as title, summary, details, issue record, action plan, and result feedback. Each section may contain specific information fields, formats, or images.

[0051] Specifically, since the obtained rich-text historical archives have complex contents and may be composed of multiple historical archive content records, their structure is generally filled in according to a fixed template, but their internal contents vary greatly due to different users, recording issues and other specific circumstances. Therefore, it is necessary to identify the obtained rich-text historical archives, obtain the overall structure of the rich text, and determine the key information and format expressed in the content of each structure.

[0052] S203. Determine detection demand-oriented information based on text usage requirements and structured multi-type rich text.

[0053] The detection demand-oriented information can be specific information directed to the detection results extracted from the rich text historical archive according to the text usage requirements; the detection demand-oriented information includes multiple record types such as pure text type, pure image type or a combination of text and image type.

[0054] Specifically, when a rich-text historical archive is generated each time a record is generated, multiple issues may be recorded in parallel at the same time, resulting in the existence of records of several other issues and actual situations that are not strongly relevant to the current text usage needs; at the same time, there may be multiple historical archive content records available for use in a rich-text historical archive; therefore, it is necessary to combine the text usage needs and structured multi-type rich text to determine the location of the key information in the structured multi-type rich text that needs to be detected, so as to extract the specific information that leads to the detection results.

[0055] S204: Determine whether there is multimedia fusion content in the detection demand-oriented information.

[0056] Multimedia fusion content can be content in other media forms besides text contained in rich text historical archives, such as a combination of text and images. Multimedia fusion content includes types such as pure image type or image-text combination type. Specifically, in order to accurately record the current problem and reproduce the problem later to formulate corresponding plans or materials, images of the current problem are usually interspersed in the text. These images play a vital role in describing the specific situation of the problem, such as an event + image that occurred at a certain time. In this process, if only the text is detected and recognized, the specific situation cannot be obtained, such as the description of the status of the news scene and many other situations that can only be determined based on the image content. Therefore, it is necessary to traverse and analyze the detection demand-oriented information. If the detection demand-oriented information contains image content, then it is determined that the structured multi-type rich text content where the detection demand-oriented information is located contains multimedia fusion content; if the detection demand-oriented information does not contain image content, then it is determined that the structured multi-type rich text content where the detection demand-oriented information is located does not contain multimedia fusion content.

[0057] S205: If it exists, determine the detection result information set based on the multimedia fusion content.

[0058] The detection result information set can be a set of information related to the detection result that is sorted out based on the text usage requirements and multimedia fusion content. Specifically, if it is detected that there is multimedia fusion content in the demand-oriented information, then based on the type and characteristics of the multimedia content, the characteristics are mainly the layout characteristics of the image and text, such as the upper and lower surround or text surround features, and the presence of image name references, the image name should be replaced with the image and its characteristics should be determined. On this basis, the corresponding detection result is determined, and the detection result is combined with the relevant text information to form a detection result information set. In other implementation methods, if the presence of multimedia fusion content in the demand-oriented information is not detected, the text in the detection demand-oriented information is directly detected to determine the detection result information set, and structured text output is performed in the subsequent process.

[0059] S206: Output the detection result information set in structured text.

[0060] The structured text may be a text format in which the test result information set is arranged in a predetermined format and organization. The predetermined format and organization may be preset in different structured forms according to different requirements so as to be used by different ports.

[0061] Specifically, in the test result information, the content recorded in the original archive content record may be complex, which makes it difficult to read or process subsequently and difficult to understand. Therefore, the text and multimedia content in the test result information set are structured to form an output format that is easy to read and understand.

[0062] Through the method provided in this embodiment, rich text historical archives are obtained in a targeted manner according to text usage requirements, thereby ensuring the accuracy and reliability of the data and reducing the interference of irrelevant data; by parsing the rich text historical archives, deeply understanding its content and structure, and determining the structured multi-type rich text, information can be extracted more specifically, improving the efficiency of information extraction, ensuring the accuracy and relevance of the detection demand-oriented information, and making the extracted detection demand-oriented information more convenient for subsequent data analysis and processing; identifying whether there is multimedia fusion content in the detection demand-oriented information, and if so, marking the multimedia fusion content, and sorting out the detection result information set, which is convenient for subsequent utilization and analysis, and improves the practical value of the information; structured text output improves the readability and comprehensibility of the detection result information set, facilitates users to quickly obtain information, and facilitates computer processing and analysis, providing support for subsequent data mining and application.

[0063] In some embodiments, the multimedia fusion content is divided into text data and image data; the text data is parsed to determine the paragraph line information and semantic information of the text data; the image data is parsed to determine the image coordinates; based on the semantic information, image coordinates and paragraph line information, the target image and target image coordinates referring to the corresponding text are determined; based on the target image and semantic information, the valid text component represented by the target image is determined; based on the valid text component, semantic information, paragraph line information and target image coordinates, a detection result information set is determined.

[0064] Text data can be the free-form rich text within rich text historical archives, including all textual information recorded by users, such as problem descriptions and solutions. Image data can be images or other visual elements included in multimedia fusion content. These images are associated with text data and serve to assist in illustrating or supplement the textual information. Paragraph and line information can be structural information about the text data, including paragraph divisions and line arrangement, used to determine the organization and layout of the text. Semantic information can be the meaning and content of the text data, including keywords, phrases, sentences, and the relationships between them, used to understand the message and intent conveyed by the text. Image coordinates can be the location information of the image data within the multimedia fusion content, including the image's starting point, size, and orientation, used to determine the relative position of the image and text data. Target images can be images within the image data associated with the textual content at the location of key information within the detected structured multi-type rich text. These images may contain visual representations or supplementary explanations of the textual content. Target image coordinates can be the specific location information of the target image within the multimedia fusion content, including its starting point, size, and orientation, used to precisely locate the target image. The effective text component can be text information extracted from the target image through image analysis technology, which is of great significance for understanding the image content or the text data associated with it.

[0065] Specifically, for multimedia fusion content, since these images are generally photographed directly or simply annotated and stored in documents, if the text and image are only processed separately, the content expressed by the image will be distorted in the separate processing method that can be achieved in the existing technology, that is, the context will be lost, and the meaning of the image being placed in that position cannot be fully and appropriately reflected, and more content cannot be interpreted in the image. The multimedia fusion content is traversed and divided according to its data type (text or image). Text extraction techniques are used to extract text data from the multimedia fusion content, and image extraction techniques are used to extract image data from the multimedia fusion content. The extracted text data is parsed to identify its paragraph and line structure and determine paragraph and line information. Natural language processing tools are used to perform semantic analysis on the text data to extract semantic information. The extracted image data is parsed to identify the image coordinates within the multimedia fusion content. Combining semantic information, image coordinates, and paragraph and line information, the referential relationship between the image and the text is analyzed to determine the target image that refers to the corresponding text and its target image coordinates within the original multimedia fusion content. The target image is further analyzed, combining semantic information to analyze the text content represented by the target image, thereby extracting the valid text components represented by the target image. Combining the valid text components, semantic information, paragraph and line information, and target image coordinates, a detection result information set is constructed. This detection result information set is structured for subsequent output and use.

[0066] Through the method provided by this embodiment, the multimedia fusion content is effectively divided into text data and image data, ensuring the accurate extraction of text data and image data, avoiding information loss, and providing a basis for subsequent processing of these two types of data separately. Parsing the text data to obtain its detailed structural information (paragraph line information) and semantic information helps to understand the text layout and provides a basis for positional relationship judgment for the subsequent determination of the relationship between text and image. Parsing the image data to obtain its coordinate information plays an important role in the subsequent extraction of effective text components and providing accurate target images and their coordinates corresponding to the text content; extracting the effective text components represented by the target image supplements the deficiencies of the text data and provides supplementary data for the subsequent construction of the detection result information set. Forming a detection result information set containing rich information such as text, images, coordinates, etc. provides convenience for subsequent output and use, meets user needs, improves the consistency and comparability of archive content records, and quickly and accurately detects, extracts and integrates required information.

[0067] In some embodiments, the association degree between each text data and each image data is calculated based on the semantic information, image coordinates, and paragraph line information, and the target image is determined according to the association degree, referring to the following formula (1): (1) Among them, is the target image; For image data data points; is a collection of image data; is the degree of association; is the correlation threshold;

[0068] According to the correlation degree, the target image coordinates are determined, referring to the following formula (2): (2)

[0069] in, is the target image coordinate; for The image in the image data pointed to by the index; is the index of the target image that refers to the corresponding text; is the index of any image in the image data set; is the degree of association; is the correlation threshold.

[0070] The degree of association can be a metric that measures the degree of correlation between a data point in text data and a data point in image data. The index of a target image can be the unique location information corresponding to an image data point in an image data set whose degree of association with a specific text data point exceeds a threshold. The threshold is determined based on experimental data or experience. The image data set can be a dataset containing multiple image data points, each of which contains specific image information and possibly metadata.

[0071] Specifically, in the search for the relationship between image and text, how to judge the correlation between image and text is crucial. For example, the text in the previous structure cannot be used to interpret the image in the current structure. This will lead to a distorted interpretation of the image content. Therefore, a certain relationship strength between the image and text is required to interpret the image based on the semantic information of the text. This means that the image needs to select a certain paragraph or paragraphs of text as the background for image interpretation. The specific implementation method is as follows: Combine semantic information, image coordinates and paragraph line information to analyze the potential relationship between text data and image data. Use the correlation calculation method to calculate the correlation between each text data and each image data to obtain a correlation matrix, where each element in the correlation matrix represents the degree of correlation between a data point of text data and a data point of image data. Traverse the correlation matrix to find the image data point whose correlation with the text data at the location of the key information in the detected structured multi-type rich text is greater than the threshold, and determine it as the target image. Use formula (1) to calculate the index of the target image, and select the corresponding target image from the image data set according to the index. Using the target image as an index, the corresponding target image coordinates are extracted from the image data set according to formula (2).

[0072] Through the method provided in this embodiment, the correlation between each text data and each image data is calculated, providing a calculation basis for the subsequent determination of the target image and the target image coordinates. According to the correlation, the target image that is most relevant to the data point of the text data at the location of the key information in the detected structured multi-type rich text is determined, providing a basis for the subsequent extraction of the target image coordinates. According to the correlation, the coordinates of the target image that is most relevant to the text data point at the location of the key information in the detected structured multi-type rich text are determined, providing accurate image position information for subsequent data analysis, information extraction and integration. Through the above calculation method, it is possible to accurately find the strong correlation between the image and the text, thereby avoiding the image from being distorted and the image interpretation failing to express the true meaning of the placement here.

[0073] In some embodiments, a keyword set is extracted from a set of text data; a keyword set is extracted from a set of image data; and the intersection size of the keyword set in the set of text data and the keyword set in the set of image data is calculated to obtain a keyword matching score, referring to the following formula (3):

[0074] (3) Among them, Score keyword matching; is a set of keywords in a set of text data; is a set of keywords in a set of image data;

[0075] Convert the set of text data into text semantic vectors; convert the set of image data into image semantic vectors; calculate the cosine similarity between the text semantic vectors and the image semantic vectors to obtain the semantic matching score, refer to the following formula (4); (4)

[0076] in; is the cosine similarity; is the dot product formula of text semantic vector and image semantic vector; is the length of the text semantic vector; is the length of the image semantic vector; calculate the proximity between the target image coordinates and the paragraph line information corresponding to the keyword set in any text data to obtain the position relationship score; according to the keyword matching score, semantic matching score and position relationship score, obtain the association degree between each text data and each image data, refer to the following formula (5):

[0077] (5)

[0078] in, is the degree of association; is the weight coefficient of the position relationship score; Score keyword matching; is the weight coefficient of semantic matching score; Score for semantic matching; is the weight coefficient of the position relationship score; Score the positional relationship.

[0079] The keyword set in a text data set can be a set of representative or important words extracted from the text data. The keyword set in an image data set can be a set of keywords further extracted from text information extracted from the image data using image recognition or optical character recognition (OCR) technology. The intersection size of the keyword sets can be the number of keywords shared by the keyword sets of the text data and the image data. The keyword matching score can be a score calculated based on the keyword intersection size and used to measure the degree of match between the text and image at the keyword level. The text semantic vector can be a vector converted from text data to represent the semantic information of the text. The image semantic vector can be a vector converted from image data to represent the semantic information of the image. The cosine similarity can measure the directional similarity between two vectors and is used to compare the similarity between text and image semantic vectors. The semantic matching score can be a score calculated based on the similarity between the text and image semantic vectors to measure their match at the semantic level. The length of the text semantic vector can be the Euclidean length of the text semantic vector in multidimensional space. The length of an image semantic vector can be the Euclidean length of the image semantic vector in multidimensional space. The dot product formula can be the sum of the multiplication of corresponding elements of two vectors, used to calculate the similarity between the vectors. The proximity can measure the positional proximity between text and images. The positional relationship score can be calculated based on the positional relationship between text and images in multimedia fusion content, used to measure their spatial relevance.

[0080] Specifically, in the process of determining the degree of association, we can start from three dimensions, namely, the degree of keyword matching between the image and text, the degree of semantic matching between the image and text, and the positional relationship between the image and text. The above three dimensions can be quantified. Specifically, the keywords in the text data set can be extracted by keyword extraction algorithm to form a keyword set of the text data set; the keywords in the image data set can be extracted by image recognition technology; the intersection size of the keyword set of each text data and the keyword set of each image data is calculated according to formula (3); and the intersection size is recorded as the keyword matching score. The text data is converted into a text semantic vector using a pre-trained text embedding model; the image data is converted into an image semantic vector using a pre-trained image embedding model; the cosine similarity between the text semantic vector and the image semantic vector is calculated according to formula (4), and the cosine similarity is recorded as the semantic matching score. The paragraph line information corresponding to the keyword set in the text data is extracted; the proximity between the target image coordinates and the paragraph line information is calculated, and the proximity is recorded as the positional relationship score. Based on the keyword matching score, semantic matching score and positional relationship score, the association degree is calculated according to formula (5).

[0081] By the method provided by this embodiment, the key information in the archival content record is extracted more accurately by combining the keyword sets of text and image, which is more comprehensive and accurate than the method that relies solely on text or image. The calculation of the correlation takes into account the keywords, semantics and positional relationships between text and image, which helps to identify the consistent information in the archival content record, thereby improving the comparability between data. By calculating the correlation between text and image, the required information can be located quickly and accurately, avoiding the tedious process of manual search and sorting in unstructured free rich text, greatly improving the efficiency of data analysis; using the above formula to calculate the correlation makes the information in the archival content record easier to integrate and utilize; by adjusting the weight coefficients corresponding to the keyword matching score, semantic matching score and positional relationship score, it can flexibly adapt to different application scenarios and needs, which makes the method very scalable and adaptable.

[0082] In some embodiments, the proximity between the target image coordinates and the paragraph line information corresponding to the keywords in the keyword set in any text data is analyzed to obtain the difference between the row numbers of the text and the image, the difference between the column numbers of the text and the image, and the surrounding situation between the text and the image; based on the surrounding situation of the image and the text, the surrounding distance of the text on the X axis of the image and the surrounding distance of the text on the Y axis of the image are obtained;

[0083] The positional relationship score is calculated based on the difference in row numbers, column numbers, X-axis wraparound distance, and Y-axis wraparound distance, using the following formula (6):

[0084] (6)

[0085] in, is the weight coefficient of the absolute value of the difference between row numbers; is the absolute value of the difference between row numbers; is the weight coefficient of the difference in column numbers; is the difference in column numbers; is the weight coefficient of the surrounding distance on the X axis; is the orbital distance on the X axis; is the weight coefficient of the surrounding distance on the Y axis; is the wrap-around distance on the Y axis.

[0086] The row number difference may be the difference between the row number of the target image and the row number of the paragraph corresponding to the keyword in the text data. The column number difference may be the difference between the column number of the target image and the column number of the paragraph corresponding to the keyword in the text data. The wrapping condition may indicate whether the text wraps around the image, as well as the direction and degree of wrapping. The wrapping distance may be the distance between the text and the image boundary when the text wraps around the image.

[0087] Specifically, when quantifying the positional relationship between an image and text, it's necessary to consider the layout characteristics of the image and text, such as wrapping or text surrounding the image. Furthermore, there are situations where images are indirectly used, such as when a figure title is referenced. Specifically, when descriptive terms such as "as shown" are used in text, the figure title should be replaced with the image before its characteristics are determined. The specific quantification process for these features can be found as follows: Extract the coordinate information of the target image, including its row and column numbers in the document. For each keyword in the text data, find the corresponding paragraph line information, including its row and column numbers. Calculate the row and column number differences between the target image coordinates and the paragraph line information corresponding to each keyword. Analyze the wrapping between the text and image, namely, whether the text wraps around the image, as well as the direction and degree of wrapping. Based on the wrapping between the text and image, determine the relative position of the text on the X and Y axes of the image. Calculate the offset of the text on the X axis relative to the image center as the wrapping distance. Calculate the offset of the text on the Y axis relative to the image center as the wrapping distance. The position relationship score is calculated using formula (6) based on the difference in row numbers, column numbers, X-axis wraparound distance, and Y-axis wraparound distance, and can be adjusted based on actual conditions.

[0088] Through the method provided by this embodiment, by carefully analyzing the relative positions of text and image in the document, including the difference in line numbers, column numbers and surrounding conditions, it is possible to more accurately locate text information closely related to the image; the calculation of the position relationship score takes into account the spatial layout relationship between text and image, which helps to identify and verify consistency information in archival content records. Even in different archival content records, as long as the spatial relationship between text and image is similar, their position relationship scores will be similar, thereby improving the comparability between data. By calculating the position relationship score, we can quickly and accurately locate text information closely related to the image, avoiding the tedious process of manual search and sorting in unstructured free rich text, and greatly improving the efficiency of data analysis. This method is not only applicable to text data, but also combines image data and supports the fusion analysis of multimedia data, which is particularly important for the analysis of image information such as photos that contain a large amount of key information in archival content records.

[0089] In some embodiments, the target image is parsed to determine whether the target image has text content; if so, the image text content of the target image is extracted; an interpretation is performed based on the image text content and the semantic information of the image screen context, and valid text components in the image are extracted through the image text content, image interpretation information and semantic information; if not, an interpretation is performed based on the semantic information of the image screen context, and the image interpretation information is used as a valid text component.

[0090] Text content can be recognizable text within an image. Image text content can be readable text extracted from a target image. Interpretation information can be an interpretation based on the image and contextual semantics when the target image lacks readable text or the text is insufficient to convey complete information.

[0091] Specifically, when detecting text content in an image, consider whether the image contains the annotations described above. Use image recognition technology to identify text areas in the image and determine whether there is recognizable text content. If text content exists, use image recognition tools to extract the text content in the target image. Preprocess the extracted text content and perform text analysis based on the extracted image text content, image interpretation, and contextual semantic information. Key information related to the archival content record is identified and extracted as valid text components. If text content does not exist, analyze the target image's image content, identify elements such as the subject, state, and scene in the image, and interpret and describe the image content based on contextual semantic information. Organize the image interpretation information into text form as valid text components.

[0092] The method provided by this embodiment not only considers the textual content that may exist in an image, but also the information conveyed by the image itself, thereby improving the comprehensiveness of information extraction. By extracting textual content from images and combining it with semantic information, unstructured free-form rich text can be converted into a more structured and understandable form, which facilitates better understanding and utilization of the information contained in archival content records. Consistent key information is extracted from different types of archival content records (including text and images), thereby improving the consistency and comparability of archival content records and facilitating comparison and analysis across multiple archival content records to identify potential issues and areas for improvement. The process of obtaining effective text components is highly applicable to image data and supports the integrated analysis of multimedia data, which is particularly important for the analysis of image information such as photographs contained in archival content records. By extracting effective textual components from images, key information in archival content records can be located more quickly and accurately, avoiding the tedious process of manually searching and sorting through unstructured free-form rich text, thereby optimizing data analysis efficiency.

[0093] In some embodiments, interpretation is performed based on the image screen in conjunction with the context to obtain interpretation information; the similarity between the interpretation information and the semantic information is calculated; if the similarity is greater than a preset threshold, the interpretation information is used as a valid text component.

[0094] The similarity can be the degree of similarity between the interpretation information and the semantic information. The preset threshold can be a condition for determining whether the similarity is satisfied, obtained from experimental data or experience, and stored in a preset database.

[0095] Specifically, image interpretation information must be relevant to the image content to avoid ambiguity, thus requiring a reverse lookup of the interpretation information's similarity. This is achieved by analyzing the target image's content and identifying its elements. Combining contextual semantic information with the image's specific meaning within the archival content record, the interpretation information is generated to describe the image's meaning and its relationship to the context. The image interpretation information and contextual semantic information are subjected to text preprocessing. A text similarity calculation algorithm is used to calculate the similarity between the interpretation information and the semantic information, resulting in a similarity score representing the degree of match between the interpretation information and the semantic information. A similarity threshold is set as the criterion for determining the validity of the interpretation information. The calculated similarity score is compared with the threshold. If the similarity score exceeds the threshold, the interpretation information is considered valid; otherwise, it is considered invalid. If the interpretation information is determined to be valid, it is included as a valid text component of the archival content record.

[0096] Through the method provided by this embodiment, the image and contextual semantic information are combined for interpretation, and the specific meaning of the image in the archival content record is understood more accurately, thereby extracting more accurate information. At the same time, not only the text content in the image is considered, but also the information conveyed by the image itself, which improves the comprehensiveness of information extraction; the unstructured free rich text and image information are converted into structured interpretation information, which facilitates subsequent data analysis and processing. The structured interpretation information improves the consistency and comparability of the archival content record, making it easier to compare and analyze between multiple archival content records; by extracting the effective text components in the image, the key information in the archival content record is more accurately located.

[0097] In some embodiments, based on the valid text components, the distribution status of the valid text components is determined according to the target image coordinates, paragraph line information, and semantic information; multimodal information fusion is performed on the valid text components and semantic information according to the distribution status to obtain a detection result information set.

[0098] The distribution state of effective text components can be the distribution and arrangement of effective information in the text. Multimodal information fusion can be the integration and fusion of information from different modalities to form a more comprehensive and richer information representation, such as text, images, etc.

[0099] Specifically, the interpretation information of an image is generally not a single word, but rather a paragraph or multiple paragraphs. Its relevance to the text, the specific placement of each paragraph in the text, and many other factors that need to match the semantic environment need to be integrated again based on the text semantic information and paragraph line information. Based on the target image coordinates, paragraph line information, and semantic information, the distribution of each valid text component in the text is determined. The valid text components are associated with the semantic information to ensure that each valid text component has a suitable corresponding position for better integration with the semantic environment. Based on the distribution of the valid text components, multimodal information fusion technology is used to fuse the valid text components with the text's semantic information into a unified information set.

[0100] By combining the effective text components, semantic information, paragraph line information, and target image coordinates, the distribution of text in an image can be more accurately determined, improving the accuracy of information extraction. Unstructured free rich text and image information are converted into a structured detection result information set, facilitating subsequent data analysis and processing. Through multimodal information fusion, the effective text components are integrated with the text's semantic information to form a more unified, rich, and comprehensive detection result information set.

[0101] In some embodiments, the detection result information set is filtered according to the query information to obtain a filtered information set; a preset structured template is obtained; the structured template is filled with information through the filtered information set to obtain a streamlined information set; and the streamlined detection result information set is output as structured text.

[0102] The filtered information set can be the information set obtained after filtering and processing the test result information set. The preset structured template can be a structured format for populating information and can include key fields of the archive content record, such as recording time, user, and event content. It can be stored in a preset database. Information filling can be the process of populating the extracted key information into the preset structured template.

[0103] Specifically, determine the specific content of the query information, including keywords, time range, event type, etc. Traverse the detection result information set and perform a matching check on each information item to see if it meets the conditions of the query information. Filter out the information items that meet the conditions to form a filtered information set. Design and preset a structured template based on the specific needs and format requirements of the archive content records. Store the preset structured template in the database for subsequent use. Traverse the filtered information set and extract the key data from each information item. Fill the extracted key data into the corresponding fields of the structured template. Repeat the above steps until all filtered information items have been processed to obtain a streamlined information set. Convert the streamlined information set into text format. Ensure that the text format meets the preset structured template requirements. Output the structured text to the specified location or file.

[0104] Through the method provided by this embodiment, the detection result information set is screened according to the query information, and the required information is quickly and accurately located, avoiding the tedious process of manual search in unstructured text, and improving the efficiency of subsequent processing. The screened information is filled with a preset structured template to ensure the unified format and standard of the output information, improve the consistency and comparability of the archive content records, make the information clearer and easier to understand, and provide convenience for subsequent data analysis. It can be directly used for report generation, data visualization or further data processing and analysis. The structured text output allows users to understand the specific circumstances and background of the event more clearly, which helps to make more accurate judgments and decisions. Since the information is screened and structured, the efficiency and accuracy of data analysis are improved.

[0105] Figure 3 A structural diagram of a text detection system provided in one embodiment of the present application is shown in FIG. Figure 3As shown, the text detection system 300 of this embodiment includes: a rich text history archive acquisition module 301, a structured multi-type rich text parsing module 302, a detection demand-oriented information determination module 303, a multimedia content judgment module 304, a detection result information set construction module 305 and a structured text output module 306.

[0106] Rich text history archive acquisition module 301, used to obtain text usage requirements; based on text usage requirements, obtain rich text history archives; structured multi-type rich text parsing module 302, used to parse rich text history archives and determine structured multi-type rich texts; detection demand-oriented information determination module 303, used to determine detection demand-oriented information based on text usage requirements and structured multi-type rich texts; multimedia content judgment module 304, used to determine whether there is multimedia fusion content in the detection demand-oriented information; detection result information set construction module 305, used to determine the detection result information set based on the multimedia fusion content if it exists; structured text output module 306, used to output the detection result information set in structured text

[0107] Optionally, the detection result information set construction module 305 is specifically used to: divide the multimedia fusion content into text data and image data; parse the text data to determine the paragraph line information and semantic information of the text data; parse the image data to determine the image coordinates; determine the target image and target image coordinates referring to the corresponding text based on the semantic information, image coordinates and paragraph line information; determine the valid text component represented by the target image based on the target image and semantic information; determine the detection result information set based on the valid text component, semantic information, paragraph line information and target image coordinates.

[0108] The detection result information set construction module 305 is specifically used to calculate the correlation between each text data and each image data based on the semantic information, image coordinates and paragraph line information, and determine the target image according to the correlation, referring to the following formula:

[0109] ;in, is the target image; For image data data points; is a collection of image data; is the degree of association; is the correlation threshold;

[0110] According to the degree of association, determine the target image coordinates, refer to the following formula: ;in, is the target image coordinate; for The image in the image data pointed to by the index; is the index of the target image that refers to the corresponding text; is the index of any image in the image data set; is the degree of association; is the correlation threshold.

[0111] The detection result information set construction module 305 is specifically used to: extract a keyword set from the text data set; extract a keyword set from the image data set; calculate the intersection size of the keyword set in the text data set and the keyword set in the image data set to obtain a keyword matching score, referring to the following formula:

[0112] ;in, Score keyword matching; is a set of keywords in a set of text data; is a set of keywords in a set of image data;

[0113] Convert the set of text data into text semantic vectors; convert the set of image data into image semantic vectors; calculate the cosine similarity between the text semantic vectors and the image semantic vectors to obtain the semantic matching score, refer to the following formula;

[0114] in; is the cosine similarity; is the dot product formula of text semantic vector and image semantic vector; is the length of the text semantic vector; is the length of the image semantic vector;

[0115] Calculate the proximity between the target image coordinates and the paragraph line information corresponding to the keyword set in any text data to obtain the position relationship score; based on the keyword matching score, semantic matching score and position relationship score, obtain the association between each text data and each image data, referring to the following formula:

[0116] ;

[0117] in, is the degree of association; is the weight coefficient of the position relationship score; Score keyword matching; is the weight coefficient of semantic matching score; Score for semantic matching; is the weight coefficient of the position relationship score; Score the positional relationship.

[0118] The detection result information set construction module 305 is specifically used to: analyze the proximity between the target image coordinates and the paragraph line information corresponding to the keywords in the keyword set in any text data, and obtain the difference between the row numbers of the text and the image, the difference between the column numbers of the text and the image, and the surrounding situation between the text and the image; based on the surrounding situation of the image and text, obtain the surrounding distance of the text on the X axis of the image and the surrounding distance of the text on the Y axis of the image; and calculate the position relationship score based on the difference in row numbers, column numbers, surrounding distance on the X axis, and surrounding distance on the Y axis, referring to the following formula:

[0119] ;

[0120] in, is the weight coefficient of the absolute value of the difference between row numbers; is the absolute value of the difference between row numbers; is the weight coefficient of the difference in column numbers; is the difference in column numbers; is the weight coefficient of the surrounding distance on the X axis; is the orbital distance on the X axis; is the weight coefficient of the surrounding distance on the Y axis; is the wrap-around distance on the Y axis.

[0121] The detection result information set construction module 305 is specifically used to: parse the target image and determine whether the target image has text content; if so, extract the image text content of the target image; interpret based on the image text content and the semantic information of the image screen context, and extract the valid text components in the image through the image text content, image interpretation information and semantic information; if not, interpret based on the semantic information of the image screen context, and use the image interpretation information as the valid text component.

[0122] The detection result information set construction module 305 is specifically used to: interpret the image based on the context to obtain interpretation information; calculate the similarity between the interpretation information and the semantic information; if the similarity is greater than a preset threshold, the interpretation information is used as a valid text component.

[0123] The detection result information set construction module 305 is specifically used to: determine the distribution status of the valid text components based on the valid text components, according to the target image coordinates, paragraph line information, and semantic information; perform multimodal information fusion on the valid text components and semantic information according to the distribution status to obtain the detection result information set.

[0124] The structured text output module 306 is specifically used to: filter the detection result information set according to the query information to obtain a filtered information set; obtain a preset structured template; fill the structured template with information through the filtered information set to obtain a streamlined information set; and output the streamlined detection result information set as structured text.

[0125] The system of this embodiment can be used to execute the method of any of the above embodiments. Its implementation principles and technical effects are similar and will not be described in detail here.

Claims

1. A text detection method, characterized in that: include: Obtaining text usage requirements; obtaining rich text historical archives based on the text usage requirements; Parsing the rich text historical archive to determine structured multi-type rich text; Determining detection demand-oriented information based on the text usage requirements and the structured multi-type rich text; Determining whether there is multimedia fusion content in the detection demand-oriented information; If so, dividing the multimedia fusion content into text data and image data; Parsing the text data to determine paragraph and line information and semantic information of the text data; parsing the image data to determine image coordinates; Calculating the degree of association between each text data and each image data based on the semantic information, the image coordinates, and the paragraph line information; Determining a target image and target image coordinates referring to the corresponding text according to the association degree; determining a detection result information set according to the semantic information, the paragraph line information, and the target image coordinates; Output the test result information set as structured text; The calculating the association degree between each text data and each image data based on the semantic information, the image coordinates and the paragraph line information includes: extracting a set of keywords from the set of text data; extracting a keyword set from the set of image data; Calculate the intersection size of the keyword set in the text data set and the keyword set in the image data set to obtain a keyword matching score, referring to the following formula: ;in, Score keyword matching; is a set of keywords in a set of text data; is a set of keywords in a set of image data; Convert a collection of text data into a text semantic vector; Convert a collection of image data into an image semantic vector; Calculate the cosine similarity between the text semantic vector and the image semantic vector to obtain the semantic matching score, refer to the following formula; ;in; is the cosine similarity; is the dot product formula of text semantic vector and image semantic vector; is the length of the text semantic vector; is the length of the image semantic vector; calculating the proximity between the target image coordinates and the paragraph line information corresponding to the keyword set in any text data to obtain a position relationship score; According to the keyword matching score, the semantic matching score and the position relationship score, the association degree between each text data and each image data is obtained, referring to the following formula: in, is the degree of association; is the weight coefficient of the position relationship score; Score keyword matching; is the weight coefficient of semantic matching score; Score for semantic matching; is the weight coefficient of the position relationship score; is a position relationship score; the calculation of the proximity between the target image coordinates and the paragraph line information corresponding to the keyword in the keyword set in any text data to obtain the position relationship score includes: Analyze the proximity between the target image coordinates and paragraph line information corresponding to keywords in a keyword set in any text data to obtain the difference between the row numbers of the text and the image, the difference between the column numbers of the text and the image, and the surrounding situation between the text and the image; According to the wrapping conditions of the image and text, obtaining the wrapping distance of the text on the X-axis of the image and the wrapping distance of the text on the Y-axis of the image; The positional relationship score is calculated based on the difference in row numbers, the difference in column numbers, the surrounding distance on the X axis, and the surrounding distance on the Y axis, with reference to the following formula: ;in, is the weight coefficient of the absolute value of the difference between row numbers; is the absolute value of the difference between row numbers; is the weight coefficient of the difference in column numbers; is the difference in column numbers; is the weight coefficient of the surrounding distance on the X axis; is the orbital distance on the X axis; is the weight coefficient of the surrounding distance on the Y axis; is the wrap-around distance on the Y axis.

2. The method according to claim 1, wherein determining the detection result information set based on the semantic information, the paragraph line information, and the target image coordinates comprises: determining, based on the target image and the semantic information, a valid text component represented by the target image; A detection result information set is determined according to the valid text components, the semantic information, the paragraph line information and the target image coordinates.

3. The method according to claim 2, characterized in that Determining the target image and the target image coordinates referring to the corresponding text according to the association degree includes: According to the correlation degree, the target image is determined, referring to the following formula: in, is the target image; For image data data points; is a collection of image data; is the degree of association; is the correlation threshold; According to the correlation degree, the target image coordinates are determined, referring to the following formula: ;in, is the target image coordinate; for The image in the image data pointed to by the index; is the index of the target image that refers to the corresponding text; is the index of any image in the image data set; is the degree of association; is the correlation threshold.

4. The method according to claim 2, characterized in that The determining, based on the target image and the semantic information, the valid text component represented by the target image includes: Parse the target image and determine whether there is text content in the target image; If it exists, extract the image text content of the target image; interpret the image text content and the semantic information of the image screen in the context, and extract the effective text components in the image through the image text content, image interpretation information and the semantic information; If it does not exist, the image interpretation information is interpreted based on the semantic information in the context of the image and the image interpretation information is used as a valid text component.

5. The method according to claim 4, characterized in that The step of interpreting the semantic information based on the context of the image and taking the image interpretation information as a valid text component includes: Interpret the image based on its context to obtain interpretation information; Calculating the similarity between the interpretation information and the semantic information; If the similarity is greater than a preset threshold, the interpretation information is used as a valid text component.

6. The method according to claim 2, characterized in that The determining of the detection result information set according to the valid text components, the semantic information, the paragraph line information, and the target image coordinates includes: Based on the effective text components, the distribution state of the effective text components is determined according to the target image coordinates, paragraph line information, and semantic information; Multimodal information fusion is performed on the valid text components and the semantic information according to their distribution states to obtain a detection result information set.

7. The method according to claim 6, characterized in that Outputting the detection result information set in structured text includes: Filtering the test result information set according to the query information to obtain a filtered information set; Obtaining a preset structured template; filling the structured template with information using the filtered information set to obtain a streamlined information set; The simplified detection result information set is output as structured text.

8. A text detection system, characterized in that: The method as claimed in any one of claims 1 to 7 comprises: A rich text history archive acquisition module is used to obtain text usage requirements; and obtain rich text history archives according to the text usage requirements; A structured multi-type rich text parsing module, used to parse the rich text historical archives and determine the structured multi-type rich text; A detection demand-oriented information determination module, configured to determine detection demand-oriented information based on the text usage requirements and the structured multi-type rich text; A multimedia content determination module, configured to determine whether the detection demand-oriented information contains multimedia fusion content; a detection result information set construction module, for determining the detection result information set based on the multimedia fusion content, if any; The structured text output module is used to output the detection result information set in structured text.

Citation Information

Patent Citations

  • Text structured processing method and device, computer equipment, medium and product

    CN114692573A

  • Souvenir generation method and device and electronic equipment

    CN114969397A