Text detection method, apparatus, device, and medium

By using multimodal analysis and text semantic version graph matching, the problem of low text detection accuracy in existing technologies has been solved, and efficient identification of subtle text information and risky texts has been achieved.

CN122113893APending Publication Date: 2026-05-29PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-14
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies lack the ability to detect subtle changes in text, resulting in low accuracy in text detection.

Method used

Text fingerprints are obtained through multimodal analysis, a text semantic version graph is constructed, and text matching is performed to determine semantic differences and assess text risk.

Benefits of technology

It improves the accuracy and reliability of text detection, and is able to identify subtle changes in information and risky text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113893A_ABST
    Figure CN122113893A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent decision-making, and discloses a text detection method, a text detection device, equipment and a medium: a to-be-detected text is acquired, multi-modal analysis is performed on the to-be-detected text, and a multi-modal text fingerprint is obtained; when a text verification result of the to-be-detected text represents that all text verifications pass, a text semantic version graph is acquired; text matching is performed on the multi-modal text fingerprint and historical text fingerprints in the text semantic version graph, and a target detection text is obtained; a difference between the to-be-detected text and the target detection text is determined as a semantic difference, and when the semantic difference is key field information and a state of the target detection text is completed, the to-be-detected text is determined as a risk text. The application is applied to a financial scene. Through the multi-modal text fingerprint and the text semantic version graph, subtle information in the to-be-detected text and overall text fingerprint detection are realized, and the accuracy and reliability of risk text detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, and in particular to a text detection method, apparatus, device, and medium. Background Technology

[0002] Text intelligence processing technology, as a core component of fintech, has been widely applied in insurance, credit, and claims processing. Current technologies primarily focus on accurately extracting text content from documents, lacking the ability to detect subtle changes in text information, resulting in low accuracy in text detection. Therefore, a text detection method is urgently needed to address these issues. Summary of the Invention

[0003] This invention provides a text detection method, apparatus, device, and medium to improve the technical problem in the prior art where the lack of detection of subtle changes in text leads to low accuracy in text detection.

[0004] A text detection method, comprising: Obtain the text to be detected, perform multimodal analysis on all the texts to be detected, and obtain the multimodal text fingerprint corresponding to each text to be detected; When the text verification results of all the texts to be detected indicate that all text verifications have passed, a text semantic version map corresponding to each of the texts to be detected is obtained; Text matching is performed on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version map to obtain the target detection text corresponding to each of the texts to be detected; The differences between each of the texts to be detected and the target text to be detected are determined as semantic differences corresponding to each of the texts to be detected. When the semantic differences are only key field information and the target text is in the state of completed, the text to be detected is determined to be a risky text.

[0005] A text detection device, comprising: The modal analysis module is used to acquire the text to be detected, perform multimodal analysis on all the texts to be detected, and obtain the multimodal text fingerprint corresponding to each text to be detected. The graph acquisition module is used to acquire the text semantic version graph corresponding to each of the texts to be detected when the text verification result of all the texts to be detected is characterized as passing all text verifications; The text matching module is used to perform text matching on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version graph to obtain the target detection text corresponding to each of the texts to be detected; The text recognition module is used to determine the differences between each of the texts to be detected and the target text to be detected as semantic differences corresponding to each of the texts to be detected, and to determine the text to be detected as risky text when the semantic differences are only key field information and the target text is in the state of completed.

[0006] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor being used to perform the text detection method described above.

[0007] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described text detection method.

[0008] The aforementioned text detection method, apparatus, device, and medium of this invention, through the text to be detected, achieve the extraction of multimodal text fingerprints, thereby enabling the verification of the text to be detected and the acquisition of a text semantic version graph. Text matching is performed on each multimodal text fingerprint and historical text fingerprint in the text semantic version graph to determine similar texts to the text to be detected, thus determining semantic differences and ultimately judging whether the text to be detected is risky, and identifying risky texts. Furthermore, through multimodal text fingerprints and text semantic version graphs, subtle information in the text to be detected and the overall fingerprint of the document are detected, thereby improving the accuracy and reliability of risky text detection. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart of a text detection method in one embodiment of the present invention; Figure 2 This is a schematic diagram of a text detection device in one embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] In one embodiment, such as Figure 1 As shown, a text detection method is provided, including the following steps: S10: Obtain the text to be detected, perform multimodal analysis on all the texts to be detected, and obtain the multimodal text fingerprint corresponding to each text to be detected.

[0013] Understandably, the text to be detected refers to a user-uploaded document. For example, in a credit scenario, this could be proof of income, bank statements, or identification documents; in an insurance claims scenario, it could be an accident liability determination letter or a vehicle repair order. Multimodal Document Fingerprint (MDF) refers to the fusion of multimodal content features from a document, including text, images, tables, formulas, etc., such as semantic features, image features, and content features.

[0014] Specifically, all texts to be detected are acquired. Then, multimodal analysis is performed on each text, which involves extracting image features, semantic features, and content features from the text. This means extracting and fusing content and morphological features from the text to obtain a multimodal text fingerprint corresponding to each text. This includes extracting layout, seal location, signature handwriting, relationships between entities, as well as Exif data, creation time, and file type.

[0015] S20: When the text verification results of all the texts to be detected indicate that all text verifications have passed, obtain the text semantic version map corresponding to each of the texts to be detected.

[0016] Understandably, text validation results refer to the results used to characterize whether all text has passed validation. A Document Semantic Version Graph (DSVG) is a graph-based representation that uses semantics as its core to structurally model the semantic evolution trajectory of text during multiple edits / iterations. Its core is to use nodes to record versions, use edges to label semantic associations and changed attributes, and add semantic vectors and metadata.

[0017] Specifically, the time and content of all texts to be tested are examined to obtain text verification results. If the time and content of all texts to be tested meet preset conditions, a text verification result indicating that all texts have passed verification is obtained. Then, by performing information matching on each text to be tested, the corresponding text semantic version graph is retrieved sequentially.

[0018] S30: Perform text matching on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version map to obtain the target detection text corresponding to each of the texts to be detected.

[0019] Understandably, historical text fingerprints refer to the multimodal text fingerprints of each historical detected text in the text semantic version graph. Target detected text refers to historical detected texts that are similar to the text to be detected.

[0020] Specifically, text matching is performed on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version graph. That is, the fingerprint similarity between the multimodal text fingerprint and the historical text fingerprint is calculated, and a preset similarity threshold corresponding to each text to be detected is obtained. The fingerprint similarity corresponding to the same text to be detected is compared with the preset similarity threshold. When the fingerprint similarity is greater than the preset similarity threshold, the historical detection text corresponding to the historical text fingerprint is determined as the target detection text corresponding to the text to be detected.

[0021] In one embodiment, when the fingerprint similarity is less than or equal to a preset similarity threshold, it is determined that there is no text similar to the text to be detected in the text semantic version graph, and the text to be detected is determined as a non-risk text, and enters the normal business process.

[0022] S40: The difference between each of the texts to be detected and the target text to be detected is determined as the semantic difference corresponding to each of the texts to be detected, and when the semantic difference is only key field information and the status of the target text to be detected is completed, the text to be detected is determined to be a risky text.

[0023] Understandably, semantic difference refers to the informational differences between the text to be detected and the target text to be detected. Key field information refers to the information corresponding to key fields in the text, such as invoice amount, invoice date, etc. Some key field information can be pre-defined. Risky text refers to text that may be tampered with or fraudulent.

[0024] Specifically, the information between the text to be detected and the target text to be detected is compared to determine the information differences between the two texts, which are then identified as semantic differences. Next, the status of the target text to be detected is checked to see if it is complete, and the semantic differences are analyzed to determine if they are only key field information. If the semantic differences are only key field information and the target text is in a completed state, the text to be detected is determined to be risky text.

[0025] In one embodiment, if the semantic difference is not only key field information, and / or the status of the target detection text is incomplete, then the text to be detected is determined to be non-risk text, and the normal business process is entered.

[0026] In one embodiment, in the financial field, the drafting, revision, and approval processes of financial contracts (such as loan contracts, asset management contracts, and trust agreements) are prone to risks such as tampering with key fields, omissions in clauses, and mixed versions (e.g., unauthorized modification of core information such as interest rates, repayment periods, and guarantee clauses). Specifically, a multimodal text fingerprint is generated for the entire process of the contract, including the initial draft, various revisions, and the final approved version. This not only extracts text character features but also integrates text format features (such as signature positions and clause levels) and semantic features (such as the expression of core rights and obligations) to form a unique multimodal fingerprint. Then, a text semantic version graph is constructed: the text fingerprints, version numbers, revisers, and approval status of all historical compliant versions are associated and stored to form a traceable version genealogy. Next, after the contract text to be tested completes basic text verification (such as format compliance and complete signatures), its multimodal fingerprint is matched with the historical compliant fingerprints (status "approved") in the graph. When a match is successful, the target matching text corresponding to the historical compliant fingerprint is obtained. If the two contracts differ only in key fields (such as the interest rate changing from 3.5% to 4.0% or the repayment period changing from 5 years to 3 years), and the target matching text is a compliant version that has already taken effect, then the contract to be tested is determined to be a risky text, triggering an alert.

[0027] The text detection method in this embodiment extracts multimodal text fingerprints from the text to be detected, thereby verifying the text and obtaining a text semantic version graph. Text matching is performed on each multimodal text fingerprint and historical text fingerprints in the text semantic version graph to identify similar texts, thus determining semantic differences and ultimately judging whether the text to be detected is risky. Furthermore, by using multimodal text fingerprints and the text semantic version graph, subtle information within the text to be detected and the overall fingerprint of the document are detected, thereby improving the accuracy and reliability of risky text detection.

[0028] In one embodiment, in step S10, performing multimodal analysis on all the texts to be detected to obtain a multimodal text fingerprint corresponding to each text to be detected includes: S101, feature extraction is performed on all the texts to be detected to obtain the current image features corresponding to each text to be detected.

[0029] S102, semantic extraction is performed on all the texts to be detected to obtain the current semantic features corresponding to each text to be detected.

[0030] S103, extract information from all the texts to be detected to obtain the current content features corresponding to each text to be detected.

[0031] S104, determine the multimodal text fingerprint corresponding to each of the texts to be detected based on the current image features, current semantic features and current content features corresponding to the same text to be detected.

[0032] Understandably, current image features refer to the features of information such as layout, stamp position, and signature handwriting in the text to be detected. Current semantic features refer to the semantic features of entities and their associated information in the text to be detected. Current content features refer to the features of information such as Exif data, creation time, and file type in the text to be detected.

[0033] Specifically, feature extraction is performed on all text to be detected. This involves using image analysis techniques to divide the text into regions based on information such as layout, stamp location, and signature handwriting, and then extracting the corresponding image features from each region to obtain the current image features corresponding to each text to be detected. Similarly, semantic extraction is performed on all text to be detected. This involves identifying entities in the text and extracting the association information between entities to obtain the current semantic features corresponding to each text to be detected. Then, information extraction is performed on all text to be detected. This involves identifying and extracting information such as Exif data, creation time, and file type from the text to obtain the current content features corresponding to each text to be detected. Next, based on the current image features, current semantic features, and current content features corresponding to the same text to be detected, a multimodal text fingerprint is determined for each text to be detected. This involves feature fusion, which concatenates the current image features, current semantic features, and current content features corresponding to the same text to obtain the multimodal text fingerprint corresponding to each text to be detected.

[0034] In one embodiment, the current image features are feature-encoded to obtain a structural fingerprint; similarly, the current semantic features are feature-encoded to obtain a semantic fingerprint; and the current content features are feature-encoded to obtain a metadata fingerprint. Then, the structural fingerprint, semantic fingerprint, and metadata fingerprint are concatenated and fused to obtain the multimodal text fingerprint corresponding to each text to be detected.

[0035] In this embodiment, feature extraction is performed on the text to be detected, which realizes the extraction of current image features, current semantic features and current content features, as well as multimodal analysis of the text to be detected, thereby realizing the determination of multimodal text fingerprint, and further realizing the extraction of text content and form.

[0036] In one embodiment, in step S20, the text verification results of all the texts to be detected include: S201, extract the time information from all the texts to be detected to obtain the text time corresponding to each text to be detected.

[0037] S202, perform timeline verification on all the text times to obtain the time verification result.

[0038] S203, perform cross-text detection on all the field information in the text to be detected, and obtain the cross-text detection results.

[0039] S204, when the time verification result indicates that all the text times have passed the verification, and the cross-text detection result indicates that the same field information is consistent, a text verification result indicating that all texts have passed the verification is obtained.

[0040] Understandably, text time refers to the time of each text to be checked; for example, a ticket includes the invoice date. Time verification results are used to characterize whether all text times have passed verification. Cross-text detection results are used to characterize whether the same field information is consistent across different texts.

[0041] Specifically, after obtaining the text to be detected, time information is extracted from all the texts, that is, the time, date, and time period in the texts are identified to determine the corresponding text time for each text. When multiple times exist in a text, semantic analysis is performed to filter out the text time that represents the text. Then, a timeline verification is performed on all text times, that is, according to preset time rules, all text times are checked to see if they conform to the timeline's occurrence sequence. If all text times conform to the timeline's occurrence sequence, a time verification result indicating that all text times have passed the verification is obtained. For example, the accident liability determination letter is earlier than the vehicle repair order information. If one or more text times do not conform to the timeline's occurrence sequence, a time verification result indicating that all text times have failed the verification is obtained.

[0042] Next, cross-text detection is performed on the field information of all texts to be detected. This involves extracting the field information of each text and comparing the same field information in different texts to determine whether the same field information is consistent across different texts, thus obtaining the cross-text detection result. If one or more field information is inconsistent, a cross-text detection result indicating inconsistency in the same field information is obtained. Finally, a text verification result indicating that all text times have passed verification, and that the cross-text detection results indicate that the same field information is consistent, is obtained. In one embodiment, each field information corresponds to one cross-text detection result; a text verification result indicating that all texts have passed verification is obtained when all cross-text detection results indicate that the same field information is consistent.

[0043] In this embodiment, by extracting text time and performing timeline verification, the validity of all texts to be detected is ensured, and the time verification results are obtained, thereby improving the accuracy of text risk identification. Cross-text detection ensures the accuracy of information in the texts to be detected, and the cross-text detection results are obtained, thus determining the text verification results and ultimately verifying all texts.

[0044] In one embodiment, in step S30, the step of performing text matching on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version graph to obtain the target detection text corresponding to each of the texts to be detected includes: S301, perform similarity processing on the current image features and the historical image features in each of the historical text fingerprints to obtain the image similarity corresponding to each of the historical image features.

[0045] S302, perform similarity processing on the current semantic feature and the historical semantic features in each of the historical text fingerprints to obtain the semantic similarity corresponding to each of the historical semantic features.

[0046] S303, perform similarity processing on the current content features and the historical content features in each of the historical text fingerprints to obtain the content similarity corresponding to each of the historical content features.

[0047] S304, based on the image similarity, semantic similarity and content similarity corresponding to the same text to be detected, filter out the target detection text corresponding to the text to be detected.

[0048] Intelligibly, each historical text fingerprint corresponds to a historical detected text. Image similarity refers to the feature similarity between the current image features of the text to be detected and the historical image features of the historical detected text. Semantic similarity refers to the feature similarity between the current semantic features of the text to be detected and the historical semantic features of the historical detected text. Content similarity refers to the feature similarity between the current content features of the text to be detected and the historical content features of the historical detected text.

[0049] Specifically, the current image features are processed to resemble the historical image features in each historical text fingerprint. This is achieved by calculating the Euclidean distance between the current and historical image features to obtain the image similarity to each historical image feature. Next, the current semantic features are processed to resemble the historical semantic features in each historical text fingerprint. This is achieved by calculating the cosine similarity between the current and historical semantic features to obtain the semantic similarity to each historical semantic feature. Then, the current content features are processed to resemble the historical content features in each historical text fingerprint. This is achieved by calculating the cosine similarity between the current and historical content features to obtain the content similarity to each historical content feature. Finally, based on the image similarity, semantic similarity, and content similarity to the same text to be detected, historical text fingerprints that simultaneously meet all three preset thresholds are selected, and the historical detection texts corresponding to these historical text fingerprints are identified as the target detection texts. Note that the similarity calculation in this embodiment is merely illustrative and not intended to be limiting; the algorithm can be adjusted according to actual circumstances.

[0050] In this embodiment, by calculating image similarity, semantic similarity, and content similarity respectively, matching of different information is achieved, thereby enabling the detection of subtle information and ultimately determining the target text.

[0051] In one embodiment, in step S304, the step of filtering out target detection texts corresponding to the text to be detected based on image similarity, semantic similarity, and content similarity with the same text to be detected includes: S3041, Filter out all image detection texts corresponding to image similarities greater than a preset image threshold from the historical detection texts corresponding to each of the historical text fingerprints.

[0052] S3042, Filter out all semantically detected texts that have a semantic similarity greater than a preset semantic threshold from all the image detected texts.

[0053] S3043, Select content detection texts with content similarity greater than a preset content threshold from all the semantic detection texts, and determine the content detection texts as target detection texts.

[0054] In essence, image-based text detection refers to selecting text that meets the requirements from all historical text detections. Semantic text detection refers to selecting text that meets the requirements from all image-based text detections. Content-based text detection refers to selecting text that meets the requirements from all semantic text detections. Each historical text fingerprint corresponds to one image similarity score, one semantic similarity score, and one content similarity score.

[0055] Specifically, after obtaining image similarity, semantic similarity, and content similarity, a preset image threshold corresponding to the image similarity is obtained. The image similarity of each historical detected text is compared with this preset image threshold, and historical detected texts with image similarities greater than the preset image threshold are selected and identified as image detected text. Next, a preset semantic threshold corresponding to the semantic similarity is obtained. The semantic similarity of each image detected text is compared with this preset semantic threshold, and all image detected texts with semantic similarities greater than the preset semantic threshold are selected and identified as semantic detected text. Then, a preset content threshold corresponding to the content similarity is obtained. The content similarity of each semantic detected text is compared with this preset content threshold, and semantic detected texts with content similarities greater than the preset content threshold are selected. These semantic detected texts are identified as content detected text, and the content detected texts are identified as target detected text.

[0056] In one embodiment, if the corresponding text is not found in a certain step, the text to be detected is determined to be non-risk text, and the normal business process is initiated.

[0057] In one embodiment, one of image similarity, semantic similarity, and content similarity is selected as the first filtering condition. Then, one of the remaining two similarities is selected as the second filtering condition. Finally, the remaining similarity is selected as the third filtering condition to obtain the target detection text. This embodiment does not limit the order of similarity selection. For example, text corresponding to content similarity can be filtered first, then text corresponding to semantic similarity can be filtered, and finally, the target detection text can be selected using image similarity. Alternatively, text corresponding to semantic similarity can be filtered first, then text corresponding to content similarity can be filtered, and finally, the target detection text can be selected using image similarity.

[0058] In this embodiment, by setting preset image thresholds, preset semantic thresholds, and preset content thresholds, the historical detection text is filtered, thereby determining the target detection text, and further determining similar texts to the text to be detected, which facilitates the detection of subtle information.

[0059] In one embodiment, in step S40, determining the difference between each of the texts to be detected and the target text to be detected as a semantic difference corresponding to each of the texts to be detected includes: S401, Extract all information from the text to be detected to obtain the current field information.

[0060] S402, extract all information from the target detection text to obtain historical field information.

[0061] S403, based on all the current field information and all the historical field information, determine the semantic difference corresponding to the detected text.

[0062] Understandably, current field information refers to information in the text to be detected, such as amount, date, etc. Historical field information refers to information in the target text to be detected.

[0063] Specifically, the identification and extraction of the text to be detected involves first identifying the key fields in the text, and then extracting the key fields and their corresponding current field information. Similarly, the identification and extraction of the target detection text involves first identifying the key fields in the target detection text, and then extracting the key fields and their corresponding historical field information. Next, the current field information and historical field information corresponding to the same key fields are compared to determine the information differences between the current field information and the historical field information, and these information differences are identified as semantic differences corresponding to the detection text.

[0064] In this embodiment, by using current field information and historical field information, the subtle differences between the text to be detected and the target text are determined, thereby realizing the determination of semantic differences and improving the accuracy of text detection.

[0065] In one embodiment, after step S40, after determining the difference between each of the texts to be detected and the target text to be detected as a semantic difference corresponding to each of the texts to be detected, the method further includes: S50, after determining that the text to be detected is a non-risk text, the text semantic version graph is updated using the text to be detected to obtain an updated text semantic graph.

[0066] Understandably, updating the text semantic graph means adding a version of the text semantic graph after the text to be detected.

[0067] Specifically, after determining that the text to be detected is a non-risk text, the text semantic version graph is updated using the text to be detected. That is, the multimodal text fingerprint, image similarity, semantic similarity, content similarity, and the text to be detected are added to the text semantic version graph to obtain an updated text semantic graph.

[0068] In this embodiment, the text semantic version graph is updated by updating the text to be detected, thereby obtaining the updated text semantic graph, enriching the text semantic version graph, and thus improving the accuracy of subsequent text detection.

[0069] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0070] In one embodiment, a text detection device is provided, which corresponds one-to-one with the text detection methods described in the above embodiments. For example... Figure 2 As shown, the text detection device includes a modality analysis module 10, a graph acquisition module 20, a text matching module 30, and a text recognition module 40. Detailed descriptions of each functional module are as follows: Modal analysis module 10 is used to receive the text to be detected from the target user, perform multimodal analysis on all the texts to be detected, and obtain the multimodal text fingerprint corresponding to each text to be detected; The graph acquisition module 20 is used to acquire the text semantic version graph corresponding to each of the texts to be detected when the text verification result of all the texts to be detected indicates that all text verifications have passed. The text matching module 30 is used to perform text matching on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version map to obtain the target detection text corresponding to each of the texts to be detected. The text recognition module 40 is used to determine the differences between each of the texts to be detected and the target text to be detected as semantic differences corresponding to each of the texts to be detected, and to determine the text to be detected as risky text when the semantic differences are only key field information and the state of the target text to be detected is completed.

[0071] In one embodiment, the map acquisition module 20 includes: The time extraction unit is used to extract time information from all the texts to be detected, and obtain the text time corresponding to each text to be detected. A timeline verification unit is used to perform timeline verification on all the text times and obtain the time verification result. The cross-text detection unit is used to perform cross-text detection on the field information in all the texts to be detected, and obtain the cross-text detection results. The verification pass unit is used to obtain a text verification result indicating that all text times have passed verification when the time verification result indicates that all the text times have passed verification and the cross-text detection result indicates that the same field information is consistent.

[0072] In one embodiment, the modal analysis module 10 includes: The feature extraction unit is used to extract features from all the texts to be detected, and obtain the current image features corresponding to each text to be detected; The semantic extraction unit is used to extract semantics from all the texts to be detected and obtain the current semantic features corresponding to each text to be detected. An information extraction unit is used to extract information from all the texts to be detected to obtain the current content features corresponding to each text to be detected. The text fingerprint unit is used to determine the multimodal text fingerprint corresponding to each of the texts to be detected based on the current image features, current semantic features, and current content features corresponding to the same text to be detected.

[0073] In one embodiment, the text matching module 30 includes: An image similarity unit is used to perform similarity processing on the current image features and the historical image features in each of the historical text fingerprints to obtain the image similarity corresponding to each of the historical image features; A semantic similarity unit is used to perform similarity processing on the current semantic feature and the historical semantic features in each of the historical text fingerprints to obtain the semantic similarity corresponding to each of the historical semantic features; The content similarity unit is used to perform similarity processing on the current content feature and the historical content features in each of the historical text fingerprints to obtain the content similarity corresponding to each of the historical content features; The target text unit is used to filter out target detection texts corresponding to the text to be detected based on image similarity, semantic similarity, and content similarity to the same text to be detected.

[0074] In one embodiment, the target text unit includes: The image text filtering subunit is used to filter out all image detection texts with image similarity greater than a preset image threshold from the historical detection texts corresponding to each of the historical text fingerprints. The semantic text filtering subunit is used to filter out all semantic texts with semantic similarity greater than a preset semantic threshold from all the image-detected texts; The content text filtering subunit is used to filter out the content detection texts with a content similarity greater than a preset content threshold from all the semantic detection texts, and to determine the content detection texts as target detection texts.

[0075] In one embodiment, the text recognition module 40 includes: The current information unit is used to extract all information from the text to be detected to obtain the current field information; The historical information unit is used to extract all information from the target detection text to obtain historical field information; The semantic difference unit is used to determine the semantic difference corresponding to the detected text based on all the current field information and all the historical field information.

[0076] In one embodiment, the device further includes: The graph update module is used to update the text semantic version graph through the text to be detected after determining that the text to be detected is non-risk text, so as to obtain an updated text semantic graph.

[0077] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor being used to perform the text detection method described above.

[0078] Specific limitations regarding the computer equipment, processor, and their various units and modules can be found in the limitations of the text detection method described above, and will not be repeated here. Each module in the aforementioned processor can be implemented entirely or partially through software, hardware, or a combination thereof. Understandably, the processor includes a processor, memory, network interface, and database connected via a device bus. Each module of the processor can be embedded in hardware or independent of the processor, or stored in memory in software form, so that the processor can call and execute the operations corresponding to each module. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores operating devices, computer programs, and a database. The internal memory provides an environment for the operation of the operating devices and computer programs in the non-volatile storage media. The database stores the data used by the text detection method in the above embodiments. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a text detection method.

[0079] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described text detection method.

[0080] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0081] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0082] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A text detection method, characterized in that, include: Obtain the text to be detected, perform multimodal analysis on all the texts to be detected, and obtain the multimodal text fingerprint corresponding to each text to be detected; When the text verification results of all the texts to be detected indicate that all text verifications have passed, a text semantic version map corresponding to each of the texts to be detected is obtained; Text matching is performed on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version map to obtain the target detection text corresponding to each of the texts to be detected; The differences between each of the texts to be detected and the target text to be detected are determined as semantic differences corresponding to each of the texts to be detected. When the semantic differences are only key field information and the target text is in the state of completed, the text to be detected is determined to be a risky text.

2. The text detection method as described in claim 1, characterized in that, The text verification results of all the texts to be detected include: Extract the time information from all the texts to be detected to obtain the text time corresponding to each text to be detected. Perform timeline verification on all the text times to obtain the time verification results; Cross-text detection is performed on all the field information in the text to be detected to obtain cross-text detection results; When the time verification result indicates that all the text times have passed the verification, and the cross-text detection result indicates that the same field information is consistent, a text verification result indicating that all texts have passed the verification is obtained.

3. The text detection method as described in claim 1, characterized in that, The step of performing multimodal analysis on all the texts to be detected to obtain a multimodal text fingerprint corresponding to each text to be detected includes: Feature extraction is performed on all the texts to be detected to obtain the current image features corresponding to each text to be detected; Semantic extraction is performed on all the texts to be detected to obtain the current semantic features corresponding to each text to be detected; Information is extracted from all the texts to be detected to obtain the current content features corresponding to each text to be detected; Based on the current image features, current semantic features, and current content features corresponding to the same text to be detected, a multimodal text fingerprint corresponding to each text to be detected is determined.

4. The text detection method as described in claim 3, characterized in that, The step of performing text matching on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version graph to obtain the target detection text corresponding to each of the texts to be detected includes: The current image features are compared with the historical image features in each of the historical text fingerprints to obtain the image similarity corresponding to each of the historical image features; The current semantic feature is similar to the historical semantic features in each of the historical text fingerprints to obtain the semantic similarity corresponding to each of the historical semantic features; The current content features are compared with the historical content features in each of the historical text fingerprints to obtain the content similarity corresponding to each of the historical content features; Based on the image similarity, semantic similarity, and content similarity with the same text to be detected, target detection texts corresponding to the text to be detected are selected.

5. The text detection method as described in claim 4, characterized in that, Based on image similarity, semantic similarity, and content similarity with the same text to be detected, target detection texts corresponding to the text to be detected are selected, including: Filter out all image detection texts corresponding to image similarities greater than a preset image threshold from the historical detection texts corresponding to each of the historical text fingerprints; Filter out all semantically detected texts from all the image-detected texts that have a semantic similarity greater than a preset semantic threshold; From all the semantic detection texts, select the content detection texts with a content similarity greater than a preset content threshold, and determine the content detection texts as the target detection texts.

6. The text detection method as described in claim 1, characterized in that, The step of determining the difference between each of the texts to be detected and the target text to be detected as the semantic difference corresponding to each of the texts to be detected includes: Extract all information from the text to be detected to obtain the current field information; Extract all information from the target detection text to obtain historical field information; Based on all the current field information and all the historical field information, determine the semantic differences corresponding to the detected text.

7. The text detection method as described in claim 1, characterized in that, After determining the differences between each of the texts to be detected and the target text to be detected as semantic differences corresponding to each of the texts to be detected, the method further includes: After determining that the text to be detected is non-risk text, the text semantic version graph is updated using the text to be detected to obtain an updated text semantic graph.

8. A text detection device, characterized in that, include: The modal analysis module is used to acquire the text to be detected, perform multimodal analysis on all the texts to be detected, and obtain the multimodal text fingerprint corresponding to each text to be detected. The graph acquisition module is used to acquire the text semantic version graph corresponding to each of the texts to be detected when the text verification result of all the texts to be detected is characterized as passing all text verifications; The text matching module is used to perform text matching on the multimodal text fingerprint corresponding to the same text to be detected and the historical text fingerprint in the text semantic version graph to obtain the target detection text corresponding to each of the texts to be detected; The text recognition module is used to determine the differences between each of the texts to be detected and the target text to be detected as semantic differences corresponding to each of the texts to be detected, and to determine the text to be detected as risky text when the semantic differences are only key field information and the target text is in the state of completed.

9. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor being used to perform the text detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the text detection method as described in any one of claims 1 to 7.