Information identification method and device
Through the combination of entity extraction and semantic relationship analysis, sensitive information is accurately identified, and the problem of inaccurate identification in the prior art is solved, achieving more efficient privacy data protection.
Patent Information
- Application Number
- CN202510728892.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art is difficult to accurately identify sensitive information, resulting in the problem of privacy data leakage.
Through entity extraction and label matching, combined with semantic relationship analysis, predefined labels and semantic field effect matrix, the confidence and probability information of text information are calculated and the target information is determined.
It improves the accuracy and comprehensiveness of sensitive information identification, reduces misjudgment and misjudgment, and ensures effective protection of privacy data.
Smart Images

Figure CN120541229A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more specifically, to an information recognition method and device. Background Art
[0002] With the advent of the big data era, the issue of leaking sensitive information, such as private data, has also arisen. Accurately identifying sensitive information and taking appropriate protection measures (such as obfuscation) is a challenge that needs to be addressed. Summary of the Invention
[0003] The first aspect of the present disclosure provides an information recognition method, comprising: performing entity extraction and label matching on input data to obtain a first recognition result of multiple text information in the data, the first recognition result including first type information of the text information and a first confidence level corresponding to the first type information; parsing semantic relationships in the data to obtain a second recognition result of multiple text information in the data, the second recognition result including second type information of the text information and a second confidence level corresponding to the second type information; determining target information based on the first confidence level and the second confidence level
[0004] According to an embodiment of the present disclosure, entity extraction and label matching are performed on input data to obtain a first recognition result of multiple text information in the data, including: obtaining the initial text of the data; splitting the initial text into multiple text blocks; and recognizing the multiple text blocks based on predefined labels to obtain a specific type of text information and a first recognition result of the text information.
[0005] According to an embodiment of the present disclosure, it also includes: generating a first label based on the first type information and the first confidence of the text information, the first label is used to characterize the first recognition result of the text information; adding the first label to the text information in the initial text to obtain the first text.
[0006] According to an embodiment of the present disclosure, semantic relationships in data are parsed to obtain second recognition results of multiple text information in the data, including: constructing a semantic field effect matrix based on the semantic relationships in the target text; wherein the target text is the initial text or the first text of the data; parsing the target text based on the semantic field effect matrix to obtain a specific type of text information and a second recognition result of the text information.
[0007] According to an embodiment of the present disclosure, it also includes: generating a second label based on the second type information and the second confidence of the text information, the second label is used to characterize the second recognition result of the text information; adding the second label to the text information in the target text to obtain a second text.
[0008] According to an embodiment of the present disclosure, target information is determined based on a first confidence level and a second confidence level, including: calculating first probability information and second probability information of the text information based on a first label, a second label, and preset parameters, the first probability information and the second probability information are both used to indicate the possibility that the text information is target information, the first probability information is calculated based on the first confidence level and the preset parameters, and the second probability information is calculated based on the second confidence level and the preset parameters; calculating the target confidence level of the text information based on the first probability information and the second probability information; determining that the text information is target information when the target confidence level of the text information meets preset conditions; wherein the preset parameters are determined based on historical recognition results of text information of this type, and the preset parameters corresponding to different types of text information are the same or different.
[0009] According to an embodiment of the present disclosure, it also includes: when the target text is the first text, obtaining the first label and the second label of the text information from the second text; or when the target text is the initial text, obtaining the first label and the second label of the text information from the first text and the second text respectively.
[0010] According to an embodiment of the present disclosure, obtaining the initial text of data includes: when the input data is text data, determining the text data as the initial text; when the input data is other types of data, converting the data into text data and determining the obtained text data as the initial text.
[0011] According to an embodiment of the present disclosure, the initial text is split into multiple text blocks, including: preprocessing the initial text based on preset rules to obtain a preprocessed text containing an initial tag, wherein the initial tag is used to characterize the pre-recognition result of the text information; splitting the preprocessed text to obtain multiple text blocks; or splitting the content in the preprocessed text that does not contain the initial tag to obtain multiple text blocks.
[0012] The second aspect of the present disclosure provides an information identification device, including: a first identification module, used to perform entity extraction and label matching on input data, and obtain a first identification result of multiple text information in the data, the first identification result includes first type information of the text information and a first confidence level corresponding to the first type information; a second identification module, used to parse semantic relationships in the data, and obtain a second identification result of multiple text information in the data, the second identification includes second type information of the text information and a second confidence level corresponding to the second type information; a determination module, used to determine target information based on the first confidence level and the second confidence level. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0014] Figure 1 The following schematically illustrates an application scenario of the information identification method and device according to an embodiment of the present disclosure;
[0015] Figure 2 The following schematically shows a flow chart of an information identification method according to an embodiment of the present disclosure;
[0016] Figure 3 A flowchart schematically illustrates a method of performing entity extraction and label matching on input data to obtain a first recognition result of multiple text information in the data according to an embodiment of the present disclosure;
[0017] Figure 4 A schematic diagram illustrating a principle of obtaining an initial text of data according to an embodiment of the present disclosure is shown;
[0018] Figure 5 The following schematically shows a principle diagram of an information identification method according to an embodiment of the present disclosure;
[0019] Figure 6 A flowchart of the information identification method according to an embodiment of the present disclosure is schematically shown;
[0020] Figure 7 A flowchart schematically illustrates a method of parsing semantic relationships in data to obtain a second recognition result of multiple text information in the data according to an embodiment of the present disclosure;
[0021] Figure 8 A flowchart schematically illustrates a method of parsing semantic relationships in data to obtain a second recognition result of multiple text information in the data according to an embodiment of the present disclosure;
[0022] Figure 9 A schematic diagram illustrating a principle diagram of an information recognition method in which a target text is a first text according to an embodiment of the present disclosure is shown;
[0023] Figure 10 The following schematically illustrates a flow chart of determining target information based on a first confidence level and a second confidence level according to an embodiment of the present disclosure;
[0024] Figure 11 The following schematically shows a structural block diagram of an information identification device according to an embodiment of the present disclosure;
[0025] Figure 12 A block diagram of an electronic device suitable for implementing the information identification method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0027] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0029] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0030] The embodiments of the present disclosure provide an information identification method and device. Before introducing the technical solutions provided by the embodiments of the present disclosure, the related technologies involved in the present disclosure are first described.
[0031] Identifying sensitive information is the basis for preventing private data from being leaked. Only by accurately identifying sensitive information can it be accurately protected, thereby effectively preventing private data from being leaked.
[0032] Traditional information recognition methods usually rely on fixed rules or machine learning methods to identify target information in texts. They are difficult to cover all information types and have poor recognition accuracy for ambiguous information, which can easily lead to information omissions or misjudgments, resulting in inaccurate and incomplete recognition of sensitive information.
[0033] Figure 1 The application scenario diagram of the information identification method and device according to the embodiment of the present disclosure is schematically shown.
[0034] like Figure 1As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0035] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0036] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0037] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0038] It should be noted that the information identification method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the information identification method apparatus provided in the embodiments of the present disclosure can generally be set in the server 105. The information identification method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the information identification method apparatus provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0039] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0040] Figure 2 The flowchart of the information identification method according to the embodiment of the present disclosure is schematically shown.
[0041] like Figure 2 As shown, the information identification method of this embodiment includes operations S110 to S130.
[0042] In operation S110, entity extraction and label matching are performed on the input data to obtain a first recognition result of multiple text information in the data, where the first recognition result includes first type information of the text information and a first confidence level corresponding to the first type information.
[0043] In some embodiments, entity extraction is performed on the input data to obtain multiple entity information with independent semantics and recognition results of each entity information. The recognition result can be the type of the entity information, for example, the type of the entity information is a person's name, a place name, etc.
[0044] Label matching is performed on multiple entity information, and text information is determined from the multiple entity information based on the label matching results. For example, the entity information with successful label matching is determined as text information, the first type information of the text information is the recognition result of the entity information, and the first confidence level is the confidence level of the entity information recognition result.
[0045] For example, entity information identification results can be matched using a predefined tagging system (or knowledge base). The predefined tagging system is related to the information identification scenario. For example, if the information identification scenario is sensitive information, the predefined tags may include multiple types of sensitive information (such as phone numbers, ID account numbers, bank card numbers, etc.). By performing tag matching, it is determined whether the entity information in the data contains sensitive information, namely text information.
[0046] In some embodiments, operation S110 may be performed by a first model, where initial text to be recognized is input into the first model. The first model analyzes the initial text, identifies entities in the initial text and the type of each entity, matches the entity type with a preset label, determines entities with the same entity type as the preset label as text information, and determines a first recognition result for the text information based on the entity recognition result. The first recognition result includes first type information of the text information and a first confidence level representing the credibility of the first type information.
[0047] Continuing with the sensitive information identification scenario as an example, the first model identifies whether there is text information of the same type as sensitive information in the input data based on a predefined label system, and determines the specific type of the text information (i.e., the first type of information) and the type of information.
[0048] For example, if the input data is "Xiaoming's mobile phone number is 123***123," the entity recognition result is: Xiaoming (name), 123***123 (number). Based on the predefined tags, the matching is performed and the number is determined to be a predefined tag. Therefore, "123***123" is determined to be text information. The first type of information in this text information is "number," and the first confidence level is 0.85.
[0049] In operation S120, the semantic relationship in the data is parsed to obtain a second recognition result of the plurality of text information in the data, where the second recognition includes second type information of the text information and a second confidence level corresponding to the second type information.
[0050] In some embodiments, certain information types may appear in specific contexts, and the semantic relationships between text in the data can be analyzed to trigger the recognition of specific types of information. For example, the numbers appearing near the content "Please contact your account manager" may be a phone number.
[0051] In some embodiments, operation S120 can be performed by the second model, and the input data is input into the second model. The second model analyzes the global context of the text, determines the role and meaning of the information in the text, captures implicit information in the context, analyzes the type of information implied by the information in the current structure, and obtains a second recognition result of the text information. The second recognition result includes the second type of information of the text information and a second confidence level used to characterize the credibility of the second type of information.
[0052] In operation S130 , target information is determined based on the first confidence level and the second confidence level.
[0053] In some embodiments, the target confidence of the text information is determined by calculating the first confidence level and the second confidence level of the text information. For example, the target confidence level can be calculated by weighted averaging, multiplication fusion, etc. If the target confidence level of the text information meets a preset condition, the text information is determined as the target information.
[0054] The information recognition method provided by this disclosure identifies text information from two dimensions: entity recognition and semantic relationship recognition. A target confidence is calculated based on the confidence of each recognition result, and the target information is determined based on the target confidence. The target confidence is calculated by combining the confidences of two independent recognition results, effectively reducing the bias or uncertainty of a single recognition method. Determining the target information based on the target confidence improves the accuracy of text recognition results, thereby improving the precision of target information recognition and preventing misidentification of the target information.
[0055] Figure 3 The flowchart schematically shows a method of performing entity extraction and label matching on input data to obtain a first recognition result of multiple text information in the data according to an embodiment of the present disclosure.
[0056] like Figure 3 As shown, in Figure 2 Based on the information identification method shown in FIG, the information identification method of this embodiment may further include operations S210 to S230. It should be noted that other implementation details of the information identification method of this embodiment can be found in FIG. Figure 2 The embodiment of the information identification method shown will not be described in detail here.
[0057] In operation S210 , an initial text of data is obtained.
[0058] In some embodiments, the data can be structured or unstructured. Structured data can be text fields extracted from a database or a webpage interface, while unstructured data can be, for example, documents (e.g., PDFs), images, or audio. For each type of data, the initial text of the data is obtained using corresponding methods.
[0059] In operation S220 , the initial text is split into a plurality of text blocks.
[0060] In some embodiments, the initial text may be divided into smaller units (ie, text blocks) according to a certain splitting strategy to facilitate subsequent processing.
[0061] Exemplary splitting strategies may include: splitting by line break or blank line, splitting by paragraph, detecting sentence boundaries and splitting by sentence, or splitting based on a fixed-length window, to obtain multiple text blocks. Multiple text blocks can be processed in parallel to improve recognition efficiency.
[0062] In some embodiments, each text block may be an entity or a text unit including multiple entities.
[0063] In operation S230 , the plurality of text blocks are recognized based on predefined tags to obtain text information of a specific type and a first recognition result of the text information.
[0064] In some embodiments, multiple text blocks can be identified based on predefined tags to obtain specific types of text information. Predefined tags may include "name, location, organization, date, phone number," etc. Entity recognition technology is used to predict multiple entities in each text block to obtain specific types of text information and a first recognition result of the text information.
[0065] Exemplarily, the specific type of text information may be text information of the same type as sensitive information. The first recognition result for the text information includes the first type information of the text information and a confidence level. For example, sensitive information types include "bank card number" and "ID card number." If the first type information of a text message is "bank card number," the text message is determined to be of the specific type.
[0066] In some embodiments, the confidence level can be determined by rule matching. For example, if the text information completely matches the preset label, the confidence level is 1.0, and if it partially matches, the confidence level is 0.8. It can also be determined based on model prediction. For example, the softmax probability mean of the entity's corresponding token is used as the first confidence level of the text information.
[0067] Splitting the initial text into blocks and recognizing them using predefined tags can effectively improve the accuracy and consistency of recognition results, making them more relevant to real-world applications. Recognizing in blocks improves recognition precision and reduces interference from irrelevant information or noise on entity recognition.
[0068] In some embodiments, when recognizing multiple text blocks, cross-block entity merging can be established. For example, if "Xiao Ming" is recognized as a person's name in text block 1, the existing recognition results for "Xiao Ming" in other text blocks can be directly inherited. This avoids duplicate recognition and improves recognition efficiency.
[0069] According to an embodiment of the present disclosure, the input data type is not limited, and the initial text is obtained from the input data in a corresponding manner.
[0070] In some embodiments, the type of input data can be determined by, for example, checking the file extension, data format, etc. of the input data, and the data can be divided into two categories, text data and other data, according to the data type.
[0071] Figure 4The schematic diagram of the principle of obtaining the initial text of data in one embodiment of the present disclosure is shown. It should be noted that other implementation details of the information identification method of this embodiment can be found in Figure 2 The embodiment of the information identification method shown will not be described in detail here.
[0072] like Figure 4 As shown, when the input data is text data, the text data is directly determined as the initial text. When the input data is other types of data, the data is converted into text data, and the obtained text data is determined as the initial text.
[0073] For example, other data may include images, audio, video, etc. Different types of other data are converted into text data using corresponding conversion methods. For example, for image data, optical character recognition technology can be used to extract text from the image. For audio data, speech recognition technology can be used to convert the audio into text.
[0074] The disclosed embodiments integrate different types of data into a unified text processing flow through data conversion, thereby performing information recognition on the text, which can effectively improve the compatibility of the information recognition method, meet different application scenarios and data input types, and improve the flexibility and scalability of information recognition.
[0075] Figure 5 The schematic diagram shows the principle diagram of the information identification method of an embodiment of the present disclosure. It should be noted that other implementation details of the information identification method of this embodiment can be found in Figure 2 The embodiment of the information identification method shown will not be described in detail here.
[0076] like Figure 5 As shown, after obtaining the initial text of the data, the initial text can be preprocessed based on preset rules to obtain a preprocessed text containing an initial label, wherein the initial label is used to represent the pre-recognition result of the text information.
[0077] The initial text can be split based on the initial tag. For example, the text containing the initial tag is considered as a text block, and the content of the preprocessed text that does not contain the initial tag is split to obtain multiple text blocks. Alternatively, the preprocessed text can be split to obtain multiple text blocks.
[0078] In some embodiments, the initial text is pre-processed based on preset rules, which may include, for example, quickly identifying target information in a standard format, such as an ID card number, bank card number, or phone number, based on regular expressions. Pre-processing allows for rapid identification of target information in a standard format, providing a reference for subsequent identification.
[0079] The disclosed embodiments progressively optimize recognition results through a multi-layered recognition mechanism, significantly enhancing the robustness of information recognition, effectively reducing misidentifications and missed recognitions, and improving the accuracy of the final recognized information. Through regular expression recognition, named entity recognition, and semantic recognition, target information is cross-validated. If the recognition results at one level are incorrect or omitted, they can be effectively corrected through subsequent levels of recognition to ensure the accuracy and comprehensiveness of the final information recognition results.
[0080] Figure 6 The flowchart of the information identification method according to the embodiment of the present disclosure is schematically shown.
[0081] like Figure 6 As shown, in Figure 2 Based on the information identification method shown in FIG, the information identification method of this embodiment may further include operations S240 to S250. It should be noted that other implementation details of the information identification method of this embodiment can be found in FIG. Figure 1 The embodiment of the information identification method shown will not be described in detail here.
[0082] In operation S240 , a first label is generated based on the first type information and the first confidence level of the text information, where the first label is used to characterize a first recognition result of the text information.
[0083] In operation S250 , a first tag is added to the text information in the initial text to obtain a first text.
[0084] In some embodiments, generating a first label based on the first type information and the first confidence level of the text information and adding the first label to the text information in the initial text includes: finding the text information to be labeled in the initial text, adding the first label near the text information, and obtaining a first text containing the first label.
[0085] Combining type information and confidence into labels can achieve structured representation of recognition results and traceability of the first recognition results of text information. By obtaining the first label in the first text, the specific type of text information in the initial text and the first recognition result of each text information can be obtained.
[0086] Figure 7 The flowchart schematically illustrates a method for parsing semantic relationships in data to obtain a second recognition result of multiple text information in the data according to an embodiment of the present disclosure.
[0087] like Figure 7 As shown, in Figure 2 Based on the information identification method shown in FIG, the information identification method of this embodiment may further include operations S121 to S122. It should be noted that other implementation details of the information identification method of this embodiment can be found in FIG. Figure 2The embodiments of the information recognition method shown are not elaborated here.
[0088] In operation S121, a semantic field effect matrix is constructed based on the semantic relationships in the target text; wherein, the target text is the initial text or the first text of the data.
[0089] In some embodiments, a matrix reflecting the context relationship is constructed by quantifying the semantic associations between words in the text.
[0090] Exemplarily, the text is segmented into multiple word segments, and the词性 of each word segment is labeled (such as verb, noun, adjective, etc.). Words without actual semantics (such as "的", "是", etc.) are removed by labeling the词性. And a semantic field effect matrix is constructed based on the obtained word segments. The semantic field effect matrix is a matrix used to represent the semantic relationships between words or concepts in the text, and shows the association strength between different words in a specific context in a quantified manner to help identify the implicit structures and relationships in the text.
[0091] Exemplarily, the rows and columns of the semantic field effect matrix usually represent words and phrases in the text, and the matrix values represent the semantic association strength between the corresponding row and column elements. The implicit semantic relationships in the text can be determined through the semantic field effect matrix.
[0092] Exemplarily, before constructing the semantic field effect matrix, the semantic scenario of the text can be determined according to the context environment or theme category of the text, and in this semantic scenario, the specific semantic relationships in the text are analyzed. For example, if the text contains verbs such as "咨询" and "办理", it is determined that the semantic scenario type corresponding to this text is service request, and a semantic field effect matrix is constructed in this semantic scenario to record the association relationships of the keywords in the scenario.
[0093] In operation S122, the target text is parsed based on the semantic field effect matrix to obtain text information of a specific type and a second recognition result of the text information.
[0094] In some embodiments, the association degrees between the words in the target text are determined based on the semantic field effect matrix, and text information of a specific type and a second recognition result of the text information are determined according to the association degrees.
[0095] It should be noted that the Chinese character "词性" in the original text seems to be an incorrect expression. It might be intended to be something like "词性标注" or other more accurate terms. But based on the translation rules, it is translated as "词性" as it is. You may need to double-check the accuracy of this part in the original context.For example, the target text is: "To handle service A, please call 123**888." Based on the terms "handle" and "consult," the semantic scenario is determined to be a "service request" type. A service request field effect matrix is constructed. In this effect matrix, the correlation between "consult" and "123**888" is 0.92. This determines that the second type of information in the text message "123**888" is a phone number, and the second confidence level corresponding to this second type of information is calculated based on the historical recognition accuracy results for this scenario.
[0096] The disclosed embodiments construct a semantic field effect matrix to identify information in target text, effectively improving the accuracy of information recognition. By quantifying the dynamic relationships between words through the matrix, semantic understanding is enhanced. Furthermore, through matrix path analysis, indirectly related entities are identified, effectively enhancing the accuracy and flexibility of information recognition. For information with ambiguous meaning or potentially multiple meanings, the semantic field effect matrix can be constructed and combined with contextual meaning to achieve accurate recognition of the information.
[0097] Figure 8 The flowchart schematically illustrates a method for parsing semantic relationships in data to obtain a second recognition result of multiple text information in the data according to an embodiment of the present disclosure.
[0098] like Figure 8 As shown, in Figure 2 Based on the information identification method shown in FIG, this embodiment analyzes the semantic relationship in the data to obtain the second recognition result of multiple text information in the data, which may further include operations S123 to S124. It should be noted that other implementation details of the information identification method of this embodiment can be found in Figure 2 The embodiment of the information identification method shown will not be described in detail here.
[0099] In operation S123, a second label is generated based on the second type information and the second confidence level of the text information, where the second label is used to represent a second recognition result of the text information.
[0100] In operation S124 , a second tag is added to the text information in the target text to obtain a second text.
[0101] In some embodiments, a second label is generated based on the second type information and the second confidence of the text information and the second label is added to the text information in the target text, including: finding the text information to be labeled in the target text, attaching the second label near the text information, and obtaining a second text containing the second label.
[0102] Figure 9The schematic diagram shows the principle of the information recognition method according to an embodiment of the present disclosure, in which the target text is the first text. It should be noted that other implementation details of the information recognition method of this embodiment can be found in Figure 2 The embodiment of the information identification method shown will not be described in detail here.
[0103] Exemplarily, if the target text is the first text, the second text includes text information with the first tag and text information with the second tag. The same text information may include the first tag and the second tag at the same time.
[0104] In some embodiments, when the target text is the first text, the target text can be dynamically adjusted based on the actual application scenario and information recognition requirements. Figure 4 The information identification method shown.
[0105] For example, for scenarios with high requirements for comprehensive recognition, it is necessary to ensure that all potential text information in the first text is recognized, with particular attention paid to the text portion of the first text that does not contain the first label. For example, text fragments not covered by the first label can be extracted, and a deep semantic analysis can be performed on the unlabeled text to obtain a second recognition result of the text information. The second label generated based on the second recognition result is added to the first text to obtain a second text. While improving the comprehensiveness of recognition, the waste of computing resources is reduced and recognition efficiency is improved.
[0106] For scenarios with high recognition accuracy requirements, it is necessary to ensure that the recognition results of all text segments are accurate. All text segments in the first text need to be scanned to ensure the reliability and accuracy of the recognition results.
[0107] Comprehensively adjusting the recognition strategy according to the needs of different scenarios can effectively improve the flexibility and dynamic adaptability of recognition and meet the information recognition needs of different scenarios.
[0108] Figure 10 The flowchart of determining target information based on the first confidence level and the second confidence level according to an embodiment of the present disclosure is schematically shown.
[0109] like Figure 10 As shown, in Figure 2 Based on the information identification method shown in FIG, the method of determining target information based on the first confidence level and the second confidence level in this embodiment may further include operations S131 to S133. It should be noted that other implementation details of determining target information based on the first confidence level and the second confidence level in this embodiment can be found in FIG. Figure 2 The embodiment of the information identification method shown will not be described in detail here.
[0110] In operation S131, first probability information and second probability information of the text information are calculated based on the first label, the second label, and the preset parameters. The first probability information and the second probability information are both used to indicate the possibility that the text information is the target information. The first probability information is calculated based on the first confidence level and the preset parameters, and the second probability information is calculated based on the second confidence level and the preset parameters.
[0111] In some embodiments, the preset parameters can be determined based on historical recognition results for that type of text information, and different types of text information may have the same or different preset parameters. The preset parameter P(H) can reflect the recognition reliability of the first recognition model and the second recognition model for a specific category.
[0112] The preset parameters can be dynamically adjusted based on the historical recognition accuracy. For example, in the historical information recognition process, if the first recognition model and the second recognition model have a high recognition consistency for this type of text information, a higher preset parameter can be assigned, for example, the value of the preset parameter can be set to 0.8.
[0113] In some embodiments, the method for obtaining the first and second labels is associated with the text recognition method. If the target text is the first text, the first and second labels of the text information can be directly obtained from the second text. If the target text is the initial text, the first and second labels of the text information are obtained from the first and second texts, respectively.
[0114] In some embodiments, calculating the first probability information based on the first confidence level in the first tag and preset parameters includes:
[0115] Take the first confidence Cm as the first likelihood probability , determine the second likelihood probability based on the first confidence , . Among them, the first likelihood probability represents the correct probability of the first recognition result, and the second likelihood probability represents the incorrect probability of the first recognition result.
[0116] The evidence probability P(E) is calculated based on the preset parameters, the first likelihood probability and the second likelihood probability. The evidence probability represents the total probability of the predicted result occurring when the model prediction result (such as the first recognition result) is observed. .in, .
[0117] The first probability information is calculated based on the evidence probability, the first probability and the preset parameters. The first probability information indicates the possibility that the text information is the target information after observing the first recognition result. Compared with the first confidence level, the first probability information is calculated by introducing prior knowledge (i.e., preset parameters), which improves the comprehensiveness and accuracy of the judgment. .
[0118] The process of calculating the second probability information based on the second confidence level and the preset parameters is similar to the process of calculating the first probability information, and will not be repeated here. By combining the confidence level of the first recognition result with prior knowledge of the business scenario, a more reliable classification result can be obtained.
[0119] In operation S132, target confidence of the text information is calculated based on the first probability information and the second probability information.
[0120] In some embodiments, the target confidence of text information can be calculated by multiplying the first probability information and the second probability information. By combining the probability information of two independent recognition results, the target confidence can reduce the bias or uncertainty of a single recognition method, effectively improving the accuracy of information recognition. The target confidence can be viewed as the joint probability of multiple recognition results for the target information. The higher the joint probability, the greater the consistency among the multiple information recognition methods in determining that the text is the target information, thereby improving the accuracy of target information recognition.
[0121] In operation S133 , if the target confidence of the text information satisfies a preset condition, the text information is determined to be target information.
[0122] In some embodiments, the preset condition may be, for example, that the target confidence of the text information is greater than a preset threshold x. When the target confidence of the text information is greater than the preset threshold x, the text information is determined to be target information.
[0123] The information recognition method disclosed in the present invention introduces preset parameters to determine the target confidence. The target confidence can effectively reflect the historical recognition performance of a specific category and improve the robustness of the recognition of target information of different categories. The target confidence integrates independent recognition results under multiple different dimensions. Compared with the confidence of independent recognition results, the target confidence can reduce the deviation or uncertainty of a single recognition result, effectively improve the accuracy of information recognition, and reduce the misjudgment rate of information recognition.
[0124] Based on the above information identification method, the present disclosure also provides an information identification device. Figure 11 The device is described in detail.
[0125] Figure 11The structural block diagram of an information identification device according to an embodiment of the present disclosure is schematically shown.
[0126] like Figure 11 As shown, the information identification device 1000 of this embodiment includes a first identification module 1010 , a second identification module 1020 and a determination module 1030 .
[0127] First recognition module 1010 is configured to perform entity extraction and label matching on the input data to obtain a first recognition result for multiple text information in the data. The first recognition result includes first type information of the text information and a first confidence level corresponding to the first type information. In one embodiment, first recognition module 1010 can be configured to perform operation S110 described above and will not be further described here.
[0128] The second recognition module 1020 is configured to analyze semantic relationships within the data and obtain second recognition results for multiple text messages within the data, wherein the second recognition includes second type information of the text messages and a second confidence level corresponding to the second type information. In one embodiment, the second recognition module 1020 may be configured to perform operation S120 described above, which will not be further described herein.
[0129] The determination module 1030 is configured to determine the target information based on the first confidence level and the second confidence level. In one embodiment, the determination module 1030 may be configured to perform the operation S130 described above, which will not be described in detail herein.
[0130] According to embodiments of the present disclosure, any multiple modules among the first identification module 1010, the second identification module 1020, and the determination module 1030 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present disclosure, at least one of the first identification module 1010, the second identification module 1020, and the determination module 1030 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or implemented in any one of software, hardware, and firmware, or any suitable combination of these. Alternatively, at least one of the first identification module 1010, the second identification module 1020, and the determination module 1030 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0131] Figure 12 A block diagram of an electronic device suitable for implementing the information identification method according to an embodiment of the present disclosure is schematically shown.
[0132] like Figure 12 As shown, the electronic device 1100 according to an embodiment of the present disclosure includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage portion 1108 into a random access memory (RAM) 1103. The processor 1101 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1101 may also include onboard memory for caching purposes. The processor 1101 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0133] Various programs and data required for the operation of the electronic device 1100 are stored in the RAM 1103. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. The processor 1101 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 1102 and / or the RAM 1103. It should be noted that the programs may also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0134] According to an embodiment of the present disclosure, electronic device 1100 may further include an input / output (I / O) interface 1105, which is also connected to bus 1104. Electronic device 1100 may also include one or more of the following components connected to I / O interface 1105: an input section 1106 including a keyboard, mouse, etc.; an output section 1107 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 1108 including a hard disk; and a communication section 1109 including a network interface card such as a LAN card or modem. Communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to I / O interface 1105 as needed. Removable media 1111, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 1110 as needed, so that computer programs read from the removable media can be installed into storage section 1108 as needed.
[0135] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0136] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 1102 and / or RAM 1103 described above, and / or one or more memories other than ROM 1102 and RAM 1103.
[0137] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the information identification method provided by the embodiments of the present disclosure.
[0138] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the computer program is executed by the processor 1101. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0139] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 1109, and / or installed from removable media 1111. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0140] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1109 and / or installed from the removable medium 1111. When the computer program is executed by the processor 1101, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0141] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0143] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.
[0144] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An information identification method, comprising: Performing entity extraction and label matching on the input data to obtain a first recognition result of multiple text information in the data, the first recognition result including first type information of the text information and a first confidence level corresponding to the first type information; parsing semantic relationships in the data to obtain second recognition results of the plurality of text information in the data, the second recognition results including second type information of the text information and a second confidence level corresponding to the second type information; Target information is determined based on the first confidence level and the second confidence level.
2. The information identification method according to claim 1, wherein: The performing entity extraction and label matching on the input data to obtain a first recognition result of multiple text information in the data includes: obtaining an initial text of the data; Splitting the initial text into multiple text blocks; The multiple text blocks are recognized based on predefined tags to obtain text information of a specific type and a first recognition result of the text information.
3. The information identification method according to claim 2, further comprising: generating a first label based on the first type information and the first confidence level of the text information, wherein the first label is used to represent a first recognition result of the text information; A first tag is added to the text information in the initial text to obtain a first text.
4. The information recognition method according to claim 3, wherein the step of analyzing the semantic relationship in the data to obtain a second recognition result of the plurality of text information in the data comprises: Constructing a semantic field effect matrix based on semantic relations in a target text; wherein the target text is an initial text or a first text of the data; The target text is parsed based on the semantic field effect matrix to obtain text information of a specific type and a second recognition result of the text information.
5. The information identification method according to claim 4, further comprising: generating a second label based on the second type information and the second confidence level of the text information, wherein the second label is used to represent a second recognition result of the text information; A second tag is added to the text information in the target text to obtain a second text.
6. The information identification method according to claim 1, wherein: The determining target information based on the first confidence level and the second confidence level includes: Calculating first probability information and second probability information of the text information based on the first label, the second label, and the preset parameters, wherein the first probability information and the second probability information are both used to indicate the likelihood that the text information is the target information, the first probability information being calculated based on the first confidence level and the preset parameters, and the second probability information being calculated based on the second confidence level and the preset parameters; Calculating a target confidence of the text information according to the first probability information and the second probability information; If the target confidence of the text information meets a preset condition, determining that the text information is target information; The preset parameters are determined based on historical recognition results of text information of this type, and the preset parameters corresponding to different types of text information may be the same or different.
7. The information identification method according to claim 6, further comprising: When the target text is a first text, obtaining a first label and a second label of the text information from a second text; or In a case where the target text is an initial text, a first label and a second label of the text information are acquired from the first text and the second text respectively.
8. The information recognition method according to claim 2, wherein obtaining the initial text of the data comprises: In the case where the input data is text data, determining the text data as initial text; In the case that the input data is other types of data, the data is converted into text data, and the obtained text data is determined as the initial text.
9. The information recognition method according to claim 2, wherein splitting the initial text into a plurality of text blocks comprises: Preprocessing the initial text based on preset rules to obtain a preprocessed text containing initial tags, wherein the initial tags are used to represent pre-recognition results of text information; Splitting the preprocessed text to obtain multiple text blocks; or The content of the preprocessed text that does not contain the initial tag is split to obtain multiple text blocks.
10. An information recognition device, comprising: a first recognition module, configured to perform entity extraction and label matching on input data to obtain a first recognition result of a plurality of text information in the data, wherein the first recognition result includes first type information of the text information and a first confidence level corresponding to the first type information; a second recognition module, configured to parse the semantic relationship in the data to obtain a second recognition result of the plurality of text information in the data, wherein the second recognition includes second type information of the text information and a second confidence level corresponding to the second type information; A determination module is used to determine target information based on the first confidence level and the second confidence level.