Data detection method, electronic device, and storage medium
By constructing a key information model and using a rule engine for analysis, data compliance risks are automatically identified, solving the problem of low efficiency in manual review and achieving efficient compliance review of data usage applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WEBANK (CHINA)
- Filing Date
- 2021-06-25
- Publication Date
- 2026-04-28
AI Technical Summary
Manually reviewing data using application forms is inefficient and prone to errors, making it difficult to meet the high requirements for data security and privacy protection in the fintech sector.
By setting up a thesaurus to extract word segments of internal proprietary terms, a key information model is constructed, and a rule engine is used to analyze user agreements to generate a set of risk warning rules, automatically identifying data compliance risks.
This improved the accuracy and efficiency of data compliance audits, ensuring that data usage application forms comply with the company's internal terminology and agreement requirements.
Smart Images

Figure CN113326699B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a data detection method, electronic device, and storage medium. Background Technology
[0002] With the development of computer technology, more and more technologies (such as big data, artificial intelligence, and blockchain) are being applied in the financial sector, and the traditional financial industry is gradually transforming into fintech. However, due to the security and real-time requirements of the financial industry, fintech also places higher demands on technology. In the fintech field, strict data classification and control standards and data authorization procedures have been established to achieve data security and privacy protection.
[0003] In related technologies, when staff need to use or view certain data for business purposes, they must submit a data usage application form within the approval system to request data access permissions. The relevant reviewer then uses their experience to determine whether the requested data access permissions violate signed confidentiality agreements or privacy protections, thereby assessing whether the data usage application poses any data compliance risks. However, manually checking the data compliance of data usage applications is not only inefficient but also prone to errors and omissions. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a data detection method, electronic device, and storage medium to solve the technical problems of low efficiency and error-proneness in the manual review of data use application forms in related technologies.
[0005] To achieve the above objectives, the technical solution of the present invention is implemented as follows:
[0006] This invention provides a data detection method, including:
[0007] Receive a first request; the first request includes a first form and first information; the first request is used to request access to the data in the first form in the first system;
[0008] A first word segmentation set is extracted from the first information; the words in the first word segmentation set are words that exist in a set word library;
[0009] Based on the field types and corresponding field values in the first form, and the field types to which the words in the first word segmentation set belong, a key information model corresponding to the first request is constructed.
[0010] Using a rule engine, the user agreement corresponding to the word segmentation representing the product name in the key information model is processed to obtain a set of risk warning rules;
[0011] If the key information model matches any of the risk warning rules in the risk warning rule set, a first warning message is output, which indicates that the first request has a data compliance risk.
[0012] This invention also provides an electronic device, comprising:
[0013] A receiving unit is configured to receive a first request; the first request includes a first form and first information; the first request is used to request access to data in the first form in a first system; the first information represents text information in the first request other than the first form.
[0014] The extraction unit is used to extract a first word segmentation set from the first information; the words in the first word segmentation set are words that exist in a set word library;
[0015] The construction unit is used to construct the key information model corresponding to the first request based on the field types and corresponding field values in the first form and the field types to which the words in the first word segmentation set belong.
[0016] The generation unit is used to process the user agreement corresponding to the word segmentation representing the product name in the key information model using the rule engine to obtain a set of risk warning rules.
[0017] The prompting unit is used to output a first prompting message when the key information model hits any of the risk prompting rules in the risk prompting rule set. The first prompting message indicates that the first request has a data compliance risk.
[0018] This invention also provides an electronic device, including: a processor and a memory for storing a computer program capable of running on the processor.
[0019] The processor is used to execute the steps of the above-described data detection method when running the computer program.
[0020] This invention also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described data detection method.
[0021] In this embodiment of the invention, since the set lexicon includes proper nouns within the enterprise, a first word segmentation set containing proper nouns can be extracted from the received first request based on the set lexicon. A key information model corresponding to the first request is constructed based on the field types and corresponding field values in the first form, and the field types to which the words in the first word segmentation set belong. A rule engine is used to analyze the user agreement related to the first request to obtain a risk warning rule set. If the key information model matches any risk warning rule in the risk warning rule set, a first warning message is output, indicating that the first request has a data compliance risk. The above solution, based on the set lexicon and natural language processing technology, processes the first request, accurately extracting the first word segmentation set from the first request, thereby constructing a key information model. Based on the word segmentation representing the product name in the key information model, the user agreement is obtained, and all user agreements related to the first request can be obtained, thus improving the accuracy and comprehensiveness of the rules in the risk warning rule set. Because a key information model and a risk warning rule set are constructed, data compliance risks related to the first request can be accurately identified. Attached Figure Description
[0022] Figure 1 A schematic diagram illustrating the implementation process of the data detection method provided in this embodiment of the invention;
[0023] Figure 2 A schematic diagram illustrating the implementation process of extracting the first word segmentation set in the data detection method provided in this embodiment of the invention;
[0024] Figure 3 A schematic diagram illustrating the implementation process of extracting the first word segmentation subset in the data detection method provided in this embodiment of the invention;
[0025] Figure 4 A schematic diagram illustrating the implementation process of extracting the second word segmentation subset in the data detection method provided in this embodiment of the invention;
[0026] Figure 5 A schematic diagram of a rule generator provided in an embodiment of the present invention;
[0027] Figure 6 A schematic diagram illustrating the implementation flow of a data detection method provided in another embodiment of the present invention;
[0028] Figure 7 This is a schematic diagram illustrating the implementation process of determining the second word segmentation set in the data detection method provided in this embodiment of the invention;
[0029] Figure 8 This is a schematic diagram illustrating the implementation process of calculating the relevance of a substring in the data detection method provided in this embodiment of the invention.
[0030] Figure 9 This is a schematic diagram illustrating the implementation process of determining the value of the joint probability factor in the data detection method provided in this embodiment of the invention.
[0031] Figure 10 A schematic diagram of the accuracy curve of the candidate word set provided in an embodiment of the present invention;
[0032] Figure 11 A schematic diagram illustrating the implementation flow of the data detection method provided in an application embodiment of the present invention;
[0033] Figure 12 This is a schematic diagram of the data processing structure provided in an embodiment of the present invention;
[0034] Figure 13 This is a schematic diagram of the hardware composition structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0035] Although Natural Language Processing (NLP) technology is widely used in fields such as speech recognition, intelligent translation, public opinion monitoring, text classification, and text recognition, NLP is based on corpora. Due to the large number of proper nouns within enterprises, it is difficult to discover new words and eliminate slang based on existing corpora. Therefore, in related technologies, manually reviewing data use applications is not only slow, but also because approvers can only remember a limited amount of confidentiality agreements or privacy protection agreements. When reviewing data use applications based on accumulated experience, errors and omissions are prone to occur.
[0036] Based on this, in various embodiments of the present invention, since the lexicon includes proprietary terms within the enterprise, a first word segmentation set containing proprietary terms can be extracted from the received first request based on the lexicon; a key information model corresponding to the first request is constructed based on the field types and corresponding field values in the first form, and the field types to which the words in the first word segmentation set belong; a risk warning rule set is obtained by analyzing the user agreement related to the first request using a rule engine; and a first warning message is output when the key information model matches any risk warning rule in the risk warning rule set, indicating that the first request has a data compliance risk. In this solution, since the lexicon includes proprietary terms within the enterprise, potential data compliance risks can be accurately identified.
[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0038] Figure 1 This is a schematic diagram illustrating the implementation flow of the data detection method provided in an embodiment of the present invention, wherein the execution subject of the flow is an electronic device such as a terminal or server. Figure 1 The data detection methods shown include:
[0039] Step 101: Receive a first request; the first request includes a first form and first information; the first request is used to request access to the data in the first form in the first system.
[0040] Here, when a user needs to request data access permissions, they fill in or select the relevant data in the interactive interface displayed on the terminal, triggering the terminal to send the first request to the electronic device.
[0041] An electronic device receives a first request sent by a terminal. The first request includes a first form and first information. The first information represents the text information in the first request, excluding the first form. The first form includes at least one field and its corresponding value. The first information includes a title and background description information about the data usage requirement. The first system generally refers to systems within an enterprise; for example, the first system could be a document issuance system, a process system, or a system for managing data usage permissions. Systems for managing data usage permissions include data hierarchical control systems and / or data authorization process systems.
[0042] In practical applications, the first form includes at least a first field, a second field, a third field, a fourth field, and a fifth field. The first field is used to write the data used in the application, the second field is used to write the name of the product associated with the data in the first field, the third field is used to write the data source environment corresponding to the data in the first field, the fourth field is used to write the target customer of the data in the first field, and the fifth field is used to write the usage scenario of the data in the first field.
[0043] Step 102: Extract the first word segmentation set from the first information; the words in the first word segmentation set are words that exist in the set word library.
[0044] Here, the electronic device performs word segmentation on the first information to obtain the word segmentation results, selects the words that exist in the set word library from the word segmentation results, and obtains the first word segmentation set.
[0045] To eliminate ambiguity and improve the accuracy of word segmentation in the first word segmentation set, in some embodiments, such as Figure 2 As shown, the step of extracting the first word segmentation set from the first information includes:
[0046] Step 201: Based on the position of the set punctuation marks, the first information is divided into multiple first sentences.
[0047] Considering that the data in the first form is structured data and there is no need to perform word segmentation on the data in the first form, the electronic device identifies the set punctuation marks included in the first information and segments the first information according to the position of the identified set punctuation marks to obtain multiple first sentences.
[0048] The punctuation marks include periods, question marks, exclamation marks, commas, pause marks, quotation marks, and semicolons.
[0049] Since punctuation marks are usually not included in the title, the entire title included in the first message is identified as a first statement.
[0050] Step 202: Based on the set maximum word segmentation length, extract the first word segmentation subset from each first sentence in order from left to right, and extract the second word segmentation subset from each first sentence in order from right to left.
[0051] Here, for each first sentence, the electronic device extracts a first subset of words from the first sentence, starting from the first character and proceeding from left to right, based on a set maximum word segmentation length. Then, based on the set maximum word segmentation length, it extracts a second subset of words from the first sentence, starting from the last character and proceeding from right to left. The number of characters in the first string extracted from the first sentence corresponds to the set maximum word segmentation length. When the set maximum word segmentation length is 5, the maximum number of characters in the extracted string is 5.
[0052] It should be noted that during the process of extracting the second word subset from the first statement, starting from the last character and proceeding from right to left, the order of the characters remains unchanged. In other words, the order of the characters in the string extracted from the first statement is consistent with their corresponding order in the first statement.
[0053] Since the processing procedure is the same for each first statement, the following uses a single first statement as an example to illustrate the process of extracting the first and second word segments from the first statement:
[0054] The process of extracting the first segmented subset from each first sentence based on the set maximum segmentation length and in order from left to right is as follows:
[0055] Here, the electronic device is based on the principle of the forward maximum matching algorithm, following the following... Figure 3 The steps shown are used to process the first statement:
[0056] Step 301: Starting from the first character of string L, extract string R from string L in left-to-right order. The length of string R is equal to the set maximum word segmentation length.
[0057] When step 301 is executed for the first time, string L corresponds to the string represented by the first statement. When step 301 is executed for the second time, string L is a new string obtained after deleting several characters.
[0058] For example, when the maximum word segmentation length is set to 5, and the first sentence is to reach customers through a targeted tweet from a public account, when step 301 is executed for the first time, string L is to reach customers through a targeted tweet from a public account, and string R extracted from string L is through a public account.
[0059] Step 302: Determine whether the string R exists in the set dictionary.
[0060] If the string R is found in the set dictionary, it indicates that the string R exists in the set dictionary, and step 303 is executed; if the string R is not found in the set dictionary, it indicates that the string R does not exist in the set dictionary, and step d is executed.
[0061] Step 303: Output string R to the first word segmentation subset, delete string R from string L to obtain a new string L, and execute step 301.
[0062] Step 304: Delete the rightmost character of string R to obtain string G.
[0063] Step 305: Determine whether the string G exists in the set dictionary.
[0064] If string G is found in the set dictionary, it indicates that string G exists in the set dictionary, and step 306 is executed; if string G is not found in the set dictionary, it indicates that string G does not exist in the set dictionary, and step 307 is executed.
[0065] Step 306: Output string G to the first word segmentation subset, delete string G from string R to obtain a new string L, and execute step 301.
[0066] Step 307: Determine whether the number of characters contained in string G is 1.
[0067] In the case where the number of characters in string G is 1, string G is output to the first word segmentation subset, string G is deleted from string R, and a new string L is obtained. Step 301 is then executed.
[0068] If the number of characters in string G is greater than 1, proceed to step 308.
[0069] Step 308: Use string G as the new string R and proceed to step 304.
[0070] The process of extracting the second segmentation subset from each first sentence based on the set maximum segmentation length and in a right-to-left order is as follows:
[0071] Here, the electronic device is based on the principle of the backward maximum matching algorithm, following the following... Figure 4 The steps shown are used to process the first statement:
[0072] Step 401: Starting from the last character of string L, extract string R from string L in a right-to-left order. The length of string R is equal to the set maximum word segmentation length.
[0073] When step 401 is executed for the first time, string L corresponds to the string represented by the first statement. When step 401 is executed for the second time, string L is a new string obtained after deleting several characters.
[0074] For example, when the maximum word segmentation length is set to 5, and the first sentence is to reach customers through targeted tweets on WeChat official accounts, when step 401 is executed for the first time, the string L is to reach customers through targeted tweets on WeChat official accounts, and the extracted string R is to reach customers through targeted tweets on WeChat official accounts.
[0075] Step 402: Determine whether the string R exists in the set dictionary.
[0076] If the string R is found in the set dictionary, it indicates that the string R exists in the set dictionary, and step 403 is executed; if the string R is not found in the set dictionary, it indicates that the string R does not exist in the set dictionary, and step d is executed.
[0077] Step 403: Output string R to the first word segmentation subset, delete string R from string L to obtain a new string L, and execute step 401.
[0078] Step 404: Delete the leftmost character of string R to obtain string G.
[0079] Step 405: Determine whether the string G exists in the set dictionary.
[0080] If string G is found in the set dictionary, it indicates that string G exists in the set dictionary, and step 406 is executed; if string G is not found in the set dictionary, it indicates that string G does not exist in the set dictionary, and step 307 is executed.
[0081] Step 406: Output string G to the first word segmentation subset, delete string G from string R to obtain a new string L, and execute step 301.
[0082] Step 407: Determine whether the number of characters contained in string G is 1.
[0083] In the case where the number of characters in string G is 1, string G is output to the first word segmentation subset, string G is deleted from string R, and a new string L is obtained. Step 401 is then executed.
[0084] If the number of characters in string G is greater than 1, proceed to step 408.
[0085] Step 408: Use string G as the new string R and proceed to step 404.
[0086] Step 203: Based on the first and second segmentation subsets corresponding to each first statement, determine the segmentation set corresponding to each first statement.
[0087] Here, after determining the first and second word segments corresponding to each first statement, the electronic device compares the first and second word segments corresponding to each first statement. Based on the comparison result, it determines the word segment set corresponding to the first statement from the first and second word segments corresponding to each first statement.
[0088] Specifically, if the comparison result indicates that the first and second segmented word subsets corresponding to the first statement are the same, the segmentation result is correct; if the comparison result indicates that the first and second segmented word subsets corresponding to the first statement are different, the segmentation result is ambiguous.
[0089] To improve the accuracy of word segmentation in the word segmentation set corresponding to the first statement, in some embodiments, when performing step 203, the method includes one of the following:
[0090] If the first segmentation subset and the second segmentation subset corresponding to the first statement are the same, the corresponding first segmentation subset or the corresponding second segmentation subset shall be determined as the corresponding segmentation set.
[0091] If the first and second word segments corresponding to the first statement are different, add the words that are the same in the first and second word segments corresponding to the first statement to the corresponding word segments set, and add the word with the fewest characters in the first and second word segments to the corresponding word segments set.
[0092] Here, if the first segmentation subset and the second segmentation subset corresponding to the first statement are the same, it indicates that the segmentation result corresponding to the first statement is correct, and the corresponding first segmentation subset or the corresponding second segmentation subset is determined as the corresponding segmentation set.
[0093] If the first and second segmentation subsets corresponding to the first statement are different, it indicates that the segmentation result corresponding to the first statement is ambiguous. In this case, the segments that are the same in the first and second segmentation subsets corresponding to the first statement are added to the corresponding segmentation set of the first statement; the segments with the fewest characters in the first and second segmentation subsets are added to the corresponding segmentation set.
[0094] For example, if the first word in the first segmentation subset is the word "application", and the second word in the second segmentation subset is the application form, then the application form is added to the corresponding segmentation set. In other words, when the first and second segmentation subsets are different, the word with fewer characters in the two subsets is more likely to be segmented correctly.
[0095] Step 204: Merge the determined word segmentation sets to obtain the first word segmentation set corresponding to the first request.
[0096] Here, having obtained the word set corresponding to each segmented first statement, the word sets corresponding to all first statements are merged to obtain the first word set corresponding to the first request.
[0097] Step 103: Based on the field types and corresponding field values in the first form, and the field types to which the words in the first word segmentation set belong, construct the key information model corresponding to the first request.
[0098] Here, upon obtaining the first word segmentation set, the electronic device can determine the field type to which the words in the first word segmentation set belong, based on the established correspondence between the set field types and the set strings. The electronic device extracts the first form from the first request, and according to the field type to which the words in the first word segmentation set belong, writes the corresponding words into the field values corresponding to the field types in the extracted first form. It then performs deduplication on the field values corresponding to each field type to obtain the key information model corresponding to the first request. The key information model represents the correspondence between field names and field values.
[0099] For example, the first form in the first request, the background description information of the data usage requirements included in the first information, the extracted first word segmentation set, and the constructed key information model are as follows:
[0100] Data fields involved ECIF number Data-related products Micro-Smart Sharing Data source environment VDI production Data target users Individual Direct Access Data use cases WeChat Official Account
[0101] First Form
[0102] Background description information of data usage requirements: Currently, it is necessary to send reminder notices for account balance withdrawals to users with remaining balances in Weihuixiang, and reach customers through targeted push messages on the official account. Now, it is applied to use the ecif numbers of Weihuixiang customers with balances to match with users who have opened accounts in the bank app, and push official account messages. The customer ecif numbers are already in the production VDI and are processed for matching through sharing to personal direct access.
[0103] The first word segmentation set includes: ecif number, Weihuixiang, official account, bank app, personal direct access, production VDI.
[0104] The key information model corresponding to the constructed first request is as follows:
[0105] Data fields involved ECIF number Data-related products Micro-Smart Sharing, Bank App Data source environment VDI production Data target users Individual Direct Access Data use cases WeChat Official Account
[0106] Step 104: Use the rule engine to process the user agreement corresponding to the word segmentation representing the product name in the key information model to obtain a set of risk warning rules.
[0107] Here, the electronic device searches in the local database or cloud database for the user agreement that matches the word segmentation representing the product name in the key information model. Exemplarily, the word segmentations representing the product name in the key information model exemplified in step 103 include Weihuixiang and bank app. That is to say, from the key information model, by obtaining the field value corresponding to the field whose field name is the product associated with the data, the word segmentation of the product name can be obtained.
[0108] The electronic device uses the rule engine to analyze the found user agreement, extracts the confidentiality agreement regarding data usage from the content of the found user agreement, and generates a set of risk warning rules based on the extracted confidentiality agreement regarding data usage.
[0109] In actual application, the rule generator includes a rule generator, which is used to sort out a set of risk warning rules according to the data usage confidentiality agreement and the found user agreement. As Figure 5 shown, the rule elements of the rule generator include: data source environment (source), data target user (target), product associated with the data (product), and data usage scenario (scene), etc.
[0110] Exemplarily, the found user agreement includes the user agreement of Product X. The confidentiality agreement regarding data usage in the user agreement of Product X is: Party B and its affiliated parties shall keep confidential the data of Party A accessed and used, and shall not provide it to any third party other than the affiliated parties of Party B's machine.
[0111] The risk warning rule generated by the rule engine is: R:[%source%]==“&&[%target%]=='Third-party organization'&&[%product%]=='X'&&[%scene%]==。 This risk warning rule indicates that there is a data usage compliance risk in the event of a data sharing requirement involving this product with a third party.
[0112] For example, when the key information model data as shown in Table 1 is constructed, the user agreements corresponding to products P1 and P2 in the constructed key information model data are found in the electronic device database, and the confidentiality agreements on data use corresponding to products P1 and P2 are extracted from the found user agreements. The extracted confidentiality agreements on data use are processed by the rule engine, and the resulting set of risk warning rules is shown in Table 2.
[0113]
[0114] Table 1
[0115] The confidentiality agreement regarding data use in the user agreement for product P1 states: Party A shall ensure that the business data generated during the service process provided by Party B to Party A's institutional users shall only be used for the purposes stipulated in this agreement and shall not be used for any other purpose.
[0116] The confidentiality agreement regarding data use in the user agreement for the P2 product states: Customer identity verification information generated by Party B in the process of providing services to Party A shall be used by Party A only for the purposes stipulated in this agreement and shall not be used for any other purpose.
[0117]
[0118]
[0119] Table 2
[0120] Step 105: If the key information model matches any risk warning rule in the risk warning rule set, output the first warning information, which indicates that the first request has a data compliance risk.
[0121] Here, the electronic device matches the key information model with the rules in the risk warning rule set generated by the rule generator to obtain the rule calculation result. When the rule calculation result represents a match with the first rule, the corresponding first warning information is output based on the rule description of the first rule.
[0122] In practical applications, when the constructed key information model hits any rule in Table 3, the rule calculation result corresponding to that rule is True, and the corresponding first prompt information is output.
[0123]
[0124] Table 3
[0125] In this embodiment of the invention, since the lexicon includes proprietary terms within the enterprise, a first word segmentation set containing proprietary terms can be extracted from the received first request based on the lexicon. A key information model corresponding to the first request is constructed based on the field types and corresponding field values in the first form, and the field types to which the words in the first word segmentation set belong. A risk warning rule set is obtained by analyzing the user agreement related to the first request using a rule engine. If the key information model matches any risk warning rule in the risk warning rule set, a first warning message is output, indicating that the first request has a data compliance risk. The above solution, based on the lexicon and natural language processing technology, processes the first request, accurately extracting the first word segmentation set from the first request, thereby constructing a key information model. Based on the word segmentation representing the product name in the key information model, the user agreement is obtained, and all user agreements related to the first request can be obtained, thus improving the accuracy and comprehensiveness of the rules in the risk warning rule set. Because a key information model and a risk warning rule set are constructed, data compliance risks related to the first request can be accurately identified.
[0126] like Figure 6 As shown, in some embodiments, prior to step 102, the method further includes:
[0127] Step 001: Perform word segmentation on each statement in the historical text file corresponding to the first system, and determine the second word segmentation set from the word segments obtained;
[0128] Step 002: Merge the second word segmentation set into the set general lexicon to obtain the set lexicon.
[0129] Here, the electronic device performs word segmentation on each sentence in each historical text file corresponding to the first system. A predefined stop dictionary is used to filter the segmented strings, removing strings containing stop words to obtain a second word segmentation set. This second word segmentation set is then merged into a predefined general lexicon to obtain a predefined lexicon. The word segments in the second word segmentation set represent newly discovered words; the predefined stop dictionary includes multiple stop words. There are multiple historical text files.
[0130] It should be noted that steps 001 and 002 can be executed before step 101, after step 101, or steps 001 and 101 can be executed simultaneously.
[0131] In actual application, during the process of word segmentation for each statement in the historical text file corresponding to the first system, subsets of 2-character strings, 3-character strings, and k-character strings corresponding to each statement are extracted. Here, in actual application, k is an integer less than or equal to 5. Exemplarily, for the statement: It is intended to apply for exporting the conjoint analysis environment data, the subset of 2-character strings extracted is: [intended to apply, apply for, for exporting, exporting the, the conjoint, conjoint analysis, analysis environment, environment data, data for, for exporting].
[0132] In actual application, in order to improve the accuracy of the word segmentation in the second word segmentation set, after filtering out the strings containing stop words in the second word segmentation result using the set stop word dictionary, the set word segmentation dictionary can also be used to filter the second word segmentation set again to filter out the word segmentations existing in the set word segmentation dictionary, and the filtered word segmentation set is output so that the user can select the correct word segmentation from the filtered word segmentation set. When the word segmentation selected by the user is detected, the word segmentation selected by the user is merged into the set general word library. In actual application, the set word segmentation dictionary can be the ICTCLAS core dictionary.
[0133] It should be noted that after obtaining the set word library, the set word library can also be updated according to the above method to update the newly added proper nouns within the enterprise to the set word library.
[0134] In actual application, the historical text file includes the first request received historically.
[0135] In this solution, a set word library is constructed based on the historical text file corresponding to the first system, and the set word library contains the proper nouns within the enterprise. Thus, the accuracy of the first word segmentation set extracted from the first request can be improved.
[0136] To improve the accuracy of the word segmentation in the second word segmentation set, in some embodiments, as Figure 7 shown, determining the second word segmentation set from the strings obtained by word segmentation includes:
[0137] Step 701: Based on the relevance corresponding to each substring in each four-character string obtained by word segmentation, determine whether the first substring in each four-character string is a seed substring to be expanded; where each substring is composed of two adjacent characters, and the first substring is composed of the middle two characters of the four-character string; the relevance characterizes the point mutual information and information richness corresponding to the substring.
[0138] Here, the electronic device performs word segmentation on each statement in each historical text file corresponding to the first system, obtaining the word segmentation result; from the word segmentation result, it determines the subset of bigram strings and subset of quadram strings corresponding to each statement in each historical text file. It should be noted that when calculating the relevance of the t-gram string, the subset of t-gram strings corresponding to each statement in each historical text file is determined from the word segmentation set.
[0139] For each quadruple string in the quadruple string subset corresponding to each statement in each historical text file, perform the following processing:
[0140] Based on the subset of binary strings corresponding to each statement in each historical text file, the left and right information entropies of each binary string in the subset of binary strings corresponding to each statement are calculated; based on the left and right information entropies of each binary string, the information richness of each binary string is calculated.
[0141] From the information richness of each binary string in the binary string subset corresponding to each statement, determine the information richness of the first, second, and third substrings of the corresponding quaternary string subset. Based on the inter-point mutual information and information richness of the first, second, and third substrings of the quaternary string, calculate the relevance of the first, second, and third substrings of the quaternary string.
[0142] Calculate the mean first relevance between the first substring and the second substring, and calculate the mean second relevance between the first substring and the third substring; the first substring consists of the middle two characters of the quadruple string; the second substring consists of the first two adjacent characters of the quadruple string; and the third substring consists of the last two adjacent characters of the quadruple string.
[0143] If the relevance of the first substring is greater than the sum of the relevance of the second substring and the average of the first relevance, and also greater than the sum of the relevance of the third substring and the average of the second relevance, then the probability that the first substring is a word or part of a word is relatively high. In this case, the first substring is determined as the seed string to be expanded; proceed to step 702. Here, relevance refers to the degree of correlation between the characters included in the substring, used to measure the cohesion of the substring.
[0144] If the relevance of the first substring is less than or equal to the sum of the relevance of the second substring and the average relevance of the first substring, or if the relevance of the first substring is less than or equal to the sum of the relevance of the third substring and the average relevance of the second substring, then the probability that the two characters representing the first substring each form a word or a word boundary is relatively high. That is, the first substring is not a seed string to be expanded. In this case, the number of occurrences of the first substring in the corresponding word segmentation results in the corresponding historical text file is reduced by 1.
[0145] For example, using w i-1 w i w i+1 w i+2 To represent any quadruple string, use AMI n When representing the relevance of substrings, the mean of the first relevance is... Second correlation mean Quaternary string w i-1 w i w i+1 w i+2 In meeting AMI n (w i ,w i+1 AMI n (w i-1 ,w i )+M1, and AMI n (w i ,w i+1 AMI n (w i+1 ,w i+2 In the case of M2+1, the substring w is represented. i w i+1 A word or part of a word has a higher probability of being a substring w. i w i+1 The seed string to be expanded; in AMI n (w i ,w i+1 )≤AMI n (w i-1 ,w i )+M1, or AMI n (w i ,w i+1 )≤AMI n (w i+1 ,w i+2 In the case of M2, the character substring w i w i+1 The character w in i and the character w i+1 The probability of each substring forming a word or a word boundary is relatively high; substring w i w i+1 This is not a seed string to be expanded.
[0146] To improve the accuracy of the calculated relevance and the accuracy of identifying words in the second word segmentation set, in some embodiments, such as Figure 8 As shown, the relevance of the substrings is calculated in the following way:
[0147] Step 801: Calculate the left information entropy and right information entropy of the substring based on the probability of occurrence of each character in the left and right neighboring character sets of the substring.
[0148] Here, the electronic device determines the k-ary character string subset corresponding to the statement to which the substring belongs based on the total number of characters in the substring, where k equals the total number of characters in the corresponding substring; from the determined k-ary character string subset, it determines the set of left and right neighboring characters of the substring, and determines the number of times each character string in the set of left and right neighboring characters of the substring appears in the corresponding word segmentation result of the corresponding historical text file.
[0149] Based on the total number of characters in the left neighbor set of the substring and the determined occurrence frequency of each character, the probability of occurrence of each character in the left neighbor set of the substring is calculated. Based on the probability of occurrence of each character in the left neighbor set of the substring, the left information entropy of the substring is calculated.
[0150] Based on the total number of characters in the right neighbor set of the substring and the determined occurrence frequency of each character, the probability of occurrence of each character in the right neighbor set of the substring is calculated; based on the probability of occurrence of each character in the right neighbor set of the substring, the right information entropy of the substring is calculated. Among these,
[0151] The set of left neighbors of a substring includes at least one left neighbor substring, which is the substring to the left of the substring in the k-ary substring subset; the set of right neighbors of a substring includes at least one right neighbor substring, which is the substring to the right of the substring in the k-ary substring subset. The total number of characters in either the left or right neighbor substring of a substring is the same as the total number of characters in the corresponding substring. The substrings in the k-ary substring subset are arranged according to the order of characters in the corresponding statement.
[0152] In practical applications, the following is adopted: The right information entropy of a substring is calculated using the formula... Calculate the left information entropy of the substring.
[0153] Where RE(w) represents the right information entropy of substring w; p(w ri The right neighboring substring w represents the substring. ri The occurrence count of substring w, where m represents the total number of substrings in the left or right neighbor set; LE(w) represents the left information entropy of substring w; p(w) li The left neighbor substring w of the substring li The number of times it appears.
[0154] For example, in the first historical text file corresponding to substring w, the right substring WA appears 5 times, the substring WB appears 2 times, and the substring WC appears 3 times. Then the right information entropy of substring w is...
[0155] It's important to note that entropy is an indicator of information content. Higher entropy means greater information content, higher uncertainty, and greater difficulty in prediction. Left and right information entropy is calculated by measuring the information entropy of the left and right sides of a string, and the value of the information entropy is used to measure the randomness of the left and right neighboring word sets of that string.
[0156] Step 802: Calculate the first parameter corresponding to the substring based on the left and right information entropy of the substring; the first parameter represents the information richness between the left and right neighboring word sets of the substring.
[0157] Here, the electronic device substitutes the left and right information entropies of the substring into a predefined formula to calculate the first parameter corresponding to the substring. This first parameter corresponds to the information richness mentioned above.
[0158] In practical applications, electronic devices are based on formulas Calculate the first parameter corresponding to the substring.
[0159] Step 803: Calculate the relevance of the substring based on the joint probability of the substring, the probability of each character in the substring appearing in the corresponding historical text file, and the first parameter; the joint probability represents the probability that all characters in the substring appear in the n-ary substring at the same time.
[0160] Here, n is also called the joint probability factor. n is a positive integer greater than or equal to 7. In practical applications, n is greater than or equal to 7 and less than 20. Of course, in practical applications, n can also be determined based on the first historical request, specifically when the word segmentation accuracy is greater than or equal to a second set threshold; the process of determining n is described in... Figure 9 The corresponding embodiments are described in detail.
[0161] The electronic device determines the number of times each character in the substring appears in the corresponding historical text file and the total number of characters in the historical text file, and calculates the probability of each character in the substring appearing in the corresponding historical text file; based on the total number of n-ary substrings in the corresponding n-ary substring subset of the historical text file, and based on the number of times all characters in the substring appear simultaneously in the n-ary substring subset, it calculates the joint probability of the substring; based on the joint probability of the substring, the probability of each character in the substring, and the first parameter, it calculates the relevance of the substring.
[0162] In practical applications, electronic devices calculate the relevance of substrings based on the following formula:
[0163]
[0164] Where t is a positive integer greater than 1, and t is less than or equal to n; AMI n (w1,…,wt The substrings w1, ..., w are represented by the substrings w1, ..., w1. t Relevance; w t p(w1) represents the t-th character in the substring; p(w1) represents the probability of the first character w1 appearing in the substring; p(w t ) represents the first character w in the substring t The probability of p occurring; n (w1,…,w t The substrings w1, ..., w are represented by the substrings w1, ..., w1. t The probability that all characters in the string appear simultaneously in the n-ary substring; L(w1,…,w t The substrings w1, ..., w are represented by the substrings w1, ..., w1. t The first parameter; L(w1,…,w t The calculation method for L(w) is similar to that for L(w), and will not be repeated here.
[0165] Step 702: If the first substring is the seed string to be expanded, expand the first substring until the number of occurrences of the expanded substring in the word segmentation result is less than the first set threshold, and output the expanded substring to the second word segmentation set.
[0166] Here, the strings in the second word segmentation set exist in the form of word chains, meaning that the strings in the second word segmentation set are arranged according to their order in the corresponding sentences. The electronic device expands the first substring according to the following steps:
[0167] Step a: Determine the first t+1 character string and the second t+1 character string corresponding to the t-character string.
[0168] When step a is executed for the first time, the t-element string corresponds to the first substring. When step a is executed for the second time or in subsequent executions, the t-element string corresponds to the first t+1-element string obtained most recently through step d, or the second t+1-element string obtained most recently through step f.
[0169] The first t+1-element string consists of the t-element string and the character adjacent to it on the left; the second t+1-element string consists of the t-element string and the character adjacent to it on the right.
[0170] For example, in the t-element string w i …w i+t-1 In the case where the first t+1 element string is w i-1 w i …w i+t-1 The second t+1 element string w i …w i+t-1 w i+t .
[0171] Step b: Calculate the relevance of the t-ary string, the relevance of the first t+1-ary string, and the relevance of the second t+1-ary string.
[0172] For example, in the t-element string w i …w i+t-1 In the case of t-element string, the relevance is AMI n (w i ,…,w i+t-1 The relevance of the first t+1-element string is AMI. n (w i-1 ,w i ,…,w i+t-1 The relevance of the second t+1-element string is the same as that of AMI. n (w i ,…,w i+t-1 ,w i+t In practical applications, based on Formula 1 above, the corresponding relevance is calculated according to steps 801 to 803; the implementation process of calculating the relevance is described in the relevant description above, and will not be repeated here.
[0173] Step c: Determine whether the relevance of the first t+1-element string is greater than the relevance of the second t+1-element string.
[0174] If the relevance of the first t+1-element string is greater than that of the second t+1-element string, it indicates that the probability of expanding the t-element string into the first t+1-element string is greater than the probability of expanding it into the second t+1-element string. In this case, step d is executed.
[0175] If the relevance of the first t+1-element string is less than or equal to the relevance of the second t+1-element string, it indicates that the probability of expanding the t-element string into the second t+1-element string is greater than the probability of expanding it into the first t+1-element string. In this case, step f is executed.
[0176] For example, in AMI n (w i-1 ,w i ,…,w i+t-1 AMI n (w i ,…,w i+t-1 ,w i+t In the case of AMI, proceed to step d; n (w i-1 ,w i ,…,w i+t-1 )≤AMI n (w i ,…,w i+t-1 ,w i+tIn the case of ), proceed to step f.
[0177] Step d: Calculate the mean third relevance between the t-element string and the first t+1-element string; if the sum of the relevance of the first t+1-element string and the mean third relevance is greater than or equal to the relevance of the t-element string, expand the t-element string into the first t+1-element string.
[0178] Step e: Determine whether the number of occurrences of the first t+1 character string in the word segmentation result is less than the first set threshold.
[0179] Specifically, if the number of occurrences of the first t+1-element string in the word segmentation result is less than a first set threshold, the first t+1-element string is output to the second word segmentation set, thus ending the expansion of the t-element string.
[0180] If the number of occurrences of the first t+1 meta-string in the word segmentation result is greater than or equal to the first set threshold, assign t to t+1 and execute step a.
[0181] In cases where the sum of the relevance of the first t+1-element string and the average of the third relevance is less than the relevance of the t-element string, the t-element string is output to the second word segmentation set, thus ending the expansion of the t-element string.
[0182] For example, in the t-element string w i …w i+t-1 In the case of the third relevance mean
[0183] When AMI n (w i-1 ,w i ,…,w i+t-1 )+M3≥AMI n (w i ,…,w i+t-1 In the case of w i …w i+t-1 Expand to w i-1 w i …w i+t-1 .
[0184] Step f: Calculate the mean fourth relevance between the t-element string and the second t+1-element string; if the sum of the relevance of the second t+1-element string and the mean fourth relevance is greater than the relevance of the t-element string, expand the t-element string into the second t+1-element string.
[0185] Specifically, if the sum of the relevance of the second (t+1)-element string and the average of the fourth relevance is less than the relevance of the t-element string, the t-element string is output to the second word segmentation set, thus ending the expansion of the t-element string.
[0186] For example, in the t-element string w i …w i+t-1 In this case,
[0187] When AMI n (w i ,…,w i+t-1 ,w i+t )+M4≥AMI n (w i ,…,w i+t-1 In the case of w i …w i+t-1 Expand to w i …w i+t-1 w i+t .
[0188] Step g: Determine whether the number of occurrences of the second t+1-element string in the word segmentation result is less than the first set threshold.
[0189] Specifically, if the number of occurrences of the second t+1-element string in the word segmentation result is less than the first set threshold, the second t+1-element string is output to the second word segmentation set, thus ending the expansion of the t-element string.
[0190] If the number of occurrences of the second t+1 meta-string in the word segmentation result is greater than or equal to the first set threshold, assign t to t+1 and execute step a.
[0191] To improve the accuracy of the determined seed string and extended string, in practical applications, the joint probability factor n corresponding to the word segmentation accuracy meeting a second set threshold is determined through the historical first request. In some embodiments, before step 001, when each sentence in the historical text file corresponding to the first system is segmented and the second word set is determined from the segmented string, the method further includes:
[0192] Step 901: Perform word segmentation on each statement in the first request received in history, and determine the candidate word set from the word segmentation string.
[0193] Here, the electronic device performs word segmentation processing on each sentence in the first request received in the past, according to... Figure 7 and Figure 8 The corresponding embodiment uses a method to determine the second word segmentation set, which determines the candidate word set.
[0194] It should be noted that the first historical request in step 901 is different from the historical text file in step 001. In other words, the first historical request used to test the word segmentation effect in step 901 is different from the historical text file used to build the set dictionary.
[0195] In accordance with Figure 8 The corresponding implementation calculates the relevance AMI of the string. n When n is a given number, the joint probability factor n takes values starting from 1.
[0196] For example, a portion of the word string in the candidate word set determined using the first request received historically is shown below:
[0197]
[0198]
[0199] In the table above, entropy is obtained by calculating the logarithm of the frequency. The frequency in the table represents the quotient of the frequency and the total number of characters in the string.
[0200] Step 902: Calculate the accuracy corresponding to the candidate word set based on the first quantity and the total number of candidate words in the candidate word set; the first quantity represents the number of words in the candidate word set that are the same as those in the tagged vocabulary set corresponding to the first request received in the past.
[0201] Here, after determining the candidate word set corresponding to the first historically received request, the determined candidate word set is compared with the tagged word set corresponding to the first historically received request to determine the number of identical words, obtaining a first count. The quotient of the first data and the total number of candidate words in the determined candidate word set is calculated to obtain the corresponding accuracy rate, and it is determined whether the determined accuracy rate is less than a second preset threshold. If the determined accuracy rate is less than the second preset threshold, step 903 is executed; if the determined accuracy rate is greater than or equal to the second preset threshold, step 904 is executed.
[0202] Step 903: If the accuracy is less than the second set threshold, adjust the value of the joint probability factor, and re-determine the candidate word set from the word segmentation string according to the adjusted value of the joint probability factor.
[0203] Here, if the determined accuracy is less than the second set threshold, the value of the joint probability factor n is adjusted, and the candidate word set is re-determined from the word string obtained by word segmentation based on the adjusted joint probability factor. Then, step 902 is returned to calculate the accuracy corresponding to the adjusted candidate word set based on the first quantity and the total number of candidate words in the adjusted candidate word set.
[0204] When adjusting the value of the joint probability factor n, it can be incremented based on the current value of n. In practical applications, n is a positive integer less than or equal to 20.
[0205] Step 904: If the accuracy is greater than or equal to the second set threshold, the value corresponding to the current joint probability factor is determined as n.
[0206] Here, if the accuracy is greater than or equal to the second set threshold, the current value of the joint probability factor is determined as the final value of the joint probability factor.
[0207] In practical applications, electronic devices can test the accuracy rate corresponding to each value of n from 1 to 20, and determine the value of n corresponding to the highest accuracy as the final joint probability factor value. For example, the accuracy curves corresponding to n values from 1 to 20 are shown below. Figure 10 As shown in the figure.
[0208] The following sets of candidate words are determined for different values of n (n≤10):
[0209]
[0210]
[0211] The table above shows that when n is greater than or equal to 7, the resulting candidate word set differs significantly from that when n = 1 or 2. When n = 1 and n = 2, the top-ranking strings all contain low-frequency characters or words, such as "add deduction," "large repayment," and "Tencent big," and the collocations of these strings are fixed. In the results where n is greater than or equal to 7, no low-frequency co-occurring strings appear, indicating that when n is greater than or equal to 7, new words can be effectively identified.
[0212] In this embodiment, the accuracy of the joint probability factor n is determined by testing the accuracy of different values of the first request received in history. The final value of n is determined by using the n corresponding to the accuracy being greater than or equal to the second set threshold, which can improve the accuracy of recognizing new words.
[0213] Figure 11 This diagram illustrates the implementation flow of the data detection method provided in an application embodiment of the present invention, as shown below. Figure 11 As shown, the data detection methods include:
[0214] Step 1101: Perform word segmentation on each statement in the first request received in history, and determine the candidate word set from the word segmentation string.
[0215] For the implementation process of step 1101, please refer to the relevant description of step 901.
[0216] Step 1102: Calculate the accuracy corresponding to the candidate word set based on the first quantity and the total number of candidate words in the candidate word set; the first quantity represents the number of words in the candidate word set that are the same as those in the tagged vocabulary set corresponding to the first request received in history.
[0217] For the implementation process of step 1102, please refer to the relevant description of step 902.
[0218] Step 1103: If the accuracy is less than the second set threshold, adjust the value of the joint probability factor, and re-determine the candidate word set from the word segmentation string according to the adjusted value of the joint probability factor.
[0219] For the implementation process of step 1103, please refer to the relevant description of step 903.
[0220] Step 1104: If the accuracy is greater than or equal to the second set threshold, the current value of the joint probability factor is determined as the final value of the joint probability factor.
[0221] For the implementation process of step 1104, please refer to the relevant description of step 904. The final value of the joint probability factor is the final value of the joint probability factor n.
[0222] Step 1105: Perform word segmentation on each statement in the historical text file corresponding to the first system, and determine the second word segmentation set from the word segments.
[0223] For the implementation process of step 1105, please refer to the relevant description of step 001.
[0224] Step 1106: Merge the second word segmentation set into the set general lexicon to obtain the set lexicon.
[0225] For the implementation process of step 1106, please refer to the relevant description of step 002.
[0226] Step 1107: Receive the first request; the first request includes a first form and first information.
[0227] For the implementation process of step 1107, please refer to the relevant description of step 101.
[0228] Step 1108: Extract the first word segmentation set from the first information.
[0229] For the implementation process of step 1108, please refer to the relevant description of step 102.
[0230] Step 1109: Based on the field types and corresponding field values in the first form, and the field types to which the words in the first word segmentation set belong, construct the key information model corresponding to the first request.
[0231] For the implementation process of step 1109, please refer to the relevant description of step 103.
[0232] Step 1110: Using a rule engine, process the user agreement corresponding to the word segmentation representing the product name in the key information model to obtain a set of risk warning rules.
[0233] For the implementation process of step 1110, please refer to the relevant description of step 104.
[0234] Step 1111: If the key information model matches any risk warning rule in the risk warning rule set, output the first warning information, which indicates that the first request has a data compliance risk.
[0235] For the implementation process of step 1111, please refer to the relevant description of step 105.
[0236] To implement the method of the embodiments of the present invention, the embodiments of the present invention also provide a terminal, such as... Figure 12 As shown, the electronic device includes:
[0237] The receiving unit 121 is configured to receive a first request; the first request includes a first form and first information; the first request is used to apply for access rights to the data in the first form in the first system; the first information represents text information in the first request other than the first form.
[0238] Extraction unit 122 is used to extract a first word segmentation set from the first information; the words in the first word segmentation set are words that exist in a set word library;
[0239] The construction unit 123 is used to construct a key information model corresponding to the first request based on the field types and corresponding field values in the first form and the field types to which the words in the first word segmentation set belong.
[0240] The generation unit 124 is used to process the user agreement corresponding to the word segmentation representing the product name in the key information model using a rule engine to obtain a risk warning rule set.
[0241] The prompting unit 125 is used to output a first prompting message when the key information model hits any of the risk prompting rules in the risk prompting rule set. The first prompting message indicates that the first request has a data compliance risk.
[0242] In some embodiments, the electronic device further includes:
[0243] The determining unit is used to perform word segmentation on each sentence in the historical text file corresponding to the first system before the extraction unit 123 extracts the first word segmentation set from the first information, and to determine the second word segmentation set from the word segmentation string.
[0244] The merging unit is used to merge the second word segmentation set into a set general lexicon to obtain the set lexicon.
[0245] In some embodiments, the extraction unit 123 is specifically used for:
[0246] Based on the relevance of each substring in each quadruple string obtained from word segmentation, it is determined whether the first substring in each quadruple string is a seed string to be expanded; where each substring consists of two adjacent characters, and the first substring consists of the middle two characters of the quadruple string; the relevance represents the mutual information and information richness between the points corresponding to the substrings;
[0247] If the first substring is the seed string to be expanded, the first substring is expanded until the number of occurrences of the expanded string in the word segmentation result is less than a first set threshold, and the expanded string is output to the second word segmentation set.
[0248] In some embodiments, the electronic device further includes:
[0249] The calculation unit is used to calculate the relevance of the substring in the following way: based on the occurrence probability of each character in the left neighbor set and right neighbor set of the substring, the left information entropy and right information entropy of the substring are calculated.
[0250] Based on the left and right information entropy of the substring, the first parameter corresponding to the substring is calculated; the first parameter represents the information richness between the left and right neighboring word sets of the substring.
[0251] Based on the joint probability of the substring, the probability of each character in the substring appearing in the corresponding historical text file, and the first parameter, the relevance of the substring is calculated; the joint probability represents the probability that all characters in the substring appear in the n-ary substring at the same time.
[0252] In some embodiments, the electronic device further includes: a training unit, configured to perform the following steps before the determining unit performs word segmentation processing on each statement in the historical text file corresponding to the first system and determines the second word segmentation set from the word segmentation string:
[0253] Each sentence in the first request received in history is segmented into words, and a set of candidate words is determined from the segmented strings;
[0254] Based on the first quantity and the total number of candidate words in the candidate word set, the accuracy corresponding to the candidate word set is calculated; the first quantity represents the number of words in the candidate word set that are the same as those in the tagged vocabulary set corresponding to the first request received in the past.
[0255] If the accuracy is less than the second set threshold, the value of the joint probability factor is adjusted, and the candidate word set is re-determined from the word segmentation string according to the adjusted value of the joint probability factor.
[0256] If the accuracy is greater than or equal to the second set threshold, the value of the current joint probability factor is determined as n.
[0257] In some embodiments, when the determining unit extracts the first word segmentation set from the first information, it is specifically used to:
[0258] Based on the position of the punctuation marks, the first information is divided into multiple first sentences;
[0259] Based on the set maximum word segmentation length, the first word segmentation subset is extracted from each first sentence in order from left to right, and the second word segmentation subset is extracted from each first sentence in order from right to left.
[0260] Based on the first and second segmentation subsets corresponding to each first statement, the segmentation set corresponding to each first statement is determined;
[0261] The determined word segmentation sets are merged to obtain the first word segmentation set corresponding to the first request.
[0262] In some embodiments, when the determining unit determines the word segmentation set corresponding to each first statement based on the first word segmentation subset and the second word segmentation subset corresponding to each first statement, it is specifically used to perform one of the following:
[0263] If the first segmentation subset and the second segmentation subset corresponding to the first statement are the same, the corresponding first segmentation subset or the corresponding second segmentation subset shall be determined as the corresponding segmentation set.
[0264] If the first and second word segments corresponding to the first statement are different, add the words that are the same in the first and second word segments corresponding to the first statement to the corresponding word segments set, and add the word with the fewest characters in the first and second word segments to the corresponding word segments set.
[0265] In practical applications, the various units included in an electronic device can be implemented by processors in the electronic device, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU), or a field-programmable gate array (FPGA).
[0266] It should be noted that the above embodiments of the electronic device are only illustrated by the division of the above program modules when performing data detection. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the electronic device and the data detection method embodiments provided above belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0267] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiments of the present invention, the embodiments of the present invention also provide an electronic device. Figure 13 This is a schematic diagram of the hardware composition structure of the electronic device according to an embodiment of the present invention, such as... Figure 13 As shown, the electronic device 13 includes:
[0268] Communication interface 131 enables information exchange with other devices, such as network devices;
[0269] The processor 132 is connected to the communication interface 131 to enable information interaction with other devices and, when running a computer program, executes the data detection method provided by one or more of the above-mentioned technical solutions. The computer program is stored in the memory 133.
[0270] Of course, in practical applications, the various components in electronic device 13 are coupled together through bus system 134. It can be understood that bus system 134 is used to realize the connection and communication between these components. In addition to a data bus, bus system 134 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in... Figure 13 The general labeled all buses as Bus System 134.
[0271] The memory 133 in this embodiment is used to store various types of data to support the operation of the electronic device 13. Examples of such data include any computer program used to operate on the electronic device 13.
[0272] It is understood that memory 133 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 133 described in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0273] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 132. The processor 132 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 132 or by instructions in the form of software. The processor 132 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 132 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the memory 133. The processor 132 reads the program in the memory 133 and completes the steps of the aforementioned method in combination with its hardware.
[0274] Optionally, when the processor 132 executes the program, it implements the corresponding processes implemented by the terminal in the various methods of the embodiments of this application. For the sake of brevity, these will not be described in detail here.
[0275] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a first memory 113 storing a computer program, which can be executed by a terminal processor 132 to complete the steps described in the aforementioned method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0276] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0277] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0278] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing module, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0279] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0280] It should be noted that the technical solutions described in the embodiments of the present invention can be combined arbitrarily without conflict.
[0281] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A data detection method, characterized in that, include: Receive the first request; The first request includes a first form and first information; The first request is used to request access to the data in the first form within the first system; The first word segmentation set is extracted from the first information; the words in the first word segmentation set are words that exist in a set word library, which includes proper nouns within the enterprise; Based on the field types and corresponding field values in the first form, and the field types to which the words in the first word segmentation set belong, a key information model corresponding to the first request is constructed. From a local database or a cloud database, find the user agreement that represents the product name in the key information model through word segmentation matching; use a rule engine to process the user agreement to obtain a set of risk warning rules; If the key information model matches any of the risk warning rules in the risk warning rule set, a first warning message is output, which indicates that the first request has a data compliance risk.
2. The method according to claim 1, characterized in that, Before extracting the first word segmentation set from the first information, the method further includes: Each sentence in the historical text file corresponding to the first system is segmented into words, and the second segmentation set is determined from the word strings obtained from the segmentation. The second word segmentation set is merged into the set general lexicon to obtain the set lexicon.
3. The method according to claim 2, characterized in that, The process of determining the second set of word segments from the word segments includes: Based on the relevance of each substring in each quadruple string obtained from word segmentation, it is determined whether the first substring in each quadruple string is a seed string to be expanded; where each substring consists of two adjacent characters, and the first substring consists of the middle two characters of the quadruple string; the relevance represents the mutual information and information richness between the points corresponding to the substrings; If the first substring is the seed string to be expanded, the first substring is expanded until the number of occurrences of the expanded string in the word segmentation result is less than a first set threshold, and the expanded string is output to the second word segmentation set.
4. The method according to claim 3, characterized in that, The method further includes: The relevance of a substring is calculated using the following method: Based on the probability of each character appearing in the left and right neighboring character sets of the substring, the left information entropy and right information entropy of the substring are calculated. Based on the left and right information entropy of the substring, the first parameter corresponding to the substring is calculated; the first parameter represents the information richness between the left and right neighboring word sets of the substring. Based on the joint probability of the substring, the probability of each character in the substring appearing in the corresponding historical text file, and the first parameter, the relevance of the substring is calculated; the joint probability represents the probability that all characters in the substring appear in the n-ary substring at the same time.
5. The method according to claim 4, characterized in that, Before performing word segmentation on each statement in the historical text file corresponding to the first system and determining the second word segmentation set from the resulting word segments, the method further includes: Each sentence in the first request received in history is segmented into words, and a set of candidate words is determined from the segmented strings; Based on the first quantity and the total number of candidate words in the candidate word set, the accuracy corresponding to the candidate word set is calculated; the first quantity represents the number of words in the candidate word set that are the same as those in the tagged vocabulary set corresponding to the first request received in the past. If the accuracy is less than the second set threshold, the value of the joint probability factor is adjusted, and the candidate word set is re-determined from the word segmentation string according to the adjusted value of the joint probability factor. If the accuracy is greater than or equal to the second set threshold, the value of the current joint probability factor is determined as n.
6. The method according to any one of claims 1-5, characterized in that, The first word segmentation set is extracted from the first information, including: Based on the position of the punctuation marks, the first information is divided into multiple first sentences; Based on the set maximum word segmentation length, the first word segmentation subset is extracted from each first sentence in order from left to right, and the second word segmentation subset is extracted from each first sentence in order from right to left. Based on the first and second segmentation subsets corresponding to each first statement, the segmentation set corresponding to each first statement is determined; The determined word segmentation sets are merged to obtain the first word segmentation set corresponding to the first request.
7. The method according to claim 6, characterized in that, When determining the word segmentation set corresponding to each first statement based on the first and second word segmentation subsets corresponding to each first statement, the method includes one of the following: If the first segmentation subset and the second segmentation subset corresponding to the first statement are the same, the corresponding first segmentation subset or the corresponding second segmentation subset shall be determined as the corresponding segmentation set. If the first and second word segments corresponding to the first statement are different, add the words that are the same in the first and second word segments corresponding to the first statement to the corresponding word segments set, and add the word with the fewest characters in the first and second word segments to the corresponding word segments set.
8. An electronic device, characterized in that, include: The receiving unit is used to receive the first request; The first request includes a first form and first information; The first request is used to request access to the data in the first form within the first system; An extraction unit is used to extract a first word segmentation set from the first information; The word segments in the first word segmentation set are those that exist in a set lexicon, which includes proper nouns within the enterprise. The construction unit is used to construct the key information model corresponding to the first request based on the field types and corresponding field values in the first form and the field types to which the words in the first word segmentation set belong. The generation unit is used to find the user agreement representing the product name in the key information model from the local database or the cloud database; and to process the user agreement using the rule engine to obtain a set of risk warning rules. The prompting unit is used to output a first prompting message when the key information model hits any of the risk prompting rules in the risk prompting rule set. The first prompting message indicates that the first request has a data compliance risk.
9. An electronic device, characterized in that, include: The processor and the memory used to store computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Safe sharing of sensitive data
CN109997143A
New word discovery method and device, computer storage medium and electronic equipment
CN112559694A