Text recognition method and device, electronic equipment, storage medium and product
By combining preset large language models and specified rules, we automatically identify sensitive information in the database, solving the problem of inefficient manual recognition, and achieving efficient and accurate screening and encryption processing of sensitive information.
Patent Information
- Application Number
- CN202510245996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, it is inefficient to manually identify sensitive information in large amounts of text, making it difficult to efficiently identify sensitive information in the database.
Using a method of combining preset large language model and specified rules, text with sensitive information characteristics is initially selected, and the preset prompt words are used to indicate further recognition of the large language model to achieve automated sensitive information recognition.
It improves the efficiency of identification of sensitive information, reduces manual intervention, improves the degree of automation and accuracy of identification, and ensures data security.
Smart Images

Figure CN120234413A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a text recognition method, device, electronic device, storage medium, and product. Background Art
[0002] To ensure user privacy, managers need to identify sensitive information in a large amount of text stored in a database. For example, the sensitive information can be the user's name, phone number, etc. Subsequently, the sensitive information can be encrypted to protect the user's sensitive information. However, the efficiency of manually identifying sensitive information in a large amount of text is not high. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide a text recognition method, device, electronic device, storage medium, and product to improve the efficiency of identifying sensitive information. The specific technical solutions are as follows:
[0004] In a first aspect, the embodiments of the present invention provide a text recognition method, the method including:
[0005] Obtain the current text to be recognized;
[0006] Determine whether the current text to be recognized conforms to the current specified rule; wherein, the current specified rule includes: the first recognition rule obtained when using a preset large language model to recognize historical text, and the first recognition rule refers to the rule that text of a specified sensitive text type conforms to;
[0007] If the current text to be recognized conforms to the current specified rule, input the current text to be recognized and a preset prompt word into the preset large language model for recognition; wherein, the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type;
[0008] Obtain the recognition result output by the preset large language model.
[0009] Optionally, the preset prompt word is further used to instruct the preset large language model to output the second recognition rule used to obtain the recognition result;
[0010] The method further includes:
[0011] When the obtained recognition result indicates that the current text to be recognized does not belong to the specified sensitive text type, obtain the second recognition rule used by the preset large language model to obtain the recognition result;
[0012] Add the second recognition rule to the current specified rule.
[0013] Optionally, the obtaining of the current text to be recognized includes:
[0014] Obtaining the text stored in the preset database as the current text to be recognized;
[0015] The method further includes:
[0016] When the obtained recognition result indicates that the current text to be recognized belongs to the specified sensitive text type, encrypting the current text to be recognized;
[0017] Deleting the current text to be recognized in the preset database and writing the encrypted result of the current text to be recognized to the storage location of the text to be recognized in the preset database to update the preset database.
[0018] Optionally, before determining whether the current text to be recognized conforms to the current specified rule, the method further includes:
[0019] Obtaining the text type to which the text to be recognized belongs, which is pre-recorded, as the text type to be processed;
[0020] According to the corresponding relationship between the preset text type and the rule set, determining the rule in the rule set corresponding to the text type to be processed as the current specified rule; wherein, the rule set corresponding to a text type includes: the first recognition rule obtained when recognizing historical texts using the preset large language model, and the first recognition rule refers to the rule that the text of this text type conforms to.
[0021] Optionally, before determining the rule in the rule set corresponding to the text type to be processed as the current specified rule according to the corresponding relationship between the preset text type and the rule set, the method further includes:
[0022] Judging whether the text type to be processed is the specified sensitive text type;
[0023] The determining the rule in the rule set corresponding to the text type to be processed as the current specified rule according to the corresponding relationship between the preset text type and the rule set includes:
[0024] When the current text type to be processed is the specified sensitive text type, determining the rule in the rule set corresponding to the text type to be processed as the current specified rule according to the corresponding relationship between the preset text type and the rule set;
[0025] The method further includes:
[0026] In the case that the type of the text to be processed is not the specified sensitive text type, perform word segmentation on the text to be recognized, and use the word segmentation result whose text type is the specified sensitive text type obtained by word segmentation as the current text to be recognized, and return to execute the step of determining the rule in the rule set corresponding to the type of the text to be processed according to the corresponding relationship between the preset text type and the rule set as the current specified rule in the case that the current type of the text to be processed is the specified sensitive text type.
[0027] Optionally, there are multiple current texts to be recognized;
[0028] The method further includes:
[0029] Obtain a specified number of texts from the updated preset database as the current texts to be verified;
[0030] If at least one of the current texts to be verified conforms to the current specified rule, return to execute the step of obtaining the text stored in the preset database as the current text to be recognized until none of the current texts to be verified conforms to the current specified rule.
[0031] Optionally, the current specified rule includes at least one of the following:
[0032] The current text to be recognized is not the specified text;
[0033] The first character of the current text to be recognized is not the specified character;
[0034] The number of characters included in the current text to be recognized belongs to the preset range.
[0035] In a second aspect, an embodiment of the present invention provides a text recognition device, and the device includes:
[0036] A first acquisition module, configured to acquire the current text to be recognized;
[0037] A first judgment module, configured to judge whether the current text to be recognized conforms to the current specified rule; wherein, the current specified rule includes: a first recognition rule obtained when using a preset large language model to recognize historical texts, and the first recognition rule refers to the rule that texts of the specified sensitive text type conform to;
[0038] An input module, configured to input the current text to be recognized and a preset prompt word into the preset large language model for recognition if the current text to be recognized conforms to the current specified rule; wherein, the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type;
[0039] A second acquisition module, configured to acquire the recognition result output by the preset large language model.
[0040] An embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0041] The memory is used to store a computer program;
[0042] The processor, when executing the program stored in the memory, implements the above-mentioned text recognition method.
[0043] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned text recognition method is implemented.
[0044] An embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned text recognition method is implemented.
[0045] A text recognition method provided by an embodiment of the present invention includes: acquiring the current text to be recognized; determining whether the current text to be recognized conforms to the current specified rule; if the current text to be recognized conforms to the current specified rule, inputting the current text to be recognized and a preset prompt word into a preset large language model for recognition; and acquiring the recognition result output by the preset large language model.
[0046] In an embodiment of the present invention, since the first recognition rule in the current specified rule is obtained by using a preset large language model to recognize historical texts, and the text of the specified sensitive text type conforms to the first recognition rule, thus, the specified rule can reflect the characteristics of the text conforming to the specified sensitive text type (collectively referred to as sensitive information characteristics in this article). By determining whether the text to be recognized conforms to the specified rule, the text to be recognized with sensitive information characteristics can be initially screened out, and the text to be recognized with sensitive information characteristics may belong to the specified sensitive text type. Since the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type, thus, inputting the text to be recognized that conforms to the specified rule and the preset prompt word into the preset large language model can further recognize whether the text to be recognized with sensitive information characteristics belongs to the specified sensitive text type. The present invention can avoid manual recognition of sensitive information in a large amount of text and improve the efficiency of recognizing sensitive information. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0048] Figure 1 Schematic flow chart of the first text recognition method provided by an embodiment of the present invention;
[0049] Figure 2 Schematic flow chart of the second text recognition method provided by an embodiment of the present invention;
[0050] Figure 3 Schematic flow chart of the third text recognition method provided by an embodiment of the present invention;
[0051] Figure 4 Schematic flow chart of the fourth text recognition method provided by an embodiment of the present invention;
[0052] Figure 5 Schematic flow chart of the fifth text recognition method provided by an embodiment of the present invention;
[0053] Figure 6 Schematic flow chart of the sixth text recognition method provided by an embodiment of the present invention;
[0054] Figure 7 Schematic principle diagram of the first text recognition method provided by an embodiment of the present invention;
[0055] Figure 8 Schematic principle diagram of the second text recognition method provided by an embodiment of the present invention;
[0056] Figure 9 Schematic structural diagram of a text recognition device provided by an embodiment of the present invention;
[0057] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0058] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0059] In a business system, the data security issue is a very important aspect. For example, the business system includes a database for storing business-related data, and the data security issue is how to protect sensitive information in the database. In the data security issue, how to quickly and efficiently identify sensitive information is a difficult problem. Among them, sensitive data is data with security risks, such as names, mobile phone numbers, ID card numbers, etc. In the prior art, sensitive information is usually screened out manually in a database with a large amount of data, which is costly and inefficient. It can be seen that how to improve the efficiency of identifying sensitive information is an urgent problem to be solved.
[0060] To improve the efficiency of identifying sensitive information, an embodiment of the present invention provides a text recognition method, apparatus, electronic device, storage medium, and product.
[0061] Among them, a text recognition method provided by an embodiment of the present invention is applied to an electronic device. The electronic device can obtain text and identify whether the obtained text belongs to a specified sensitive text type.
[0062] A text recognition method provided by an embodiment of the present invention may include the following steps:
[0063] Obtain the current text to be recognized;
[0064] Determine whether the current text to be recognized conforms to the current specified rule; among them, the current specified rule includes: the first recognition rule obtained when using a preset large language model to recognize historical text, and the first recognition rule refers to the rule that the text of the specified sensitive text type conforms to;
[0065] If the current text to be recognized conforms to the current specified rule, input the current text to be recognized and a preset prompt word into the preset large language model for recognition; among them, the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type;
[0066] Obtain the recognition result output by the preset large language model.
[0067] In an embodiment of the present invention, since the first recognition rule in the current specified rule is obtained when using a preset large language model to recognize historical text, and the text of the specified sensitive text type conforms to the first recognition rule, thus, the specified rule can reflect the characteristics of the text that conforms to the specified sensitive text type (collectively referred to as sensitive information characteristics in this article). By whether the text to be recognized conforms to the specified rule, the text to be recognized with sensitive information characteristics can be initially screened out, and the text to be recognized with sensitive information characteristics may belong to the specified sensitive text type. Since the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type, thus, inputting the text to be recognized that conforms to the specified rule and the preset prompt word into the preset large language model can further recognize whether the text to be recognized with sensitive information characteristics belongs to the specified sensitive text type. The present invention can avoid manual identification of sensitive information in a large amount of text and improve the efficiency of identifying sensitive information.
[0068] The following introduces a text recognition method provided by an embodiment of the present invention with reference to the accompanying drawings. See Figure 1 , Figure 1 is a schematic flowchart of a text recognition method provided by an embodiment of the present invention, and this method may include steps S101 - S104.
[0069] S101. Obtain the current text to be recognized.
[0070] It can be understood that the electronic device can obtain the text to be recognized, and the text to be recognized can be a text containing at least one character. Exemplarily, the electronic device can obtain the text to be recognized from a local database or a database of other devices. For example, the database stores the historical video browsing records of users, and the electronic device can obtain the text under the field of the historical video browsing records in the database, and recognize the obtained text to determine whether there is sensitive information in the historical video browsing records. Another example is that there is a field of mobile phone numbers in the database, and the mobile phone numbers are sensitive information. The electronic device can obtain the text under the field of the mobile phone numbers in the database and recognize the obtained text to determine whether the mobile phone numbers in the database are encrypted.
[0071] S102. Determine whether the current text to be recognized conforms to the current specified rule.
[0072] Among them, the current specified rule includes: the first recognition rule obtained when using a preset large language model to recognize historical texts. The first recognition rule refers to the rule that the text of the specified sensitive text type conforms to.
[0073] It can be understood that the electronic device can determine whether the text to be recognized conforms to the first recognition rule in the specified rule. The current specified rule can include at least one first recognition rule. The first recognition rule in the current specified rule is obtained by the preset large language model recognizing historical texts. The rules in the specified rule can reflect the sensitive information characteristics that the text belonging to the specified sensitive text type conforms to. Optionally, the electronic device locally stores the current specified rule, that is, the electronic device stores a local rule library.
[0074] Among them, the specified sensitive text type can be one or more of user personal privacy information such as name, mobile phone number, ID number, etc. If the specified sensitive text type is one, it is necessary to determine whether the current text to be recognized belongs to the specified sensitive text type; if the specified sensitive text type is multiple, it is necessary to determine whether the current text to be recognized belongs to one of the multiple specified sensitive text types.
[0075] Optionally, when there is no first recognition rule output by the preset large language model in the specified rule, the specified rule can include manually preset recognition rules, and the specified rule can be updated during the recognition process of the preset large language model.
[0076] Since the text specifying the sensitive text type conforms to the first recognition rule in the specified rules, thus, when the text to be recognized does not conform to the first recognition rule in the specified rules, it can be determined that the text to be recognized does not have the characteristics of sensitive information; when the text to be recognized conforms to the first recognition rule in the specified rules, it can be determined that the text to be recognized has the characteristics of sensitive information, that is, it is initially determined that the text to be recognized may belong to the specified sensitive text type, and further confirmation by the pre-set large language model is required later.
[0077] Optionally, in one implementation, the current specified rules include at least one of the following:
[0078] The current text to be recognized is not the specified text;
[0079] The first character of the current text to be recognized is not the specified character;
[0080] The number of characters contained in the current text to be recognized belongs to a preset range.
[0081] It can be understood that the specified text, the specified character, and the preset range can be summarized from the recognition of historical texts that do not belong to the specified sensitive type by the pre-set large language model. The specified text can be a text that does not belong to the specified sensitive text type, the specified character can be the first character of the specified text, or the first character of a text that does not belong to the specified sensitive text type; the preset range can be the range in which the number of characters of a text belonging to the specified sensitive type is located. For example, the specified sensitive text type can be a name, "misunderstanding" has a special semantics and will not be used as a name, the specified text can be "misunderstanding"; "mis" is not used as a surname, the specified character can be "mis"; the names of Chinese people are usually 2 - 4 characters, and the preset range can be 2 - 4 characters. Another example, the specified sensitive text type can be a mobile phone number, "12345678900" will not be used as a mobile phone number, the specified text can be "12345678900"; "2", "3", "4", "5", "6", "7", "8", "9", "0" will not be used as the first character of a mobile phone number either, the specified character can be "2", "3", "4", "5", "6", "7", "8", "9", "0", and a mobile phone number is usually 11 characters, so the preset range is 11 characters.
[0082] When the current text to be recognized meets all the rules in the current specified rules, it is determined that the current text to be recognized conforms to the current specified rules. If there are multiple first recognition rules in the current specified rules, when the current text to be recognized only meets some of the rules in the current specified rules, it cannot be determined that the current text to be recognized conforms to the current specified rules.
[0083] Exemplarily, the first recognition rules included in the current specified rules are: the current text to be recognized is not "misunderstanding", the first character of the current text to be recognized is not "mis", and the number of characters included in the current text to be recognized does not belong to 2 - 4 characters. If the current text to be recognized is "misunderstanding", it is determined that the current text to be recognized does not conform to the current specified rules; if the current text to be recognized is "mis a", it is determined that the current text to be recognized does not conform to the current specified rules; if the current text to be recognized is "jia yi bing ding wu", it is determined that the current text to be recognized does not conform to the current specified rules; if the current text to be recognized is "jia yi", it is determined that the current text to be recognized conforms to the current specified rules.
[0084] In the embodiments of the present invention, by restricting the specified text, specified characters, and preset range, specific specified rules are set, which can exclude most of the texts to be recognized that do not belong to the specified sensitive text type. While accurately screening out the texts that conform to the characteristics of sensitive information, the number of texts to be recognized for subsequent recognition by the preset large language model is reduced, and the efficiency of recognition by the preset large language model is improved.
[0085] S103, if the current text to be recognized conforms to the current specified rules, input the current text to be recognized and the preset prompt word into the preset large language model for recognition.
[0086] Among them, the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type.
[0087] It can be understood that the preset large language model is a deep learning model trained with a large number of sample texts and can handle various natural language tasks, such as text classification, question answering, and dialogue, etc. The structure of the preset large language model includes a Transformer network.
[0088] Since the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type, thus, when the preset large language model receives the current text to be recognized and the preset prompt word, it can recognize the current text to be recognized and output a recognition result indicating that the current text to be recognized belongs to the specified sensitive text type.
[0089] Exemplarily, the preset prompt word can indicate: determine whether the current text to be recognized belongs to the specified sensitive text type. For example, the specified sensitive type can be a name. The text to be recognized input into the preset large language model is "misunderstanding", and the preset large language model can recognize that "misunderstanding" is not a name and "misunderstanding" does not belong to the specified sensitive text type.
[0090] Optionally, in one implementation, there are multiple specified sensitive text types, and the electronic device can pre-store multiple preset prompt words, with each preset prompt word corresponding to a specified sensitive text type. For example, the specified sensitive text types include name and mobile phone number. The name has its corresponding preset prompt word, and the preset prompt word corresponding to the name indicates: determining whether the current text to be recognized belongs to a name; the mobile phone number has its corresponding another preset prompt word, and the preset prompt word corresponding to the mobile phone number indicates: determining whether the current text to be recognized belongs to a mobile phone number.
[0091] It can be understood that for each specified sensitive text type, the electronic device can input the preset prompt word corresponding to the specified sensitive text type and the text to be recognized into the preset large language model to determine whether the text to be recognized belongs to the specified sensitive text type.
[0092] For example, the specified sensitive text types include name and mobile phone number. The electronic device can input the text to be recognized and the preset prompt word corresponding to the name into the preset large language model to obtain a recognition result indicating whether the text to be recognized belongs to a name; input the text to be recognized and the preset prompt word corresponding to the mobile phone number into the preset large language model to obtain a recognition result indicating whether the text to be recognized belongs to a mobile phone number.
[0093] In the embodiments of the present invention, different preset prompt words correspond to different specified sensitive text types. By using the preset prompt words corresponding to the specified sensitive text types to instruct the preset large language model, the recognition result of whether the text to be recognized belongs to the specified sensitive text type can be accurately obtained.
[0094] S104, obtain the recognition result output by the preset large language model.
[0095] It can be understood that the electronic device can obtain the recognition result output by the preset large language model for further processing of the obtained recognition result subsequently. Exemplarily, the recognition result can be stored in a database, or the recognition result can be displayed to the user to inform the user whether the current text to be recognized belongs to the specified sensitive text type.
[0096] In the embodiments of the present invention, since the first recognition rule in the current specified rule is obtained when using a preset large language model to recognize historical texts, and the text of the specified sensitive text type conforms to the first recognition rule, thus, the specified rule can reflect the characteristics of the text that conforms to the specified sensitive text type (collectively referred to as sensitive information characteristics in this article). By determining whether the text to be recognized conforms to the specified rule, the text to be recognized with sensitive information characteristics can be initially screened out, and the text to be recognized with sensitive information characteristics may belong to the specified sensitive text type. Since the preset prompt is used to instruct the preset large language model to output the recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type, thus, inputting the text to be recognized that conforms to the specified rule and the preset prompt into the preset large language model can further recognize whether the text to be recognized with sensitive information characteristics belongs to the specified sensitive text type. The present invention can avoid manual recognition of sensitive information in a large amount of text and improve the efficiency of identifying sensitive information.
[0097] In one embodiment, based on Figure 1 the text recognition method shown, the preset prompt is further used to instruct the second recognition rule used by the preset large language model to obtain the recognition result; as Figure 2 shown, the method further includes steps S201 - S202.
[0098] S201, when the obtained recognition result indicates that the current text to be recognized does not belong to the specified sensitive text type, obtain the second recognition rule used by the preset large language model to obtain the recognition result.
[0099] S202, add the second recognition rule to the current specified rule.
[0100] It can be understood that the preset prompt is not only used to instruct the preset large language model to output the recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type, but also used to instruct the second recognition rule used by the preset large language model to obtain the recognition result.
[0101] Exemplarily, the preset prompt words can indicate: determining whether the current text to be recognized belongs to a specified sensitive text type and giving the recognition rules based on which the determination is made. For example, the specified sensitive text type can be a name, the current text to be recognized is "misunderstanding", and the recognition rule in the specified rules is that the number of characters of the current text to be recognized is between 2 and 4, and the current text to be recognized conforms to the current specified rule; inputting the current text to be recognized and the preset prompt words into the preset large language model for recognition, the preset large language model can analyze that the meaning of the word "misunderstanding" is incorrect understanding, which is a word with a specific meaning, while a name needs to have high recognition and uniqueness, "misunderstanding" does not have the sensitive information characteristics of a name, and "wu" is not a surname. Therefore, "misunderstanding" is not a name. That is to say, the preset large language model can output that "misunderstanding" does not belong to a name, and the recognition rule used is: "misunderstanding" is not a name, and "wu" is not the first character of a name.
[0102] It can be understood that if the preset large language model recognizes that the text to be recognized does not belong to the specified sensitive text type, the second recognition rule used by the preset large language model for recognition can be obtained, and then the second recognition rule can be added to the specified rules to update the specified rules. When recognizing the next text to be recognized, the updated specified rules can be used to screen out the text that was recognized by the preset large language model as not belonging to the specified sensitive text type last time. For example, the current text to be recognized is "misunderstanding", and it can be determined through the preset large language model that the current text to be recognized does not belong to the specified sensitive text type, and the second recognition rule used is: "misunderstanding" is not a name. Then, after adding the second recognition rule to the specified rules; the next text to be recognized is also "misunderstanding". When recognizing the next text to be recognized, "misunderstanding" can be screened out using the updated specified rules, avoiding subsequent recognition of "misunderstanding" by the preset large language model and improving the efficiency of text recognition.
[0103] Exemplarily, the electronic device locally stores a dictionary set, a rule set, an experience set, and an artificial strategy library. An artificial person can supplement rules to the artificial strategy library at any time. During the process of the electronic device adding recognition rules to the specified rules, the specified text that is not of the specified sensitive text type in the recognition rules can be added to the dictionary set, the specified character that is not the first character of the specified sensitive text type can be added to the experience set, and the new recognition rule can be added to the rule set. For example, the recognition rules output by the preset large language model are: "misunderstanding" is not a name and "wu" is not a surname. Then the electronic device can add "misunderstanding" is not a name to the dictionary set and add the first character being "wu" to the experience set.
[0104] Since the pre-set large language model can think in a way similar to the human brain under the instruction of the pre-set prompt words, by leveraging the capabilities of the pre-set large language model, the pre-set large language model can be used to replace manual analysis. When the obtained recognition result indicates that the current text to be recognized does not belong to the specified sensitive text type, adding the second recognition rule output by the pre-set large language model to the current specified rule can incorporate the analysis process of the pre-set large language model into the text recognition method, continuously improve the specified rule, and the text to be recognized can be recognized through the specified rule. Subsequently, only a relatively small number of texts to be recognized need to be input into the pre-set large language model for recognition, which can reduce the usage cost of the pre-set large language model and improve the efficiency and automation of identifying sensitive data while ensuring a high correct rate.
[0105] Optionally, in one embodiment, based on the Figure 1 text recognition method shown, as Figure 3 shown, step S101 includes step S1011, and the method further includes steps S301 - S302.
[0106] S1011, Obtain the text stored in the preset database as the current text to be recognized.
[0107] It can be understood that the preset database can be a local database or a database of other devices outside the electronic device. Obtaining the text stored in the preset database can be obtaining the text under the specified field name in the preset database. For example, the text under the specified field name can be the text under the name field or the text under the historical video viewing information field. The electronic device obtaining the text under the specified field can reduce the number of texts to be recognized and improve the efficiency of identifying sensitive information. Optionally, the electronic device can obtain all the data in the preset database, and the electronic device can determine whether all the data in the preset database belongs to the specified sensitive type to ensure the security of sensitive information.
[0108] Among them, the number of preset databases can be one or more, and the electronic device, as an intermediate device, can manage the data in multiple preset databases. Exemplarily, the electronic device can proxy the preset database to encrypt and delete the text in the preset database.
[0109] S301, When the obtained recognition result indicates that the current text to be recognized belongs to the specified sensitive text type, encrypt the current text to be recognized.
[0110] It can be understood that when the recognition result output by the preset large language model indicates that the current text to be recognized belongs to the specified sensitive text type, the current text to be recognized is sensitive information, and the electronic device can encrypt the current text to be recognized. Exemplarily, the text to be recognized can be encrypted by a preset encryption algorithm. Among them, the encryption algorithm can be an asymmetric encryption algorithm or a symmetric encryption algorithm. Subsequently, if a person with viewing permission needs to view the encrypted result, the encrypted result can be decrypted according to the preset encryption algorithm to obtain the text to be recognized. It should be noted that the encrypted text to be recognized does not have the characteristics of sensitive information, that is, the encrypted result does not conform to the current specified rules, and the recognition result of the preset large language model for the encrypted result indicates that the encrypted result does not belong to the specified sensitive text type.
[0111] S302, delete the current text to be recognized in the preset database, and write the encrypted result of the current text to be recognized to the storage location of the text to be recognized in the preset database to update the preset database.
[0112] It can be understood that after it is determined that the current text to be recognized obtained from the preset database is sensitive information, the current text to be recognized in the preset database can be deleted, and the encrypted result of the current text to be recognized can be written to the storage location where the text to be recognized was originally in the preset database to complete the update of the preset database.
[0113] In the embodiments of the present invention, the text to be recognized as sensitive information in the preset database can be deleted, and the encrypted text to be recognized can be written to the preset database, that is, the sensitive information recognized in the preset database is encrypted to ensure that the sensitive information in the preset database is in an encrypted state and improve the security of the sensitive information.
[0114] Optionally, in one embodiment, on the basis of Figure 3 the text recognition method shown, as Figure 4 shown, the method further includes steps S401 - S402.
[0115] S401, obtain the text type to which the text to be recognized belongs that is pre-recorded as the text type to be processed.
[0116] It can be understood that the text type to which the text to be recognized belongs can be pre-recorded, and the electronic device can obtain the text type to which the text to be recognized belongs as the text type to be processed. Exemplarily, when the electronic device obtains the text to be recognized from the preset database, it can read the table structure information of the data table where the text to be recognized is located, and obtain the field name to which the text to be recognized belongs as the text type to be processed for the text to be recognized. For example, if the text to be recognized is the text in the personal information table in the preset database, the electronic device can obtain the table structure information of the personal information table and obtain the field name to which the text to be recognized belongs as the text type to be processed.
[0117] S402. According to the corresponding relationship between the preset text type and the rule set, determine the rule in the rule set corresponding to the text type to be processed as the current specified rule.
[0118] Among them, the rule set corresponding to a text type includes: the first recognition rule obtained when using the preset large language model to recognize historical texts, and the first recognition rule refers to the rule that the text of this text type conforms to.
[0119] It can be understood that the electronic device locally stores multiple rule sets, and each rule set corresponds to a text type. For example, the text types are name and mobile phone number. The name has its corresponding rule set, and the mobile phone number has its corresponding another rule set. The electronic device can use the rules in the rule set corresponding to the text type to be processed as the current specified rule and make a judgment according to the rules in the rule set corresponding to the text type to be processed. For example, the first recognition rules in the rule set corresponding to the name are: "misunderstanding" is not a name; the first recognition rules in the rule set corresponding to the mobile phone number are: "12345678900" is not a mobile phone number; the electronic device can judge whether the current text to be recognized is "misunderstanding" to determine whether the current text to be recognized conforms to the rules in the rule set corresponding to the name; judge whether the current text to be recognized is "12345678900" to determine whether the current text to be recognized conforms to the rule set corresponding to the mobile phone number.
[0120] Optionally, there are multiple specified sensitive text types, and each specified sensitive text type corresponds to a rule set. In order to recognize whether the text to be recognized belongs to each specified sensitive text type, at this time, the rules in the rule sets corresponding to each specified sensitive text type can be used to determine whether the text to be recognized belongs to one of the multiple specified sensitive text types. The text to be recognized only needs to meet the specified rules in one of the multiple rule sets to determine that the text to be recognized conforms to the current specified rule.
[0121] When the text to be recognized meets one of multiple rule sets, the specified sensitive text type corresponding to the satisfied rule set can be used as the recognition target that the preset large language model needs to recognize, and the preset prompt words corresponding to the specified sensitive text type are input into the preset large language model to obtain the recognition result indicating whether the text to be recognized belongs to the specified sensitive text type and the second recognition rule. For example, if the text to be recognized belongs to a name, the preset prompt words corresponding to the name are input into the preset large language model to obtain the recognition result indicating whether the text to be recognized belongs to the name and the second recognition rule. At the same time, when it is necessary to add the second recognition rule output by the preset large language model to the specified sensitive text type, the second recognition rule can be added to the rule set of the specified sensitive text type, that is, the second recognition rule is added to the rule set corresponding to the recognition target that obtains the second recognition rule.
[0122] In the embodiments of the present invention, different rule sets correspond to different text types. By judging according to the rule set corresponding to the text type to be processed, it can be accurately judged whether the text to be recognized conforms to the specified rules.
[0123] Optionally, in one embodiment, based on Figure 4 the text recognition method shown, as Figure 5 shown, the method further includes steps S501-S502, and step S402 includes step S4021.
[0124] S501, judge whether the text type to be processed is a specified sensitive text type.
[0125] It can be understood that the specified sensitive text type can be a preset text type. For example, names, mobile phone numbers, and ID card numbers can be preset as the specified sensitive text types in advance. If the text type to be processed is one of the above text types, it can be determined that the text type to be processed is a specified sensitive text type.
[0126] S4021, when the current text type to be processed is a specified sensitive text type, determine the rules in the rule set corresponding to the text type to be processed according to the corresponding relationship between the preset text type and the rule set as the current specified rules.
[0127] It can be understood that if the text type to be processed is a specified sensitive text type, the rules in the rule set corresponding to the text type to be processed can be directly used for judgment.
[0128] S502, when the text type to be processed is not a specified sensitive text type, perform word segmentation on the text to be recognized, and use the word segmentation result whose text type is a specified sensitive text type obtained by word segmentation as the current text to be recognized. Then execute step S4021.
[0129] It can be understood that if the text type to be processed is not the specified sensitive text type, the text to be recognized may still contain text of the specified sensitive text type, and the text to be recognized may be text under a field that does not belong to sensitive information. Exemplarily, the text to be recognized may be a text containing several characters. For example, if the specified sensitive text type is name, and the text to be recognized is under the field of chat record or note content in the preset database, rather than the field of name, then the text type of the text to be recognized is not the specified sensitive text type, and the text to be recognized is a text containing several characters. Among them, the several characters may be Chinese characters, or characters combined with Chinese and numbers. For this, the embodiments of the present invention only give examples and do not make specific limitations.
[0130] The electronic device can use a preset word segmentation tool to segment the text to be recognized, divide a text containing several Chinese characters into multiple word segmentation results, and identify whether the text type of each word segmentation result is the specified sensitive text type. Then, use the word segmentation results whose text type is the specified sensitive text type obtained by word segmentation as the current text to be recognized, return to step S4021, determine the rules in the rule set corresponding to each specified sensitive text type as the current specified rules, and judge whether the word segmentation results belonging to the specified sensitive text type conform to the current specified rules; if they conform, input the word segmentation results and the preset prompt words into the preset large language model to obtain the recognition result and the second recognition rule. Among them, the preset word segmentation tool can be a tool designed based on natural language technology, which can divide the words in a text and label the word types.
[0131] In the embodiments of the present invention, it is possible to segment a text containing several characters, judge whether the word segmentation results belonging to the specified sensitive text type conform to the rules, and perform recognition by the preset large language model, and segment the text containing sensitive information to identify the sensitive information hidden in a text, so as to improve the security of sensitive information.
[0132] Optionally, in one embodiment, based on the Figure 3 text recognition method shown, if there are multiple current texts to be recognized, as Figure 6 shown, the method further includes steps S601 - S602.
[0133] S601, obtain a specified number of texts from the updated preset database as the current texts to be verified.
[0134] It can be understood that after the electronic device finishes identifying each text to be recognized, it can be considered that the text in the preset database has been recognized and encrypted once, and the preset database has been updated. The text to be verified is the text in the database after the above encryption process, that is, the text in the updated preset database.
[0135] However, during the recognition process, the rules in the specified rules are constantly increasing and the specified rules are constantly being improved; the specified rules used in the early stage of recognition are not perfect enough, and there may be sensitive information that is missed in the text screened out. To ensure the security of sensitive information in the preset database, the data in the updated preset database can be verified to determine whether there is text of the specified sensitive text type that has not been encrypted in the updated preset database. The embodiment of the present invention adopts a sampling inspection method to obtain a specified number of texts from the updated preset database as the current texts to be verified. Among them, the specified number of texts can be used as sampling examples, and the specified number can be at least one. Optionally, the specified number can also be the number of all data in the preset database, that is, each data in the preset database is checked.
[0136] S602, if at least one of the current texts to be verified conforms to the current specified rule, then execute step S1011 until none of the current texts to be verified conforms to the current specified rule.
[0137] It can be understood that if at least one of the current texts to be verified conforms to the current specified rule, it can be considered that at least one of the sampling examples may be a text of the specified sensitive text type, that is to say, there is text of the specified sensitive text type in the preset database. At this time, the electronic device will return to step S1011, continue to obtain the text in the preset database, recognize the obtained text, and encrypt the recognized text of the specified sensitive text type.
[0138] Since the text belonging to the specified sensitive text type after encryption is ciphertext and does not have the sensitive information characteristics of the text belonging to the specified sensitive text type, that is to say, the text belonging to the specified sensitive text type after encryption does not conform to the specified rule; thus, if none of the current texts to be verified conforms to the current specified rule, it can be considered that none of the sampling examples belongs to the specified sensitive text type, that is to say, the sampling examples may be texts that are not encrypted and do not belong to the specified sensitive text type, or may be texts that belong to the specified sensitive text type after encryption. When none of the sampling examples conforms to the current specified rule, it can be determined that there is no text of the specified sensitive text type in the preset database, and the encryption of the text of the specified sensitive text type in the preset database is completed.
[0139] In the embodiments of the present invention, sensitive information in a preset database is encrypted to obtain an updated preset database, and the data in the updated preset database is verified to determine whether there is any text of a specified sensitive text type that is missing in the updated preset database. If so, the text identifying the specified sensitive text type and the process of encrypting the text of the specified sensitive text type are returned, which can reduce the unencrypted sensitive information in the preset database and improve the security of sensitive information.
[0140] To more clearly understand the text recognition method provided by the embodiments of the present invention, the following is an introduction in conjunction with the schematic diagram. As Figure 7 shown, the electronic device may include a policy fusion model and a large model comprehensive analysis module.
[0141] The electronic device can obtain the current text to be recognized, where the text to be recognized can be a large amount of data in a preset database. Then, the electronic device can determine whether the current text to be recognized conforms to the current specified rules through the policy fusion module for a round of screening.
[0142] Among them, the specified rules are locally stored in the electronic device, and the specified rules include at least one of the following: the current text to be recognized is not the specified text; the first character of the current text to be recognized is not the specified character; the number of characters included in the current text to be recognized belongs to a preset range. The specified text can be the text in the dictionary set locally stored in the electronic device, the specified character can be the character in the experience set locally stored in the electronic device, and the recognition rule that the number of characters included in the text to be recognized belongs to a preset range can be the recognition rule in the rule set locally stored in the electronic device. In addition, the operator can supplement the policies to the local artificial policy library of the electronic device. For example, the operator can set the specified text. If the current text to be recognized is not the specified text, it can be determined that the current text to be recognized may belong to the text of the specified sensitive text type. Among them, in addition to applying the rule set, the experience set, the dictionary set, and the artificial policy library, the policy fusion model can also include other functions, such as using a preset word segmentation tool to segment a piece of text.
[0143] If the current text to be recognized conforms to the current specified rules, the electronic device can input the current text to be recognized and the preset prompt word into a preset large language model for recognition, that is, deliver a large amount of data for large model analysis.
[0144] The large model comprehensive analysis module is used to identify the current text to be identified through at least one preset large language model, obtain the identification result output by the preset large language model, and the second identification rule used by the preset large language model to obtain the identification result (i.e., obtain the result). For example, the current text to be identified can be identified through large model 1, large model 2, large model 3, …, to obtain the identification result. That is, two-round screening.
[0145] In the case that the obtained identification result indicates that the current text to be inspected does not belong to the specified sensitive text type, add the second identification rule output by the preset large language model to the current specified rule, that is, extract features according to the result to improve the strategy.
[0146] Figure 8 It is a schematic flowchart of another text recognition method provided by an embodiment of the present invention.
[0147] Such as Figure 8 As shown, during the process of storing source data in the database, the electronic device can obtain the text stored in the preset database as the current text to be identified; the preset database can be at least one database storing sensitive information, and the source data is the text in the preset database. This process can also be called the scanning stage.
[0148] During the process of data recognition by combining the local rule library with the large model, the electronic device can determine whether the current text to be identified conforms to the current specified rule; if the current text to be identified conforms to the current specified rule, input the current text to be identified and the preset prompt word into the preset large language model for identification; obtain the identification result output by the preset large language model, and the second identification rule used by the preset large language model to obtain the identification result. This process can be called the discovery stage.
[0149] During the encryption process of the electronic device proxy database, the electronic device can encrypt the current text to be identified in the case that the obtained identification result indicates that the current text to be processed belongs to the specified sensitive text type; and delete the current text to be identified in the preset database, and write the encryption result of the current text to be identified to the storage location of the text to be identified in the preset database to update the preset database. Specifically, the electronic device, as the proxy middleware of the preset database, manages the preset database, and can write each encryption result to the original storage location of the encryption result through the data double-writing method, and delete the text to be identified at the original storage location, so as to realize switching the text to be identified at the original location to an encrypted column. This process can be called the encryption stage.
[0150] During the data security identification and verification process, after the electronic device finishes identifying each text to be identified currently, it can obtain a specified number of texts from the updated preset database as the current texts to be verified; if each of the current texts to be verified does not conform to the current specified rules, it is determined that there is no text of the specified sensitive text type in the preset database; if at least one of the current texts to be verified conforms to the current specified rules, it returns to the discovery stage and the encryption stage until each of the current texts to be verified does not conform to the current specified rules. This process can be called the verification stage.
[0151] In the embodiments of the present invention, the judgment process of the specified rules and the identification process of the preset large language model are combined to form a system, which cooperate with each other to carry out data security work, avoid manually screening out sensitive information, liberate a large amount of manpower in data security governance, improve the identification efficiency and accuracy at the same time, and carry out data security work better and faster. Moreover, for the protection of sensitive information, it can meet the compliance requirements and security requirements of the business.
[0152] The embodiments of the present invention also provide a text recognition device, as Figure 9 shown, the device includes:
[0153] A first acquisition module 910, configured to acquire the current text to be identified;
[0154] A first judgment module 920, configured to judge whether the current text to be identified conforms to the current specified rules; wherein, the current specified rules include: the first identification rules obtained when using the preset large language model to identify historical texts, and the first identification rules refer to the rules that the texts of the specified sensitive text type conform to;
[0155] An input module 930, configured to input the current text to be identified and the preset prompt words into the preset large language model for identification if the current text to be identified conforms to the current specified rules; wherein, the preset prompt words are used to instruct the preset large language model to output an identification result indicating whether the current text to be identified belongs to the specified sensitive text type;
[0156] A second acquisition module 940, configured to acquire the identification result output by the preset large language model.
[0157] Optionally, the preset prompt words are further used to instruct the preset large language model to output the second identification rules used when obtaining the identification result;
[0158] The device further includes:
[0159] A third acquisition module, configured to, when the acquired recognition result indicates that the current text to be recognized does not belong to the specified sensitive text type, acquire the second recognition rule used by the preset large language model to obtain the recognition result;
[0160] An addition module, configured to add the second recognition rule to the current specified rule
[0161] Optionally, the first acquisition module 910 is specifically configured to acquire the text stored in the preset database as the current text to be recognized;
[0162] The device further includes:
[0163] An encryption module, configured to encrypt the current text to be recognized when the acquired recognition result indicates that the current text to be recognized belongs to the specified sensitive text type;
[0164] A writing module, configured to delete the current text to be recognized in the preset database and write the encryption result of the current text to be recognized to the storage location of the text to be recognized in the preset database to update the preset database.
[0165] Optionally, the device further includes:
[0166] A fourth acquisition module, configured to acquire the text type to which the text to be recognized belongs as recorded in advance as the text type to be processed;
[0167] A first determination module, configured to determine, according to the corresponding relationship between the preset text type and the rule set, the rule in the rule set corresponding to the text type to be processed as the current specified rule; wherein, the rule set corresponding to a text type includes: the first recognition rule obtained when using the preset large language model to recognize historical texts, and the first recognition rule refers to the rule that the text of this text type conforms to.
[0168] Optionally, the device further includes:
[0169] A second judgment module, configured to judge whether the text type to be processed is the specified sensitive text type;
[0170] The first determination module is specifically configured to, when the current text type to be processed is the specified sensitive text type, determine, according to the corresponding relationship between the preset text type and the rule set, the rule in the rule set corresponding to the text type to be processed as the current specified rule;
[0171] The device further includes:
[0172] A word segmentation module, configured to segment the text to be recognized when the type of the text to be processed is not the specified sensitive text type, and use the word segmentation result whose text type is the specified sensitive text type obtained by word segmentation as the current text to be recognized, and return to execute the step of determining, according to the corresponding relationship between the preset text type and the rule set, the rule in the rule set corresponding to the type of the text to be processed as the current specified rule when the current type of the text to be processed is the specified sensitive text type.
[0173] Optionally, there are multiple current texts to be recognized;
[0174] The device further includes:
[0175] A fifth acquisition module, configured to acquire a specified number of texts from the updated preset database as the current texts to be verified;
[0176] A verification module, configured to, if at least one of the current texts to be verified conforms to the current specified rule, return to execute the step of acquiring the texts stored in the preset database as the current texts to be recognized until none of the current texts to be verified conforms to the current specified rule.
[0177] Optionally, the current specified rule includes at least one of the following:
[0178] The current text to be recognized is not the specified text;
[0179] The first character of the current text to be recognized is not the specified character;
[0180] The number of characters included in the current text to be recognized belongs to a preset range.
[0181] In the embodiments of the present invention, since the first recognition rule in the current specified rule is obtained when using a preset large language model to recognize historical texts, and the text of the specified sensitive text type conforms to the first recognition rule, thus, the specified rule can reflect the characteristics of the text conforming to the specified sensitive text type (collectively referred to as sensitive information characteristics in this article). By determining whether the text to be recognized conforms to the specified rule, the text to be recognized with sensitive information characteristics can be initially screened out, and the text to be recognized with sensitive information characteristics may belong to the specified sensitive text type. Since the preset prompt is used to instruct the preset large language model to output the recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type, thus, inputting the text to be recognized conforming to the specified rule and the preset prompt into the preset large language model can further recognize whether the text to be recognized with sensitive information characteristics belongs to the specified sensitive text type. The present invention can avoid manual recognition of sensitive information in a large number of texts and improve the efficiency of recognizing sensitive information.
[0182] An embodiment of the present invention further provides an electronic device, such as Figure 10 As shown, it includes a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete mutual communication through the communication bus 1004.
[0183] The memory 1003 is used to store a computer program.
[0184] When the processor 1001 is used to execute the program stored on the memory 1003, the above-mentioned text recognition method is implemented.
[0185] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0186] The communication interface is used for communication between the above terminal and other devices.
[0187] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0188] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0189] In another embodiment provided by the present invention, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the text recognition method described in any of the above embodiments is implemented.
[0190] In another embodiment provided by the present invention, a computer program product containing instructions is further provided. When it runs on a computer, the computer is caused to execute the text recognition method described in any of the above embodiments.
[0191] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0192] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0193] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0194] The above description is only for the preferred embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A text recognition method, characterized in that: The method comprises: Get the current text to be recognized; Determine whether the current text to be recognized complies with the current specified rules; wherein the current specified rules include: a first recognition rule obtained when the historical text is recognized using a preset large language model, wherein the first recognition rule refers to a rule that the text of the specified sensitive text type complies with; If the current text to be recognized meets the current specified rule, the current text to be recognized and the preset prompt word are input into the preset large language model for recognition; wherein the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type; Obtain a recognition result output by the preset large language model.
2. The method according to claim 1, characterized in that The preset prompt word is also used to indicate a second recognition rule used when the preset large language model outputs the recognition result; The method further comprises: When the obtained recognition result indicates that the current text to be recognized does not belong to the specified sensitive text type, obtaining a second recognition rule used by the preset large language model to obtain the recognition result; The second identification rule is added to the current specified rules.
3. The method according to claim 2, characterized in that The step of obtaining the current text to be recognized includes: Obtain the text stored in the preset database as the current text to be recognized; The method further comprises: When the obtained recognition result indicates that the current text to be recognized belongs to the specified sensitive text type, encrypting the current text to be recognized; The current text to be recognized in the preset database is deleted, and the encryption result of the current text to be recognized is written into the storage location of the text to be recognized in the preset database to update the preset database.
4. The method according to claim 3, characterized in that Before determining whether the current text to be recognized conforms to the current specified rule, the method further includes: Acquire the pre-recorded text type to which the to-be-recognized text belongs as the to-be-processed text type; According to the correspondence between the preset text type and the rule set, the rule in the rule set corresponding to the text type to be processed is determined as the current designated rule; wherein the rule set corresponding to a text type includes: a first recognition rule obtained when the historical text is recognized using the preset large language model, and the first recognition rule refers to the rule that the text of the text type conforms to.
5. The method according to claim 4, characterized in that Before determining the rule in the rule set corresponding to the to-be-processed text type as the current designated rule according to the preset correspondence between the text type and the rule set, the method further includes: Determine whether the type of the text to be processed is the specified sensitive text type; The step of determining a rule in the rule set corresponding to the text type to be processed as the current designated rule according to the preset correspondence between the text type and the rule set includes: In the case where the current type of text to be processed is the designated sensitive text type, according to the preset correspondence between the text type and the rule set, a rule in the rule set corresponding to the type of text to be processed is determined as the current designated rule; The method also includes: when the type of the text to be processed is not the specified sensitive text type, segmenting the text to be recognized, and taking the segmentation result of the text type obtained by segmentation as the specified sensitive text type as the current text to be recognized, and returning to execute the step of determining, when the current type of the text to be processed is the specified sensitive text type, a rule in the rule set corresponding to the text type to be processed according to the correspondence between the preset text type and the rule set, as the current specified rule step.
6. The method according to any one of claims 3 to 5, characterized in that: There are multiple texts to be recognized at present; The method further comprises: Obtain a specified number of texts from the updated preset database as the current texts to be verified; If at least one of the current texts to be verified meets the current specified rule, the process returns to the step of obtaining the text stored in the preset database as the current text to be recognized, until none of the current texts to be verified meets the current specified rule.
7. The method according to any one of claims 3 to 5, characterized in that: The current specified rule includes at least one of the following: the current text to be recognized is not the specified text; The first character of the current text to be recognized is not the specified character; The number of characters contained in the current text to be recognized belongs to the preset range.
8. A text recognition device, characterized in that: The device comprises: The first acquisition module is used to acquire the current text to be recognized; A first judgment module is used to judge whether the current text to be recognized conforms to the current specified rule; wherein the current specified rule includes: a first recognition rule obtained by using a preset large language model to recognize the historical text, wherein the first recognition rule refers to a rule that the text of the specified sensitive text type conforms to; An input module, used for inputting the current text to be recognized and a preset prompt word into the preset large language model for recognition if the current text to be recognized meets the current specified rule; wherein the preset prompt word is used to instruct the preset large language model to output a recognition result indicating whether the current text to be recognized belongs to the specified sensitive text type; The second acquisition module is used to obtain the recognition result output by the preset large language model.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-7 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when being executed by a processor.