Data processing method and device, electronic equipment and storage medium
By replacing and restoring sensitive information input from the public network model locally, the problem of information leakage when the public network model uses the local database is solved, and information security and high-quality task execution is achieved.
Patent Information
- Application Number
- CN202510120866.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-20
AI Technical Summary
When using a public network model to perform AIGC tasks, if the local database is called, it will cause confidential information in the local resources to be leaked, causing information security issues.
By replacing sensitive information from the data input to the public network model locally, the desensitized data is generated and processed, and sensitive information is restored before the real data is finally output.
It ensures that local sensitive information will not be leaked when performing tasks on the public network model, and at the same time, it ensures that the data generated by the public network model meets user needs. The entire process is transparent to users, has a good user experience, and ensures the information security of local data.
Smart Images

Figure CN120179770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] When using a public network large model to perform AIGC (Artificial Intelligence Generated Content) tasks, if a local database needs to be called, the confidential information in the local resources will be leaked, causing information security problems for users. Summary of the Invention
[0003] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium to at least solve the above technical problems existing in the prior art.
[0004] In a first aspect of the present disclosure, a data processing method is provided, and the method includes:
[0005] Obtain a first prompt word to be input into the public network large model;
[0006] Retrieve at least one corresponding first text data from at least one local text library according to the first prompt word;
[0007] Identify and replace sensitive information in each of the first text data to obtain corresponding second text data;
[0008] Generate a second prompt word according to all the second text data and the first prompt word, and input the second prompt word into the public network large model;
[0009] Obtain third text data generated by the public network large model according to the second prompt word, perform sensitive information restoration on the third text data to obtain fourth text data, and output the fourth text data.
[0010] Among them, the replacement of sensitive information in the first text data includes:
[0011] Determine a sensitive information library corresponding to the local text library to which the first text data belongs, where the sensitive information library includes sensitive information and corresponding replacement rules;
[0012] Determine the sensitive information included in the first text data according to the sensitive information library, and perform sensitive information replacement on the first text data according to the replacement rules corresponding to the sensitive information.
[0013] Among them, the replacement rules corresponding to the sensitive information include one or more of the following:
[0014] Replace the sensitive information with the corresponding set information;
[0015] Obtain similar information of the sensitive information from the non-sensitive information library, and replace the sensitive information with the similar information;
[0016] Replace the sensitive information according to the set replacement algorithm.
[0017] Wherein, when replacing the sensitive information in the first text data, the method further includes: recording the sensitive information and the corresponding replacement information;
[0018] Correspondingly, the restoring the sensitive information in the third text data to obtain the fourth text data includes: restoring the replacement information in the third text data to the sensitive information according to the recorded sensitive information and the corresponding replacement information to obtain the fourth text data.
[0019] Wherein, the method further includes: constructing the sensitive information library corresponding to the local text library, including:
[0020] Obtain the new text data of the local text library, and determine the weight of each word in the new text data;
[0021] Determine the words with weights meeting the set threshold as suspected sensitive information;
[0022] According to the review result of the suspected sensitive information, determine whether the suspected sensitive information is sensitive information;
[0023] Construct the replacement rule of the sensitive information, and add the sensitive information and the corresponding replacement rule to the sensitive information library.
[0024] Wherein, the retrieving in at least one local text library according to the first prompt word includes:
[0025] Perform intention recognition on the first prompt word, determine the corresponding local text library according to the recognized intention type, and retrieve in the determined local text library according to the first prompt word.
[0026] Wherein, the performing intention recognition on the first prompt word includes:
[0027] Obtain the feature vector of each word in the first prompt word;
[0028] Determine the probability of the word under the set intention type according to the feature vector;
[0029] Determine the probability of the first prompt word under the set intention type according to the probability corresponding to each word in the first prompt word;
[0030] Determine the intent type of the first prompt word according to the probability of the first prompt word under each set intent type.
[0031] A second aspect of the present disclosure provides a data processing device, which is applied to a local electronic device and includes:
[0032] An interaction module, configured to obtain a first prompt word to be input into a public network large model;
[0033] A retrieval module, configured to retrieve at least one corresponding first text data from at least one local text library according to the first prompt word;
[0034] A desensitization module, configured to identify and replace sensitive information in each of the first text data to obtain corresponding second text data;
[0035] The interaction module is further configured to generate a second prompt word according to all the second text data and the first prompt word, input the second prompt word into the public network large model, and obtain third text data generated by the public network large model according to the second prompt word;
[0036] The desensitization module is further configured to restore sensitive information in the third text data to obtain fourth text data;
[0037] The interaction module is further configured to output the fourth text data.
[0038] A third aspect of the present disclosure provides an electronic device, including
[0039] A processor;
[0040] A memory for storing executable instructions of the processor;
[0041] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the data processing method described above.
[0042] A fourth aspect of the present disclosure provides a computer-readable storage medium, where the storage medium stores a computer program, and the computer program is used to execute the data processing method described above.
[0043] In the above solution, all data is processed locally before the final prompt is input into the public network large model. In particular, sensitive information is desensitized. Therefore, when the public network large model performs tasks, even if local data is used, it will not cause the leakage of sensitive information. Finally, the sensitive information in the data generated by the public network large model is restored to obtain real text data that meets the user's needs. The entire process is transparent to the user. The user only needs to input information once, and the user experience is good. It also enables the public network large model to use sufficient data to perform tasks, ensuring the quality of task execution and the information security of local data. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 FIG. shows a schematic flowchart of a data processing method according to an example of the present disclosure;
[0045] Figure 2 FIG. shows a schematic flowchart of a sensitive information recognition process according to an example of the present disclosure;
[0046] Figure 3 FIG. shows a data processing process according to another example of the present disclosure;
[0047] Figure 4 FIG. shows a schematic structural diagram of a data processing device according to an example of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0049] To prevent the problem of local sensitive information leakage when using the public network large model to perform tasks, sensitive information can be desensitized locally and then provided to the public network large model for interaction so that the public network large model can perform related tasks. For this purpose, as Figure 1 shown, the present disclosure provides a data processing method, including:
[0050] Step 101, obtain a first prompt for input into the public network large model.
[0051] The first prompt is used to guide the public network large model to perform corresponding tasks. The tasks that the large model can perform include but are not limited to: text classification, sentiment analysis, information extraction, text generation, machine translation, etc.
[0052] In an example of this solution, a service can be deployed locally to execute the processes of this solution. Among them, the local service can provide an interaction interface to users for obtaining user input information and generating a first prompt word according to the user input information. The local service can determine the content of the prompt word required by the public network large model according to the type of the public network large model, and guide the user to input relevant information through the interaction interface.
[0053] Step 102: Retrieve corresponding at least one first text data from at least one local text library according to the first prompt word.
[0054] In an example of this disclosure, some data required for the public network large model to execute the corresponding task indicated by the first prompt word is saved locally. Therefore, it is necessary to retrieve locally according to the first prompt word to obtain relevant data.
[0055] In one example, multiple local text libraries can be maintained, and each text library includes multiple texts. Retrieve from the local text library according to the first prompt word to obtain corresponding one or more texts. For the convenience of description, each retrieved text is recorded as the first text data.
[0056] The methods for retrieving based on the first prompt word in the local text library include but are not limited to: using the query function of a database management system (DBMS), through SQL query statements, specific text data in the database can be quickly located and extracted; based on natural language processing (NLP) technology, NLP can convert the first prompt word into structured data, and then perform retrieval based on text retrieval algorithms such as the vector space model (VSM) and TF-IDF model.
[0057] Step 103: Identify and replace sensitive information in each first text data to obtain corresponding second text data.
[0058] Since the first text data is local data, it may contain sensitive information. Therefore, it is necessary to identify sensitive information in each first text data and replace the identified sensitive information. After the sensitive information in the first text data is replaced, it becomes the second text data.
[0059] Step 104: Generate a second prompt word according to all the second text data and the first prompt word, and input the second prompt word into the public network large model.
[0060] Merge all the second text data and the first prompt word to generate a second prompt word. This second prompt word is the final prompt word input to the public network large model. Since the sensitive information in the local text data it contains has been replaced and will not cause information leakage, the local service can directly input the second prompt word to the public network large model to execute the corresponding task.
[0061] Step 105: Obtain the third text data generated by the public network large model according to the second prompt word, restore the sensitive information in the third text data to obtain the fourth text data, and output the fourth text data.
[0062] The public network large model generates the third text data according to the second prompt word, and the local service obtains this third text data. Since some information in the third text data was desensitized in Step 103 and is not actual data, therefore, this part of the data needs to be restored with sensitive information to obtain the fourth text data that meets the user's needs.
[0063] In the above solution, before the final prompt word is input into the public network large model, all data is processed locally. In particular, sensitive information is desensitized. Therefore, when the public network large model executes tasks, even if local data is used, it will not cause the leakage of sensitive information. Finally, the sensitive information in the data generated by the public network large model is restored to obtain the real text data that meets the user's needs. The whole process is transparent to the user. The user only needs to input information once, and the user experience is better. It also enables the public network large model to use sufficient data to execute tasks, ensuring the quality of task execution and at the same time ensuring the information security of local data.
[0064] Next, the process of replacing sensitive information in the first text data in the present disclosure example will be described in detail.
[0065] First, introduce the local text library: The user can maintain multiple text libraries locally. Each text library contains multiple texts, and each text library has a corresponding type, which can be determined according to the user's needs. For example, a text library corresponding to product information, and each text contained in this text library records product information; another example is a text library corresponding to procurement information, and each text contained in this text library records the corresponding procurement information. According to needs, the user can create various types of text libraries, and the present disclosure does not limit this. During use, operations such as adding, deleting, and modifying files in the local text library can be performed.
[0066] The content of the texts saved in different types of local text libraries has its own characteristics. It is necessary to construct a corresponding sensitive information library for the local text library. The sensitive information library records the sensitive information and replacement rules in the corresponding text library.
[0067] When constructing the sensitive information library corresponding to the local text library, some sensitive information and corresponding replacement rules can be initialized. When a new text is added to the local text library, the sensitive information in this text can be identified and the replacement rules corresponding to this sensitive information can be created.
[0068] In one example, when a new text is added to the local text library, the process of identifying sensitive information, such as Figure 2 is as follows:
[0069] Step 201: Obtain the newly added text data of the local text library and determine the weight of each word in the newly added text data.
[0070] Specifically, segment the newly added text, and then determine the weight of each word. In this example, the weight of each word in the newly added text can be calculated based on the TF-IDF algorithm:
[0071] First, calculate the term frequency (TF) of each word in the newly added text: For any word in the newly added text, count the number of times the word appears in the newly added text and the total number of words in the newly added text. The ratio of the two is the term frequency of the word.
[0072] Then, calculate the inverse document frequency (IDF) of each word in the newly added text: For any word in the newly added text, count the total number of texts in the text library (all texts before the newly added text) and the number of texts containing the word. Calculate the inverse document frequency of the word through the following formula:
[0073]
[0074] Finally, multiply the term frequency and the inverse document frequency of the word to obtain the weight of the word.
[0075] Step 202: Determine the words with weights meeting the set threshold as suspected sensitive information.
[0076] In this example, a set threshold can be configured, and the words with weights greater than or equal to the set threshold are determined as suspected sensitive information.
[0077] Step 203: Determine whether the suspected sensitive information is sensitive information according to the review result of the suspected sensitive information.
[0078] To ensure the accuracy of sensitive information identification, in the example of the present disclosure, the suspected sensitive information can be provided to the user for review through an interactive interface, and the corresponding review result can be obtained, so as to determine which words are the final sensitive information.
[0079] Step 204: Construct a replacement rule for sensitive information, and add the sensitive information and the corresponding replacement rule to the sensitive information library.
[0080] For the modified text in the local text library, the modified text can be regarded as the newly added text, and the above-mentioned process of identifying sensitive information is also executed.
[0081] For each word identified as sensitive information, a replacement rule for the word can be constructed. Then, the sensitive information and the corresponding replacement rule are added to the sensitive information library. When a word in the newly added text is identified as sensitive information, if the sensitive information is already recorded in the sensitive information library, it is possible to continue to compare whether the new replacement rule for the sensitive information is the same as the old replacement rule. If they are different, the new replacement rule can be used to overwrite the old replacement rule, or the new replacement rule can be discarded and the old replacement rule can be used, or both the old replacement rule and the new replacement rule can be retained. Specifically, it can be set according to requirements. If it is set to use the old replacement rule, the step of constructing the replacement rule for the existing sensitive information can be skipped; if the sensitive information is not recorded in the sensitive information library, the sensitive information and the corresponding replacement rule are directly added to the sensitive information library.
[0082] Furthermore, for the convenience of subsequent use, a text identifier to which the sensitive information belongs can also be added to the sensitive information. When a word in the newly added text is identified as sensitive information, if the sensitive information is already recorded in the sensitive information library, the text identifier of the newly added text can be added to the sensitive information. At the same time, if the replacement rule constructed for the sensitive information this time is different from the existing replacement rule, the new replacement rule can be used to overwrite the old replacement rule, or the new replacement rule can be discarded and the old replacement rule can be used, or both the old replacement rule and the new replacement rule can be retained. Specifically, it can be set according to requirements. If it is set to use the old replacement rule, the step of constructing the replacement rule for the existing sensitive information can be skipped; if the sensitive information is not recorded in the sensitive information library, the sensitive information and the corresponding replacement rule are directly added to the sensitive information library, and the text identifier to which the sensitive information belongs is marked for the sensitive information.
[0083] In the examples of the present disclosure, the replacement rules for sensitive information include one or more of the following:
[0084] 1. Replace the sensitive information with the corresponding set information.
[0085] For example, for the model number, price, etc. of a product, a fixed information can be set for replacement, such as replacing it with a string of meaningless characters: XXXX, 00000, etc.
[0086] 2. Obtain the similar information of the sensitive information from the non-sensitive information library and replace the sensitive information with the similar information.
[0087] A thesaurus of non-sensitive information can be maintained, and for sensitive information, the non-sensitive information with the highest similarity can be selected from the non-sensitive information library for replacement.
[0088] 3. Replace the sensitive information according to the set replacement algorithm.
[0089] The replacement algorithm can be an encryption algorithm, which is used to perform encryption operations on sensitive information to obtain replacement information; the replacement algorithm can also be a series of rules for word replacement, such as: synonym replacement, antonym replacement, hypernym and hyponym replacement, etc.
[0090] In addition, since sensitive information is usually time-sensitive, a lifecycle can also be set for the sensitive information in the sensitive information library, and the expired sensitive information is deleted from the sensitive information library. Similarly, the replacement rules for sensitive information can be updated regularly.
[0091] For the words deleted from the sensitive information library, they can be added to the above-mentioned non-sensitive information library. Correspondingly, if the words added to the sensitive information library also exist in the non-sensitive information library, they need to be deleted from the non-sensitive information library. In addition, when identifying sensitive information, the information identified as non-sensitive can also be added to the non-sensitive information library, and the words in the non-sensitive information library can also come from the network. The present disclosure does not limit the construction of the non-sensitive information library.
[0092] Based on the sensitive information library constructed in the above manner, the process of replacing sensitive information in the first text data in the present disclosure includes:
[0093] Determine the sensitive information library corresponding to the local text library to which the first text data belongs. The sensitive information library includes sensitive information and corresponding replacement rules; determine the sensitive information included in the first text data according to the sensitive information library, and perform sensitive information replacement on the first text data according to the replacement rules corresponding to the sensitive information.
[0094] For each retrieved first text data, first determine the local text library to which it belongs, and then search for the sensitive information included in the first text data from the sensitive information library corresponding to the local text library. As can be seen from the foregoing content, sensitive information identification has been performed when the first text data is added to the local text library. Therefore, by comparing each word in the first text data with each word in the sensitive information library, the sensitive information included in the first text data can be determined. Further, if the sensitive information in the sensitive information library has a corresponding text identifier, then by using the text identifier of the first text data as an index for retrieval, the sensitive information included in the first text data can also be determined.
[0095] After finding the sensitive information included in the first text data, corresponding replacement information can be generated according to its corresponding replacement rules for replacement. For subsequent restoration of sensitive information, the sensitive information and the corresponding replacement information can be saved at this time. Correspondingly, after obtaining the third text data generated by the public network large model, according to the sensitive information and the corresponding replacement information, the replacement information in the third text data is restored to sensitive information to obtain the fourth text data output to the user.
[0096] To further ensure information security, the first prompt word can also be replaced with sensitive information. The sensitive information included in the first prompt word can be determined according to the sensitive information library corresponding to the local text library to which the retrieved first text data belongs. Specifically: the key information in the first prompt word can be obtained, and the key information can be compared with each word in the sensitive information library to determine the sensitive information included in the first prompt word, so as to use the replacement rule corresponding to the sensitive information. Of course, after comparison, there may be a situation where the first prompt word does not contain sensitive information.
[0097] In an example of the present disclosure, after obtaining the first prompt word, it can be retrieved in one or more specified local text libraries. This method is applicable to the situation where the type of the local text library can be directly determined according to the key information of the first prompt word. For example, according to the first prompt word: Please generate a promotional copy for notebook products, one of the key information in the first prompt word is notebook products, then the local text library corresponding to its product type can be directly determined; in another example of the present disclosure, when the type of the local text library cannot be determined according to the key information of the first prompt word, the first prompt word can be subjected to intention recognition, and the corresponding local text library can be determined according to the recognized intention type, and the first prompt word can be retrieved in the determined local text library. Here, the intention type is the type of the local text library.
[0098] The process of intention recognition of the first prompt word in the example of the present disclosure includes:
[0099] Step 301, obtain the feature vector of each word in the first prompt word.
[0100] In the example of the present disclosure, the generation of the word feature vector of the first prompt word can be realized by the Word2Vec algorithm. Among them, the Word2Vec algorithm can be implemented by using the CBOW model or the Skip-gram model. In addition to the CBOW model or the Skip-gram model, the feature vector of each word in the first prompt word can also be determined by the BERT (Bidirectional Encoder Representations from Transformers) model, the GloVe (Global Vectors for Word Representation) model, etc.
[0101] Since the extraction of sensitive information depends on the text to which it belongs and is related to its context, correspondingly, the generation of the feature vector of the first prompt word can adopt the above-mentioned Word2Vec algorithm for calculating word feature vectors based on context. Taking the CBOW model as an example: the input of the CBOW model is a set of context words (usually represented by the feature vectors of words, such as one-hot encoding), and the output is the feature vector of the central word. Suppose there is a sentence "The cat is sitting on the mat", and the goal is to generate the feature vector of each word. For the first word "cat" in it, when it is used as the central word, c words after it and c words before it can be taken as the context. Suppose c is 2, then the 2 words after "cat" are "sitting" and "on", and since there are no words before "cat", set words can be used to supplement. In this way, to generate the feature vector of "cat", 2c context words are needed. Input the feature vectors (represented by one-hot encoding) of the 2c context words of the central word into the CBOW model, and the feature vector of the central word can be generated.
[0102] Step 302: Determine the probability of the word under the set intention type according to the feature vector of the word.
[0103] Here, there are multiple set intention types that correspond one-to-one with the types of the local text library. For example, the types of the local text library are products and supply chain, and the corresponding set intention types also include products and supply chain.
[0104] For any set intention type, determine the probability of each word in the first prompt word, denoted as P(x i |c), where x i is the feature vector of the i-th word in the first prompt word, and c is the set intention type.
[0105] For example, for the intention type - products, calculate the probability of each word in the first prompt word under this intention type. For the intention type - supply chain, calculate the probability of each word in the first prompt word under this intention type.
[0106] Step 303: Determine the probability of the first prompt word under the set intention type according to the probability corresponding to each word in the first prompt word, expressed as: where d is the number of words in the first prompt word, and P(c) is the prior probability of c.
[0107] For example, for the intention type - products, after obtaining the probability of each word through step 302, according to the formula calculate the probability of this first prompt word under the intention type - products, that is, the probability that the intention of the first prompt word is products, for example, it is 0.8;
[0108] For the intention type - supply chain, after obtaining the probability of each word through step 302, according to the formula Calculate the probability of the first prompt word in the intent type - supply chain, that is, the probability that the intent of the first prompt word is supply chain, for example, it is 0.6.
[0109] Thus, determine the probability of the first prompt word under each intent type.
[0110] Step 304: Determine the intent type of the first prompt word according to the probability of the first prompt word under each set intent type.
[0111] Specifically, after determining the probability of the first prompt word under each intent type, the intent type with the highest probability can be selected as the final intent type, so as to determine the local text library corresponding to the first prompt word. For example, if the probability that the intent of the above first prompt word is product is 0.8 and the probability that the intent of the first prompt word is supply chain is 0.6, then product is determined as the intent of the first prompt word, and the type of the corresponding local text library is also determined as product.
[0112] The intent recognition in steps 301 - 304 can be implemented through an intent recognition model, and its recognition process can be divided into two parts: 1. Generate word vectors through a word vector generation model (such as the above CBOW model, Skip - gram model, BERT model, GloVe model); 2. Calculate the probability of the set intent according to the word vectors.
[0113] Among them, the training process of the intent recognition model is as follows:
[0114] Generate word vectors for the sample data through the word vector generation model, calculate the probability of the sample data under each set intent type according to the word vectors, and take the intent type with the highest probability as the final intent type of this prediction according to the size of the probability of each set intent type to complete a training process. If the predicted intent type does not meet the expectation, it means that the accuracy of the word vectors is not high. Then, after optimizing the word vector generation model, the above - mentioned training process is executed again until the prediction result meets the expectation.
[0115] Taking the CBOW model as an example, optimizing the CBOW model actually means optimizing the objective function of the CBOW model. Its objective function can be expressed as:
[0116]
[0117] Among them, D is the vocabulary (representing the set of all words in the CBOW model training dataset), l w is the size of the context window (in the example of step 301 above, 2c is the size of the context window), is an indicator variable. If the word at the j - th position in the context window is the target word w, then Otherwise σ is the sigmoid function, and X w is the context vector (obtained by averaging or summing the word feature vectors in the context window), is the weight, i.e., the parameter of the CBOW model.
[0118] For the optimization of the objective function, optimization algorithms such as gradient descent can be used. During the optimization process, the gradient of the objective function with respect to the model parameters is calculated, and the parameters of the CBOW model are updated according to the gradient. The gradient represents the directionality and rate of change of the loss function of the CBOW model on each parameter. By adjusting the parameters in the opposite direction of the gradient, the global minimum of the loss function can be gradually approached. The parameters corresponding to the global minimum of the loss function are the optimized parameters of the CBOW model.
[0119] Regarding the parameters of the CBOW model its gradient p(w∣context(w)) is calculated according to the following formula:
[0120]
[0121] where context(w) represents the words included in the context window.
[0122] The following uses a specific example to illustrate the data processing process of the present disclosure. As Figure 3 shown, the process includes:
[0123] Step 401: Generate the first prompt word.
[0124] Suppose the information obtained by the local service through the interaction interface is: flower series notebooks, publicity, copywriting. The local service generates the first prompt word, for example: Please generate publicity copy for the flower series products of the notebook.
[0125] Step 402: Perform intent recognition on the first prompt word to determine the local text library.
[0126] Suppose the intent type of the first prompt word is determined to be type 1 - product through the above intent recognition method, thereby determining that the type of the local text library is product.
[0127] Step 403: Retrieve from the local text library according to the first prompt word and output the first text data.
[0128] According to the first prompt word, multiple texts can be retrieved from the local text library of the product type. Each text records the information of the flower series products. Suppose one of the text data records the information of this series of products as follows: 16 inches, I9 processor, with air cooling function, etc.
[0129] Step 404: Desensitize the first text data using the sensitive information library corresponding to the local text library, and output the second text data and a record of sensitive information - replacement information.
[0130] Search for the sensitive information contained in each first text data from the sensitive information library corresponding to the local text library of type 1 - products. Assume that the above 16, I9, and air cooling are sensitive information, and the replacement information generated according to their respective replacement rules are: XX, YY, and natural cooling.
[0131] The product information contained in the desensitized second text data is as follows: XX inches, YY processor, with natural cooling function, etc.
[0132] And record the sensitive information - replacement information: 16 - XX, I9 - YY, air cooling - natural cooling.
[0133] Step 405: Merge all the second text data and the first prompt to obtain the second prompt.
[0134] For example, the second prompt at least includes: Please generate a promotional copy for the notebook flower series products, XX inches, YY processor, with natural cooling function.
[0135] Step 406: Input the second prompt into the public network large model, and the public network large model outputs the third text data.
[0136] The local service calls the interface provided by the public network large model, inputs the second prompt into the public network large model, and the large model generates the third text data, that is, the promotional copy. For example:
[0137] Notebook Flower Series - The peak of performance, victory in calmness!
[0138] The trend - setting Notebook Flower Series is coming again with a bang! The brand - new products launched this time are tailor - made for you who pursue extreme performance and a calm experience.
[0139]
Extra - large view, all in the XX - inch giant screen
[0140] An immersive visual feast starts from a XX - inch high - definition large screen. Whether it is immersive gaming or efficient office work, it can let you enjoy a broad view and experience an unprecedented visual shock.
[0141]
YY Processor, performance monster
[0142] Equipped with the top - notch YY processor, whether it is multi - tasking or running large - scale games, it can handle them with ease. Its powerful performance is only to meet your extreme pursuit of speed and efficiency.
[0143]
Natural cooling, calm as ever
[0144] Innovatively adopt natural cooling technology, which can achieve efficient heat dissipation without relying on traditional fans. Even under long-term high-load operation, the body can remain cool, allowing you to enjoy high performance while staying away from the troubles of noise and heat.
[0145] The Flower series is not just a laptop, but also a powerful assistant for you to conquer challenges and explore the unknown. Join us now and start a new chapter of coexisting performance and calmness!
[0146] Step 407: Restore the sensitive information in the third text data to obtain the fourth text data.
[0147] According to the sensitive information-replacement information mapping relationship recorded in step 404: 16-XX, I9-YY, air-cooled heat dissipation-natural cooling, restore the third text data to the fourth text data:
[0148] The Flower series of laptops - peak performance, calm wins!
[0149] The trend-setting Flower series of laptops is coming back with a bang! The newly launched products this time are tailor-made for you who pursue extreme performance and a calm experience.
[0150]
Extra-large view, all in a 16-inch giant screen
[0151] An immersive visual feast starts with a 16-inch high-definition large screen. Whether it is immersive gaming or efficient office work, it can let you enjoy a wide field of vision and experience an unprecedented visual shock.
[0152]
I9 processor, a performance monster
[0153] Equipped with the top-notch I9 processor, it can easily handle multitasking and running large games with ease. Its powerful performance is just to meet your extreme pursuit of speed and efficiency.
[0154]
Air-cooled heat dissipation, always calm
[0155] Innovatively adopt air-cooled heat dissipation technology, which can achieve efficient heat dissipation without relying on traditional fans. Even under long-term high-load operation, the body can remain cool, allowing you to enjoy high performance while staying away from the troubles of noise and heat.
[0156] The Flower series is not just a laptop, but also a powerful assistant for you to conquer challenges and explore the unknown. Join us now and start a new chapter of coexisting performance and calmness!
[0157] Step 408: Output the fourth text data to the user.
[0158] The local service presents the fourth text data to the user through the interaction interface. If the user believes that the content of the fourth text data does not meet the expectations, the user can issue an instruction through the interaction interface to cause the public network large model to regenerate the copywriting. After the local service obtains this instruction, it executes steps 406 - 407 again until it meets the user's expectations.
[0159] If the user modifies the input information, steps 401 - 407 are executed again.
[0160] To implement the above data processing method, as Figure 4 shown, the present disclosure also provides a data processing apparatus, which is applied to the above local service and can be deployed in a local electronic device. The apparatus includes:
[0161] An interaction module 10 for obtaining a first prompt word to be input to the public network large model;
[0162] A retrieval module 20 for retrieving at least one corresponding first text data from at least one local text library according to the first prompt word;
[0163] A desensitization module 30 for identifying and replacing sensitive information in each of the first text data to obtain corresponding second text data;
[0164] The interaction module 10 is further configured to generate a second prompt word according to all the second text data and the first prompt word, input the second prompt word into the public network large model, and obtain third text data generated by the public network large model according to the second prompt word;
[0165] The desensitization module 30 is further configured to restore sensitive information in the third text data to obtain fourth text data;
[0166] The interaction module 10 is further configured to output the fourth text data.
[0167] The above interaction module 10 provides an interaction interface for the user and can call the interface provided by the public network large model when facing the public network large model.
[0168] In one example, when replacing sensitive information in the first text data, the desensitization module 30 is further configured to determine a sensitive information library corresponding to the local text library to which the first text data belongs. The sensitive information library includes sensitive information and corresponding replacement rules; determine the sensitive information included in the first text data according to the sensitive information library, and perform sensitive information replacement on the first text data according to the replacement rules corresponding to the sensitive information.
[0169] When replacing sensitive information in the first text data, the desensitization module 30 is further configured to record the sensitive information and the corresponding replacement information; correspondingly, the desensitization module 30 is further configured to restore the replacement information in the third text data to sensitive information according to the recorded sensitive information and the corresponding replacement information, so as to obtain the fourth text data.
[0170] The desensitization module 30 is further configured to obtain the newly added text data of the local text library and determine the weight of each word in the newly added text data;
[0171] Determine the words whose weights meet the set threshold as suspected sensitive information;
[0172] Determine whether the suspected sensitive information is sensitive information according to the review result of the suspected sensitive information;
[0173] Construct the replacement rule of the sensitive information and add the sensitive information and the corresponding replacement rule to the sensitive information library.
[0174] The device further includes an intention recognition module 40 (not shown in the figure), which is configured to recognize the intention of the first prompt word, determine the corresponding local text library according to the recognized intention type, and notify the interaction module 10 to provide the first prompt word and the corresponding local text library type to the retrieval module 20, so that the retrieval module 20 retrieves according to the first prompt word in the local text library of the corresponding type.
[0175] When recognizing the intention of the first prompt word, the intention recognition module 40 is further configured to: obtain the feature vector of each word in the first prompt word; determine the probability of the word under the set intention type according to the feature vector; determine the probability of the first prompt word under the set intention type according to the probability corresponding to each word in the first prompt word; determine the intention type of the first prompt word according to the probability of the first prompt word under each set intention type.
[0176] Exemplarily, the present disclosure further provides an electronic device, including:
[0177] A processor;
[0178] A memory for storing the executable instructions of the processor;
[0179] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above data processing method.
[0180] Exemplarily, the present invention further provides a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the above data processing method.
[0181] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present application described in the "Exemplary Methods" section above of this specification.
[0182] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0183] Furthermore, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present application described in the "Exemplary Methods" section above of this specification.
[0184] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0185] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the above-disclosed specific details are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present application to necessarily adopt the above specific details for implementation.
[0186] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc. are open-ended terms meaning "including but not limited to" and can be used interchangeably with each other. The words "or" and "and" as used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" as used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0187] It should also be noted that in the devices, equipment, and methods of this application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this application.
[0188] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0189] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and subcombinations thereof.
Claims
1. A data processing method, the method comprising: Obtain the first prompt word to be input into the public network large model; Searching at least one local text library according to the first prompt word to obtain corresponding at least one first text data; Identify and replace sensitive information in each of the first text data to obtain corresponding second text data; Generate a second prompt word according to all the second text data and the first prompt word, and input the second prompt word into the public network big model; The third text data generated by the public network big model according to the second prompt word is obtained, sensitive information of the third text data is restored to obtain fourth text data, and the fourth text data is output.
2. The data processing method according to claim 1, wherein the step of replacing sensitive information of the first text data comprises: Determine a sensitive information library corresponding to the local text library to which the first text data belongs, wherein the sensitive information library includes sensitive information and corresponding replacement rules; The sensitive information included in the first text data is determined according to the sensitive information library, and the sensitive information is replaced in the first text data according to a replacement rule corresponding to the sensitive information.
3. According to the data processing method of claim 2, the replacement rules corresponding to the sensitive information include one or more of the following: Replace sensitive information with corresponding setting information; Obtain similar information of the sensitive information from a non-sensitive information database, and replace the sensitive information with the similar information; Replace sensitive information according to the set replacement algorithm.
4. According to any one of claims 1 to 3, when the sensitive information of the first text data is replaced, the method further comprises: Record sensitive information and corresponding replacement information; Correspondingly, the sensitive information restoration of the third text data to obtain the fourth text data includes: based on the recorded sensitive information and corresponding replacement information, restoring the replacement information in the third text data to the sensitive information to obtain the fourth text data.
5. The data processing method according to claim 2, further comprising: Constructing the sensitive information library corresponding to the local text library includes: Acquire new text data from the local text library, and determine the weight of each word in the new text data; Determine words whose weights meet the set threshold as suspected sensitive information; Determining whether the suspected sensitive information is sensitive information according to the review result of the suspected sensitive information; Construct replacement rules for the sensitive information, and add the sensitive information and the corresponding replacement rules to the sensitive information library.
6. The data processing method according to claim 1, wherein searching at least one local text library according to the first prompt word comprises: The first prompt word is subjected to intention recognition, a corresponding local text library is determined according to the recognized intention type, and a search is performed in the determined local text library according to the first prompt word.
7. The data processing method according to claim 6, wherein the performing intention recognition on the first prompt word comprises: Obtaining a feature vector of each word in the first prompt word; Determining the probability of the word under a set intent type according to the feature vector; Determining the probability of the first prompt word under the set intention type according to the probability corresponding to each word in the first prompt word; The intent type of the first prompt word is determined according to the probability of the first prompt word under each set intent type.
8. A data processing device, applied to a local electronic device, comprising: An interactive module, used for obtaining a first prompt word to be input into the public network large model; A retrieval module, used to search at least one local text library according to the first prompt word to obtain corresponding at least one first text data; A desensitization module, used for identifying and replacing sensitive information in each of the first text data to obtain corresponding second text data; The interaction module is further configured to generate a second prompt word according to all the second text data and the first prompt word, input the second prompt word into the public network big model, and obtain third text data generated by the public network big model according to the second prompt word; The desensitization module is further used to restore sensitive information of the third text data to obtain fourth text data; The interaction module is further configured to output the fourth text data.
9. An electronic device comprising processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program is used to execute the data processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Language model cue word sensitivity evaluation and training method based on symbolic interaction
CN122264132A