Model construction method and device, data risk identification method and device, medium and product
By using the preset seed training data set and general language model to train and verify the intent recognition model, the time-consuming and labor-intensive problem of manual labeling is solved, and efficient and accurate intent recognition is achieved.
Patent Information
- Application Number
- CN202510122483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The construction of existing intent identification model training data relies on manual annotation, which is time-consuming, labor-intensive, cost-effective, and prone to inconsistent labeling results and errors, resulting in low recognition accuracy.
By obtaining the preset seed training data set, the offline intent recognition model is trained, the trained model is used to predict the intent tag of the training data to be marked, and input it into the common language model for verification. Finally, the online intent recognition model is trained based on data of the same intent tag.
It realizes the automated construction of rich and diverse training data, reduces the demand for manual labeling, improves the labeling efficiency, reduces the cost of manpower labeling, and ensures the quality and consistency of the training data through dual verification, and improves the recognition accuracy of the intention recognition model.
Smart Images

Figure CN120045708A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular, to a method, device, medium, and product for model construction and data risk identification. Background Art
[0002] In the field of Natural Language Processing (NLP), intent recognition is a fundamental and important technology. Its core task is to accurately identify the intent of the user from the text data input by the user, so as to provide guidance for subsequent services or operations.
[0003] Currently, the construction method of intent recognition model training data mainly relies on manual annotation of a large amount of text data. This not only takes time and effort and requires extremely high labor costs, but also easily leads to inconsistent and incorrect annotation results due to factors such as the understanding deviation and fatigue of the annotators, resulting in a low recognition accuracy of the intent recognition model. Summary of the Invention
[0004] To overcome the problems existing in the related art, the embodiments of this specification provide a method, device, medium, and product for model construction and data risk identification.
[0005] According to the first aspect of any embodiment of this specification, a model construction method is provided. The method includes:
[0006] Obtain a preset seed training data set, and use the seed training data set to train a preset offline intent recognition model; wherein, the seed training data in the seed training data set is configured with correct intent labels selected from multiple preset intent labels;
[0007] Obtain a training data set to be labeled, and use the trained offline intent recognition model to predict the first intent label of the training data to be labeled in the training data set to be labeled;
[0008] Construct prompt content including the training data to be labeled and the multiple preset intent labels, and input it into a general language model, so that the general language model can identify the second intent label of the training data to be labeled among the multiple preset intent labels;
[0009] Based on the training data to be labeled with the same first intent label and second intent label, train a preset online intent recognition model.
[0010] According to the second aspect of any embodiment of this specification, a data risk identification method is provided. The method includes:
[0011] Obtain the current user attribute information and the current writing content of the current user;
[0012] Match the current user attribute information with the high-risk user attribute information stored in a preset database;
[0013] Use the online intent recognition model constructed by the method described in any embodiment of the first aspect above to identify the current intent label of the current writing content;
[0014] Determine the risk result of the current writing content according to the matching result and the current intent label.
[0015] According to the third aspect of any embodiment of this specification, a data risk identification method is provided. The method is applied to a client, and the method includes:
[0016] Obtain the current writing content written by the current user;
[0017] Obtain the risk result of the current writing content; wherein, the risk result is determined based on the current intent label identified by a preset online intent recognition model for the current writing content; the labels of the training data of the online intent recognition model are obtained after being identified by a trained offline intent recognition model and a general language model respectively.
[0018] According to the fourth aspect of any embodiment of this specification, an electronic device is provided, including:
[0019] A processor;
[0020] A memory for storing instructions executable by the processor;
[0021] Wherein, the processor realizes the method described in any embodiment of the first aspect and the second aspect above, or the method described in any embodiment of the third aspect above by running the executable instructions.
[0022] According to the fifth aspect of any embodiment of this specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the method described in any embodiment of the first aspect and the second aspect above, or the method described in any embodiment of the third aspect above is realized.
[0023] According to the sixth aspect of any embodiment of this specification, a computer program product is provided, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the method described in any embodiment of the first aspect and the second aspect above, or the method described in any embodiment of the third aspect above is realized.
[0024] The technical solutions provided by the embodiments of this specification may include the following beneficial effects:
[0025] According to the above embodiments, by obtaining a preset seed training dataset, training a preset offline intent recognition model using the seed training dataset to obtain a to-be-labeled training dataset, predicting the first intent label of the to-be-labeled training data in the to-be-labeled training dataset using the trained offline intent recognition model, constructing a prompt content including the to-be-labeled training data and multiple preset intent labels and inputting it into a general language model, so that the general language model can identify the second intent label of the to-be-labeled training data among the multiple preset intent labels, and training a preset online intent recognition model based on the to-be-labeled training data with the same first intent label and second intent label. Using the generative labeling method of the general language model, more abundant and diverse training data can be automatically constructed, reducing the need for manual annotation, greatly improving the annotation efficiency, thereby reducing the labor annotation cost. Moreover, by combining the offline intent recognition model and the general language model, double verification of the training data can be achieved, ensuring the quality and consistency of the training data, and improving the recognition accuracy of the online intent recognition model.
[0026] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the embodiments of this specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and constitute a part of the embodiments of this specification, showing the embodiments that conform to this specification, and are used together with the specification to explain the principles of the embodiments of this specification.
[0028] Figure 1 is a schematic diagram of an intent recognition scenario shown according to an exemplary embodiment of this specification;
[0029] Figure 2 is a flowchart of a model construction method shown according to an exemplary embodiment of this specification;
[0030] Figure 3A is a schematic diagram of constructing seed training data shown according to an exemplary embodiment of this specification;
[0031] Figure 3B is a schematic diagram of a takeaway interaction intent shown according to an exemplary embodiment of this specification;
[0032] Figure 4 is a flowchart of a data risk recognition method shown according to an exemplary embodiment of this specification;
[0033] Figure 5AIt is a schematic diagram of a data risk assessment shown in accordance with an exemplary embodiment of this specification;
[0034] Figure 5B It is a schematic diagram of a user behavior link shown in accordance with an exemplary embodiment of this specification;
[0035] Figure 6 It is a flowchart of another data risk identification method shown in accordance with an exemplary embodiment of this specification;
[0036] Figure 7 It is an interaction diagram of a data risk identification method shown in accordance with an exemplary embodiment of this specification;
[0037] Figure 8 It is a schematic diagram of the structure of an electronic device shown in accordance with an exemplary embodiment of this specification;
[0038] Figure 9 It is a block diagram of a model construction device shown in accordance with an exemplary embodiment of this specification;
[0039] Figure 10 It is a block diagram of a data risk identification device shown in accordance with an exemplary embodiment of this specification;
[0040] Figure 11 It is a block diagram of another data risk identification device shown in accordance with an exemplary embodiment of this specification. Detailed implementation manners
[0041] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of the embodiments of this specification as detailed in the appended claims.
[0042] The terms used in the embodiments of this specification are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of this specification. The singular forms "a", "the", and "said" used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0043] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0044] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0045] In the intent recognition scenario, the server can recognize the user's intent based on the input content of the user on the client. Please refer to Figure 1 , Figure 1 is a schematic diagram of an intent recognition scenario shown in this specification according to an exemplary embodiment. This intent recognition scenario may include a server 10, a client 11, and a client 12.
[0046] Among them, the server can be a program deployed in the background device to provide services for users. The background device can be a server, which can be an independent server or a server cluster composed of multiple servers, and can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence sampling point devices.
[0047] The client can be a program deployed in the user device to provide services for users, and can be an application program, a Web page, a small program, etc. The user device can be a personal computer, a laptop, a smartphone, a tablet computer, an Internet of Things device, a wearable device, a personal digital assistant, etc.
[0048] The server can provide various business services for users based on the client, such as: commodity delivery service, transaction service, evaluation service, etc. The role information of the users corresponding to the client 11 and the client 12 can both be consumers, the role information of the users corresponding to the client 11 and the client 12 can both be merchants, or the role information of the user corresponding to the client 11 can be a consumer, and the role information of the user corresponding to the client 12 can be a merchant.
[0049] It can be understood that Figure 1 The clients 11 and 12 in the illustrated intent recognition scenario are merely examples, and there are no restrictions on the number and role information of the clients in the actual intent recognition scenario.
[0050] When a user uses a client to input content, the server can perform intent recognition on the content input by the user. In related technologies, a model can be trained based on manually labeled training data, and this model is used to predict the user's intent. However, manually labeled training data requires extremely high labor costs and is also prone to labeling errors, resulting in a low recognition accuracy of the intent recognition model.
[0051] To solve the above problems, an embodiment of this specification proposes a model construction method. To further illustrate the embodiments of this specification, the following embodiments are provided:
[0052] Please refer to Figure 2 , Figure 2 which is a flowchart of a model construction method shown in this specification according to an exemplary embodiment. This model construction method can be applied to an intent recognition scenario. For example, it can be executed by the server 10 shown in Figure 1 or can be executed by devices in other application scenarios. The embodiments of this specification do not limit this.
[0053] As shown in Figure 2 , this model construction method may include the following steps:
[0054] Step 201: Obtain a preset seed training data set, and use the seed training data set to train a preset offline intent recognition model; wherein, the seed training data in the seed training data set is configured with correct intent labels selected from multiple preset intent labels.
[0055] In this step, the interaction behavior records of users in each business scenario can be extracted from the actual business scenario in advance, or other methods such as web crawling can be used to construct the seed training data. The seed training data is configured with correct intent labels selected from multiple preset intent labels by means of manual annotation or automated annotation tools. Based on the seed training data configured with correct intent labels, a seed training data set is constructed.
[0056] Among them, the seed training data set is a data set including several pieces of seed training data. The seed training data is representative initial sample data, and the seed training data includes at least the text content that needs to perform intent recognition.
[0057] The preset intent tags represent the intents that the user may wish to express. For example: urging the delivery progress, querying pick-up information, modifying the delivery address, booking services, making complaints, etc. They can be set and adjusted based on the actual business scenario. The correct intent tag represents the intent tag that the user actually wishes to express.
[0058] A preset seed training dataset can be obtained, and the preset offline intent recognition model is trained using the seed training dataset. Among them, the offline intent recognition model is a machine learning or deep learning model, or a pre-trained large language model, which can extract features from the input data and judge the intent of the input data based on these features.
[0059] The embodiments of this specification do not limit the specific training method of the offline intent recognition model. Exemplarily, each seed training data in the seed training dataset can be input into the offline intent recognition model, the prediction result is calculated through forward propagation, then the loss function is used to calculate the error between the prediction result and the correct intent tag, and the weight parameters of the offline intent recognition model are updated through backpropagation to minimize the error. Through multiple iterative trainings until the performance of the offline intent recognition model on the preset validation set meets the expectations.
[0060] Step 202: Obtain the training dataset to be labeled, and use the trained offline intent recognition model to predict the first intent tag of the training data to be labeled in the training dataset to be labeled.
[0061] In this step, a training dataset to be labeled can be obtained. The training dataset to be labeled can include a large number of training data to be labeled without the correct intent tag being labeled. The training data to be labeled in the training dataset to be labeled can be extracted from the historical business data of the actual business scenario, or can be constructed in other ways.
[0062] The trained offline intent recognition model can be used to perform intent recognition on the training data to be labeled in the training dataset to be labeled, and predict the first intent tag of the training data to be labeled.
[0063] Exemplarily, each training data to be labeled in the training dataset to be labeled can be input into the offline intent recognition model. The offline intent recognition model performs intent recognition on the training data to be labeled and outputs the predicted first intent tag. The first intent tag can represent the intent that the offline intent recognition model believes the user wants to express.
[0064] Step 203: After constructing the prompt content including the training data to be labeled and multiple preset intent tags, input it into the general language model so that the general language model can recognize the second intent tag of the training data to be labeled among the multiple preset intent tags.
[0065] In this step, a prompt containing the training data to be labeled and multiple preset intent labels can be constructed. The prompt can be used as context information for general language models such as Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformer (GPT), helping the general language model better understand the intent of the training data to be labeled.
[0066] The prompt can be input into the general language model. The general language model uses its own text generation and understanding capabilities to identify the second intent label of the training data to be labeled among multiple preset intent labels. The second intent label can represent the intent that the general language model believes the user wants to express.
[0067] Step 204: Train the preset online intent recognition model based on the training data to be labeled with the same first intent label and second intent label.
[0068] In this step, the first intent label and the second intent label can be compared. For each piece of training data to be labeled, if the first intent label predicted by the offline intent recognition model is the same as the second intent label recognized by the general language model, it can be determined that the intent recognition result for this training data to be labeled is reliable.
[0069] The preset online intent recognition model can be trained based on the training data to be labeled with the same first intent label and second intent label. The training method of the online intent recognition model can also include steps such as data input, forward propagation, loss calculation, backpropagation, and weight update, which are not limited in the embodiments of this specification.
[0070] Among them, the online intent recognition model can quickly and accurately identify the user's intent in the actual business scenario. The online intent recognition model and the online intent recognition model can be two different models, or the trained offline intent recognition model can be used as the initial online intent recognition model, and the initial online intent recognition model can be trained with the training data to be labeled with the same first intent label and second intent label.
[0071] In one embodiment, for the training data to be labeled with different first intent labels and second intent labels, the correct intent label of the training data to be labeled can be identified through manual review and other means. If the correct intent label is consistent with the second intent label, the training data to be labeled configured with the second intent label can be added to the seed training data set to continue training the offline intent recognition model, thereby improving the accuracy of the offline intent recognition model.
[0072] The model construction method of this embodiment obtains a preset seed training data set, uses the seed training data set to train a preset offline intent recognition model to obtain a training data set to be labeled, uses the trained offline intent recognition model to predict the first intent label of the training data to be labeled in the training data set to be labeled, constructs a prompt content including the training data to be labeled and multiple preset intent labels, and inputs it into a general language model, so that the general language model can identify the second intent label of the training data to be labeled among multiple preset intent labels. Based on the training data to be labeled with the same first intent label and second intent label, the preset online intent recognition model is trained. Using the generative labeling method of the general language model, more abundant and diverse training data can be automatically constructed, the need for manual annotation can be reduced, the annotation efficiency can be greatly improved, thereby reducing the human annotation cost. Moreover, by combining the offline intent recognition model and the general language model, double verification of the training data can be achieved, ensuring the quality and consistency of the training data, and improving the recognition accuracy of the online intent recognition model.
[0073] In one embodiment, please refer to Figure 3A , Figure 3A which shows a schematic diagram of constructing seed training data. Historical business data in the historical business data set can be generated by extracting and analyzing information such as log files of business systems, user behavior records, and customer service chat records. The historical business data includes, but is not limited to: historical scenarios, role information of historical users, and historical written content.
[0074] Among them, the historical scenario is the business scenario where the historical user performs the operation of inputting historical written content. For example: instant delivery scenario, medical consultation scenario, rider communication group scenario, fan group scenario, purchasing and delivery scenario, etc. The role information of the historical user is the user identity in which the historical user performs the operation of inputting written content. For example: consumer, merchant, rider, intermediate platform, etc. The historical written content is the input content of the historical user under the historical scenario and role information.
[0075] Contents with highly similar texts may have different intents in different scenarios or different roles. For example, "If I take this medicine now and go to sleep, will I probably not be able to get up this afternoon?" When it is an inquiry initiated by the ordering user to the doctor in the "medical consultation scenario", it is very likely to be an inquiry about whether a certain drug may have drowsiness side effects, representing a benign intent; while in the "chat group scenario", a statement made by a user in the group may have a non-benign intent. Therefore, the white rule data designed in the embodiments of this specification can distinguish benign written content through scenarios and roles, and can construct training data that accurately represents benign intent.
[0076] Based on various methods such as business experience, user behavior analysis, and expert review in advance, data analysis and rule abstraction can be performed on the benign content in historical business data to generate white rule data, which can accurately reflect the benign intentions of historical users in different scenarios.
[0077] Among them, the white rule data can include benign writing content corresponding to the labels representing benign intentions under different scenarios and different role information. The benign intention label is a label indicating the positive intention expressed by the user under different scenarios and different role information.
[0078] The benign writing content is the specific text content used by the user to express benign intentions under different scenarios and different role information, which can accurately reflect the expression methods and language habits of the user in different situations.
[0079] Taking the instant delivery scenario as an example, the white rule data can be the delivery requirements of the user for the rider, such as: no need for phone communication, leave the goods at the door, etc., and can also be the arrival notice of the rider to the user, such as: the goods have been delivered, please pick up the meal, etc.
[0080] The preset white rule data can be obtained, and the historical business data set can be screened according to the white rule data. Each historical business data in the historical business data set is compared with the white rule data, and the historical business data that matches the white rule data is screened from the historical business data set, and the matching historical business data is output to the configurator for the configurator to configure the correct intention label for the historical business data, so as to construct seed training data.
[0081] As described above, by obtaining the historical business data set and obtaining the preset white rule data, since the white rule data can reflect the benign intentions of historical users in different scenarios, the historical business data that matches the white rule data is screened from the historical business data set to construct seed training data, and the historical business data that is highly relevant and accurate to the benign intention can be screened to ensure the accuracy and relevance of the seed training data, which helps to improve the pre-labeling ability of the offline intention recognition model, and can significantly reduce the workload of manual labeling, thereby reducing costs.
[0082] In one embodiment, please continue to refer to Figure 3A , in the historical business data set, there are some historical business data that fail to match the white rule data due to reasons such as iterative updates in business development. These historical business data represent new intentions not covered by the white rules. In order to effectively utilize these historical business data, a clustering method based on text similarity can be used to label the unmatched historical business data.
[0083] For each piece of historical business data in the historical business data set that does not match the white rule data, the text similarity between each piece of historical business data can be calculated using text similarity algorithms such as cosine similarity, Jaccard similarity, and edit distance.
[0084] Clustering algorithms such as K-means clustering algorithm (K-means), hierarchical clustering, density-based spatial clustering of applications with noise (DBSCAN) and other clustering algorithms are used to cluster historical business data whose text similarity meets the preset similarity conditions to obtain multiple clusters, each of which can contain a group of similar historical business data.
[0085] After clustering is completed, multiple clusters are output to the configuration personnel, who can configure the correct intent label for each piece of historical business data based on multiple clusters, combined with historical scenarios, role information, historical writing content and other factors in the historical business data. Seed training data is constructed based on each piece of historical business data and the corresponding correct intent label.
[0086] As described above, by clustering each piece of historical business data that does not match the white rule data in the historical business data set based on the text similarity between each piece of historical business data, multiple clusters are output to the configuration personnel so that the configuration personnel can configure the correct intent label for each piece of historical business data, and construct seed training data based on each piece of historical business data and the correct intent label. This can reduce the amount of data that needs to be labeled, and can effectively utilize the unmatched historical business data to explore the potential value and information therein, enrich the diversity of seed training data, thereby continuously optimizing and updating the offline intent recognition model and the online intent recognition model to keep pace with business development.
[0087] In one embodiment, please continue to refer to Figure 3A , the seed training data set constructed by white rule matching and clustering can be used to train the offline intent recognition model to obtain a trained offline intent recognition model.
[0088] The preset intent labels corresponding to the seed training data set may include benign intent labels and non-benign intent labels. Benign intent labels include at least one of the following: intent labels related to orders, intent labels related to products, intent labels related to meal pickup and delivery, intent labels related to after-sales service, and intent labels related to user interaction.
[0089] The non-benign intention label is a label that represents the negative intention expressed by the user under different scenarios and different role information, and can be other labels than the benign intention label.
[0090] Exemplarily, taking the instant delivery scenario as the specific takeaway scenario as an example, as Figure 3B shown, it is a schematic diagram of the takeaway interaction intention shown in the embodiments of this specification. In the instant delivery scenario, the benign intention labels and non-benign intention labels corresponding to the seed training data set can be as shown in Table 1:
[0091] Table 1 Benign Intention Labels and Non-Benign Intention Labels
[0092]
[0093] It can be understood that the benign intention labels and non-benign intention labels shown in Table 1 are only examples, and can be adjusted and set based on the actual application scenario and actual business. The embodiments of this specification do not limit this.
[0094] Based on the latest historical business data, a training data set to be labeled can be constructed. Input the training data to be labeled in the training data set to be labeled into the offline intention recognition model, and obtain the first intention label output by the offline intention recognition model.
[0095] According to the training data to be labeled and multiple preset intention labels, the prompt content of the general language model can be constructed. Exemplarily, the prompt content can be as follows: "
[0096] Your task is to perform intention classification on the text content. Please refer to the given scenario and role information, and select the category that best matches the text content. The predefined preset intention labels are as follows:
[0097] 1. Takeout Delivery - Order Assignment Requirement: The user's delivery requirement for the rider, such as no need for phone communication, leave the goods at the door, etc.
[0098] 2. Takeout Delivery - Arrival Notice: The rider's arrival notice to the user, such as delivered, please pick up the meal, etc.
[0099] ……
[0100] n. Others: Do not belong to the above benign intention categories, marked as the "Others" category.
[0101] The reference scenarios include instant delivery, new retail shopping, etc. The role information of the content sender includes riders, consumers, merchants, etc.
[0102] The output is required to be in the standard JSON format: {"result": "The name of the determined category"}
[0103] The given scenario is {scene_name}, the role information of the sender is {sender}, and the content to be determined is "{content}".
[0104] Among them, the prompt content includes:
[0105] ① The historical scenario of the historical writing content, the role information of the historical user, and the historical writing content.
[0106] ② Each benign intention label and non - benign intention label; for each label, there is also a rule describing the label; for example, the labels such as "Take - out delivery - order assignment requirement" and "Take - out delivery - arrival notice" in the 1st to n - 1th are benign intention labels, and the other labels in the nth are non - benign intention labels. In the above "Take - out delivery - order assignment requirement: The user's delivery requirement for the rider, such as no need for phone communication, leave the goods at the door, etc.", "Take - out delivery - order assignment requirement" represents the benign intention label, and "The user's delivery requirement for the rider, such as no need for phone communication, leave the goods at the door, etc." represents the rule of this label.
[0107] ③ The format requirements for the model output. For example, in the above example, the JSON format is required, and other formats can also be used in actual applications. This embodiment does not limit this. Among them, result represents the second intention label, scene_name represents the historical scenario in the historical business data, sender represents the role information of the historical user in the historical business data, and content represents the historical writing content in the historical business data.
[0108] The constructed prompt content can be input into the general - purpose language model to obtain the second intention label recognized by the general - purpose language model among multiple preset intention labels. In the post - processing stage of the general - purpose language model, the general - purpose language model can also analyze the historical writing content, judge whether the historical writing content contains risk content, adjust the score of the benign intention label, and output the second intention label of the non - benign intention.
[0109] Based on the training data to be labeled with the same first intention label and second intention label, train the online intention recognition model. The trained online intention recognition model can perform intention recognition on the user writing content in the instant delivery scenario.
[0110] As described above, by using the trained online intention recognition model to perform intention recognition on the user writing content in the instant delivery scenario, the intention can be quickly and accurately recognized, which helps to quickly understand the user's needs, improve the service response speed and customer satisfaction. By referring to the benign intention label and non - benign intention label for labeling, the general - purpose language model can accurately identify whether the user writing content belongs to the positive intention or negative intention, which helps to discover and handle potential risks in a timely manner.
[0111] To build a risk control mechanism for data security, please refer to Figure 4 , Figure 4 which is a flowchart of a data risk identification method shown in this specification according to an exemplary embodiment. This data risk identification method can be applied to scenarios such as instant delivery. For example, it can be executed by the server 10 shown in Figure 1 , or can be executed by devices in other application scenarios. The embodiments of this specification do not limit this.
[0112] As Figure 4 shown, this data risk identification method may include the following steps:
[0113] Step 401: Obtain the current user attribute information and the current written content of the current user.
[0114] In this step, the server may receive the current written content input by the current user from the client. Analyze information such as the account attributes, historical behaviors, and user backgrounds of the current user to obtain the user attribute information of the current user.
[0115] Among them, the current user attribute information includes but is not limited to: the user's identity information (such as username, account ID, registration time, etc.), historical behavior records (such as historical interaction records, violation records, etc.), etc. The current written content is the text content that the user is currently inputting or editing, and there may be potential risks in the current written content.
[0116] Step 402: Match the current user attribute information with the high-risk user attribute information stored in the preset database.
[0117] In this step, the server compares and matches the current user attribute information obtained in step 401 with the high-risk user attribute information in the preset database. Through a matching algorithm or rule engine, compare the current user attribute information with the high-risk user attribute information to determine the matching result.
[0118] Among them, the preset database is a data set that is constructed and stored in advance according to historical risk data and rules and contains high-risk user attribute information. The high-risk user attribute information is a characteristic or attribute indicating that the user may have high-risk behaviors, which is associated with historical risk events or behaviors, and the high-risk user attribute information can be obtained based on historical data analysis.
[0119] The high-risk user attribute information may include high-risk user account information, high-risk behavior indicators, etc. The high-risk user account information is the account information indicating that the user has high-risk characteristics, mainly focusing on the registration information of the user account, the account usage behavior, and other information related to the account. For example, the account information with high risks such as abnormal registration information and abnormal account usage behavior.
[0120] High - risk behavior indicators are behavioral indicators that indicate users' engagement in high - risk activities, mainly focusing on abnormal behaviors demonstrated by users during service usage. For example, there are historical behaviors such as violating content posting rules, abnormal interaction behaviors, and transaction - related risks.
[0121] Step 403: Use the online intent recognition model constructed by the aforementioned model construction method to identify the current intent label of the current written content.
[0122] In this step, the server uses the online intent recognition model constructed by the aforementioned model construction method to perform intent recognition on the current written content of the current user, and identifies the current intent label of the current written content. The data input into the online intent recognition model can specifically be a text containing: current scenario information, current user role information, and the current written content. The online intent recognition model can analyze information such as the semantics of the current written content, the role information of the current user, and the scenario, and identify whether the current user is engaged in normal communication, seeking help, or launching a malicious attack, etc., so as to determine which one of the aforementioned multiple preset intent labels the current written content of the current user belongs to.
[0123] Step 404: Determine the risk result of the current written content based on the matching result and the current intent label.
[0124] In this step, the server comprehensively considers the matching result of the current user attribute information in step 402 and the current intent label in step 403, and through a preset risk assessment algorithm or rule engine, performs a risk assessment on the current written content, and takes the evaluated comprehensive risk score or risk level as the risk result.
[0125] Based on the risk result of the current written content, it can be determined whether the current written content has risks, and corresponding handling measures can be taken, such as: directly intercepting, marking for review, sending a warning, etc.
[0126] The data risk identification method of this embodiment, by obtaining the current user attribute information and the current written content of the current user, matching the current user attribute information with the high - risk user attribute information stored in the preset database, using the online intent recognition model to identify the current intent label of the current written content, and determining the risk result of the current written content based on the matching result and the current intent label, while considering the intent of the current written content itself, deeply excavates the current user attribute information of the current user, and comprehensively evaluates the real risk of the current written content from multiple dimensions, can improve the comprehensiveness and accuracy of risk identification, contribute to maintaining the healthy purity of the platform content ecosystem, and enhance user experience and satisfaction.
[0127] In one embodiment, when determining the risk result of the current written content, if the matching result is a failed match and the current intent label is a benign intent label, it indicates that the current user is an ordinary user and the current written content is benign content, and the risk result can be determined as low risk.
[0128] If the matching result is a successful match and / or the current intent label is a non-benign intent label, it indicates that the current user is a high-risk user and / or the current written content is non-benign content, and the risk result can be determined as high risk.
[0129] It can be understood that the above process of determining the risk result is only an example. It is also possible to conduct a more in-depth feature analysis of the current written content, combine information such as the extracted keywords, sentiment analysis, and semantic understanding for risk determination, and it is also possible to refer to factors such as the context and time node of the current written content for risk determination. The embodiments of this specification do not limit this.
[0130] As described above, by determining the risk result as low risk if the matching result is a failed match and the current intent label is a benign intent label, and determining the risk result as high risk if the matching result is a successful match and / or the current intent label is a non-benign intent label, the risk result can be quickly determined in a short time, improving the efficiency of risk assessment.
[0131] In one embodiment, in the case where the risk result is determined to be high risk, the current written content can be matched with the risk words stored in the preset word library, and the keywords or phrases in the current written content can be compared with the risk words. If there are no matching risk words in the current written content, further analysis or other risk assessment measures can be taken.
[0132] If there are matching risk words, it can be determined that the current written content is high-risk content, and an interception operation can be performed on the current written content to prevent the spread of high-risk content. Among them, the interception operation includes but is not limited to: preventing content publication, sending a warning message to the current user, marking the previous written content as pending review, etc.
[0133] As described above, by matching the current written content with the risk words stored in the preset word library in the case where the risk result is determined to be high risk, high-risk content can be more accurately identified, reducing unnecessary intervention and false alarms, improving the efficiency and accuracy of risk prevention and control. If there are matching risk words, an interception operation is performed on the current written content to prevent the spread of high-risk content, which helps prevent the spread of risk events.
[0134] In one embodiment, please refer to Figure 5A , Figure 5AShows a schematic diagram of data risk assessment. Exemplarily, the preset database may include: an account database storing high-risk user account information, and a behavior database storing high-risk behavior metrics.
[0135] Historical user behaviors can be pre-analyzed and collected, and the historical data can be mined and analyzed to identify high-risk behavior metrics. For example Figure 5B , is a schematic diagram of a user behavior link shown in an embodiment of this specification. According to the user behavior link from login, account, search, browsing, placing an order, interaction, evaluation, dispute to risk control, different behavior types can be used to classify the high-risk behavior metrics in the user behavior link feature set and store them in the behavior database in a classified manner.
[0136] Exemplarily, the high-risk behavior metrics in the behavior database can be as shown in Table 2:
[0137] Table 2 High-risk behavior metrics in the behavior database
[0138]
[0139]
[0140] It can be understood that the behavior types and corresponding high-risk behavior metrics shown in Table 2 are only examples and can also be adjusted and set based on the actual application scenario. For example, in the e-commerce scenario, high-risk behavior metrics of behavior types such as placing an order, dispute, and risk control can be focused on; on the social media platform, high-risk behavior metrics of behavior types such as search, interaction, and evaluation can be focused on. The embodiments of this specification do not limit this.
[0141] The current user account information in the current user attribute information and one or more behavior information of the current user can be obtained. The current user account information is matched with the high-risk user account information in the account database to determine whether the current user belongs to a known high-risk user, and one or more behavior information of the current user is matched with the high-risk behavior metrics in the behavior database to determine whether the behavior of the current user is risky, thereby obtaining a matching result.
[0142] Please continue to refer to Figure 5A , the currently obtained writing content, the role information and scenario of the current user can be input into the online intent recognition model. The online intent recognition model combines the role information and scenario to perform intent recognition on the current writing content and obtain the current intent label output by the online intent recognition model. The decision-making layer in the server can determine whether the current intent label belongs to a benign intent label or a non-benign intent label. Then, combining the matching result and the judgment result of the current intent label, the risk result of the current writing content is determined.
[0143] As described above, by matching the current user account information in the current user attribute information with the high-risk user account information in the account database, and matching one or more behavior information of the current user in the current user attribute information with the high-risk behavior indicators in the behavior database, potential high-risk users and risk behaviors can be identified more accurately, which helps to reduce false positives and false negatives and improve the efficiency and accuracy of risk identification.
[0144] In one embodiment, the current written content may include at least one of the following: instant chat messages, order evaluation content, and inquiry content with the intelligent customer service. Among them, the instant chat message is the text content for instant dialogue communication between users, the order evaluation content is the evaluation content of the goods or services after the user completes the transaction, and the inquiry content is the text content that the user asks the intelligent customer service system.
[0145] The online intent recognition model can be pre-trained using the training data sets to be labeled in application scenarios such as instant chat, order evaluation, and intelligent customer service conversations, so that the online intent recognition model can accurately identify the current intent labels in different application scenarios for instant chat messages, order evaluation content, and inquiry content with the intelligent customer service.
[0146] It can be understood that the above-mentioned instant chat messages, order evaluation content, and inquiry content with the intelligent customer service are only some examples of the current written content, and it can also be social content, product description content, user feedback content, etc. The embodiments of this specification do not limit this.
[0147] As described above, by performing data risk identification on at least one of the instant chat messages, order evaluation content, and inquiry content with the intelligent customer service, the security prevention and control in different application scenarios are ensured, and the needs and intentions of users are better understood, so as to provide a more personalized data security decision-making scheme.
[0148] Please refer to Figure 6 , Figure 6 which is a flowchart of another data risk identification method shown in an exemplary embodiment of this specification. This data risk identification method can be executed by the Figure 1 client shown, or can be executed by devices in other application scenarios. The embodiments of this specification do not limit this.
[0149] As Figure 6 shown, this data risk identification method may include the following steps:
[0150] Step 601: Obtain the current written content written by the current user.
[0151] In this step, the client can detect whether the current user is performing a writing operation and obtain the current writing content written by the current user through methods such as text input monitoring and file upload detection.
[0152] Step 602: Obtain the risk result of the current writing content; wherein, the risk result is determined based on the current intent label identified by the preset online intent recognition model for the current writing content; the labels of the training data of the online intent recognition model are obtained after being respectively recognized by the trained offline intent recognition model and the general language model.
[0153] In this step, in response to detecting the current writing content written by the current user, the client sends the current writing content to the server, and the server can determine the risk result of the current writing content based on the current intent label identified by the preset online intent recognition model for the current writing content.
[0154] The client can obtain the risk result determined by the server and process the current writing content according to the risk result. Exemplarily, if the risk result is a positive result, the current writing content can be normally displayed; if the risk result is a negative result, the client can display warning information for the current writing content, prompt information such as asking the user to re-edit, etc.
[0155] In the data risk identification method of this embodiment, by obtaining the current writing content written by the current user and obtaining the risk result of the current writing content, since the risk result is determined based on the current intent label identified by the preset online intent recognition model for the current writing content, the current user's intent for the current writing content itself can be deeply mined, the accuracy of risk identification can be improved, which helps to maintain the health and purity of the platform content ecosystem and enhance the user experience and satisfaction; since the labels of the training data of the online intent recognition model are obtained after being respectively recognized by the trained offline intent recognition model and the general language model, more abundant and diverse training data can be automatically constructed, the need for manual annotation can be reduced, the annotation efficiency can be greatly improved, thereby reducing the human annotation cost, and moreover, by combining the offline intent recognition model and the general language model, double verification of the training data can be realized, ensuring the quality and consistency of the training data and improving the recognition accuracy of the online intent recognition model.
[0156] To further introduce the interaction process of data risk identification, Figure 7 An interaction diagram of a data risk identification method is shown. This method can be applied to the instant delivery scenario, and this method can include the following steps:
[0157] Step 701: The server trains the offline intent recognition model using the seed training dataset.
[0158] In this step, the server can obtain the historical business dataset in the instant delivery scenario. Screen the historical business data in the historical business dataset that matches the preset white rule data, so that the configurator can configure the correct intention label for the historical business data from multiple preset intention labels including benign intention labels and non-benign intention labels, and construct seed training data.
[0159] It is also possible to cluster each piece of historical business data in the historical business dataset that does not match the white rule data to obtain multiple clusters, so that the configurator can configure the correct intention label for each piece of historical business data and construct seed training data. The constructed seed training dataset can be used to train the offline intention recognition model.
[0160] Step 702: The server uses the offline intention recognition model to predict the first intention label of the training data to be labeled.
[0161] In this step, the server uses the trained offline intention recognition model to perform intention recognition on the training data to be labeled in the training dataset to be labeled, and predicts the first intention label of the training data to be labeled.
[0162] Step 703: The server uses the general language model to recognize the second intention label of the training data to be labeled.
[0163] In this step, the server can construct a prompt content including the training data to be labeled and multiple preset intention labels, input the prompt content into the general language model, and obtain the second intention label of the training data to be labeled recognized by the general language model.
[0164] Step 704: The server trains the online intention recognition model based on the training data to be labeled with the same first intention label and second intention label.
[0165] In this step, the server compares the first intention label and the second intention label, and trains the online intention recognition model based on the training data to be labeled with the same first intention label and second intention label, so that the online intention recognition model learns to perform intention recognition in the instant delivery scenario.
[0166] Step 705: The client sends the current written content to the server.
[0167] In this step, the client obtains the current written content written by the current user and sends the current written content to the server.
[0168] Step 706: The server matches the current user attribute information with the high-risk user attribute information stored in the preset database.
[0169] In this step, the server can match the current user account information in the current user attribute information with the high-risk user account information in the account database, and match one or more behavior information in the current user attribute information with the high-risk behavior indicators in the behavior database to obtain a matching result of successful or failed matching.
[0170] Step 707: The server uses the online intent recognition model to recognize the current intent label of the current writing content.
[0171] In this step, the server obtains the role information of the current user and the scenario of the current writing content, inputs the current writing content, role information and scenario into the online intent recognition model, and obtains the current intent label output by the online intent recognition model.
[0172] Step 708: If the matching result is a failed match and the current intent label is a benign intent label, determine that the risk result is a low risk.
[0173] In this step, if the matching result is a failed match and the current intent label is a benign intent label, the server can determine that the risk result is a low risk.
[0174] Step 709: If the matching result is a successful match and / or the current intent label is a non-benign intent label, determine that the risk result is a high risk.
[0175] In this step, if the matching result is a successful match and / or the current intent label is a non-benign intent label, the server can determine that the risk result is a high risk. The server can also match the current writing content with the risk words stored in the preset word library. If there are matching risk words, perform an interception operation on the current writing content.
[0176] Step 710: The server sends the risk result to the client.
[0177] Step 711: The client processes the current writing content according to the risk result.
[0178] In this step, the client can obtain the risk result of the current writing content sent by the server. If the risk result is a low risk, the client can normally display the current writing content. If the risk result is a high risk, the client can display prompt information such as a warning to the current user.
[0179] It can be seen from the above embodiments that this embodiment designs an online intent recognition model, which can effectively distinguish the enthusiasm of the content posted by users and classify it into intent labels, such as urging the progress of meal delivery, querying meal pickup information, modifying the delivery address, etc., and finally falling into classification labels. This model can reduce the average time consumption from the original 100MS to an average of 16.77MS. This part is currently used in instant chat scenarios. It can be based on different attribute groups (three-party call groups, users & doctors, fan groups, rider association groups, etc.), and user identities (users, riders, merchants), based on business rules (white rules such as picking up and delivering meals, commodities, etc.) combined with the semantic recognition and classification capabilities of the model itself, and finally return the intent label.
[0180] At the same time, a general language model, i.e. a pre-trained model in the field of natural language processing, is introduced. With its deep learning ability on large-scale text data, it can automatically label input samples efficiently and accurately. Relying on the model's learning of the prompt content, it can adapt to specific tasks through fine-tuning with a small amount of labeled data; generative labeling provides a flexible and detailed labeling method; and through continuous learning, it continuously optimizes the labeling effect of the training data to be labeled.
[0181] Instant delivery scenarios such as takeout interactions have very rich intentions. In order to achieve accurate intent recognition, this embodiment expands and builds a large amount of training sample data by labeling sample data with different rules, and then improves the model effect through evaluation and continuous iteration. The use of automated intent label generation technology can achieve efficient and automated processing of large amounts of real historical business data on the platform, significantly reducing the need for manual labeling, thereby greatly improving labeling efficiency and effectively reducing labor costs.
[0182] This embodiment also introduces a preset database when identifying data risks. It serves as a data set that centrally manages known or suspected bad content features, including account attributes, behavior patterns, etc. This feature collection acts as an intelligent filter to identify suspected black and gray information and filter to the next level. It achieves rapid detection through different feature classifications, covering multi-dimensional scenarios such as instant messaging (IM), evaluation, and artificial intelligence generated content (AIGC). It strives to limit and narrow the risk flow while reducing direct dependence on the risk management strategy of the middle-end service layer, but it can also ensure that risks are not missed and reduce the misjudgment of high-risk content.
[0183] In this embodiment, the intent recognition model is used to identify intent tags and match high-risk user attribute information. Different from the content of small models for text and image recognition, it not only considers the intuitive text information of the currently written content, but also combines factors such as the user's behavior pattern, account history, role information, and scenario to comprehensively evaluate the true intent behind a piece of content.
[0184] In this embodiment, based on the identified true intent, risk decision-making for data is realized, filtering out the healthy remarks of trusted users and content without malicious intent, without repeatedly identifying the text and image levels of the content at the middle platform service layer, greatly reducing the expenditure on risk control costs. For high-risk content, it is monitored through features such as accounts and behaviors. For example, for potentially risky spam accounts or abnormal behavior characteristics, more stringent monitoring measures are taken, such as reducing the algorithm tolerance and increasing the review frequency, to ensure that high-risk content is not leaked externally. Through the content recognition algorithm of the intent recognition model and the combined method of deeply analyzing and diagnosing the user's identity, account, purpose, emotion, and behavior, two-way parallel security control is achieved.
[0185] Based on the risk recognition and filtering capabilities of the online intent recognition model and the preset database in parallel, this embodiment can generate account information, user identity, and intent tags in a short time. Based on the intent tags generated by the decision-making layer and combined with the user's account information and behavior information, the content is classified into risk decisions, divided into two categories: low risk and high risk, which can effectively screen and release trusted content that does not need to trigger risk control, thus reducing unnecessary resource consumption.
[0186] Through the parallel risk recognition strategy of the preset database and the online intent recognition model, this embodiment can adapt to applications in different scenarios such as IM scenarios, evaluation scenarios, and AIGC scenarios. In the applications of various different scenarios, it can help the platform better understand the needs and intents of users, so as to provide more personalized solutions.
[0187] Figure 8 It is a schematic structural diagram of an electronic device shown in this specification according to an exemplary embodiment. The electronic device can be, for example, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a personal digital assistant, a server, a smart home appliance, a vehicle-mounted computer, etc. Refer to Figure 8, at the hardware level, the electronic device includes a processor 802, an internal bus 804, a network interface 806, a memory 808, and a non-volatile memory 810. Of course, it may also include other hardware required for other services. The processor 802 reads the corresponding computer program from the non-volatile memory 810 into the memory 808 and then runs it, forming a model construction device / data risk identification device at the logical level. Of course, in addition to the software implementation, the embodiments of this specification do not exclude other implementation manners, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and may also be hardware or a logic device.
[0188] Figure 9 is a block diagram of a model construction device shown according to an exemplary embodiment of this specification. Refer to Figure 9 , the device may include: an offline training module 901, a first identification module 902, a second identification module 903, and an online training module 904, where:
[0189] The offline training module 901 is configured to obtain a preset seed training data set, and use the seed training data set to train a preset offline intent recognition model; wherein, the seed training data in the seed training data set is configured with correct intent labels selected from multiple preset intent labels;
[0190] The first identification module 902 is configured to obtain a to-be-labeled training data set, and use the trained offline intent recognition model to predict the first intent label of the to-be-labeled training data in the to-be-labeled training data set;
[0191] The second identification module 903 is configured to construct prompt content including the to-be-labeled training data and the multiple preset intent labels and input the prompt content into a general language model, so that the general language model identifies the second intent label of the to-be-labeled training data among the multiple preset intent labels;
[0192] The online training module 904 is configured to train a preset online intent recognition model based on the to-be-labeled training data with the same first intent label and second intent label.
[0193] In one example, the offline training module 901 is further configured to obtain a historical business data set, where the historical business data in the historical business data set includes: historical scenarios, role information of historical users, and historical writing content; obtain preset white rule data, where the white rule data includes: benign writing content representing benign intent labels corresponding to different scenarios and different role information; screen historical business data matching the white rule data from the historical business data set, and construct the seed training data.
[0194] In one example, the offline training module 901 is further configured to cluster, based on the text similarity between the historical service data that does not match the white rule data in the historical service dataset, each piece of historical service data that does not match the white rule data, to obtain a plurality of clusters; output the plurality of clusters to a configurator for the configurator to determine correct intent labels corresponding to the respective pieces of historical service data based on the plurality of clusters; and use the respective pieces of historical service data and the correct intent labels corresponding to the respective pieces of historical service data as the seed training data.
[0195] In one example, the online intent recognition model is used to recognize the intent of the content written by a user in an instant delivery scenario; the preset intent labels include benign intent labels and non-benign intent labels, and the benign intent labels include at least one of the following: intent labels related to orders, intent labels related to goods, intent labels related to food pickup and delivery, intent labels related to after-sales, and intent labels related to user interaction.
[0196] Figure 10 It is a block diagram of a data risk recognition device shown according to an exemplary embodiment of the present specification. Refer to Figure 10 This device may include: a data acquisition module 1001, an information matching module 1002, an intent recognition module 1003, and a risk determination module 1004, where:
[0197] The data acquisition module 1001 is configured to acquire the current user attribute information and the current written content of the current user.
[0198] The information matching module 1002 is configured to match the current user attribute information with the high-risk user attribute information stored in a preset database.
[0199] The intent recognition module 1003 is configured to recognize the current intent label of the current written content by using the online intent recognition model constructed by the foregoing model construction method.
[0200] The risk determination module 1004 is configured to determine the risk result of the current written content according to the matching result and the current intent label.
[0201] In one example, when the risk determination module 1004 is configured to determine the risk result of the current written content according to the matching result and the current intent label, it includes: if the matching result is a matching failure and the current intent label is a benign intent label, determining that the risk result is a low risk; if the matching result is a matching success and / or the current intent label is a non-benign intent label, determining that the risk result is a high risk.
[0202] In one example, the risk determination module 1004 is further configured to, when determining that the risk result is a high risk, match the current writing content with risk words stored in a preset word library; if there is a matching risk word, perform an interception operation on the current writing content.
[0203] In one example, the preset database includes an account database storing high-risk user account information and a behavior database storing high-risk behavior metrics; the current user attribute information includes current user account information and one or more behavior information of the current user; when the information matching module 1002 is configured to match the current user attribute information with the high-risk user attribute information stored in the preset database, it includes: matching the current user account information with the high-risk user account information; matching one or more behavior information of the current user with the high-risk behavior metrics.
[0204] In one example, the current writing content includes at least one of the following: instant messaging messages, order evaluation content, and inquiry content with an intelligent customer service.
[0205] Figure 11 It is a block diagram of another data risk identification device shown in the present specification according to an exemplary embodiment. Refer to Figure 11 , the data risk identification device may include: a content acquisition module 1101 and a result acquisition module 1102, where:
[0206] The content acquisition module 1101 is configured to acquire the current writing content written by the current user;
[0207] The result acquisition module 1102 is configured to acquire the risk result of the current writing content; wherein, the risk result is determined based on the current intent label identified by a preset online intent recognition model for the current writing content; the labels of the training data of the online intent recognition model are obtained by being respectively recognized by a trained offline intent recognition model and a general language model.
[0208] In one example, the risk result is determined by combining the matching result and the current intent label; the matching result is obtained by matching the current user attribute information of the current user with the high-risk user attribute information stored in the preset database.
[0209] In one example, the risk result includes low risk and high risk; wherein, the low risk is determined by the matching result being a match failure and the current intent label being a benign intent label, and the high risk is determined by the matching result being a match success and / or the current intent label being a non-benign intent label.
[0210] In one example, the high risk is used to intercept the current writing content if there is a risk word in the current writing content that matches a risk word stored in a preset word library.
[0211] In one example, the preset database includes an account database storing high-risk user account information and a behavior database storing high-risk behavior metrics; the current user attribute information includes current user account information and one or more behavior information of the current user; the matching result is obtained by matching the current user account information with the high-risk user account information and matching one or more behavior information of the current user with the high-risk behavior metrics.
[0212] In one example, the current writing content includes at least one of the following: instant chat messages, order evaluation content, and inquiry content with an intelligent customer service.
[0213] In one example, the label is determined by the same first intention label and second intention label; wherein, the first intention label is predicted for the training data to be labeled by using an offline intention recognition model; the second intention label is recognized for the training data to be labeled from multiple preset intention labels by using a general language model; the offline intention recognition model is trained by using a seed training dataset for a preset offline intention recognition model, and the seed training data in the seed training dataset is configured with correct intention labels selected from the multiple preset intention labels.
[0214] In one example, the seed training data is constructed by screening historical business data that matches white rule data from a historical business dataset; wherein, the historical business data in the historical business dataset includes: historical scenarios, role information of historical users, and historical writing content; the white rule data contains: benign writing content representing benign intention labels corresponding to different scenarios and different role information.
[0215] In one example, the seed training data is constructed based on each piece of historical business data that does not match the white rule data and correct intention labels; the correct intention labels are configured for multiple clusters obtained by clustering based on the text similarity between each piece of historical business data.
[0216] In one example, the online intent recognition model is used to recognize the intent of the content written by users in the instant delivery scenario; the preset intent tags include benign intent tags and non-benign intent tags, and the benign intent tags include at least one of the following: intent tags related to orders, intent tags related to goods, intent tags related to pick-up and delivery, intent tags related to after-sales, and intent tags related to user interaction.
[0217] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0218] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0219] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by the processor of the model construction device / data risk identification device to implement the method described in any one of the above embodiments.
[0220] Among them, the non-transitory computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc., and the embodiments of this specification do not limit this.
[0221] In an exemplary embodiment, a computer program product including computer programs / instructions is also provided. The above computer programs / instructions can be executed by the processor of the model construction device / data risk identification device to implement the method described in any one of the above embodiments.
[0222] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0223] Those skilled in the art will readily conceive of other embodiments of the embodiments of this specification after considering the specification and practicing the invention claimed herein. The embodiments of this specification are not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the embodiments of this specification is only limited by the appended claims.
[0224] The above are only the preferred embodiments of this specification and are not intended to limit the embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of this specification shall be included within the scope of protection of the embodiments of this specification.
Claims
1. A model building method, the method comprising: Obtaining a preset seed training data set, and using the seed training data set to train a preset offline intent recognition model; wherein the seed training data in the seed training data set is configured with a correct intent label selected from a plurality of preset intent labels; Obtain a training data set to be labeled, and use the trained offline intent recognition model to predict a first intent label of the training data to be labeled in the training data set to be labeled; Constructing prompt content including the training data to be labeled and the plurality of preset intent labels and inputting the prompt content into a general language model, so that the general language model can identify the second intent label of the training data to be labeled among the plurality of preset intent labels; Based on the to-be-labeled training data having the same first intent label and the same second intent label, the preset online intent recognition model is trained.
2. According to the method of claim 1, the seed training data set is obtained by: Acquire a historical business data set, where the historical business data in the historical business data set includes: Historical scenarios, role information of historical users, and historical writing content; Acquire preset white rule data, wherein the white rule data includes: benign writing content representing benign intention labels corresponding to different scenarios and different role information; The historical business data matching the white rule data is selected from the historical business data set to construct the seed training data.
3. The method according to claim 2, further comprising: For each piece of historical business data in the historical business data set that does not match the white rule data, clustering is performed based on text similarity between the pieces of historical business data to obtain multiple clusters; Outputting the multiple clusters to a configuration personnel, so that the configuration personnel can configure the correct intent label for each piece of historical business data based on the multiple clusters; The seed training data is constructed based on the various pieces of historical business data and the correct intent labels.
4. A data risk identification method, the method comprising: Get the current user attribute information and current writing content of the current user; Matching the current user attribute information with high-risk user attribute information stored in a preset database; Using the online intent recognition model constructed by the method according to any one of claims 1 to 3, identifying the current intent label of the currently written content; The risk result of the currently written content is determined according to the matching result and the current intention label.
5. According to the method of claim 4, the preset database includes an account database storing high-risk user account information and a behavior database storing high-risk behavior indicators; the current user attribute information includes the current user account information and one or more behavior information of the current user; The matching of the current user attribute information with the high-risk user attribute information stored in a preset database includes: Matching the current user account information with the high-risk user account information; Matching one or more behavior information of the current user with the high-risk behavior indicator.
6. A data risk identification method, the method being applied to a client, the method comprising: Get the current content written by the current user; Obtaining the risk result of the currently written content; wherein the risk result is determined based on the current intent label identified by a preset online intent recognition model for the currently written content; the label of the training data of the online intent recognition model is obtained after being identified by a trained offline intent recognition model and a general language model respectively.
7. According to the method of claim 6, the risk result is determined by combining the matching result and the current intention label; the matching result is obtained by matching the current user attribute information of the current user with the high-risk user attribute information stored in a preset database.
8. An electronic device comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 5 or the method according to any one of claims 6 to 7 by running the executable instructions.
9. A computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the method according to any one of claims 1 to 5 or the method according to any one of claims 6 to 7 is implemented.
10. A computer program product having a computer program / instruction stored thereon, wherein the computer program / instruction, when executed by a processor, implements the method according to any one of claims 1 to 5 or the method according to any one of claims 6 to 7.
Citation Information
Cited By
Risk control message intelligent matching system and method based on data analysis
CN121598319A