Method, apparatus, medium, and program product for generating annotations
By acquiring and expanding information on violation cases and combining it with background knowledge rules to input into a large language model, we solved the problem of labeling sensitive information in the training of large generative machine-reviewed text models, achieved efficient violation labeling and detailed explanations, and improved the model's recall capability.
Patent Information
- Application Number
- CN202411564061.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-04
AI Technical Summary
In the training of large generative machine-reviewed text models, labeling of sensitive information is difficult to achieve, and the demand for labelers is high, making it difficult to meet data requirements.
By obtaining violation case information that matches sensitive query information, case expansion processing is performed, and background knowledge rule information is input into the target large language model to generate accurate violation labeling information.
The recall capability of the generative machine-reviewed text model is improved, the data requirements for model training are met, and accurate violation labeling information is generated.
Smart Images

Figure CN119474338B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device, computer-readable medium, and computer program product for generating annotations. Background Art
[0002] Based on existing technology solutions, in the training of large generative machine-reviewed text models, the labeling of sensitive information becomes extremely difficult due to its obscurity. Labelers are required to have a complete knowledge background, and their number is relatively small, making it difficult to meet the data requirements of model training. Summary of the Invention
[0003] Various aspects of the present application provide a method, apparatus, electronic device, computer-readable medium, and computer program product for generating annotations.
[0004] In one aspect of the present application, a method for generating annotations is provided, wherein the method comprises:
[0005] According to preset sensitive query information, obtaining first violation case information that matches the sensitive query information;
[0006] Obtaining second violation case information by performing case expansion processing on the first violation case information;
[0007] inputting background knowledge rule information matching the first violation case information and / or the second violation case information into the target large language model;
[0008] For the violation sample data, the target large language model is used to perform labeling processing based on the first violation case information, the second violation case information and the corresponding background knowledge rule information to generate corresponding violation labeling information.
[0009] In one aspect of the present application, a device for generating annotations is provided, wherein the device comprises:
[0010] means for obtaining, based on preset sensitive query information, first violation case information that matches the sensitive query information;
[0011] means for obtaining second violation case information by performing case expansion processing on the first violation case information;
[0012] means for inputting background knowledge rule information matching the first violation case information and / or the second violation case information into a target large language model;
[0013] A device for labeling violation sample data using the target large language model based on first violation case information, second violation case information and corresponding background knowledge rule information to generate corresponding violation labeling information.
[0014] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the present application.
[0015] In another aspect of the present application, a computer program product is provided, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.
[0016] In the solution provided by the embodiment of the present application, violation cases related to sensitive data are obtained as first case information, and the first case information is expanded to obtain more violation cases as second case information, and the background knowledge and rules matching the violation cases are informed to the generative large language model, so that it can generate accurate annotations and detailed explanations for the sample data based on the first case information, the second case information, the background knowledge and the rules, thereby improving the recall ability of the generative machine-reviewed text large model and effectively meeting the training needs of the generative machine-reviewed text large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0019] Figure 1 A schematic diagram of a process for generating annotations according to an embodiment of the present application is shown;
[0020] Figure 2 A schematic structural diagram of a device for generating annotations provided in an embodiment of the present application is shown;
[0021] Figure 3 A structural diagram of a device suitable for implementing the solution in the embodiments of the present application is shown.
[0022] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0023] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPUs), input / output interfaces, network interfaces and memories.
[0025] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0026] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer program instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0027] Figure 1 A flow chart of a method for generating annotations provided in an embodiment of the present application is shown, wherein the method comprises at least steps S101, S102, S103 and S104.
[0028] In practical scenarios, the execution subject of this method can be a network device or an application running on a network device. The network device includes, but is not limited to, a network host, a single network server, a set of multiple network servers, or a collection of computers based on cloud computing, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a group of loosely coupled computers forming a virtual computer.
[0029] In the machine review scenario, the method of the embodiment of the present application is implemented by a generative large language model. The embodiment of the present application obtains violation cases related to sensitive data, expands the violation cases to obtain more violation cases, and informs the generative large language model of the background knowledge and rules that match the violation cases. Based on the violation cases, background knowledge and rules, the generative large language model generates accurate annotations and detailed explanations for the sample data, effectively meeting the training requirements of the generative machine review text large model.
[0030] The concepts involved in the embodiments of this application are explained below.
[0031] Large Language Model (LLM): Large language models are commonly used in natural language processing (NLP) to handle a variety of natural language tasks, such as text classification, question answering, and conversation. They are used to generate natural language text or understand the meaning of text. Large language models are also a general term for deep learning models trained using large amounts of text data. Models such as GPT-3, PaLM, Galactica, and LLaMA are all commonly used by those skilled in the art.
[0032] Reference Figure 1 In step S101, according to preset sensitive query information, first violation case information matching the sensitive query information is obtained.
[0033] The sensitive query information includes various information that can be used to query violation cases.
[0034] The violation case information may include various types of text information, such as a text description of the violation.
[0035] Optionally, violation cases in the embodiments of the present application include but are not limited to cases where reported comments, reported barrages, and reported tags are judged to be violations based on preset rules or review agencies.
[0036] Optionally, violation cases may also include source information or type information. Source information may include, but is not limited to, all source cases, system-banned cases, etc. Type information may include, but is not limited to, all types of cases, comment cases, barrage cases, tag cases, or submission cases.
[0037] Optionally, the sensitive query information includes preset sensitive words and violation tags. In step S101, a query is performed in a predetermined case library based on the preset sensitive words and violation tags, and the matching violation cases obtained by the query are used as the first violation case information.
[0038] Optionally, the acquired violation case is processed according to a predetermined format, and the processed case information in the predetermined format is used as the first violation case information.
[0039] For example, for violation cases of comment barrage, the corresponding video manuscript information (video title, tags, introduction) is added, and this information and the text information of the illegal comment barrage content are used as the first violation case information.
[0040] In step S102, second violation case information is obtained by performing case expansion processing on the first violation case information.
[0041] The methods for expanding the first violation case information include but are not limited to:
[0042] 1) Expanding the case by rewriting the first violation case information;
[0043] Specifically, step S102 according to this method includes step S1021.
[0044] In step S1021 , a predetermined number of violation cases similar to the first violation case are generated as second violation case information by rewriting part of the violation case information.
[0045] There are many ways to rewrite. For example, some keywords in the text of the violation case can be replaced with their synonyms or near synonyms through synonym replacement. Another example is to use Internet meme replacement to replace the words or phrases in the original text with popular Internet memes or expressions commonly used in specific communities. Another example is to change the background or environment of the violation case through context conversion.
[0046] 2) Case expansion processing by defining personas in the target large language model;
[0047] Specifically, according to this method, step S102 includes step S1022 and step S1023.
[0048] According to one embodiment, in step S1022, a target persona is defined in the target large language model. Optionally, the target persona includes various people who are good at processing text, such as a writer, a network expert, etc.
[0049] Character design refers to the personality traits, appearance, behavior, and character traits that are set for a character during the creation, writing, and role-playing of a virtual character. Character design can be achieved by specifying various factors, such as the character's gender, age, occupation, hobbies, and language style.
[0050] For example, in a large language model, the character's personality traits can be gradually revealed by inputting command text during the conversation, making it more in line with user needs during the conversation.
[0051] In step S1023 , based on the target persona, a predetermined number of violation cases that are similar to the first violation case and conform to the target persona are generated in the target large language model as second violation case information.
[0052] 3) Case expansion is performed by combining the persona defined in the target large language model and rewriting the first violation case information;
[0053] Specifically, according to this method, step S102 includes step S1024 and step S1025.
[0054] According to one embodiment, in step S1024 , a target persona is defined in the target large language model.
[0055] In step S1025, by inputting instructions to the target large language model, the target large language model is caused to rewrite part of the information of the first violation case based on the target persona, and generate a predetermined number of violation cases similar to the first violation case as second violation cases.
[0056] According to the first example of this application, in the scenario of machine review of comments on video manuscripts, this example screens the case library based on preset sensitive words and violation tags to obtain matching violation cases, and uses a large language model to generate more violation cases based on the obtained violation cases. The violation cases in this example include the text information of the video manuscript comments that are judged to be illegal.
[0057] This example inputs prompts into the large language model, causing it to generate similar violation comments to the violation comments in the violation case, and then generates a new violation case based on the obtained similar violation comments.
[0058] The specific rules are as follows: Create a seed comment group (the comments themselves are illegal). Then, for the same article, if there are more than 5 comments on the article, add all the comments to the sample prompt for generating a violation case as shown below. If the article has less than 5 comments, randomly select comments from the seed comment group to make up the total 5 comments.
[0059] In this example, the prompt structure for generating a new violation case is as follows:
[0060] The following is a piece of background material: {Fill in the text content of the background material}.
[0061] The background material is used to explain what types of speech are considered illegal in the online environment, such as attacks or satires targeting specific issues.
[0062] Now there is a video from the xx partition.
[0063] The title is: {fill in the title text content};
[0064] The label is: {fill in the label text content};
[0065] The introduction is: {fill in the introduction text};
[0066] The root comment is: {fill in the root comment text};
[0067] An audience member made the following comment:
[0068] Comment 1: {fill in the text content of comment 1};
[0069] Comment 2: {fill in the text content of comment 2};
[0070] Comment 3: {fill in the text content of comment 3};
[0071] Comment 4: {fill in the text content of comment 4};
[0072] Comment 5: {fill in the text content of Comment 5}.
[0073] In this example, the large language model generates similar illegal comments in the following ways:
[0074] 1) Define a surfing expert persona in the large language model and have it generate similar violating comments based on this persona and the input contextual information. Specifically, the input is: "Based on the contextual information, the above five comments require further review. If you are a surfing expert who likes to use memes, please write 10 other violating comments based on the context."
[0075] 2) Define a surfing expert persona in the large language model, and use the large language model to generate similar illegal comments based on the persona, input background information, and variant descriptions of online memes.
[0076] By inputting "Based on the background material, the above 5 comments all need further review" into the large language model, if you are a surfing expert, your ability is to be good at playing with memes, and you will describe the Internet memes with variants. For example, the variant of '永永的神' can be 'yyds'. You can also describe other Internet memes with variants by yourself. Please write down the other 10 comments that violate the rules based on the background material", the large language model will output new illegal comments.
[0077] Continue to refer to the following Figure 1 To illustrate, in step S103 , background knowledge rule information matching the first violation case information and / or the second violation case information is input into the target large language model.
[0078] The background knowledge rule information includes various types of background knowledge information and rule information that can be used by the target large language model to judge violations, such as text data such as material information, factual knowledge, etc. from different sources such as news, blogs, and forums.
[0079] Optionally, the method analyzes the first violation case and / or the second violation case to determine one or more sensitive elements involved, and then collects data related to the sensitive elements as background knowledge rule information matching the first violation case and the second violation case.
[0080] For example, for violations involving sensitive content, relevant historical data and recent developments are collected. Based on this data, existing rules are refined to ensure they accurately cover different types of violations. Furthermore, a comprehensive background knowledge base is established and maintained, encompassing information on political figures, events, organizations, and related terminology. This knowledge base serves as the foundation for identifying and assessing violations.
[0081] According to one embodiment, the method periodically obtains background knowledge rule information matching the first violation case information and / or the second violation case information, and updates the background knowledge rule information input into the target large language model based on the obtained information.
[0082] In step S104 , the target large language model is used to perform labeling processing on the violation sample data based on the first violation case information, the second violation case information and the corresponding background knowledge rule information to generate corresponding violation labeling information.
[0083] The violation marking information includes a violation determination result indicating whether a violation exists or does not exist. Optionally, the violation marking information also includes explanatory information describing the corresponding violation point. The explanatory information may include a detailed description of the violation point and the basis for determining the violation.
[0084] The data annotated by the target large language model can be used as training data to train the large model used for machine review.
[0085] Optionally, data annotation requirements are predefined in the target large language model.
[0086] The data annotation requirements are used to indicate the aspects and information from which the target large language model needs to make judgments, and what conclusions it ultimately gives.
[0087] Optionally, the method defines data annotation requirements to indicate the data format or data structure of the illegal annotation information that the target large language model needs to output. Optionally, the data format or data structure of the illegal annotation information is presented in a format that is easy for manual review and modification, facilitating manual verification and annotation.
[0088] Continuing with the first example, a large amount of sample data is input into a large language model, which generates violation annotation information for each comment based on violation cases and corresponding background knowledge / rules. The violation annotation information includes the violation annotation results and an explanation of the violation. Violation cases include those derived based on preset sensitive words and violation labels, as well as new violation cases generated by the large language model. The annotated data can be used to train the generative large model.
[0089] According to the method of the embodiment of the present application, by obtaining violation cases related to sensitive data as first case information, and expanding the first case information to obtain more violation cases as second case information, and informing the generative large language model of the background knowledge and rules matching the violation cases, it generates accurate annotations and detailed explanations for the sample data based on the first case information, the second case information, the background knowledge and the rules, thereby improving the recall capability of the generative machine-reviewed text large model and effectively meeting the training needs of the generative machine-reviewed text large model.
[0090] In addition, the present invention also provides a device for generating annotations, the structure of which is as follows: Figure 2 As shown. The device includes: a device for obtaining first violation case information matching the sensitive query information according to preset sensitive query information (hereinafter referred to as "case obtaining device 101"), a device for obtaining second violation case information by performing case expansion processing on the first violation case information (hereinafter referred to as "case expansion device 102"), a device for inputting background knowledge rule information matching the first violation case information and / or the second violation case information into a target large language model (hereinafter referred to as "knowledge rule input device 103"), and a device for using the target large language model to perform annotation processing on violation sample data based on the first violation case information, the second violation case information and the corresponding background knowledge rule information to generate corresponding violation annotation information (hereinafter referred to as "annotation generating device 104").
[0091] Reference Figure 2 The case acquisition device 101 acquires first violation case information matching the preset sensitive query information based on the preset sensitive query information.
[0092] The sensitive query information includes various information that can be used to query violation cases.
[0093] The violation case information may include various types of text information, such as a text description of the violation.
[0094] Optionally, violation cases in the embodiments of the present application include but are not limited to cases where reported comments, reported barrages, and reported tags are judged to be violations based on preset rules or review agencies.
[0095] Optionally, violation cases may also include source information or type information. Source information may include, but is not limited to, all source cases, system-banned cases, and so on. Type information includes, but is not limited to, all types of cases, comment cases, barrage cases, tag cases, or submission cases.
[0096] Optionally, the sensitive query information includes preset sensitive words and violation labels. The case acquisition device 101 searches a predetermined case library based on the preset sensitive words and violation labels, and uses the matching violation cases obtained from the search as the first violation case information.
[0097] Optionally, the case acquisition device 101 processes the acquired violation case according to a predetermined format, and uses the processed case information in the predetermined format as the first violation case information.
[0098] For example, for violation cases of comment barrage, the corresponding video manuscript information (video title, tags, introduction) is added, and this information and the text information of the illegal comment barrage content are used as the first violation case information.
[0099] The case expansion device 102 obtains second violation case information by performing case expansion processing on the first violation case information.
[0100] The method in which the case expansion device 102 expands the first violation case information includes but is not limited to:
[0101] 1) Expanding the case by rewriting the first violation case information;
[0102] Specifically, the case expansion device 102 according to this method rewrites part of the information of the violation case to generate a predetermined number of violation cases similar to the first violation case as the second violation case information.
[0103] Among them, various ways can be used for rewriting, for example, some keywords in the text content of the violation case are replaced by their synonyms or near synonyms through synonym replacement. For example, by replacing the network segment, the words or phrases in the original text are replaced with popular network segments or expressions commonly used in a particular community. For example, by changing the background or environment in which the violation case occurs,
[0104] 2) Case expansion processing is performed by defining a persona in the target large language model;
[0105] Specifically, the case expansion device 102 according to this mode defines a target persona in the target large language model. Optionally, the target persona includes various personas skilled in processing text, such as text workers and network experts.
[0106] Among them, the persona refers to the personality characteristics, appearance description, behavior habits, character traits, etc. set for the role in the process of creating virtual characters, writing, role-playing, etc. Defining a persona can be achieved by specifying the gender, age, occupation, interests, language style, etc. of the role.
[0107] For example, in a large language model, the characteristics of the role can be gradually revealed by inputting instruction text during the conversation process, so that it is more in line with user needs during the conversation process.
[0108] Next, the case expansion device 102 generates a predetermined number of violation cases similar to the first violation case in the target large language model based on the target persona as the second violation case information.
[0109] 3) Case expansion processing is performed by combining the definition of the persona in the target large language model and rewriting the first violation case information;
[0110] Specifically, the case expansion device 102 according to this mode defines a target persona in the target large language model. Next, by inputting instructions to the target large language model, the target large language model is instructed to rewrite part of the information of the first violation case based on the target persona, and a predetermined number of violation cases similar to the first violation case are generated as the second violation case.
[0111] According to the first example of the present application, in the scene of machine auditing of video article comments, the present example is based on the pre-set sensitive words and violation labels to screen in the case library to obtain the matched violation case (case), and uses the large language model to generate more violation cases based on the obtained violation case. The violation case of the present example includes the video article comment text information judged to be in violation.
[0112] In this example, the case expansion device 102 inputs prompts to the large language model, causing the large language model to generate similar violation comments to the violation comments in the violation case, and then generates a new violation case based on the obtained similar violation comments.
[0113] The specific rules are as follows: Create a seed comment group (the comments themselves are illegal). Then, for the same article, if there are more than 5 comments on the article, add all the comments to the sample prompt for generating a violation case as shown below. If the article has less than 5 comments, randomly select comments from the seed comment group to make up the total 5 comments.
[0114] In this example, the prompt structure for generating a new violation case is as follows:
[0115] The following is a piece of background material: {Fill in the text content of the background material}.
[0116] The background material is used to explain what types of speech are considered illegal in the online environment, such as attacks or satires targeting specific issues.
[0117] Now there is a video from the xx partition.
[0118] The title is: {fill in the title text content};
[0119] The label is: {fill in the label text content};
[0120] The introduction is: {fill in the introduction text};
[0121] The root comment is: {fill in the root comment text};
[0122] An audience member made the following comment:
[0123] Comment 1: {fill in the text content of comment 1};
[0124] Comment 2: {fill in the text content of comment 2};
[0125] Comment 3: {fill in the text content of comment 3};
[0126] Comment 4: {fill in the text content of comment 4};
[0127] Comment 5: {fill in the text content of Comment 5}.
[0128] In this example, the large language model generates similar illegal comments in the following ways:
[0129] 1) Define a surfing expert persona in the large language model and have it generate similar violating comments based on this persona and the input contextual information. Specifically, the input is: "Based on the contextual information, the above five comments require further review. If you are a surfing expert who likes to use memes, please write 10 other violating comments based on the context."
[0130] 2) Define a surfing expert persona in the large language model, and use the large language model to generate similar illegal comments based on the persona, input background information, and variant descriptions of online memes.
[0131] By inputting "Based on the background material, the above 5 comments all need further review" into the large language model, if you are a surfing expert, your ability is to be good at playing with memes, and you will describe the Internet memes with variants. For example, the variant of '永永的神' can be 'yyds'. You can also describe other Internet memes with variants by yourself. Please write down the other 10 comments that violate the rules based on the background material", the large language model will output new illegal comments.
[0132] Continue to refer to the following Figure 2 To illustrate, the knowledge rule input device 103 inputs background knowledge rule information that matches the first violation case information and / or the second violation case information into the target large language model.
[0133] The background knowledge rule information includes various types of background knowledge information and rule information that can be used by the target large language model to judge violations, such as text data such as material information, factual knowledge, etc. from different sources such as news, blogs, and forums.
[0134] Optionally, the device analyzes the first violation case and / or the second violation case to determine one or more sensitive elements involved, and then collects data related to the sensitive elements as background knowledge rule information matching the first violation case and the second violation case.
[0135] For example, for violations involving sensitive content, relevant historical data and recent developments are collected. Based on this data, existing rules are refined to ensure they accurately cover different types of violations. Furthermore, a comprehensive background knowledge base, encompassing information on political figures, events, organizations, and related terminology, is established and maintained. This knowledge base serves as the foundation for identifying and assessing violations.
[0136] According to one embodiment, the apparatus periodically obtains background knowledge rule information that matches the first violation case information and / or the second violation case information, and updates the background knowledge rule information input into the target large language model based on the obtained information.
[0137] The annotation generating device 104 uses the target large language model to perform annotation processing on the violation sample data based on the first violation case information, the second violation case information and the corresponding background knowledge rule information to generate corresponding violation annotation information.
[0138] The violation marking information includes a violation determination result indicating whether a violation exists or does not exist. Optionally, the violation marking information also includes explanatory information describing the corresponding violation point. The explanatory information may include a detailed description of the violation point and the basis for determining the violation.
[0139] The data annotated by the target large language model can be used as training data to train the large model used for machine review.
[0140] Optionally, data annotation requirements are predefined in the target large language model.
[0141] The data annotation requirements are used to indicate the aspects and information from which the target large language model needs to make judgments, and what conclusions it ultimately gives.
[0142] Optionally, the device defines data annotation requirements to indicate the data format or data structure of the illegal annotation information that the target large language model needs to output. Optionally, the data format or data structure of the illegal annotation information is presented in a format that is easy for manual review and modification, facilitating manual verification and annotation.
[0143] Continuing with the first example, the knowledge rule input device 103 inputs a large amount of sample data into the large language model, which then generates violation annotation information for each comment based on the violation cases and the corresponding background knowledge / rules. The violation annotation information includes the violation annotation results and an explanation of the violation point. Violation cases include those derived based on preset sensitive words and violation labels, as well as new violation cases generated by the large language model. The annotated data can be used to train the generative large model.
[0144] According to the device of the embodiment of the present application, by obtaining violation cases related to sensitive data as first case information, and expanding the first case information to obtain more violation cases as second case information, and informing the generative large language model of the background knowledge and rules matching the violation cases, it generates accurate annotations and detailed explanations for the sample data based on the first case information, the second case information, the background knowledge and the rules, thereby improving the recall capability of the generative machine-reviewed text large model and effectively meeting the training needs of the generative machine-reviewed text large model.
[0145] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. The method corresponding to the electronic device may be the method for generating annotations in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.
[0146] The electronic device may be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablets, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer collections, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computers.
[0147] Figure 3 The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown. The device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 into the random access memory (RAM) 1203. Various programs and data required for system operation are also stored in RAM 1203. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through a bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.
[0148] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, and the like; an output section 1207 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, and a speaker; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, and a semiconductor memory; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 1209 performs communication processing via a network such as the Internet.
[0149] In particular, the methods and / or embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above-mentioned functions defined in the method of the present application are performed.
[0150] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.
[0151] Specifically, the present embodiment can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0152] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0153] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0154] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0155] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.
[0156] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, device and unit described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0157] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or page components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0158] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0159] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.
[0160] The integrated unit realized in the form of software functional unit can be stored in a computer readable storage medium. The software functional unit is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0162] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.
Claims
1. A method for generating annotations, wherein: The method comprises: According to preset sensitive query information, obtaining first violation case information that matches the sensitive query information; Obtaining second violation case information by performing case expansion processing on the first violation case information; inputting background knowledge rule information matching the first violation case information and / or the second violation case information into the target large language model; For the violation sample data, the target large language model is used to perform labeling processing based on the first violation case information, the second violation case information and the corresponding background knowledge rule information to generate corresponding violation labeling information; The method further comprises: Analyze the first violation case and / or the second violation case to determine one or more sensitive elements involved; Data related to the sensitive element is collected as background knowledge rule information matching the first violation case and the second violation case.
2. The method according to claim 1, wherein The second violation case information obtained by performing case expansion processing on the first violation case information includes: Define the target persona in the target large language model; Based on the target persona, a predetermined number of violation cases that are similar to the first violation case and conform to the target persona are generated in the target large language model as second violation case information.
3. The method according to claim 2, wherein: The generating, in the target large language model based on the target persona, a predetermined number of violation cases that are similar to the first violation case and conform to the target persona as second violation case information includes: By inputting instructions into the target large language model, the target large language model is caused to rewrite part of the information of the first violation case based on the target persona, and generate a predetermined number of violation cases similar to the first violation case as second violation cases.
4. The method according to claim 1, wherein The second violation case information obtained by performing case expansion processing on the first violation case information includes: A predetermined number of violation cases similar to the first violation case are generated as second violation case information by rewriting part of the violation case information.
5. The method according to claim 1, wherein The method further comprises: Periodically acquiring background knowledge rule information that matches the first violation case information and / or the second violation case information; The background knowledge rule information input to the target large language model is updated based on the acquired information.
6. The method according to claim 1, wherein The sensitive query information includes preset sensitive words and violation labels, and obtaining first violation case information matching the sensitive query information according to the preset sensitive query information includes: A query is performed in a predetermined case library based on preset sensitive words and violation tags, and matching violation cases obtained from the query are used as first violation case information.
7. A device for generating annotations, wherein: The device comprises: means for obtaining, based on preset sensitive query information, first violation case information that matches the sensitive query information; means for obtaining second violation case information by performing case expansion processing on the first violation case information; means for inputting background knowledge rule information matching the first violation case information and / or the second violation case information into a target large language model; A device for labeling the violation sample data using the target large language model based on the first violation case information, the second violation case information, and the corresponding background knowledge rule information to generate corresponding violation labeling information; Wherein, the device is also used for: Analyze the first violation case and / or the second violation case to determine one or more sensitive elements involved; Data related to the sensitive element is collected as background knowledge rule information matching the first violation case and the second violation case.
8. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Text data sample determination method and text data processing model determination method
CN117056521A
Multi-modal large model fine tuning method and device, computer equipment and storage medium
CN117095257A