Sample construction method and device

By filtering and dividing dialogue samples in multiple historical dialogue sequences, the problem of low accuracy of dialogue data compliance detection in the prior art is solved, and more efficient dialogue content quality inspection is achieved.

CN115712712BActive Publication Date: 2025-08-19BEIJING ZHENGUANYU TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211465617.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-08-19
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

In the prior art, dialogue data compliance detection relies on manual reading and keyword retrieval, resulting in high human resources consumption and low accuracy, single sample, high probability of error recall, and low prediction accuracy.

Method used

By obtaining multiple historical conversation sequences, filtering the initial conversation sequence containing keywords, generating the initial conversation sample, and dividing it into the first positive conversation sample and the second negative conversation sample according to the attribute information, and storing it in the corresponding set, improving sample diversity.

Benefits of technology

The prediction accuracy of the detection model is improved, and the accuracy of the quality inspection of dialogue content is improved through diversity sample training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115712712B_ABST
    Figure CN115712712B_ABST
Patent Text Reader

Abstract

This specification provides a sample construction method and device, wherein the sample construction method includes: obtaining multiple historical dialogue sequences, taking at least two dialogue sequences containing keywords in the multiple historical dialogue sequences as initial dialogue sequences, and screening a first negative dialogue sequence from the multiple historical dialogue sequences; generating initial dialogue samples corresponding to the at least two initial dialogue sequences, and a first negative dialogue sample corresponding to the first negative dialogue sequence; dividing the at least two initial dialogue samples into a first positive dialogue sample and a second negative dialogue sample based on attribute information of the at least two initial dialogue samples, wherein the first positive dialogue sample and the second negative dialogue sample both contain keywords; storing the first negative dialogue sample and the second negative dialogue sample in a negative dialogue sample set, and storing the first positive dialogue sample in a positive dialogue sample set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and more particularly to a sample construction method, a sample construction apparatus, a computing device, and a computer-readable storage medium. Background Art

[0002] With the development of internet technology, online services are becoming increasingly integrated into people's lives and learning. This online communication model generates a large amount of conversation data. By monitoring this conversation data, we can determine whether service providers are using non-compliant service methods or language when providing consulting, problem-solving, and other services.

[0003] Conventional methods typically rely on manual reading of conversation data and keyword retrieval for compliance testing. However, manual reading consumes significant human resources and has low accuracy. Keyword retrieval, which directly detects keywords based on conversation data, uses a relatively limited sample size and has significant limitations, resulting in a high probability of false recall and low prediction accuracy. Therefore, a sample construction method is urgently needed to address these issues. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide a sample construction method, a sample construction apparatus, a computing device, and a computer-readable storage medium to address the technical deficiencies in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a sample construction method is provided, including:

[0006] Acquire multiple historical dialogue sequences, select at least two dialogue sequences containing the keyword from the multiple historical dialogue sequences as initial dialogue sequences, and select a first negative dialogue sequence from the multiple historical dialogue sequences;

[0007] generating initial dialogue samples corresponding to at least two initial dialogue sequences, and a first negative dialogue sample corresponding to the first negative dialogue sequence;

[0008] Dividing the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample according to attribute information of the at least two initial conversation samples, wherein both the first positive conversation sample and the second negative conversation sample contain keywords;

[0009] The first negative conversation sample and the second negative conversation sample are stored in a negative conversation sample set, and the first positive conversation sample is stored in a positive conversation sample set.

[0010] According to a second aspect of the embodiments of this specification, there is provided a sample construction device, including:

[0011] an acquisition module configured to acquire a plurality of historical dialogue sequences, select at least two dialogue sequences containing the keyword from the plurality of historical dialogue sequences as initial dialogue sequences, and screen a first negative dialogue sequence from the plurality of historical dialogue sequences;

[0012] a generating module configured to generate initial dialogue samples corresponding to at least two initial dialogue sequences, and a first negative dialogue sample corresponding to the first negative dialogue sequence;

[0013] a division module configured to divide the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample based on attribute information of the at least two initial conversation samples, wherein both the first positive conversation sample and the second negative conversation sample contain keywords;

[0014] The storage module is configured to store the first negative conversation sample and the second negative conversation sample into a negative conversation sample set, and store the first positive conversation sample into a positive conversation sample set.

[0015] According to a third aspect of an embodiment of this specification, a computing device is provided, including:

[0016] memory and processor;

[0017] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the sample construction method.

[0018] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the sample construction method are implemented.

[0019] The sample construction method provided in this specification obtains multiple historical dialogue sequences, takes at least two dialogue sequences containing keywords in the multiple historical dialogue sequences as initial dialogue sequences, and screens a first negative dialogue sequence from the multiple historical dialogue sequences; generates initial dialogue samples corresponding to the at least two initial dialogue sequences, and a first negative dialogue sample corresponding to the sample-constructed first negative dialogue sequence; divides the at least two initial dialogue samples into a first positive dialogue sample and a second negative dialogue sample based on attribute information of the at least two initial dialogue samples, wherein the sample-constructed first positive dialogue sample and the sample-constructed second negative dialogue sample both contain keywords; stores the sample-constructed first negative dialogue sample and the sample-constructed second negative dialogue sample in a negative dialogue sample set, and stores the sample-constructed first positive dialogue sample in a positive dialogue sample set.

[0020] One embodiment of this specification implements the selection of at least two initial conversation sequences containing keywords from multiple historical conversation sequences, and then classifies the initial conversation samples into first positive conversation samples and second negative conversation samples based on the attribute information of the initial conversation samples corresponding to each initial conversation sequence. By screening the first negative conversation sequences from the multiple historical conversation sequences to generate first negative conversation samples, the diversity of the samples is increased. Subsequently, a detection model is trained based on the first negative conversation samples, the first positive conversation samples, and the second negative conversation samples, thereby improving the prediction accuracy of the detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a sample construction diagram of a sample construction method provided in an embodiment of this specification;

[0022] Figure 2 This is a flow chart of a sample construction method provided in an embodiment of this specification;

[0023] Figure 3 is a schematic diagram of a sample construction method provided in an embodiment of this specification;

[0024] Figure 4 This is a processing flow chart of a sample construction method applied to conversation data provided in an embodiment of this specification;

[0025] Figure 5 This is a schematic structural diagram of a sample construction device provided in one embodiment of this specification;

[0026] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. DETAILED DESCRIPTION

[0027] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0028] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0029] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0030] First, the terms involved in one or more embodiments of this specification are explained.

[0031] BERT (Bidirectional Encoder Representation from Transformers): A pre-training technology for natural language processing. BERT uses a large amount of unsupervised data to pre-train a neural network stacked with Transformers, which is then applied to downstream tasks. Transformers can encode bidirectional information between words, enabling better text understanding.

[0032] Conversation content quality control: Natural language processing technology is used to determine whether there are any violations in the conversation. The main quality control content includes illegal words, illegal behaviors, service attitude, etc.

[0033] Focal Loss: Focal loss (focus loss function) is mainly used to solve the problem of serious imbalance in the ratio of positive and negative samples in supervised machine learning scenarios. By designing a new loss function, the model can automatically allocate sample weights during training to achieve the purpose of balancing positive and negative samples.

[0034] Positive samples and negative samples: In quality inspection tasks, illegal samples are recorded as positive samples, and compliant samples are recorded as negative samples.

[0035] In this specification, a sample construction method is provided. This specification also relates to a sample construction device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0036] With the advancement of computer technology, online services are becoming increasingly popular, allowing service providers to communicate with users through voice, text, and other means. Application scenarios for online services include, but are not limited to, merchandise trading, communication between teachers and students, communication between teachers and parents, consulting services, and rental services. Online services generate a large amount of conversation data. To ensure service quality, conversation content can be quality-checked. Specifically, natural language processing technology can be used to determine whether any violations are present. Key quality-check areas include, but are not limited to, illegal language, illegal behavior, and service attitude.

[0037] One embodiment of this specification implements the selection of at least two initial conversation sequences containing keywords from multiple historical conversation sequences, and then classifies the initial conversation samples into first positive conversation samples and second negative conversation samples based on the attribute information of the initial conversation samples corresponding to each initial conversation sequence. By screening the first negative conversation sequences from the multiple historical conversation sequences to generate first negative conversation samples, the diversity of the samples is increased. Subsequently, a detection model is trained based on the first negative conversation samples, the first positive conversation samples, and the second negative conversation samples, thereby improving the prediction accuracy of the detection model.

[0038] Figure 1 This is a sample construction diagram of a sample construction method provided in an embodiment of this specification. In scenarios such as communication between teachers and students, parents, and customer service personnel and customers, a large amount of dialogue data in the form of text or voice will be generated. Figure 1 As shown, a one-to-one or one-to-many conversation between a teacher and a student or parent is taken as a historical conversation sequence. A historical conversation sequence can also be all the conversation data generated in a chat group within a period of time. Then the conversation between a teacher and multiple parents or students is a plurality of historical conversation sequences. After obtaining the conversation in voice form, the voice is converted into a conversation sequence in text form. The historical conversation sequences are matched according to a preset keyword table to determine at least one initial conversation sequence containing keywords in the multiple historical conversation sequences, wherein the keywords refer to illegal words, including but not limited to uncivilized language, words with bad attitudes, etc. The historical conversation sequences are randomly sampled to obtain a first negative conversation sequence. Each conversation sequence in the initial conversation sequence is integrated to obtain an initial conversation sample, and the first negative conversation sequence is integrated to obtain a first negative conversation sample.

[0039] Based on the attribute information of the initial conversation sample, the initial conversation sample is divided into a first positive conversation sample and a second negative conversation sample. The first positive conversation sample contains keywords and is a violation sample, while the second negative conversation sample contains keywords but is not a violation sample. For example, "private contact information" is a violation word, i.e., a keyword. The sample "Please add my private contact information" is a violation sample containing keywords. Correspondingly, the sample "We do not allow teachers to provide private contact information" is a compliance sample containing keywords. The first positive conversation sample is stored in the first positive conversation sample set, and the first negative conversation sample and the second negative conversation sample are stored in the first negative conversation sample set. The conversation samples to be processed are then extracted from the first positive conversation sample set and the first negative conversation sample set for model training. This enables the model to detect keywords and determine whether a sample is a violation sample, thereby achieving conversation content quality inspection and improving the accuracy of conversation content quality inspection.

[0040] Figure 2 A flow chart of a sample construction method provided according to an embodiment of this specification is shown, which specifically includes the following steps:

[0041] Step S202: Acquire multiple historical dialogue sequences, use at least two dialogue sequences containing keywords from the multiple historical dialogue sequences as initial dialogue sequences, and screen a first negative dialogue sequence from the multiple historical dialogue sequences.

[0042] Specifically, a historical conversation sequence refers to conversation data in the form of voice, audio recording, or text generated by communication between a service provider and a service recipient. Service providers include but are not limited to teachers, sellers, merchants, customer service personnel, etc., and corresponding service recipients are consumers such as students, parents, and buyers. A historical conversation sequence can be conversation data generated by one-on-one communication between a service provider and a service recipient over a period of time, or it can be conversation data generated by communication between a service provider and multiple service recipients through a communication group. Keywords refer to illegal words, including but not limited to uncivilized language and words with bad attitudes. An initial conversation sequence refers to a historical conversation sequence containing keywords among multiple historical conversation sequences, and a first negative conversation sequence refers to a historical conversation sequence randomly selected from multiple historical conversation sequences.

[0043] Based on this, when a customer service representative and a user generate conversation data, the one-on-one conversation data between the customer service representative and the user is used as a historical conversation sequence. Multiple historical conversation sequences are obtained, and keyword matching is performed on each historical conversation sequence based on a pre-built keyword table. At least two conversation sequences containing the keyword in the multiple historical conversation sequences are used as the initial conversation sequence. At least two historical conversation sequences are randomly selected from the multiple historical conversation sequences as the first negative conversation sequence.

[0044] In practical applications, when acquiring multiple historical conversation sequences, conversation data generated between multiple customer service personnel and corresponding users within a certain timeframe can be used as the historical conversation sequence. Alternatively, conversation data generated between a single customer service personnel and multiple users can be used as the historical conversation sequence. For example, within a day or week, all conversation data generated between all customer service personnel and multiple users within a single day can be acquired and used as the multiple historical conversation sequences. Alternatively, conversation data generated between customer service personnel A and all users within a week can be used as the multiple historical conversation sequences. When selecting the first negative conversation sequence from the multiple historical conversation sequences, the multiple historical conversation sequences can be randomly sampled. Since the number of conversation sequences containing the keyword is relatively small, the first negative conversation sequences obtained from the random sampling of the multiple historical conversation sequences are all considered compliant conversation sequences. By performing keyword matching on the multiple historical conversation sequences, violations can be detected for specific scenarios. This can then be used to expand the conversation data of the customer service personnel and detect violations in the conversation data between the customer service personnel and other users.

[0045] Furthermore, considering that there are a large number of keywords and that the definition of keywords may vary due to different conversation scenarios, it is necessary to set up a keyword table in advance and store words that can be used as keywords in the keyword table so that historical conversation sequences can be matched according to the keyword table. The specific implementation is as follows:

[0046] Keyword matching is performed on each historical dialogue sequence based on a preset keyword table, and at least two dialogue sequences containing keywords in the preset keyword table are used as initial dialogue sequences.

[0047] Specifically, the preset keyword table refers to a pre-set keyword set consisting of keywords. Whether a word is a keyword is determined based on the semantics and emotional color of the word. If the word is stored in the keyword table, the keyword refers to an illegal word. In the scenario of communication between teachers and students, "playing games", "adding private contact information", "husband" and words with insulting or uncivilized meanings are illegal words; in the scenario of communication between teachers and parents, "withdrawing from class", "complaining", "adding private contact information" and words with insulting or uncivilized meanings are illegal words. The preset keyword table can be expanded, deleted, modified and adjusted according to actual conversation data.

[0048] Based on this, keyword matching is performed on each historical conversation sequence based on a preset keyword table. At least two conversation sequences from the multiple historical conversation sequences that contain keywords from the preset keyword table are used as initial conversation sequences. The initial conversation sequences contain at least one keyword. When the number of historical conversation sequences is large, a certain ratio of historical conversation sequences containing keywords to historical conversation sequences not containing keywords is likely to exist in the multiple historical conversation sequences. The multiple historical conversation sequences represent at least two historical conversation sequences.

[0049] For example, in scenarios where salespeople communicate with buyers, to ensure service quality and optimize salespeople's communication methods, quality control is often performed on the conversations generated during the conversations. This process detects whether the salesperson's speech contains any offensive terms. These include terms that lead buyers to file complaints against the seller and terms that contain insults or other uncivilized language. A keyword table is constructed from these offensive terms, and the conversation data is then tested for the presence of these keywords. Conversations containing these keywords are then identified as offensive conversation sequences.

[0050] To sum up, by presetting a keyword table, keyword matching is performed on each historical dialogue sequence based on the preset keyword table, and then at least two dialogue sequences containing keywords in the preset keyword table are used as initial dialogue sequences, thereby improving the accuracy and standardization of the determination of the initial dialogue sequence.

[0051] Furthermore, considering that there are a large number of historical dialogue sequences and that the proportion of historical dialogue sequences containing keywords is relatively small among a large number of historical dialogue sequences, a random sampling method can be used to randomly select a historical dialogue sequence from multiple historical dialogue sequences as the first negative dialogue sequence. The specific implementation is as follows:

[0052] Perform random sampling on multiple historical dialogue sequences to obtain the first negative dialogue sequence.

[0053] Based on this, random sampling refers to random selection, that is, randomly selecting a portion of historical conversation sequences from multiple historical conversation sequences as the first negative conversation sequence. After obtaining multiple historical conversation sequences, since relatively few of these historical conversation sequences contain keywords, random sampling can be used to randomly select a portion of the historical conversation data as the first negative conversation sequence. Alternatively, one can first determine which historical conversation sequences do not contain keywords from the multiple historical conversation sequences, and then select a certain number of these historical conversation sequences as the first negative conversation sequence.

[0054] Continuing with the above example, when there are 40,000 historical dialogue sequences, 20,000 of them are randomly selected as the first negative dialogue sequences. Alternatively, 38,000 of the 40,000 historical dialogue sequences that do not contain keywords are first identified, and 20,000 of these 38,000 are randomly selected as the first negative dialogue sequences. The easy negative samples are then generated based on these first negative dialogue sequences.

[0055] In summary, the first negative dialogue sequence is determined from multiple historical dialogue sequences by random sampling, thereby improving the randomness of the determination of the first negative dialogue sequence and making the first negative dialogue sequence more representative.

[0056] Step S204: generating initial dialogue samples corresponding to at least two initial dialogue sequences and a first negative dialogue sample corresponding to the first negative dialogue sequence.

[0057] Specifically, after acquiring multiple historical conversation sequences, selecting at least two conversation sequences containing keywords from the multiple historical conversation sequences as initial conversation sequences, and screening out a first negative conversation sequence from the multiple historical conversation sequences, initial conversation samples can be generated based on the at least two initial conversation sequences, and a first negative conversation sample can be generated based on the first negative conversation sequence. Initial conversation samples refer to conversation samples generated after data cleaning and preprocessing of the initial conversation sequences, and can be used for model training. Correspondingly, first negative conversation samples refer to conversation samples generated after data cleaning and preprocessing of the first negative conversation sequence, and can be used for model training.

[0058] Based on this, after determining the initial dialogue sequence and the first negative dialogue sequence based on multiple historical dialogue sequences, the initial dialogue sequence and the first negative dialogue sequence can be processed. Specifically, emoticons and spaces contained in the initial dialogue sequence and the first negative dialogue sequence are deleted, and links are replaced with [url]. The processed dialogues are then sequentially spliced together to obtain the initial dialogue sample corresponding to the initial dialogue sequence and the first negative dialogue sample corresponding to the first negative dialogue sequence.

[0059] In practical applications, when generating the initial conversation sample and the first negative conversation sample, the initial conversation sequence and the first negative conversation sequence can be cleaned first, that is, noise data such as emoticons and spaces are deleted, and links are replaced with special symbols to remove the noise data in the initial conversation sequence. The conversation sequence with the noise data removed is then spliced into a piece of text in the order in which the conversations are generated to generate the initial conversation sample. Correspondingly, the first negative conversation sequence can also be processed using the above method to generate the first negative conversation sample.

[0060] Furthermore, since the initial conversation sequence can be all the conversation data generated between two people within a day or a week, the amount of conversation data is relatively large. If there is only one conversation sentence containing a keyword in the conversation data, processing all the conversation data to generate the initial conversation sample will result in unnecessary waste of resources. Therefore, the conversation sentence containing the keyword can be used as the central conversation sentence, and then the initial conversation sample is determined based on the central conversation sentence. The specific implementation is as follows:

[0061] A central dialogue sentence containing a keyword is determined in the initial dialogue sequence; and an initial dialogue sample containing the central dialogue sentence is generated based on the initial dialogue sequence.

[0062] Based on this, the central dialogue sentence refers to the dialogue sentence containing keywords in the initial dialogue sequence. When multiple dialogue sentences in the initial dialogue sequence contain keywords, the dialogue sentences containing keywords can be used as central dialogue sentences respectively, and then the initial dialogue samples containing the central dialogue sentences can be determined in the initial dialogue sequence based on the central dialogue sentences.

[0063] In summary, the dialogue sentences containing keywords in the initial dialogue sequence are used as the central dialogue sentences, and then the initial dialogue samples are determined based on the central dialogue sentences. In this way, the subsequent model training can be combined with the context of the central dialogue sentences to improve the model's prediction accuracy.

[0064] Furthermore, considering that the initial dialogue sequence contains many dialogue sentences, it cannot be directly used as a sample for model training. Considering the impact of semantic features on model prediction accuracy, some dialogue sentences in the initial dialogue sequence can be selected as initial dialogue samples. The specific implementation is as follows:

[0065] Selecting a preceding dialogue text and a subsequent dialogue text corresponding to the central dialogue sentence in the initial dialogue sequence; combining the preceding dialogue text, the subsequent dialogue text and the central dialogue sentence to obtain an initial dialogue sample.

[0066] Specifically, the preceding dialogue text refers to the dialogue sentences that are arranged in chronological order before the central dialogue sentence in the initial dialogue sequence, and the number of sentences in the preceding dialogue text can be set according to actual needs; correspondingly, the subsequent dialogue text refers to the dialogue sentences that are arranged in chronological order after the central dialogue sentence in the initial dialogue sequence, and the number of sentences in the subsequent dialogue text can be set according to actual needs.

[0067] Based on this, a set number of preceding dialogue sentences and a set number of subsequent dialogue sentences corresponding to the central dialogue sentence are selected from the initial dialogue sequence. The preceding dialogue text is composed of the set number of preceding dialogue sentences within a certain time range, and the subsequent dialogue text is composed of the set number of subsequent dialogue sentences. The preceding dialogue text, the subsequent dialogue text, and the central dialogue sentence are combined to obtain the initial dialogue sample.

[0068] In practical applications, the first 10 sentences and the last 10 sentences corresponding to the central dialogue sentence in a dialogue sequence can be selected, and the sentences generated within 12 hours can be used as sentences associated with the central dialogue sentence. The initial dialogue sample is then composed of the central dialogue sentence, the first 10 sentences, and the last 10 sentences.

[0069] Continuing with the previous example, a salesperson communicates with a buyer, generating the conversation data "Salesperson A: 1.******; Buyer B: 1.******; ...Salesperson A: 30.******; Buyer B: 30.******." If Salesperson A's 15th sentence contains the keyword "complaint," we use Salesperson A's 15th sentence as the central sentence, and the 10 sentences preceding and following it as related sentences. The initial conversation sample is then composed of Salesperson A's 15th sentence, the 10 sentences preceding it, and the 10 sentences following it.

[0070] In summary, the preceding dialogue text and the subsequent dialogue text corresponding to the central dialogue sentence are determined, and the initial dialogue sample is composed of the central dialogue sentence, the preceding dialogue text and the subsequent dialogue text, thereby obtaining the initial dialogue sample.

[0071] Step S206 : dividing the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample according to the attribute information of the at least two initial conversation samples, wherein both the first positive conversation sample and the second negative conversation sample contain keywords.

[0072] Specifically, after generating the initial dialogue samples corresponding to at least two initial dialogue sequences and the first negative dialogue sample corresponding to the first negative dialogue sequence, the at least two initial dialogue samples can be divided into a first positive dialogue sample and a second negative dialogue sample according to the attribute information of the initial dialogue samples, wherein the first positive dialogue sample refers to a dialogue sample in which the initial dialogue sample contains keywords, and the sentences containing keywords in the initial dialogue sample are illegal sentences; correspondingly, the second negative dialogue sample refers to a dialogue sample in which the initial dialogue sample contains keywords, and the sentences containing keywords in the initial dialogue sample are non-illegal sentences, and the attribute information refers to the semantic information of the sentences containing keywords in the initial dialogue sample, as well as the semantic information combined with the context.

[0073] Based on this, after generating initial dialogue samples corresponding to at least two initial dialogue sequences and a first negative dialogue sample corresponding to the first negative dialogue sequence, the at least two initial dialogue samples are divided into a first positive dialogue sample and a second negative dialogue sample based on their attribute information. Both the first positive dialogue sample and the second negative dialogue sample contain keywords. The first positive dialogue sample contains keywords, and sentences containing keywords are considered violating sentences. The second negative dialogue sample contains keywords, but sentences containing keywords are considered non-violating sentences. For example, if the violating word "playing games" is present in the initial dialogue sample and the sentence containing the keyword is "Let's play games later," then this sentence is considered a violating sentence containing the keyword. If the sentence containing the keyword is "Playing games during work hours," then this sentence is considered a non-violating sentence containing the keyword.

[0074] In actual applications, when determining whether the initial conversation sample is the first positive conversation sample or the second negative conversation sample, since the judgment is based on the semantic information of the sentence containing the keyword, the initial conversation sample can be divided into samples by manual judgment, or by training a neural network model. This embodiment does not impose any restrictions on this. Combining the semantics of the context of the sentence containing the keyword can relatively easily determine whether the sentence containing the keyword is an illegal sentence.

[0075] Step S208: storing the first negative conversation sample and the second negative conversation sample into a negative conversation sample set, and storing the first positive conversation sample into a positive conversation sample set.

[0076] Specifically, after dividing the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample based on the attribute information of the at least two initial conversation samples, the first positive conversation sample and the second negative conversation sample can be stored in corresponding sample sets respectively, wherein the positive conversation sample set is used to store positive conversation samples, that is, samples containing keywords, and the sentences containing keywords are illegal sentences; the negative conversation sample set is used to store negative samples, that is, samples containing keywords, but the sentences containing keywords are non-illegal sentences, and samples that do not contain keywords. Since negative conversation samples can be determined by random sampling, there can be samples containing keywords, and the sentences containing keywords are illegal sentences in the negative conversation sample set.

[0077] Based on this, after dividing the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample according to the attribute information of the at least two initial conversation samples, the first negative conversation sample and the second negative conversation sample can be stored in the negative conversation sample set, and the first positive conversation sample can be stored in the positive conversation sample set.

[0078] In actual applications, before storing the first negative conversation sample and the second negative conversation sample in the negative conversation sample set, and storing the first positive conversation sample in the positive conversation sample set, the positive conversation sample set and the negative conversation sample set can be existing sample sets that have been created and contain positive samples / negative samples. It is also possible that after determining the first negative conversation sample and the second negative conversation sample, a negative conversation sample set is generated based on the first negative conversation sample and the second negative conversation sample, and after determining the first positive conversation sample, a positive conversation sample set is generated based on the first positive conversation sample.

[0079] Furthermore, after generating the first negative dialogue sample, the second negative dialogue sample, and the first positive dialogue sample, considering that the first negative dialogue sample, the second negative dialogue sample, and the first positive dialogue sample are all generated based on the historical dialogue sequence, there will be a lot of non-standard data in the historical dialogue sequence. It is necessary to further adjust the dialogue samples before they can be stored in the corresponding sample set. The specific implementation is as follows:

[0080] The first negative dialogue sample and the second negative dialogue sample are adjusted and stored in a negative dialogue sample set, and the first positive dialogue sample is adjusted and stored in a positive dialogue sample set.

[0081] Based on this, after generating the first negative dialogue sample, the second negative dialogue sample, and the first positive dialogue sample, the first negative dialogue sample and the second negative dialogue sample are adjusted and processed respectively and stored in the negative dialogue sample set, and the first positive dialogue sample is adjusted and processed and stored in the positive dialogue sample set.

[0082] In practical applications, since conversation samples are conversation data generated during actual conversations, and each person's expression methods and habits are different, the conversation data will contain noise data such as expressions, symbols, and links. Before model training, the conversation samples need to be cleaned to remove the noise data.

[0083] To sum up, the first negative conversation sample and the second negative conversation sample are adjusted and processed and then stored in the negative conversation sample set, and the first positive conversation sample is adjusted and processed and then stored in the positive conversation sample set, thereby improving the standardization of the conversation samples stored in the sample set and reducing the difficulty of subsequent model training.

[0084] Furthermore, considering that the generated dialogue samples contain a lot of noisy data and the length of each dialogue sentence in the dialogue samples is different, in order to facilitate subsequent model training, the dialogue samples need to be denoised and integrated. The specific implementation is as follows:

[0085] Noise data contained in the first negative conversation sample and the second negative conversation sample are deleted or modified respectively to obtain a first negative denoised conversation sample and a second negative denoised conversation sample, and the first negative denoised conversation sample and the second negative denoised conversation sample are integrated and processed respectively, and the processing results are stored in the negative conversation sample set; noise data contained in the first positive conversation sample is deleted or modified to obtain a first positive denoised conversation sample, and the first positive denoised conversation sample is integrated and processed, and the processing results are stored in the positive conversation sample set.

[0086] Specifically, in this embodiment, noise data refers to data such as expressions, symbols, special characters, links, pictures, and repeated conversation content in the conversation data; different processing methods correspond to different noise data. Expressions, symbols, spaces, empty sentences, short conversations, and sentences that appear repeatedly and whose repetition rate reaches a preset threshold in the conversation sample can be deleted. For links and line breaks, the link can be modified to [url] and the line break can be replaced with a space; integration processing refers to the splicing processing of multiple sentences in the conversation sample, that is, each sentence is spliced in turn according to the arrangement order of the sentences in the conversation sample or the order in which the sentences are generated to form a piece of text.

[0087] Based on this, after determining the first negative dialogue sample, the second negative dialogue sample and the first positive dialogue sample, the noise data contained in the first negative dialogue sample, the second negative dialogue sample and the first positive dialogue sample are detected respectively, and the noise data are deleted and modified accordingly, and then subsequent sentence integration is performed to obtain the first negative denoised dialogue sample, the second negative denoised dialogue sample and the first positive denoised dialogue sample, and store them in the corresponding sample sets.

[0088] In practical applications, the noise data contained in the first negative conversation sample and the second negative conversation sample are deleted or modified to obtain a first negative denoised conversation sample corresponding to the first negative conversation sample, and a second negative denoised conversation sample corresponding to the second negative conversation sample. The first negative denoised conversation sample and the second negative denoised conversation sample are then integrated and processed, and the processing results are stored as negative conversation samples in the negative conversation sample set. The noise data contained in the first positive conversation sample is deleted or modified to obtain a first positive denoised conversation sample. The first positive denoised conversation sample is then integrated and processed, and the processing results are stored as positive conversation samples in the positive conversation sample set.

[0089] Continuing with the previous example, when a salesperson communicates with a buyer, the conversation data generated may contain symbols, emoticons, images, hyperlinks, spaces, empty sentences, and short sentences (sentences containing only one character or word). These elements, such as emoticons and symbols, can negatively impact subsequent model training. Therefore, we can clean the conversation samples using methods such as regular expression matching to remove noise and obtain high-quality conversation samples. The cleaned sentences are then concatenated into segments in chronological order, using [CLS] as the start marker and [SEP] as the sentence segment and end marker.

[0090] To summarize, the noise data contained in the conversation samples is deleted and modified to achieve data cleaning, and then the conversation samples after data cleaning are integrated into a conversation sample, which facilitates subsequent model training.

[0091] Furthermore, after storing the first negative conversation sample and the second negative conversation sample in the negative conversation sample set and storing the first positive conversation sample in the positive conversation sample set, the existing positive conversation sample set and negative conversation sample set are expanded. Then, model training can be performed based on the expanded sample set to obtain the target conversation detection model. The specific implementation is as follows:

[0092] Extracting a target dialogue sample from the negative dialogue sample set and the positive dialogue sample set, wherein the target dialogue sample includes a target positive dialogue subsample and a target negative corresponding subsample; and training a dialogue detection model based on the target dialogue sample until a target dialogue detection model that meets a training stop condition is obtained.

[0093] Specifically, the target dialogue sample refers to the dialogue sample extracted from the negative dialogue sample set and the positive dialogue sample set for model training; the dialogue samples in both the negative dialogue sample set and the positive dialogue sample set are used to train the dialogue detection model. The dialogue detection model refers to an untrained neural network model used to detect illegal sentences contained in the dialogue data. In this embodiment, the dialogue detection model can be an untrained neural network model such as the BERT model. Correspondingly, the target dialogue detection model refers to a trained dialogue detection model that can be directly used to detect dialogue data.

[0094] Based on this, conversation samples from the negative and positive conversation sample sets are used to train the conversation detection model. During model training, target conversation samples are extracted from the negative and positive conversation sample sets, and the conversation detection model is trained based on the target conversation samples until a target conversation detection model that meets the training stop criteria is obtained. The target conversation samples include target positive conversation subsamples from the positive conversation sample set and target negative conversation subsamples from the negative conversation sample set. In this embodiment, the training stop criteria may include the model completing a preset number of training cycles, the model's prediction accuracy reaching a preset accuracy threshold, or the model's training time reaching a preset time range.

[0095] Continuing with the previous example, target conversation samples are extracted from the negative and positive conversation sample sets, and the conversation detection model is then trained based on these target conversation samples. Conversation samples that do not contain keywords are extracted from the negative conversation sample set and fed into the conversation detection model for training, enabling the model to predict negative conversation samples. Conversation samples that contain keywords are extracted from the positive conversation sample set and fed into the conversation detection model for training, enabling the model to predict positive conversation samples.

[0096] In summary, the conversation detection model is trained based on the conversation samples in the negative conversation sample set and the positive conversation sample set until the target conversation detection model that meets the training stop condition is obtained. This realizes model training based on specific conversation samples and improves the efficiency of model training.

[0097] Furthermore, when training the model based on the conversation samples in the positive and negative conversation sample sets, considering that the conversation samples contain a lot of noise data such as symbols and expressions, and the conversation samples contain a lot of sentences, it is necessary to denoise the extracted conversation samples to be processed and perform sentence splicing. The specific implementation is as follows:

[0098] Extracting a conversation sample to be processed from the negative conversation sample set and the positive conversation sample set, wherein the conversation sample to be processed includes a positive conversation sub-sample to be processed and a negative conversation sub-sample to be processed; deleting or modifying noise data included in the conversation sample to be processed to obtain a denoised conversation sample; integrating the denoised conversation sample to obtain an initial conversation sample; and labeling the initial conversation sample to obtain a target conversation sample.

[0099] Specifically, the dialogue samples to be processed refer to dialogue samples that can be used for model training. The dialogue samples to be processed may come from negative dialogue samples or positive dialogue samples. In this embodiment, noise data refers to data such as expressions, symbols, special characters, links, pictures, and repeated dialogue content in the dialogue data. Different processing methods correspond to different noise data. Expressions, symbols, spaces, empty sentences, short dialogues, and sentences that appear repeatedly and whose repetition rate reaches a preset threshold in the dialogue samples can be deleted. For links and line breaks, links can be modified to [url] and line breaks can be replaced with spaces. Integration processing refers to the splicing processing of multiple sentences in the denoised dialogue samples, that is, each sentence is spliced in sequence according to the order of arrangement of the sentences in the dialogue samples or the order in which the sentences are generated to form a text segment. Annotation processing refers to the annotation processing of each sentence in the initial dialogue sample, that is, assigning a label to each sentence in the initial dialogue sample, assigning a center label to the sentences containing keywords in the initial dialogue sample, and assigning a non-center label to the sentences other than the sentences containing keywords in the initial dialogue sample.

[0100] Based on this, we extract the unprocessed conversation samples from the negative and positive conversation sample sets. These unprocessed conversation samples consist of the unprocessed positive conversation subsamples from the positive conversation sample set and the unprocessed negative conversation subsamples from the negative conversation sample set. We then remove or modify the noise data contained in the unprocessed conversation samples to obtain denoised conversation samples. These denoised conversation samples are then integrated to obtain the initial conversation samples. The initial conversation samples are annotated, with each conversation sentence contained in the initial conversation samples annotated to obtain the target conversation samples.

[0101] In practical applications, the deletion or modification of noise data, the integration of denoised conversation samples, and the labeling of initial conversation samples are performed on each conversation sample in the negative conversation sample set and the positive conversation sample set. That is, after extracting the conversation samples to be processed from the negative conversation sample set and the positive conversation sample set, the above processing method is applied to the conversation samples to be processed, and they are processed into target conversation samples before model training.

[0102] Continuing with the above example, the 15th sentence of Sales A is the sentence to be predicted. After determining the initial conversation sample consisting of 21 sentences consisting of the 15th sentence of Sales A, the 10 sentences before the 15th sentence of Sales A, and the 10 sentences after the 15th sentence of Sales A, the symbols, expressions, pictures, hyperlinks, spaces, empty sentences, short sentences (sentences containing only one character or one word) contained in these 21 sentences are cleaned to remove noise data and obtain high-quality conversation samples. The sentences after data cleaning are then spliced into a section according to the person corresponding to the sentence and the time sequence in which the sentence was generated, and [CLS] is used as the start mark and [SEP] is used as the sentence segmentation mark and end mark. Each sentence is then labeled to obtain the following Figure 3 The initial conversation sample shown in the figure, that is, the central sentence containing the keyword is marked as 1, and the other sentences are marked as 0. The generated conversation sample is input into the BERT model for training to obtain the vector representation of each sentence, that is, the text representation.

[0103] In summary, the noise data contained in the conversation samples to be processed is deleted or modified, the denoised conversation samples are integrated, and the integrated initial conversation samples are labeled to obtain the target conversation samples, so as to facilitate subsequent model training based on the processed conversation samples, thereby improving the prediction efficiency and accuracy of the model and reducing the impact of noise data on the prediction effect.

[0104] Furthermore, when inputting the target dialogue sample into the dialogue detection model for detection, considering that the model trained with fewer dialogue samples may not achieve the required prediction accuracy, multiple target dialogue samples can be used for model training until the training stop condition is met to obtain the target dialogue detection model. The specific implementation is as follows:

[0105] The target dialogue sample is input into the dialogue detection model for detection to obtain a detection probability of the target dialogue sample; the dialogue detection model is trained based on the detection probability and a loss function until a target dialogue detection model that satisfies a training stop condition is obtained, wherein the loss function formula is as follows:

[0106]

[0107] Among them, L represents the loss value; N represents the total number of samples; (1-P t ) γ represents the regulatory factor; P t Represents the probability of correct classification; α t represents the class weight, γ represents the focus parameter, and α t and γ are hyperparameters of the loss function.

[0108] Specifically, the detection probability refers to the violation probability of the target conversation sample obtained by inputting the target conversation sample into the conversation detection model for prediction; the loss function is used to train the conversation detection model. In this embodiment, the loss function can be a Focal Loss (focus loss function) or other loss functions, and this embodiment does not impose any limitation on this.

[0109] Based on this, the target dialogue sample is input into the dialogue detection model for detection to obtain the detection probability of the target dialogue sample. The dialogue detection model is trained based on the detection probability and the loss function until the target dialogue detection model that meets the training stop conditions is obtained.

[0110] In practical applications, the constructed target dialogue sample can be fed into the BERT model for processing to obtain the word vector h for each word in the target dialogue sample. ij ∈R L , and then use average pooling to obtain the vector representation of each sentence: h i =avgpool(h ij ) Take the final text representation h i After a linear transformation and then an activation function, the probability of violation is obtained: p(c|h i )=sigmoid(Wh i ) where W is the parameter matrix to be learned. Finally, the conversation detection model is trained using Focal Loss as the loss function.

[0111] A sample processing method provided in one embodiment of this specification selects at least two initial conversation sequences containing keywords from multiple historical conversation sequences, and then divides the initial conversation samples into first positive conversation samples and second negative conversation samples based on the initial conversation sample attribute information corresponding to each initial conversation sequence. By screening the multiple historical conversation sequences for the first negative conversation sequence to generate the first negative conversation samples, the diversity of the samples is increased. Subsequently, a detection model is trained based on the first negative conversation samples, the first positive conversation samples, and the second negative conversation samples, thereby improving the prediction accuracy of the detection model.

[0112] The following combined Figure 4 , taking the application of the sample construction method provided in this specification in dialogue quality inspection as an example, the sample construction method is further explained. Figure 4 A processing flow chart of a sample construction method for conversation quality inspection provided in an embodiment of this specification is shown, which specifically includes the following steps:

[0113] Step S402: Perform keyword matching on each historical conversation sequence based on the keyword table, and use at least two conversation sequences containing keywords in the keyword table as initial conversation sequences.

[0114] When teachers communicate with students / parents, they typically use online voice, video, or text. This generates a large amount of conversation data. To ensure communication quality, this conversation data can be quality-checked to identify any inappropriate words or behaviors, as well as to assess service attitude and feedback speed, thereby improving service quality. The conversation data between teachers and parents / students is treated as a historical conversation sequence. A historical conversation sequence can also be conversation data generated within a certain timeframe. The supervisory team specifies a keyword list and uses keyword matching methods to search for clues to violations within the large-scale conversation data. At least two conversation sequences containing the keyword are used as the initial conversation sequence.

[0115] In practical applications, keywords can be used to detect violations in specific scenarios, and then expand the content of the teacher's conversation, search for violations in the conversation data between the teacher and other parents / students, and continue to discover violations in other scenarios.

[0116] Step S404: determining a central dialogue sentence containing a keyword in the initial dialogue sequence, and selecting a preceding dialogue text and a subsequent dialogue text corresponding to the central dialogue sentence in the initial dialogue sequence to form an initial dialogue sample.

[0117] In this embodiment, keywords can be uncivilized words, words that over-promise, or words that lead users to complain or withdraw from a course. For example, keywords such as "playing games," "withdrawing from a course," "complaint," and "private contact information" are used as the central conversation sentence in the initial conversation sequence. A set number of preceding and subsequent conversation texts corresponding to the central conversation sentence are then selected from the initial conversation sequence. The first 10 and last 10 sentences of the central conversation sentence, which are no more than 12 hours old, can be selected to form the initial conversation sample. If there are fewer than 10 sentences within 12 hours, all sentences generated within 12 hours can be selected to form the initial conversation sample.

[0118] Step S406 : dividing the at least two initial dialogue samples into a first positive dialogue sample and a first negative dialogue sample according to the attribute information of the at least two initial dialogue samples.

[0119] Since sentences containing keywords may or may not violate regulations, after determining the initial conversation sample containing keywords, we need to determine whether it contains any violations based on its semantic information. For example, if the sentence "You can add my private contact information" indicates that the user is adding the teacher's private contact information, it is considered a violation and a positive example. However, if the sentence "We do not allow teachers to add students' private contact information" indicates that the teacher refuses to add private contact information, it is considered a non-violation sentence and a hard negative example.

[0120] Step S408: Randomly sample multiple historical dialogue sequences to generate second negative dialogue samples.

[0121] Randomly sample some sentences other than keywords from the full conversation data. Since violations are rare, these sentences are considered compliant and recorded as likely negative samples.

[0122] Step S410 : storing the first negative conversation sample and the second negative conversation sample into a negative conversation sample set, and storing the first positive conversation sample into a positive conversation sample set.

[0123] Negative samples are composed of difficult negative samples and easy negative samples, which are stored in the existing negative sample set. Positive samples are stored in the positive sample set. Then, the samples in the positive sample set and the negative sample set are divided into training set and test set respectively, so as to facilitate subsequent model training and testing based on the negative sample set and the positive sample set.

[0124] Step S412: extracting dialogue samples to be processed from the negative dialogue sample set and the positive dialogue sample set.

[0125] When training the model, dialogue samples from the negative dialogue sample set and the positive dialogue sample set are extracted as training samples, and then the model training is performed.

[0126] Step S414: Delete or modify the noise data contained in the dialogue sample to be processed to obtain a denoised dialogue sample, and integrate the denoised dialogue sample to obtain an initial dialogue sample.

[0127] The conversation samples used for model training are de-noised, including but not limited to data cleaning of symbols, emoticons, images, hyperlinks, spaces, empty sentences, and short sentences (sentences containing only one character or word). This removes noise and produces high-quality conversation samples. The cleaned sentences are then concatenated into segments, based on the corresponding person and the time of their production. [CLS] is used as the start marker, and [SEP] is used as the sentence segmentation and end marker to produce the initial conversation sample.

[0128] Step S416: annotate the initial dialogue sample to obtain a target dialogue sample, input the target dialogue sample into the dialogue detection model for detection, and obtain the detection probability of the target dialogue sample.

[0129] The sentences in the initial conversation sample are labeled, with the sentences to be predicted marked as 1 and the other sentences marked as 0 to obtain the target conversation sample. The target conversation sample is then input into the conversation detection model to obtain the vector representation of each word. Average pooling is then used to obtain the vector representation of each sentence, and the final text representation is obtained. After a linear transformation and an activation function, the probability of violation is obtained.

[0130] Step S418: Train the conversation detection model based on the detection probability and the loss function until a target conversation detection model that meets the training stop condition is obtained.

[0131] Use Focal Loss as the loss function to train the conversation detection model until the target conversation detection model meets the training stop conditions.

[0132] In summary, by selecting at least two initial conversation sequences containing keywords from multiple historical conversation sequences, and then classifying the initial conversation samples into first positive conversation samples and second negative conversation samples based on the initial conversation sample attribute information corresponding to each initial conversation sequence, the method screens the multiple historical conversation sequences for the first negative conversation sequence to generate the first negative conversation samples, thereby increasing sample diversity. Subsequently, the detection model is trained based on the first negative conversation samples, the first positive conversation samples, and the second negative conversation samples, thereby improving the prediction accuracy of the detection model.

[0133] One embodiment of this specification selects at least two initial conversation sequences containing keywords from multiple historical conversation sequences, and then divides the initial conversation samples into first positive conversation samples and second negative conversation samples based on the initial conversation sample attribute information corresponding to each initial conversation sequence. By screening the first negative conversation sequences from the multiple historical conversation sequences to generate first negative conversation samples, the diversity of the samples is increased. Subsequently, a detection model is trained based on the first negative conversation samples, the first positive conversation samples, and the second negative conversation samples, thereby improving the prediction accuracy of the detection model.

[0134] Corresponding to the above method embodiment, this specification also provides a sample construction device embodiment, Figure 5 FIG1 shows a schematic diagram of the structure of a sample construction device provided in an embodiment of this specification. Figure 5 As shown, the device includes:

[0135] An acquisition module 502 is configured to acquire a plurality of historical conversation sequences, select at least two conversation sequences containing the keyword from the plurality of historical conversation sequences as initial conversation sequences, and filter a first negative conversation sequence from the plurality of historical conversation sequences;

[0136] A generating module 504 is configured to generate initial dialogue samples corresponding to at least two initial dialogue sequences, and a first negative dialogue sample corresponding to the first negative dialogue sequence;

[0137] a division module 506 configured to divide the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample based on attribute information of the at least two initial conversation samples, wherein both the first positive conversation sample and the second negative conversation sample contain keywords;

[0138] The storage module 508 is configured to store the first negative conversation sample and the second negative conversation sample into a negative conversation sample set, and store the first positive conversation sample into a positive conversation sample set.

[0139] In an optional embodiment, the generating module 504 is further configured to:

[0140] A central dialogue sentence containing a keyword is determined in the initial dialogue sequence; and an initial dialogue sample containing the central dialogue sentence is generated based on the initial dialogue sequence.

[0141] In an optional embodiment, the generating module 504 is further configured to:

[0142] Selecting a preceding dialogue text and a subsequent dialogue text corresponding to the central dialogue sentence in the initial dialogue sequence; combining the preceding dialogue text, the subsequent dialogue text and the central dialogue sentence to obtain an initial dialogue sample.

[0143] In an optional embodiment, the storage module 508 is further configured to:

[0144] The first negative dialogue sample and the second negative dialogue sample are adjusted and stored in a negative dialogue sample set, and the first positive dialogue sample is adjusted and stored in a positive dialogue sample set.

[0145] In an optional embodiment, the storage module 508 is further configured to:

[0146] Noise data contained in the first negative conversation sample and the second negative conversation sample are deleted or modified respectively to obtain a first negative denoised conversation sample and a second negative denoised conversation sample, and the first negative denoised conversation sample and the second negative denoised conversation sample are integrated and processed respectively, and the processing results are stored in the negative conversation sample set; noise data contained in the first positive conversation sample is deleted or modified to obtain a first positive denoised conversation sample, and the first positive denoised conversation sample is integrated and processed, and the processing results are stored in the positive conversation sample set.

[0147] In an optional embodiment, the storage module 508 is further configured to:

[0148] Extracting a target dialogue sample from the negative dialogue sample set and the positive dialogue sample set, wherein the target dialogue sample includes a target positive dialogue subsample and a target negative corresponding subsample; and training a dialogue detection model based on the target dialogue sample until a target dialogue detection model that meets a training stop condition is obtained.

[0149] In an optional embodiment, the storage module 508 is further configured to:

[0150] Extracting a conversation sample to be processed from the negative conversation sample set and the positive conversation sample set, wherein the conversation sample to be processed includes a positive conversation sub-sample to be processed and a negative conversation sub-sample to be processed; deleting or modifying noise data included in the conversation sample to be processed to obtain a denoised conversation sample; integrating the denoised conversation sample to obtain an initial conversation sample; and labeling the initial conversation sample to obtain a target conversation sample.

[0151] In an optional embodiment, the storage module 508 is further configured to:

[0152] The target dialogue sample is input into the dialogue detection model for detection to obtain a detection probability of the target dialogue sample; and the dialogue detection model is trained based on the detection probability and the loss function until a target dialogue detection model that meets the training stop condition is obtained.

[0153] The loss function formula is as follows:

[0154]

[0155] Among them, L represents the loss value; N represents the total number of samples; (1-P t ) γ represents the regulatory factor; P t Represents the probability of correct classification; α t represents the class weight, γ represents the focus parameter, and α t and γ are hyperparameters of the loss function.

[0156] In an optional embodiment, the acquisition module 502 is further configured to:

[0157] Keyword matching is performed on each historical dialogue sequence based on a preset keyword table, and at least two dialogue sequences containing keywords in the preset keyword table are used as initial dialogue sequences.

[0158] In an optional embodiment, the acquisition module 502 is further configured to:

[0159] Perform random sampling on multiple historical dialogue sequences to obtain the first negative dialogue sequence.

[0160] A sample processing device provided in one embodiment of this specification selects at least two initial conversation sequences containing keywords from multiple historical conversation sequences, and then divides the initial conversation samples into first positive conversation samples and second negative conversation samples based on the initial conversation sample attribute information corresponding to each initial conversation sequence. By screening the first negative conversation sequences from the multiple historical conversation sequences to generate first negative conversation samples, the device increases sample diversity. Subsequently, a detection model is trained based on the first negative conversation samples, the first positive conversation samples, and the second negative conversation samples, thereby improving the prediction accuracy of the detection model.

[0161] The above is a schematic diagram of a sample construction device according to this embodiment. It should be noted that the technical solution of the sample construction device and the technical solution of the sample construction method described above are based on the same concept. For details not described in detail in the technical solution of the sample construction device, please refer to the description of the technical solution of the sample construction method described above.

[0162] Figure 6 6 shows a block diagram of a computing device 600 according to an embodiment of the present disclosure. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0163] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0164] In one embodiment of the present application, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 6 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.

[0165] The computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 600 can also be a mobile or stationary server. The processor 620 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the sample construction method described above.

[0166] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the sample construction method described above are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the sample construction method described above.

[0167] An embodiment of the present specification further provides a computer-readable storage medium storing computer instructions, which implement the steps of the above-mentioned sample construction method when executed by a processor.

[0168] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the sample construction method described above are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the sample construction method described above.

[0169] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0170] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0171] It should be noted that for the aforementioned method embodiments, for ease of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this specification is not limited to the order of the actions described, because according to this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this specification.

[0172] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0173] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A sample construction method, characterized in that: include: Acquire multiple historical conversation sequences, select at least two conversation sequences containing a keyword from the multiple historical conversation sequences as initial conversation sequences, and screen a first negative conversation sequence from the multiple historical conversation sequences, wherein the keyword is an illegal word; generating initial dialogue samples corresponding to at least two initial dialogue sequences, and a first negative dialogue sample corresponding to the first negative dialogue sequence; dividing the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample according to attribute information of the at least two initial conversation samples, wherein both the first positive conversation sample and the second negative conversation sample contain keywords, and wherein the attribute information is determined by semantic information of the initial conversation samples; The first negative conversation sample and the second negative conversation sample are stored in a negative conversation sample set, and the first positive conversation sample is stored in a positive conversation sample set, wherein the negative conversation sample set and the positive conversation sample set are used to train a conversation detection model for detecting conversation data.

2. The method according to claim 1, characterized in that Determining the initial conversation sample corresponding to any one of the at least two initial conversation sequences includes: Determining a central dialogue sentence containing a keyword in the initial dialogue sequence; An initial dialogue sample including the central dialogue sentence is generated based on the initial dialogue sequence.

3. The method according to claim 2, characterized in that Generating an initial dialogue sample including the central dialogue sentence based on the initial dialogue sequence includes: Selecting a preceding dialogue text and a subsequent dialogue text corresponding to the central dialogue sentence in the initial dialogue sequence; The preceding dialogue text, the subsequent dialogue text, and the central dialogue sentence are combined to obtain an initial dialogue sample.

4. The method according to claim 1, wherein The storing the first negative conversation sample and the second negative conversation sample into a negative conversation sample set, and storing the first positive conversation sample into a positive conversation sample set, comprises: The first negative dialogue sample and the second negative dialogue sample are adjusted and stored in a negative dialogue sample set, and the first positive dialogue sample is adjusted and stored in a positive dialogue sample set.

5. The method according to claim 4, characterized in that The adjusting and processing the first negative conversation sample and the second negative conversation sample and storing them in a negative conversation sample set, and adjusting and processing the first positive conversation sample and storing them in a positive conversation sample set, includes: Deleting or modifying noise data included in the first negative conversation sample and the second negative conversation sample to obtain a first negative denoised conversation sample and a second negative denoised conversation sample, integrating the first negative denoised conversation sample and the second negative denoised conversation sample, and storing the processing results in a negative conversation sample set; Noise data included in the first positive conversation sample is deleted or modified to obtain a first positive denoised conversation sample, the first positive denoised conversation sample is integrated and processed, and the processing result is stored in a positive conversation sample set.

6. The method according to claim 1, characterized in that After the step of storing the first negative conversation sample and the second negative conversation sample into a negative conversation sample set and the step of storing the first positive conversation sample into a positive conversation sample set is performed, the method further includes: Extracting a target dialogue sample from the negative dialogue sample set and the positive dialogue sample set, wherein the target dialogue sample includes a target positive dialogue subsample and a target negative corresponding subsample; The dialogue detection model is trained based on the target dialogue sample until a target dialogue detection model that meets a training stop condition is obtained.

7. The method according to claim 6, characterized in that Extracting a target dialogue sample from the negative dialogue sample set and the positive dialogue sample set includes: Extracting a to-be-processed conversation sample from the negative conversation sample set and the positive conversation sample set, wherein the to-be-processed conversation sample includes a to-be-processed positive conversation sub-sample and a to-be-processed negative conversation sub-sample; Deleting or modifying noise data contained in the conversation sample to be processed to obtain a denoised conversation sample, and integrating the denoised conversation sample to obtain an initial conversation sample; The initial dialogue sample is labeled to obtain a target dialogue sample.

8. The method according to claim 7, characterized in that The step of training the dialogue detection model based on the target dialogue sample until a target dialogue detection model that satisfies a training stop condition is obtained includes: Inputting the target conversation sample into the conversation detection model for detection to obtain a detection probability of the target conversation sample; The conversation detection model is trained based on the detection probability and the loss function until a target conversation detection model that meets the training stopping condition is obtained.

9. The method according to claim 8, characterized in that The loss function formula is as follows: Among them, L represents the loss value; N represents the total number of samples; (1-P t ) γ represents the regulatory factor; P t Indicates the probability of correct classification; represents the class weight, γ represents the focusing parameter, and and γ are hyperparameters of the loss function.

10. The method according to claim 1, characterized in that The step of taking at least two dialogue sequences containing keywords from the plurality of historical dialogue sequences as initial dialogue sequences includes: Keyword matching is performed on each historical dialogue sequence based on a preset keyword table, and at least two dialogue sequences containing keywords in the preset keyword table are used as initial dialogue sequences.

11. The method according to claim 1, wherein The screening of the first negative dialogue sequence from the plurality of historical dialogue sequences includes: Perform random sampling on multiple historical dialogue sequences to obtain the first negative dialogue sequence.

12. A sample construction device, characterized in that: include: an acquisition module configured to acquire a plurality of historical conversation sequences, select at least two conversation sequences containing a keyword from the plurality of historical conversation sequences as initial conversation sequences, and screen a first negative conversation sequence from the plurality of historical conversation sequences, wherein the keyword is an illegal word; a generating module configured to generate initial dialogue samples corresponding to at least two initial dialogue sequences, and a first negative dialogue sample corresponding to the first negative dialogue sequence; a division module configured to divide the at least two initial conversation samples into a first positive conversation sample and a second negative conversation sample based on attribute information of the at least two initial conversation samples, wherein both the first positive conversation sample and the second negative conversation sample contain keywords, and wherein the attribute information is determined by semantic information of the initial conversation samples; The storage module is configured to store the first negative conversation sample and the second negative conversation sample in a negative conversation sample set, and store the first positive conversation sample in a positive conversation sample set, wherein the negative conversation sample set and the positive conversation sample set are used to train a conversation detection model for detecting conversation data.

13. A computing device, characterized in that It comprises a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the sample construction method according to any one of claims 1 to 11.

14. A computer-readable storage medium storing computer instructions, characterized in that: When the instruction is executed by a processor, the steps of the sample construction method described in any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Online dialogue log violation detection method and system based on BERT model

    CN112199480A

  • Sensitive word detection method and device, electronic equipment and storage medium

    CN114417881A