A text sample processing method, device, equipment and medium

By clustering and sample de-correcting of the query text in the initial training sample of the knowledge question and answer system model, the quality of the model training sample is improved, the problem of rough selection of negative samples is solved, and the output accuracy of the text matching model is improved.

CN113408301BActive Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110785709.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-12
Publication Date
2025-05-23
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

During the training process of the knowledge question and answer system model, the selection of negative samples is too rough, resulting in low sample quality and affecting the model learning results.

Method used

By clustering the query text in the initial training sample of the preset text matching model, and deduplication and correction of the negative samples based on the clustering results and sample timestamps, the target model training sample is obtained.

Benefits of technology

The quality of the model training samples is improved, and the output results of the text matching model are more accurate, solving the problems of negative sample label errors and high repetition rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113408301B_ABST
    Figure CN113408301B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a text sample processing method, device, equipment and medium, wherein the method comprises: obtaining an initial training sample of a preset text matching model, and clustering the query text in the initial training sample, wherein the query text is a keyword input into the preset text matching model; according to the result of the clustering processing and the timestamp of each initial training sample, deduplication and correction are performed on the negative samples in the initial training sample to obtain a target model training sample. The problem of low sample data quality caused by incorrect negative sample labels and high repetition rate in the training sample data of the preset text matching model collected in the prior art is solved, and the sample deduplication is achieved according to the query text similarity and sample timestamp in the initial training sample, thereby improving the quality of the training sample of the preset text matching model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a text sample processing method, device, equipment and medium. Background Art

[0002] In the knowledge question answering system, based on the query text content input by the user, the user will be provided with text content related to the query text, such as multiple articles related to keywords, for the user to click and read. Among them, the ranking of the articles fed back to the user by the knowledge question answering system will directly affect the user's experience of using the knowledge question answering system.

[0003] The knowledge question answering system uses the behavior of users entering query text and then clicking to view related articles fed back by the system as positive and negative samples for training the knowledge question answering system model.

[0004] However, in the process of implementing the present invention, it was found that there are at least the following technical problems in the prior art: during the training of the knowledge question and answer system model, the selection of negative samples is too rough. In some cases, the articles fed back by the knowledge question and answer system based on the query text are not clicked and are not negative samples. The quality of model training samples needs to be improved, and the knowledge question and answer system model whose learning results depend on the quality of sample data needs to be further optimized. Summary of the invention

[0005] The embodiments of the present invention provide a text sample processing method, device, equipment and medium to improve the quality of model training samples, enable the text matching model to be better learned, and the output results of the trained text matching model to be more accurate.

[0006] In a first aspect, an embodiment of the present invention provides a text sample processing method, the method comprising:

[0007] Acquire an initial training sample of a preset text matching model, and perform clustering processing on query texts in the initial training sample, wherein the query texts are keywords input into the preset text matching model;

[0008] According to the result of clustering processing and the timestamp of each initial training sample, the negative samples in the initial training samples are deduplicated and corrected to obtain the target model training samples.

[0009] In a second aspect, an embodiment of the present invention further provides a text sample processing device, the device comprising:

[0010] A text clustering module, used to obtain an initial training sample of a preset text matching model, and perform clustering processing on the query text in the initial training sample, wherein the query text is a keyword input into the preset text matching model;

[0011] The sample processing module is used to remove duplicates and correct negative samples in the initial training samples according to the results of clustering processing and the timestamps of each initial training sample to obtain target model training samples.

[0012] In a third aspect, an embodiment of the present invention further provides a computer device, the computer device comprising:

[0013] one or more processors;

[0014] A memory for storing one or more programs;

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement a text sample processing method provided by any embodiment of the present invention.

[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a text sample processing method as provided in any embodiment of the present invention.

[0017] The embodiments of the above invention have the following advantages or beneficial effects:

[0018] The embodiment of the present invention performs clustering processing on the query text in the initial training sample of the preset text matching model, that is, the keywords input into the preset text matching model; then, the clustered query text is deduplicated and corrected according to the category and the corresponding sample timestamp, that is, multiple initial training samples generated within a certain period of time are deduplicated or the labels corresponding to the samples are corrected, and finally the target model training samples with better sample data quality are obtained. The problem of low sample data quality caused by the error of negative sample labels and high repetition rate in the training sample data of the preset text matching model collected in the prior art is solved, and the sample deduplication is realized according to the query text similarity and sample timestamp in the initial training sample, so as to improve the quality of the training samples of the preset text matching model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of a text sample processing method provided by Embodiment 1 of the present invention;

[0020] Figure 2 It is a text query record data graph provided by the first embodiment of the present invention;

[0021] Figure 3 This is a text matching result display diagram of a text query provided by the first embodiment of the present invention;

[0022] Figure 4is a flow chart of a text sample processing method provided by Embodiment 2 of the present invention;

[0023] Figure 5 This is a schematic diagram of a query text clustering analysis process provided by Embodiment 2 of the present invention;

[0024] Figure 6 is a structural schematic diagram of a text sample processing device provided in Embodiment 3 of the present invention;

[0025] Figure 7 It is a structural diagram of a computer device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0026] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0027] Embodiment 1

[0028] Figure 1 This is a flowchart of a text sample processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to the case of constructing high-quality training samples of text matching models / question-answering models. The method can be executed by a text sample processing device, which can be implemented in software and / or hardware and integrated into a computer device with application development functions.

[0029] like Figure 1 As shown, the text sample processing method includes the following steps:

[0030] S110, obtaining an initial training sample of a preset text matching model, and performing clustering processing on a query text in the initial training sample, wherein the query text is a keyword input into the preset text matching model.

[0031] Among them, the preset text matching model can be a question-and-answer model of a knowledge question-and-answer system, which matches answers to the queried questions; or an article matching model, which matches the input keywords or key words with relevant text content. Accordingly, the query text, as a keyword input into the preset text matching model, can be a single word, a word or phrase composed of multiple words, or a text of different lengths such as a sentence. Usually, the results matched with the query text output by the preset text matching model are sorted, and the accuracy of the sorting of the output results depends on the CTR (Click-Through-Rate) technology. The good samples of the text matching model whose output results need to be sorted are usually the texts of the texts that are clicked or not clicked after the query as the model training samples. Assuming that after a user enters a query text in a preset text matching model such as a knowledge question-and-answer system, the question-and-answer system feedbacks 20 relevant texts, and the user clicks on one of the texts, then 20 samples can be collected in the end, including 1 positive sample and 19 negative samples, that is, the behavior of clicking to view the feedback text in a query is a positive sample, and the behavior of querying but not clicking the text in the same text query is a negative sample.

[0032] However, the negative samples obtained by the above sample collection method have large noise. Figure 2 As shown in the text query data table, a sample contains information such as time, sample serial number, query text, user ID and click text ID. A user enters 4 different query texts within a few seconds. The letters A, B, C, D and E corresponding to the query text represent different words. The user did not click on any text in the first 3 queries, and the text ID was nan (indicating null). In the 4th query, the text content with the ID 21508 was clicked. Then the 4 query records will be regarded as 3 negative samples and 1 positive sample. However, further, refer to Figure 3 In the text matching results displayed according to the query text, the four query records all display the articles that the user identified as 21508 is interested in, and they are all ranked first. However, for some reasons, the user did not click to view the articles after the first few queries. That is to say, the text lists displayed by the four query behaviors are not exactly the same, but the article ranked first is the same, but the user did not click it. The query texts of these four queries are similar, and can even be considered as one query. The negative samples determined by the above table are not necessarily true negative samples. The reason for this phenomenon may be that although the user did not click on any matching text after entering the query text, he has obtained the desired information in the partial summary content displayed in the text list and no longer continues to view it; or due to network jams, the user repeatedly enters different but similar query texts to find the corresponding text content.

[0033] In this embodiment, in order to reduce the noise in the negative samples, the query texts are first clustered, so that the negative samples in the initial training samples are semantically corrected and deduplicated according to the similarity of the query texts.

[0034] Specifically, the clustering of query texts can select an algorithm suitable for the text classification scenario from commonly used clustering methods such as partition method, hierarchical method, density-based method, grid-based method or model-based method. Exemplarily, the K-means clustering method in the partition method can be used to construct K groups in N query texts, and finally obtain K cluster categories, where N and K are both natural numbers, and K is less than. In the process of cluster analysis, the cluster centers of the K groups can be randomly or according to preset classification rules, and clustering can be performed to obtain clustering results.

[0035] S120, according to the result of clustering processing and the timestamp of each initial training sample, the negative samples in the initial training samples are deduplicated and corrected to obtain the target model training samples.

[0036] Specifically, when deduplication and correction of negative samples are performed, in addition to considering the semantic similarity of the query text, the collection time of each initial training sample must also be considered. Only multiple queries within a similar time period can be considered as one query. For query texts belonging to the same category in the clustering results, each initial training sample can be grouped according to the timestamp of the initial training sample corresponding to each query text. That is, a time window is pre-set, and text queries and matches within the same time window can be considered as one query. In the implementation process, the initial training samples can be sorted according to the time sequence of the timestamps of the initial training samples corresponding to each query text. Assuming that the length of the time window is 20 seconds, starting from the initial training sample ranked first, the initial training texts corresponding to the timestamps whose difference with the timestamp of the initial training sample ranked first is within 20 seconds can be divided into one group, and the same sample will not be divided into two sample groups.

[0037] Furthermore, when the initial training samples in the same group include both positive samples and negative samples, the negative samples in the group are corrected to positive samples, and the corrected positive samples are deduplicated to form one positive sample; when the initial training samples in the same group are all negative samples, the negative samples are deduplicated to form one negative sample. Figure 2 Taking the four query record samples in as an example, the four samples can be grouped together, and the four samples can be deduplicated into one positive sample. Through the processing of this step, the accuracy of the model training sample data set is improved, and at the same time, the redundancy of the data set can be reduced, achieving the effect of optimizing the quality of the model training samples.

[0038] The technical solution of this embodiment is to cluster the query text in the initial training sample of the preset text matching model, that is, the keywords input into the preset text matching model; then, the clustered query text is deduplicated and corrected according to the category and the corresponding sample timestamp, that is, multiple initial training samples generated within a certain period of time are deduplicated or the labels corresponding to the samples are corrected, and finally the target model training samples with better sample data quality are obtained. The problem of low sample data quality caused by the wrong negative sample labels and high repetition rate in the training sample data of the preset text matching model collected in the prior art is solved, and the sample deduplication is realized according to the query text similarity and sample timestamp in the initial training sample, so as to improve the quality of the training samples of the preset text matching model.

[0039] Embodiment 2

[0040] Figure 2 This is a flowchart of a text sample processing method provided in Embodiment 2 of the present invention. This embodiment and the text sample processing method in the above embodiment belong to the same inventive concept, and further describes the text clustering process of sample processing. The method can be executed by a text sample processing device, which can be implemented by software and / or hardware and integrated into a computer device with application development function.

[0041] like Figure 2 As shown, the text sample processing method includes the following steps:

[0042] S210: Obtain an initial training sample of a preset text matching model, and convert the query text in the initial training sample into a text vector.

[0043] In this step, the initial training sample data is preprocessed to convert the query text into a text vector, which enables the computer to directly and effectively understand the meaning of the text. Specifically, the word2vec tool can be used to convert the text vector, that is, Figure 5 The process of vectorizing the original text into a text vector.

[0044] S220, selecting a preset number of text vectors from the text vectors as clustering centers based on a genetic algorithm, performing text vector clustering processing, and completing the clustering processing when the clustering effect meets a preset condition.

[0045] When performing cluster analysis on query text, i.e., the corresponding text vector, whether the selection of cluster centers is reasonable will affect the convergence of the cost function in the clustering algorithm. Compared with randomly selecting a certain number of cluster centers for cluster analysis, we still hope to find cluster centers that can make the classification results more accurate and the clustering effect better, so as to complete the final cluster analysis.

[0046] Therefore, in this embodiment, multiple groups of different cluster centers are determined through population iteration of the genetic algorithm, and multiple clustering processes are performed to determine the best cluster center and the best clustering effect. Figure 5 The clustering process of text vectors is shown:

[0047] First, k vectors are randomly selected from the vectorized text vectors as the initial population points of the first generation of the genetic algorithm and the initial clustering centers in the clustering algorithm "k-means algorithm".

[0048] In a genetic algorithm, the number of iterations can be set in advance, and the randomly selected k vectors are the initial population of the first generation. The initial population is transformed into a new population through population mutation, population crossover, calculation of cost function and roulette selection process. When the new population no longer reduces the cost function value of the genetic algorithm model, the optimal population is obtained, and the initial population point of the next generation can be determined based on the optimal population.

[0049] In the clustering algorithm "k-means algorithm", k randomly selected vectors are used as the initial clustering center of the first clustering, and then the initial population points of each generation in the genetic algorithm are used as the initial clustering center of the next clustering, and multiple clustering analysis processes are performed. In each clustering analysis process, the distances from the text vectors of the non-clustering center to the text vectors of each initial clustering center are calculated, and the text vectors of each non-clustering center are classified into the initial clustering center with the smallest distance; then, the center points of each cluster after classification are recalculated, and when the center points no longer change, a clustering process is completed. At the end of each clustering analysis process, the sum of the distances from each text vector in the cluster to the corresponding clustering center and the sum of the distances of the clustering centers of each cluster in the clustering result are also counted as clustering costs to judge the clustering effect. Because a good clustering model requires small spacing within the cluster and large spacing between clusters. After the query text is represented by text vectorization, it is more suitable to use Euclidean distance to measure the clustering effect. The clustering cost result after each clustering analysis will also affect the value of the cost function in the genetic algorithm. In one possible implementation, the cost function in the genetic algorithm is the same as the clustering cost function in the k-means clustering process, and both use the Euclidean distance value of the vector within and between clusters as the numerical standard for judgment.

[0050] With multiple iterations in the genetic algorithm, the k-means clustering analysis process will be performed for a corresponding number of times. Thus, the calculation results of multiple cost functions will be obtained. When the cost function in the clustering algorithm and the cost function in the genetic algorithm meet the convergence conditions at the same time, the clustering operation and the iterative process of the genetic algorithm can be terminated to determine the final clustering result. Among them, the cost function meets the convergence condition, which can be that in the clustering result, the intra-cluster distance has reached the minimum value, or when the intra-cluster distance and the inter-cluster distance are comprehensively considered, the value reaches an optimal result, and the cost function is considered to have converged.

[0051] In this embodiment, based on the genetic algorithm and taking into account the influence between generations, different initial population points (i.e., initial clustering centers) are iteratively selected to optimize the clustering effect. This can enable each text vector to achieve a better semantic classification effect and further improve the quality of the model training samples.

[0052] S230, according to the result of clustering processing and the timestamp of each initial training sample, the negative samples in the initial training samples are deduplicated and corrected to obtain the target model training samples.

[0053] The technical solution of this embodiment is to cluster the query text in the initial training sample of the preset text matching model, that is, the keywords input into the preset text matching model, based on the genetic algorithm; then, the clustered query text is deduplicated and corrected according to the category and the corresponding sample timestamp, that is, multiple initial training samples generated within a certain period of time are deduplicated or the labels corresponding to the samples are corrected, and finally the target model training samples with better sample data quality are obtained. The problem of low sample data quality caused by the wrong negative sample labels and high repetition rate in the training sample data of the preset text matching model collected in the prior art is solved, and the sample deduplication is realized according to the query text similarity and sample timestamp in the initial training sample, so as to improve the quality of the training samples of the preset text matching model.

[0054] Furthermore, after processing the training samples of the preset text matching model, the optimized samples can be used for model training to enable the model to learn better, thereby obtaining a target text matching model. When using the target text matching model, the obtained text matching keywords can be input as query text into the target text matching model to obtain the target text matching result.

[0055] In a specific example, a knowledge question-and-answer system was used for testing, and it was verified that after optimizing the model training samples using the text sample processing method under a 30S time window length, the text matching effectiveness was increased from 81% to 85%. The customer service inquiry experience was improved. The following is an embodiment of a text sample processing device provided in an embodiment of the present invention. The device and the text sample processing methods of the above-mentioned embodiments belong to the same inventive concept and can implement the text sample processing methods of the above-mentioned embodiments. For details not described in detail in the embodiment of the text sample processing device, please refer to the embodiment of the above-mentioned text sample processing method.

[0056] Embodiment 3

[0057] Figure 6 This is a schematic diagram of the structure of the text sample processing device provided in Example 3 of the present invention. This embodiment is applicable to the case of constructing high-quality training samples of text matching models / question-answering models. The device can be implemented by software and / or hardware and integrated into a computer device with application development capabilities.

[0058] like Figure 6 As shown, the text sample processing device includes: a text clustering module 310 and a sample processing module 320 .

[0059] The text clustering module 310 is used to obtain the initial training samples of the preset text matching model and perform clustering processing on the query text in the initial training samples, wherein the query text is the keyword input into the preset text matching model; the sample processing module 320 is used to deduplicate and correct the negative samples in the initial training samples according to the results of the clustering processing and the timestamps of each initial training sample to obtain the target model training samples.

[0060] The technical solution of this embodiment is to cluster the query text in the initial training sample of the preset text matching model, that is, the keywords input into the preset text matching model; then, the clustered query text is deduplicated and corrected according to the category and the corresponding sample timestamp, that is, multiple initial training samples generated within a certain period of time are deduplicated or the labels corresponding to the samples are corrected, and finally the target model training samples with better sample data quality are obtained. The problem of low sample data quality caused by the wrong negative sample labels and high repetition rate in the training sample data of the preset text matching model collected in the prior art is solved, and the sample deduplication is realized according to the query text similarity and sample timestamp in the initial training sample, so as to improve the quality of the training samples of the preset text matching model.

[0061] Optionally, the text clustering module 310 specifically includes:

[0062] A vector conversion submodule, used for converting the query text into a text vector;

[0063] The text clustering submodule is used to select a preset number of text vectors from the text vectors as clustering centers based on a genetic algorithm to perform text vector clustering processing; when the clustering effect meets the preset conditions, the clustering processing is completed.

[0064] Optionally, the text clustering submodule is specifically used for:

[0065] A preset number of text vectors are randomly selected as the initial population points of the first generation of the genetic algorithm and the cluster center points in the first clustering to perform genetic calculations and cluster analysis;

[0066] The initial population points of each iteration in the genetic algorithm are used as the cluster center points of each cluster analysis.

[0067] Optionally, the text clustering submodule is further used for:

[0068] When the cost function in the clustering algorithm and the cost function in the genetic algorithm simultaneously meet the convergence condition, the clustering operation and the iterative process of the genetic algorithm are terminated.

[0069] Optionally, the sample processing module 320 is specifically used for:

[0070] For the text vectors belonging to the same category in the clustering results, the initial training samples are grouped according to the timestamps of the initial training samples corresponding to the text vectors;

[0071] When the initial training samples in the same group include both positive samples and negative samples, the negative samples in the group are corrected to positive samples, and the corrected positive samples are deduplicated to form one positive sample;

[0072] When all initial training samples in the same group are negative samples, the negative samples are deduplicated into one negative sample.

[0073] Optionally, the sample processing module 320 is further used for:

[0074] Sort the initial training samples according to the time sequence of the timestamps of the initial training samples corresponding to each text vector;

[0075] The initial training samples that are sorted and belong to the same preset length time window are taken as a group of samples.

[0076] Optionally, the text sample processing device further includes:

[0077] The model training module is used to perform model training on the preset text matching model through the target model training sample to obtain the target text matching model.

[0078] Optionally, the text sample processing device further includes:

[0079] The text matching module is used to obtain text matching keywords; input the text matching keywords into the target text matching model to obtain target text matching results.

[0080] The text sample processing device provided in the embodiment of the present invention can execute the text sample processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0081] Embodiment 4

[0082] Figure 7 A schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Figure 7 A block diagram of an exemplary computer device 12 suitable for use in implementing embodiments of the present invention is shown. Figure 7 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as an intelligent controller and server, a mobile phone and other terminal devices.

[0083] like Figure 7 As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).

[0084] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0085] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0086] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 7 not shown, usually called a "hard drive"). Although Figure 7 Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The system memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.

[0087] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.

[0088] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the computer device 12 via the bus 18. It should be understood that although Figure 7 Not shown, other hardware and / or software modules may be used in conjunction with computer device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0089] The processing unit 16 executes various functional applications and data processing by running the program stored in the system memory 28, for example, implementing the text sample processing method provided in the embodiment of the present invention, which includes:

[0090] Acquire an initial training sample of a preset text matching model, and perform clustering processing on query texts in the initial training sample, wherein the query texts are keywords input into the preset text matching model;

[0091] According to the result of clustering processing and the timestamp of each initial training sample, the negative samples in the initial training samples are deduplicated and corrected to obtain the target model training samples.

[0092] Embodiment 5

[0093] The fifth embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the text sample processing method provided in any embodiment of the present invention is implemented, including:

[0094] Acquire an initial training sample of a preset text matching model, and perform clustering processing on query texts in the initial training sample, wherein the query texts are keywords input into the preset text matching model;

[0095] According to the result of clustering processing and the timestamp of each initial training sample, the negative samples in the initial training samples are deduplicated and corrected to obtain the target model training samples.

[0096] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0097] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0098] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0099] Computer program code for performing the operation of the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0100] It should be understood by those skilled in the art that the modules or steps of the present invention described above can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, optionally, they can be implemented by a program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0101] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A text sample processing method, It is characterized in that The method comprises: Acquire an initial training sample of a preset text matching model, and perform clustering processing on query texts in the initial training sample, wherein the query texts are keywords input into the preset text matching model; According to the result of clustering processing and the timestamp of each initial training sample, the negative samples in the initial training samples are deduplicated and corrected to obtain the target model training samples; Deduplication and correction of negative samples in the initial training samples according to the clustering processing result and the timestamp of each initial training sample include: For the text vectors belonging to the same category in the clustering results, the initial training samples are grouped according to the timestamps of the initial training samples corresponding to the text vectors; When the initial training samples in the same group include both positive samples and negative samples, the negative samples in the group are corrected to positive samples, and the corrected positive samples are deduplicated to form one positive sample; When all initial training samples in the same group are negative samples, the negative samples are deduplicated into one negative sample.

2. The method according to claim 1, It is characterized in that The clustering process of all query texts in the initial training sample includes: Convert the query text into a text vector; Selecting a preset number of text vectors from the text vectors as cluster centers based on a genetic algorithm to perform text vector clustering processing; When the clustering effect meets the preset conditions, the clustering process is completed.

3. The method according to claim 2, It is characterized in that The method of selecting a preset number of text vectors from the text vectors as cluster centers based on the genetic algorithm and performing text vector clustering processing includes: A preset number of text vectors are randomly selected as the initial population points of the first generation of the genetic algorithm and the cluster center points in the first clustering to perform genetic calculations and cluster analysis; The initial population points of each iteration in the genetic algorithm are used as the cluster center points of each cluster analysis.

4. The method according to claim 3, It is characterized in that When the clustering effect meets the preset condition, the clustering process is completed, including: When the cost function in the clustering algorithm and the cost function in the genetic algorithm simultaneously meet the convergence condition, the clustering operation and the iterative process of the genetic algorithm are terminated.

5. The method according to claim 1, It is characterized in that The grouping of the initial training samples according to the timestamps of the initial training samples corresponding to the text vectors includes: Sort the initial training samples according to the time sequence of the timestamps of the initial training samples corresponding to each text vector; The initial training samples that are sorted and belong to the same preset length time window are taken as a group of samples.

6. The method according to claim 1, It is characterized in that The method further comprises: The preset text matching model is trained through the target model training samples to obtain a target text matching model.

7. The method according to claim 6, It is characterized in that The method further comprises: Get text matching keywords; The text matching keyword is input into the target text matching model to obtain a target text matching result.

8. A text sample processing device, It is characterized in that The device comprises: A text clustering module, used to obtain an initial training sample of a preset text matching model, and perform clustering processing on the query text in the initial training sample, wherein the query text is a keyword input into the preset text matching model; A sample processing module, used to remove duplicates and correct negative samples in the initial training samples according to the results of clustering processing and the timestamps of each initial training sample, so as to obtain target model training samples; The sample processing module is specifically used for: For the text vectors belonging to the same category in the clustering results, the initial training samples are grouped according to the timestamps of the initial training samples corresponding to the text vectors; When the initial training samples in the same group include both positive samples and negative samples, the negative samples in the group are corrected to positive samples, and the corrected positive samples are deduplicated to form one positive sample; When all initial training samples in the same group are negative samples, the negative samples are deduplicated into one negative sample.

9. A computer device, It is characterized in that The computer device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the text sample processing method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the sample processing according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Search method and device based on artificial intelligence

    CN106354852A

  • Search recall method and device, server and storage medium

    CN107491518A