A training text data acquisition method and device, electronic equipment and storage medium
By constructing a hypercube of text vectors and dividing it into subcubes, determining the initial centroids for clustering, the problem of insufficient quality and professionalism of training text data in existing technologies is solved, and high-quality training text data acquisition is achieved.
Patent Information
- Application Number
- CN202411935419.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing technologies cannot generate highly specialized and high-quality training text data, thus failing to meet the quality and professional requirements of artificial intelligence training in various professional fields.
By constructing text vectors corresponding to each candidate document, drawing a hypercube, and dividing it into multiple sub-cubes on an average basis, determining the number of clusters and the initial centroid, and performing clustering based on the number of clusters and the number of text vectors within the sub-cubes, training text data is obtained.
It improves the quality of initial centroid selection, reduces sensitivity, enhances the stability and consistency of clustering, optimizes text clustering results, and improves the quality and professionalism of training text data.
Smart Images

Figure CN119886129B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a training text data acquisition method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the continuous development of Internet technology, technologies such as big data and cloud computing are continuously accumulated and refined. Internet-based public data resource acquisition channels and a large number of enterprise internal data resources have accumulated massive data resources. These data have a significant effect on the training of artificial intelligence models in various fields, including the financial field. Pre-training and fine-tuning of large models using a large amount of general data obtained from public channels combined with professional field data accumulated in enterprise internal can significantly improve the application performance of large models in professional fields.
[0003] However, due to the uneven quality of data obtained from public channels, it is difficult to meet the quality and professional requirements of artificial intelligence training in various professional fields, and therefore there is an urgent need for a training text generation scheme that can generate training text data with strong professional and high quality. SUMMARY
[0004] The present application provides a training text data acquisition method, device, electronic device, and storage medium, which can generate training text data with strong professional and high quality.
[0005] In a first aspect, the present application provides a training text data acquisition method, comprising:
[0006] establishing text vectors corresponding to each candidate document, and drawing a hypercube comprising each text vector;
[0007] dividing the hypercube into a plurality of sub-cubes on average;
[0008] determining the number of clusters for clustering and determining the initial centroid based on the number of clusters and the number of text vectors in each sub-cube; and
[0009] clustering the text vectors based on the number of clusters and the initial centroid to obtain a plurality of clustering result clusters, and determining training text data based on the plurality of clustering result clusters.
[0010] In a second aspect, the present application provides a training text data acquisition device, comprising:
[0011] a hypercube acquisition module for establishing text vectors corresponding to each candidate document, and drawing a hypercube comprising each text vector;
[0012] a sub-cube acquisition module for dividing the hypercube into a plurality of sub-cubes on average;
[0013] a cluster number acquisition module configured to determine a cluster number for clustering;
[0014] a clustering result and training text data acquisition module configured to cluster the individual text vectors based on the cluster number and the initial centroids to obtain a plurality of clustering result clusters, and determine training text data based on the plurality of clustering result clusters.
[0015] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the training text data acquisition method according to any of the embodiments of the present application when executing the program.
[0016] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program executable on a processor, and the program implements the training text data acquisition method according to any of the embodiments of the present application when executed on the processor.
[0017] The training text data acquisition method, device, electronic device, and storage medium provided by the embodiments of the present application can effectively improve the selection quality of the initial centroid, reduce the sensitivity to the selection of the initial centroid, by establishing a text vector corresponding to each candidate document, drawing a hypercube including the individual text vectors, and then determining the initial centroid for clustering based on the number of text vectors in the sub-hypercube obtained by dividing the hypercube. The stability and consistency of clustering can be improved, and the efficiency of capturing text semantic information can be improved, and the text clustering effect can be optimized, by clustering the individual text vectors based on the initial centroid. The efficiency of obtaining training text data can be improved, and the quality and professionalism of the obtained training text data can be improved, by determining the training text data based on the clustering result clusters. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be considered as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained without creative labor.
[0019] Figure 1 is a flowchart of the training text data acquisition method provided by the embodiments of the present application;
[0020] Figure 2 is another flowchart of the training text data acquisition method provided by the embodiments of the present application;
[0021] Figure 3 is another flowchart of the training text data acquisition method provided by an embodiment of the present application;
[0022] Figure 4 is another flowchart of the training text data acquisition method provided by an embodiment of the present application;
[0023] Figure 5 is a flowchart of a training text data acquisition device provided by an embodiment of the present application;
[0024] Figure 6 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present application.
[0026] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] Figure 1 is a flowchart of a training text data acquisition method, device, electronic device and storage medium provided by an embodiment of the present application. The present embodiment can be applied to the scene of acquiring training data of an artificial intelligence model in the financial field. The method can be executed by a training text data acquisition device provided by an embodiment of the present application. The device can be realized in the form of software and / or hardware. In a specific embodiment, the device can be integrated in an electronic device, such as a computer, a server, etc. The following embodiments will be described by taking the device integrated in an electronic device as an example. Referring to Figure 1 , the method can specifically include the following steps:
[0028] Step 101, a text vector corresponding to each candidate document is established, and a hypercube including each text vector is drawn. This step can facilitate the average division of the hypercube to obtain a plurality of sub-cubes, and then determine the initial centroid of the cluster based on the number of texts in each sub-cube.
[0029] Specifically, before step 101, initial text data can be prepared, and each initial document in the initial text data can be pre-processed to obtain the above-mentioned each candidate document. Specifically, the above-mentioned initial text data can be obtained from an open source Chinese data set or crawled from a web page.
[0030] Specifically, the process of pre-processing the initial text data can include: cleaning messy characters and punctuation marks, filtering harmful character texts, and screening and removing colloquial texts and general texts. The above-mentioned harmful character texts, for example, involve yellow and social phobia texts. The above-mentioned colloquial texts, for example, you, I, and the like. The above-mentioned general texts, for example, de, and the like.
[0031] Specifically, the process of establishing a text vector corresponding to each candidate document can be based on algorithms such as Bag-of-Words (BoW), Term Frequency-Inverse Document Frequency (TF-IDF), or Word Embedding.
[0032] Specifically, the process of drawing a hypercube including each text vector can include: determining the dimension of the hypercube based on the above-mentioned each text vector; determining the vertex of the hypercube based on the dimension of the hypercube and the element value of the above-mentioned each text vector; and establishing the above-mentioned hypercube based on the vertex of the hypercube.
[0033] Step 102, the hypercube is evenly divided into a plurality of sub-cubes. This step can facilitate the determination of the initial centroid of the cluster based on the number of texts in each sub-cube.
[0034] Optionally, the process of evenly dividing the hypercube into a plurality of sub-cubes includes: determining the number of sub-cubes, and evenly dividing the above-mentioned hypercube based on the number of sub-cubes to obtain the above-mentioned plurality of sub-cubes.
[0035] Specifically, the number of sub-cubes can be set based on empirical data.
[0036] In one specific example, for a hypercube of m dimensions, when the number of sub-cubes is M, and M=k1×k2...k j ...k m ,k iis an integer, j = 1, 2,..., m. The division can be made along m dimensional directions, and the length of each dimension is divided into k i parts, so that M sub-hypercubes can be obtained. The length of each dimension of each sub-hypercube is
[0037] In step 103, the number of clusters for clustering is determined, and initial centroids are determined based on the number of clusters and the number of text vectors in each sub-cube. Compared with the prior art of randomly selecting initial centroids, the selection quality of the initial centroids can be effectively improved, the sensitivity to the selection of the initial centroids is reduced, the stability and consistency of clustering are improved, the efficiency of capturing semantic information of the text is improved, and the effect of text clustering is optimized.
[0038] Specifically, in the clustering algorithm, a cluster refers to a set of data points, and the number of clusters refers to the number of different groups into which the data set is divided.
[0039] Optionally, the process of determining the number of clusters for clustering includes determining the number of clusters for clustering based on the elbow method.
[0040] Specifically, the process of determining the number of clusters for clustering can also be based on experience or based on the silhouette coefficient method or the like.
[0041] Optionally, the process of determining the initial centroids based on the number of clusters and the number of text vectors in each sub-cube includes determining the probability of each sub-cube containing an initial centroid based on the weighted average probability of the number of text vectors in the corresponding sub-cube, and determining the initial centroids based on the probability of each sub-cube containing an initial centroid and the number of clusters.
[0042] Specifically, the weighted average probability of the number of text vectors in each sub-cube can be calculated based on the following formula:
[0043]
[0044] where N represents the total number of text vectors, H i represents the i-th sub-cube, represents the number of text vectors in the hypercube H i .
[0045] Specifically, the weighted average probability of the number of text vectors in the corresponding sub-cube can be determined as the probability of each sub-cube containing an initial centroid.
[0046] Specifically, the corresponding weighted average probability can also be subjected to other common processing, for example, multiplied by a corresponding correction coefficient determined based on other influencing factors, to obtain the probability of each sub-cube containing an initial centroid.
[0047] The process of determining the initial centroid based on the probability of each sub-cube containing the initial centroid and the cluster number includes: obtaining a target sub-cube from the cluster number of sub-cubes with the maximum probability of containing the initial centroid; and determining the geometric center of the target sub-cube as the initial centroid.
[0048] Specifically, the initial centroid probability of each sub-cube is determined based on the weighted average probability of the number of text vectors in the corresponding sub-cube, and the center of the target sub-cube obtained from the cluster number of sub-cubes with the maximum initial centroid probability is determined as the initial centroid, which can help to ensure the rationality of the selection of the initial centroid.
[0049] Specifically, the number of candidate sub-cubes exceeding the cluster number can also be obtained, and then the initial centroid is determined based on the geometric center of the candidate sub-cube.
[0050] Specifically, other points in the corresponding target sub-cube can also be determined as the initial centroid.
[0051] Step 104, based on the cluster number and the initial centroid, the text vectors are clustered to obtain a plurality of clustering result clusters, and the training text data is determined based on the plurality of clustering result clusters. Based on steps 101 to 103, by establishing the text vector corresponding to each candidate document, and drawing a hypercube including each text vector, and then dividing the hypercube, the number of text vectors in the sub-cube obtained by dividing the hypercube is determined to determine the initial centroid of clustering, which can effectively improve the selection quality of the initial centroid, reduce the sensitivity of the selection of the initial centroid, and improve the stability and consistency of clustering by clustering each text vector based on the initial centroid. The efficiency of capturing text semantic information is improved, and the text clustering effect is optimized; further, by determining the training text data based on the clustering result cluster, the efficiency of obtaining the training text data is improved, and the quality and professionalism of the obtained training text data are improved.
[0052] Optionally, the process of determining the training text data based on the plurality of clustering result clusters includes: establishing a clustering result data set corresponding to each clustering result cluster based on the candidate document corresponding to each clustering result cluster; and determining the clustering result data set most relevant to the target field in each clustering result data set as the training text data.
[0053] Specifically, the target field can be one target field or multiple target fields.
[0054] Optionally, the process of determining the clustering result data set most relevant to the target field in each clustering result data set as the training text data includes:
[0055] For a current target field in the plurality of target fields, obtain a current target data set from a clustering result data set most relevant to the current target field in each clustering result data set, and determine the current target data set as training text data corresponding to the current target field.
[0056] Optionally, the method for obtaining training text data provided by the embodiment of the present application further comprises: performing dimension reduction on the hypercube and each text vector obtained in step 101 to obtain a hypercube and text vectors after dimension reduction, and performing steps 102 to 104 based on the hypercube and text vectors after dimension reduction.
[0057] Specifically, the method for obtaining training text data provided by the embodiment of the present application can decompose a corresponding computing task into a plurality of subtasks by using a distributed computing framework, and parallelly process the plurality of subtasks on a plurality of computing nodes to accelerate the computing speed and improve the efficiency of obtaining training text data.
[0058] The method for obtaining training text data provided by the embodiment of the present application is further introduced as follows, specifically as shown in the following steps. Figure 2 Figure 1 Step 101 can comprise the following steps.
[0059] Step 1011, determining a text vector corresponding to each candidate document based on a term frequency-inverse document frequency algorithm.
[0060] Specifically, the process of determining a text vector corresponding to each candidate document based on the term frequency-inverse document frequency (TF-IDF) algorithm can comprise: traversing all text data, extracting words appearing therein to construct a vocabulary; for each candidate document, calculating the term frequency of each word in the vocabulary in the current candidate document, and calculating the inverse document frequency of each word; for each candidate document, calculating the term frequency-inverse document frequency value of each word corresponding to the current candidate document based on the term frequency and the corresponding inverse document frequency of each word corresponding to the current candidate document; and constructing a text vector of the current candidate document based on the term frequency-inverse document frequency value of each word corresponding to the current candidate document.
[0061] Step 1012, drawing a hypercube based on the number of dimensions of each text vector and the extreme value of each text vector in each dimension.
[0062] Optionally, the process of determining the number of dimensions of each text vector comprises: determining the number of dimensions of each text vector as the number of dimensions of the hypercube.
[0063] Optionally, the process of drawing the hyper-cube by the extreme values of each text vector in each dimension includes: determining the extreme values of each text vector in each dimension as the vertices of the hyper-cube.
[0064] It can be understood that, since the number of dimensions of each text vector determined based on the term frequency-inverse document frequency algorithm is equal to the number of dimensions of the constructed vocabulary, the dimensions of each text vector are determined as the number of dimensions of the hyper-cube, and the extreme values of each text vector in each dimension are determined as the vertices of the hyper-cube, so that the hyper-cube including each text vector is drawn.
[0065] Specifically, the values obtained by adding a correction value to the extreme values of each text vector in each dimension can be determined as the vertices of the hyper-cube, so as to ensure that the hyper-cube completely includes each text vector.
[0066] The embodiment of the present application can simply and conveniently draw the hyper-cube meeting the requirements of the embodiment of the present application, and can facilitate the subsequent candidate document clustering and the acquisition of training text data.
[0067] The method for acquiring training text data provided by the embodiment of the present application will be further introduced below, and specifically as shown in the following Figure 3 , step 104 in the above Figure 1 may include the following steps:
[0068] Step 1041: based on the number of clusters and the initial centroid, performing this round of clustering on each text vector to obtain a plurality of this round of result clusters.
[0069] Step 1042: based on the sum of squared errors within each this round of result cluster and the within-cluster sum of squares, verifying the effect of the this round of clustering result.
[0070] Specifically, the sum of squared errors within a cluster (SSE) and the within-cluster sum of squares (WCSS) can be used to evaluate the clustering effect, and the smaller the sum of squared errors within a cluster and the within-cluster sum of squares, the closer the corresponding data points to the corresponding centroid, and the better the clustering effect.
[0071] Optionally, the process of verifying the effect of the this round of clustering result based on the sum of squared errors within each this round of result cluster and the within-cluster sum of squares includes:
[0072] The system checks whether the sum of squared intra-cluster errors and squared intra-cluster errors of each cluster in this round of results are less than a preset threshold. If the sum of squared intra-cluster errors and squared intra-cluster errors of each cluster in this round of results are less than the preset threshold, the clustering results of this round are deemed to have passed the effectiveness verification; otherwise, the clustering results of this round are deemed to have failed the effectiveness verification.
[0073] Step 1043: If the clustering result of this round fails the effect verification, restart the execution to divide the hypercube into multiple sub-cubes on an average basis, and perform the next round of clustering on each text vector until the clustering result of this round passes the effect verification.
[0074] Step 1044: When the results of this round of clustering pass the effect verification, the multiple clusters of this round of results are determined as multiple clustering result clusters.
[0075] Optional, such as Figure 4 As shown, step 1041 may include the following steps:
[0076] Step 1041A: Based on the number of clusters and the initial centroid, perform this round of clustering on each text vector to obtain multiple current process clusters.
[0077] Specifically, the process of performing this round of clustering on each text vector based on the number of clusters and the initial centroid to obtain multiple current-round clusters includes:
[0078] Calculate the Euclidean distance between each text vector and each initial centroid, and assign each text vector to the cluster containing the initial centroid with the closest Euclidean distance; recalculate the centroid of each cluster and reassign each text vector based on the new centroid until each centroid no longer changes significantly, and determine the corresponding clusters as the process clusters for this round.
[0079] Step 1041B: Calculate the contour coefficient of each text vector based on multiple current process clusters, and verify the rationality of the number of clusters based on the contour coefficient of each text vector.
[0080] Optionally, the above process of verifying the reasonableness of the number of clusters based on the contour coefficients of each text vector includes: calculating the difference between 1 and the contour coefficients of each text vector; if the corresponding difference is less than a preset difference threshold, the number of clusters is determined to pass the reasonableness verification; otherwise, the number of clusters is determined to fail the reasonableness verification.
[0081] Understandably, the silhouette coefficient is used to measure the closeness of a data point to other data points within its own cluster, as well as the degree of separation between data points in other clusters. The silhouette coefficient value is between -1 and 1. The closer the silhouette coefficient is to 1, the more reasonable the sample clustering is. Therefore, the reasonableness of the number of clusters can be verified by whether the difference between 1 and the silhouette coefficient of each text vector is less than a preset difference threshold.
[0082] Step 1041C, when the cluster number does not pass the rationality verification, re-determine the cluster number, and re-start to execute the determination of the initial centroid based on the cluster number and the number of text vectors in each sub-cube, until the cluster number passes the rationality verification.
[0083] Step 1041D, when the cluster number passes the rationality verification, determine the multiple current round process clusters as the current round result clusters.
[0084] It can be understood that when performing text clustering, the phenomenon of too wide or too narrow distribution often occurs, for example, the text of the financial field, the sports field and the entertainment English related vocabulary, which is often clustered into a cluster in actual clustering. The embodiments of the present application can obtain the optimal clustering result by performing multiple rounds of clustering, verifying the rationality of the cluster number in each round of clustering, verifying the effect of the corresponding clustering result after each round of clustering is completed, adjusting the cluster number based on the rationality verification result, and adjusting the sub-cube and the cluster number based on the effect verification result, so as to obtain the training text data with the highest quality.
[0085] Figure 5 is a structural diagram of a training text data acquisition device provided by the embodiments of the present application, and the device is suitable for executing the training text data acquisition method provided by the embodiments of the present application. As shown in the figure, Figure 5 The device can specifically include:
[0086] The hypercube acquisition module 501 is configured to establish a text vector corresponding to each candidate document, and draw a hypercube including each text vector. After the hypercube is evenly divided into multiple sub-cubes, the initial centroid of clustering can be determined based on the number of texts in each sub-cube.
[0087] Optionally, the hypercube acquisition module 501 can be specifically configured to determine the text vector corresponding to each candidate document based on the term frequency-inverse document frequency algorithm; and
[0088] Draw the hypercube based on the number of dimensions of each text vector and the extreme value of each text vector in each dimension.
[0089] The sub-cube acquisition module 502 is configured to evenly divide the hypercube into multiple sub-cubes. The initial centroid of clustering can be determined based on the number of texts in each sub-cube.
[0090] The cluster number and centroid obtaining module 503 is configured to determine the cluster number for clustering and determine the initial centroid based on the cluster number and the number of text vectors in each sub-cube. The selection quality of the initial centroid can be effectively improved, the sensitivity to the selection of the initial centroid is reduced, the stability and consistency of clustering are improved, the efficiency of capturing semantic information of the text is improved, and the effect of text clustering is optimized.
[0091] Optionally, the cluster number and centroid obtaining module 503 can be specifically configured to determine the cluster number for clustering based on the elbow method.
[0092] The probability of each sub-cube containing the initial centroid is determined based on the weighted average probability of the number of text vectors in the corresponding sub-cube.
[0093] The initial centroid is determined based on the probability of each sub-cube containing the initial centroid and the cluster number.
[0094] Optionally, the cluster number and centroid obtaining module 503 can be specifically configured to obtain a target sub-cube from the cluster number of sub-cubes with the maximum probability of containing the initial centroid.
[0095] The geometric center of the target sub-cube is determined as the initial centroid.
[0096] The clustering result and training text data obtaining module 504 is configured to cluster each text vector based on the cluster number and the initial centroid to obtain a plurality of clustering result clusters, and determine training text data based on the plurality of clustering result clusters. In combination with the modules 501 to 503, the text vector corresponding to each candidate document can be established, the hypercube including each text vector can be drawn, the initial centroid for clustering can be determined based on the number of text vectors in the sub-cube after the hypercube is evenly divided, the selection quality of the initial centroid can be effectively improved, the sensitivity to the selection of the initial centroid is reduced, the stability and consistency of clustering can be improved, the efficiency of capturing semantic information of the text is improved, and the effect of text clustering is optimized. Further, the training text data can be determined based on the clustering result clusters, the efficiency of obtaining the training text data is improved, and the quality and professionalism of the obtained training text data are improved.
[0097] Optionally, the clustering result and training text data obtaining module 504 can be specifically configured to cluster each text vector based on the cluster number and the initial centroid to obtain a plurality of this-round result clusters.
[0098] The effect of the this-round clustering result is verified based on the within-cluster sum of squares and the inter-cluster sum of squares of each this-round result cluster.
[0099] When the current clustering result fails to pass the effectiveness verification, restarting execution of dividing the hypercube into multiple sub-cubes, performing next round clustering on each text vector, until the current clustering result passes the effectiveness verification; and
[0100] When the current clustering result passes the effectiveness verification, determining the multiple current round result clusters as the multiple clustering result clusters.
[0101] Optionally, the clustering result and training text data obtaining module 504 can be specifically configured to perform current round clustering on each text vector based on the cluster number and the initial centroid to obtain multiple current round process clusters.
[0102] Based on the multiple current round process clusters, calculating the silhouette coefficient of each text vector, and verifying the rationality of the cluster number based on the silhouette coefficient of each text vector.
[0103] When the cluster number fails to pass the rationality verification, re-determining the cluster number, and restarting execution of determining the initial centroid based on the cluster number and the number of text vectors in each sub-cube, until the cluster number passes the rationality verification; and
[0104] When the cluster number passes the rationality verification, determining the multiple current round process clusters as the current round result clusters.
[0105] Optionally, the clustering result and training text data obtaining module 504 can be specifically configured to,
[0106] Based on each clustering result cluster corresponding to a candidate document, establishing a clustering result data set corresponding to each clustering result cluster; and
[0107] Determining the clustering result data set most relevant to the target field in each clustering result data set as the training text data.
[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the above described functional modules can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0109] The embodiment of the present application further provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the training text data obtaining method provided in any of the foregoing embodiments when executing the program.
[0110] The embodiment of the present application further provides a computer readable medium, which has stored thereon a computer program, and the program is executed by a processor to implement the training text data acquisition method provided by any of the above embodiments.
[0111] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the training text data acquisition method according to any of the embodiments of the present application.
[0112] Reference is made below in conjunction with Figure 6 which shows a structural schematic diagram of a computer system 600 of an electronic device suitable for implementing the embodiment of the present application. Figure 6 The electronic device shown is merely an example, and should not bring any limitation to the functions and use range of the embodiment of the present application.
[0113] As shown in Figure 6 , the computer system 600 comprises a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage portion 608 to a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, the ROM 602 and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0114] The following components are connected to the I / O interface 605: an input portion 606 comprising a keyboard, a mouse, etc.; an output portion 607 comprising a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 comprising a hard disk, etc.; and a communication portion 609 comprising a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as necessary. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 610 as necessary, so that a computer program read therefrom is installed in the storage portion 608 as necessary.
[0115] In particular, the processes described above with reference to the flow charts can be implemented as a computer software program in accordance with the embodiments disclosed herein. For example, embodiments disclosed herein include a computer program product which includes a computer program tangibly embodied on a computer readable medium, the computer program including program code for executing the methods illustrated by the flow charts. In such embodiments, the computer program can be downloaded and installed from a network via the communication portion 609 and / or installed from the removable media 611. When the computer program is executed by the central processing unit (CPU) 601, the above-described functions defined in the system of the present application are executed.
[0116] It should be noted that the computer readable medium shown in the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer readable signal medium can include a data signal carried in a baseband or as a carrier wave in a propagated data signal, in which the computer readable program code is carried. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that can send, propagate or transfer a program for use by or in connection with an instruction execution system, device or apparatus. The program code contained on the computer readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0117] The computer program product of the present application can be a computer program embodied on a computer readable medium. The computer program product can be stored on a computer readable medium, such as a floppy disk, a compact disk, a DVD, a Blu-ray disk, a hard disk, a memory, a memory card, a ROM, a PROM, an EPROM, an EEPROM, a FLASH memory, or the like. The computer program product can also be transmitted on a computer network, such as the Internet, a local network, a wide area network, a wireless network, a cellular network, a telephone network, or the like.
[0118] The modules and / or units described in the embodiments of the present application can be implemented by software or by hardware. The described modules and / or units can also be set in a processor, for example, can be described as: a processor includes a hypercube obtaining module, a sub-cube obtaining module, a cluster number and a centroid obtaining module, and a clustering result and training text data obtaining module. In some cases, the names of these modules do not constitute a limitation on the modules themselves.
[0119] As another aspect, the present application also provides a computer readable medium, which can be included in the device described in the above embodiments, or can exist independently without being assembled into the device. The computer readable medium carries one or more programs, when the one or more programs are executed by the device, the device is caused to implement: establishing a text vector corresponding to each candidate document, and drawing a hypercube including each text vector; dividing the hypercube into a plurality of sub-cubes averagely; determining a cluster number for clustering and determining an initial centroid based on the cluster number and the number of text vectors in each sub-cube; and clustering each text vector based on the cluster number and the initial centroid to obtain a plurality of clustering result clusters, and determining training text data based on the plurality of clustering result clusters.
[0120] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made depending on design requirements and other factors. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method for acquiring training text data, characterized in that, include: Create the text vectors corresponding to each candidate document and draw a hypercube that includes each text vector; The hypercube is divided into multiple sub-cubes on an equal basis; The number of clusters to be clustered is determined, and the initial centroid is determined based on the number of clusters and the number of text vectors within each sub-cube; as well as Based on the number of clusters and the initial centroid, the text vectors are clustered to obtain multiple clustering result clusters, and the training text data is determined based on the multiple clustering result clusters.
2. The training text data acquisition method according to claim 1, characterized in that, The step of establishing the text vectors corresponding to each candidate document and drawing a hypercube including each text vector includes: The text vector corresponding to each candidate document is determined based on the term frequency-inverse document frequency algorithm; and The hypercube is drawn based on the number of dimensions of each text vector and the extreme values of each text vector in each dimension.
3. The training text data acquisition method according to claim 1, characterized in that, The step of determining the number of clusters for clustering and determining the initial centroid based on the number of clusters and the number of text vectors within each sub-cube includes: The number of clusters to be clustered is determined based on the elbow method; Based on the weighted average probability of the number of text vectors within the corresponding sub-cube, the probability that each sub-cube contains an initial centroid is determined; and The initial centroid is determined based on the probability that each sub-cube contains an initial centroid and the number of clusters.
4. The training text data acquisition method according to claim 3, characterized in that, The determination of the initial centroid based on the probability that each sub-cube contains an initial centroid and the number of clusters includes: The target sub-cube is obtained by selecting the number of sub-cubes containing the cluster with the highest initial centroid probability from the plurality of sub-cubes; and The geometric center of the target sub-cube is determined as the initial centroid.
5. The training text data acquisition method according to claim 1, characterized in that, The process of clustering the text vectors based on the number of clusters and the initial centroids to obtain multiple clustering result clusters includes: Based on the number of clusters and the initial centroid, the current round of clustering is performed on each text vector to obtain multiple current round result clusters; The effectiveness of this round of clustering results is verified based on the sum of squared errors within each cluster and the sum of squared errors within each cluster in this round. If the clustering result in this round fails the effectiveness verification, the process of dividing the hypercube into multiple sub-cubes on an average basis and performing the next round of clustering on each text vector is restarted until the clustering result in this round passes the effectiveness verification; and When the results of this round of clustering pass the effectiveness verification, the multiple clusters of results from this round are determined as the multiple clustering result clusters.
6. The training text data acquisition method according to claim 5, characterized in that, The current round of clustering of the text vectors based on the number of clusters and the initial centroid yields multiple current round result clusters, including: Based on the number of clusters and the initial centroid, multiple current process clusters are obtained by performing current-round clustering on each text vector; The contour coefficients of each text vector are calculated based on the multiple current process clusters, and the reasonableness of the number of clusters is verified based on the contour coefficients of each text vector. If the number of clusters fails the validity verification, the number of clusters is re-determined, and the process of determining the initial centroid based on the number of clusters and the number of text vectors within each sub-cube is restarted until the number of clusters passes the validity verification; and When the number of clusters passes the reasonableness verification, the multiple current process clusters are determined as the current result clusters.
7. The training text data acquisition method according to claim 1, characterized in that, The process of determining training text data based on the multiple clustering result clusters includes: Based on the candidate documents corresponding to each clustering result cluster, a clustering result dataset corresponding to each clustering result cluster is established; and The clustering result dataset that is most relevant to the target domain among all the clustering result datasets is determined as the training text data.
8. A training text data acquisition device, characterized in that, include: The hypercube acquisition module is used to build the text vectors corresponding to each candidate document and draw a hypercube that includes each text vector. A sub-cube acquisition module is used to divide the hypercube into multiple sub-cubes on an average basis. The cluster number and centroid acquisition module is used to determine the number of clusters to be clustered and to determine the initial centroid based on the number of clusters and the number of text vectors in each sub-cube; as well as The clustering result and training text data acquisition module is used to cluster each text vector based on the number of clusters and the initial centroid to obtain multiple clustering result clusters, and to determine the training text data based on the multiple clustering result clusters.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the training text data acquisition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the training text data acquisition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Initial clustering center optimization selection method based on RFID data intensity-time distribution
CN107437093A
Distributed satellite group optimization design method for tracking and observation of space passage
CN109656133A