Method and device for generating word vector
By obtaining and analyzing the field description information and enumeration values of short text, combining the correspondence between data types and features, participle words and calculating subtext weights, the problem of short text word vectors is solved, and more accurate natural language processing is achieved.
Patent Information
- Application Number
- CN202111444646.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-11-30
AI Technical Summary
In natural language processing, word vectors of short texts are difficult to accurately express their semantic information, which affects the accuracy of natural language processing.
By obtaining the field description information and field enumeration values of the target field, combining the correspondence between the predefined data type and data characteristics, the target data type and matching probability value are determined, and the field description information is segmented, the subtext weight is determined based on the part of speech and data type of the subtext, and finally the word vector is calculated to generate the word vector.
This method can more accurately calculate the word vector of field description information, reflecting its semantic center of gravity, thereby improving the accuracy of natural language processing.
Smart Images

Figure CN114139537B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a method and device for generating word vectors. Background Art
[0002] In the process of natural language processing, the text needs to be vectorized first to obtain the word vector corresponding to the text, and then the word vector is calculated to obtain the intrinsic semantic relationship of the natural language, so that the computer can understand the meaning of the natural language.
[0003] Generally speaking, when determining a word vector for a text, the text's context can be analyzed to ensure that the word vector fully reflects the text's characteristic information. However, when processing short texts, due to their sparse vocabulary and discrete semantics, the word vectors obtained by analyzing the context are difficult to accurately express the semantic information of the short text, affecting the accuracy of natural language processing. Summary of the Invention
[0004] In view of this, the present application provides a method and device for generating word vectors.
[0005] Specifically, this application is implemented through the following technical solutions:
[0006] According to the first aspect of the present application, a method for generating a word vector is proposed, comprising:
[0007] Get the field description information and field enumeration value corresponding to the target field;
[0008] Determine the target data type and the corresponding target matching probability value corresponding to the data feature that best matches the field enumeration value based on the predefined correspondence between the data type and the data feature;
[0009] Segmenting the field description information to obtain at least one subtext, and determining a subtext weight of each subtext based on the part of speech of each subtext, the target data type, and a predefined correspondence between the data type, part of speech, and weight;
[0010] The word vector of each subtext is calculated according to the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information.
[0011] According to the second aspect of the present application, a device for generating a word vector is proposed, comprising:
[0012] The acquisition unit is used to obtain the field description information and field enumeration value corresponding to the target field;
[0013] A first determining unit is configured to determine a target data type and a corresponding target matching probability value corresponding to a data feature that best matches the field enumeration value based on a predefined correspondence between data types and data features;
[0014] a second determining unit, configured to segment the field description information to obtain at least one subtext, and determine a subtext weight of each subtext based on the part of speech of each subtext, the target data type, and a predefined correspondence between the data type, part of speech, and weight;
[0015] The first calculation unit is used to calculate the word vector of each subtext according to the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information.
[0016] According to a third aspect of the present application, an electronic device is provided, including:
[0017] processor;
[0018] a memory for storing processor-executable instructions;
[0019] The processor implements the method described in the embodiment of the first aspect above by running the executable instructions.
[0020] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in the embodiment of the first aspect are implemented.
[0021] It can be seen from the technical solution provided by the above application that the application can more accurately calculate the word vector of the field description information by comprehensively considering the data type and part of speech of each sub-text in the field description information, and make the word vector more accurately reflect the semantic center of gravity of the field description information. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0023] Figure 1 This is a flowchart of a method for generating a word vector according to an exemplary embodiment of the present application;
[0024] Figure 2 This is a flowchart of a method for generating a word vector according to an exemplary embodiment of the present application;
[0025] Figure 3 1 is a schematic diagram of an electronic device for generating a word vector according to an exemplary embodiment of the present application;
[0026] Figure 4 It is a block diagram of a device for generating a word vector according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0027] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0028] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0030] Next, the embodiments of the present application are described in detail.
[0031] With the development of computer science and technology, more and more data appears on the Internet in the form of short texts. Similarity calculation based on short text word vectors can promote many natural language processing tasks and has wide applications in search engines, recommendation systems, machine translation, automatic question answering, named entity recognition, spelling correction and other fields.
[0032] In natural language processing, word vectors refer to using a high-dimensional vector to represent the meaning of a word. In related technologies, word vectors can usually be obtained by taking a weighted average or sum of the word vectors of the individual characters that make up the word, or by using an attention mechanism to fully interact with the context of the text and integrate the complete context for calculation. However, when calculating the word vectors of short texts, since short texts are usually composed of several subtexts, the field length is short and the amount of information contained is small, the above two methods cannot determine the semantic focus of the short text, and the calculated word vector cannot accurately express the semantic information of the short text.
[0033] As an example to solve the above problem, this application provides a method for generating word vectors. Figure 1 FIG. 1 is a flow chart of a method for generating a word vector according to an exemplary embodiment of the present application. Figure 1 As shown, the following steps may be included:
[0034] Step 102: Obtain the field description information and field enumeration value corresponding to the target field.
[0035] Among them, the field description information is the short text for which the word vector needs to be generated, and the field enumeration value is the sample data corresponding to the field description information, which is used to illustrate the field description information. For example, when the field description information is "gender", the corresponding field enumeration value can be "male" and / or "female"; when the field description information is "birthday", the corresponding field enumeration value can be "1989.11.23" and / or "March 11, 1996", etc. When the target field is a standard data field, since standard data generally has a corresponding dictionary table that records the dictionary value of each field, the dictionary value corresponding to the field can also be used as the field enumeration value of the field.
[0036] In one embodiment, the target field can be any field in the target data table. During the data governance process, the data table that has undergone data exploration typically contains the value range distribution information of the original data table. Therefore, by performing data exploration on the target data table, a data exploration table corresponding to the target data table can be obtained. This data exploration table can contain field description information and field enumeration values corresponding to each field in the target data table. By searching this data exploration table, the field description information and field enumeration values corresponding to the target field can be determined.
[0037] Step 104: According to the predefined correspondence between data types and data features, determine the target data type corresponding to the data feature that best matches the field enumeration value and the corresponding target matching probability value.
[0038] In the technical solution of the present application, a correspondence between data types and data features is pre-set, wherein data features may include features such as length, format, and character type. For example, according to common knowledge in the art, eight-digit numbers similar to the format of "2020-07-01" or "2020 / 07 / 01", or data containing specific characters such as year, month, and day similar to "July 1, 2020" are usually used to represent dates, so "date formats such as July 1, 2020, 2020-07-01, 2020 / 07 / 01, or full digital representations with a length of 8" can be used as data features of the data type "date". Therefore, in this embodiment, the correspondence between data types and data features can be summarized by those skilled in the art based on common sense or experience, or can be set based on the GB / T unified standard, and this application does not impose any restrictions on this.
[0039] On this basis, the field enumeration value corresponding to the field description information of the target field can be compared with the data features corresponding to each data type to determine whether the format, length, etc. of the field enumeration value conform to the data features in the predefined correspondence. The data feature that best matches the field enumeration value is determined based on the comparison result. The data type corresponding to the data feature is the target data type that best matches the field description information corresponding to the field enumeration value, and the matching probability value is the target matching probability value of the field description information matching the target data type. For example, two data types, "date" and "code", are predefined. The data characteristics corresponding to the "date" data type are "data similar to the format of July 1, 2020, 2020-07-01, 2020 / 07 / 01, or all numbers with a length of 8." The data characteristics corresponding to the "code" data type are "all numbers and a length greater than 11; or containing Chinese, English, and numbers and a length less than 15, and not in the format of July 1, 2020; or all letters and a length greater than or equal to 5." If the field enumeration value of the target field is "13082300470501", then this field enumeration value best matches the characteristic of "all numbers and a length greater than 11". It can be determined that the data type that best matches the field description information corresponding to the target field is the "code" corresponding to the characteristic of "all numbers and a length greater than 11".
[0040] In one embodiment, a multi-classification model can be pre-built based on the correspondence between predefined data types and data features. After obtaining the field description information and field enumeration value corresponding to the target field, the field enumeration value can be input into the trained multi-classification model. The multi-classification model processes the input field enumeration value and outputs a matching probability value for each data feature that matches it. The data type corresponding to the maximum value among the matching probability values is determined as the target data type that best matches the field description information corresponding to the target field, and the maximum value is determined as the target matching probability value for the field description information matching the target data type.
[0041] In another embodiment, in some cases, the expression of field enumeration values may not be expressed in a standard format. For example, when the field description information is "gender", the field enumeration value should generally be "male" or "female", but in some fields it may be represented by "0" or "1". If the field enumeration value is matched only based on the correspondence between predefined data types and data features, errors may occur. And because the word vectors corresponding to semantically similar words are spatially close, when building a multi-classification model, in addition to the correspondence between field types and field features, it can also be trained based on the correspondence between predefined word vectors and data types. When determining the target data type corresponding to the data feature that best matches the field enumeration value and the corresponding target matching probability value, the original word vector of the field description information can be obtained first according to the method of generating word vectors in the relevant technology. For example, an original word vector model is trained based on a large amount of Chinese or industry-related corpus, and the field description information is input into the original word vector model to obtain the corresponding original word vector. Although the original word vector is not accurate enough, it can be input into the above-mentioned trained multi-classification model together with the field enumeration value to obtain the matching probability value of each data feature that matches the field enumeration value and the original word vector. The data type corresponding to the maximum value of each matching probability value is determined as the target data type that best matches the field description information corresponding to the target field, and the maximum value is determined as the target matching probability value of the field description information matching the target data type. By increasing the consideration of the original word vector of the character description information, the deviation caused by the non-standard field enumeration value can be avoided, and the accuracy of determining the data type to which the field description information belongs can be improved.
[0042] Furthermore, when obtaining the original word vector of the field description information, since the original word vector model pre-trained in the related art can process multiple input field descriptions at the same time, it is necessary to first fill the field description information during the processing so that the length of each field description information is unified to the maximum length of the input field description information. Therefore, the various field description information can be sorted according to the field length first, and the various field description information can be grouped according to the sorting result to divide the preset number of adjacent field description information into the same text group. For example, if there are 6 field description information to be processed, and their lengths are 5, 3, 22, 25, 9, and 12 respectively, if the pre-trained original word vector model can process 3 field description information at a time, compared to directly dividing the 3 field description information with lengths of 5, 3, and 22 into the first text group, the field description information with lengths of 25, 9, and 12 is divided into the second text group, so that the original word vector model fills the length of each field description information in the first text group to 22 when processing the first text group, and fills the length of each field description information in the second text group to 25. In this embodiment, the fields can be sorted according to their length first, and the three field description information with lengths of 3, 5, and 12 can be divided into the first text group, and the field description information with lengths of 13, 22, and 25 can be divided into the second text group. In this way, when the original word vector model processes the first text group, it only needs to fill the length of each field description information therein to 12, and fill the length of each field description information in the second text group to 25, thereby improving the efficiency of the original word vector model in processing the first text group.
[0043] Step 106: Segment the field description information to obtain at least one subtext, and determine the subtext weight of each subtext based on the part of speech of each subtext, the target data type, and the predefined correspondence between data type, part of speech, and weight.
[0044] Among them, word segmentation can be to divide the text content in the field description information into multiple sub-texts, wherein the field description information can be divided according to the context semantics when dividing, so that the word segmentation of the field description information is more accurate. For example, "equipment in use status" can be divided into four words: "equipment", "in", "use" and "status". When the field description information is segmented, the sub-texts after the division can be tagged with parts of speech. For example, "equipment", "in", "use" and "status" can be tagged with "noun", "preposition", "preposition" and "noun" respectively. Among them, the word segmentation of the field description information and the part-of-speech tagging of the sub-texts after word segmentation can refer to the relevant content in the prior art, and this application does not limit this.
[0045] In this application, the correspondence between parts of speech and weights is pre-set for each data type. Different parts of speech within the same data type have different weights, and the same part of speech can also have different weights in different data types. The corresponding correspondence between parts of speech and weights can be determined based on the target data type determined above, and the subtext weight corresponding to each subtext can be determined based on the part of speech of each subtext determined above.
[0046] Step 108: Calculate the word vector of each subtext according to the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information.
[0047] In one embodiment, after determining the subtext weight corresponding to the target data type, the target matching probability value can be multiplied by the subtext weight corresponding to each subtext to obtain a modified weight, and the modified weight is normalized to obtain a normalized weight corresponding to each subtext. Based on the normalized weight, the word vectors of each subtext are weighted and summed to calculate the word vector corresponding to the field description information. The word vector of each subtext can be obtained by inputting each subtext into a word vector model trained based on a large amount of Chinese or industry-specific corpus.
[0048] Based on the above-mentioned word vector generation method, the word vectors of the field description information and the standard metadata can be calculated respectively, and the semantic similarity between the field description information and the standard metadata can be calculated through the word vectors of the field description information and the standard metadata, so as to determine the standard data element with the highest semantic similarity with the word vector of the field description information as the target data source associated with the field description information, so that the recommended association between data items and standard data elements can be realized in the process of data governance.
[0049] It can be seen from the technical solution provided by the present application that the present application determines the target data type to which the field description information belongs and the probability of belonging to the target data type through the data features of each data type summarized in advance, and calculates the word vector of the subtext based on the probability and the part-of-speech weight corresponding to the subtext after the field description information is segmented to obtain the word vector of the field description information. The generated word vector can more accurately represent the key content in the field description information, so as to further realize the natural language processing task based on the word vector. Figure 3 Detailed description is given. Figure 2 This is a flow chart of a method for generating a word vector according to an exemplary embodiment of the present application. Figure 2 As shown, the word vector generation method may include the following steps:
[0050] Step 201: Obtain a target data table.
[0051] Step 202: Perform data exploration on the target data table to obtain a data exploration table.
[0052] Table 1 is a data exploration table obtained after analyzing a dry beach equipment table.
[0053]
[0054]
[0055] Table 1
[0056] Step 203: Obtain field description information and field enumeration values.
[0057] The field description information and field enumeration values corresponding to each field in the target data table can also be obtained from the data exploration table shown in Table 2. For example, for the field description information of "Device Number", its corresponding field enumeration values are "13082300470501", "42030200010103", etc.
[0058] Step 204: Determine the target data type and target matching probability value.
[0059] Table 2 shows the correspondence between predefined data types and data features. In this embodiment, a multi-classification model is pre-constructed based on the correspondence between data types and data features shown in Table 2.
[0060]
[0061]
[0062] Table 2
[0063] Input the field description information in Table 1 into the multi-classification model and obtain the matching probability value output by the multi-classification model that matches each field description information with each data feature. Table 3 shows the matching probability value corresponding to each field description information and each data type.
[0064]
[0065] Table 3
[0066] Determine the data type corresponding to the maximum value among the respective matching probability values as the target data type that most closely matches the field description information corresponding to the target field, and determine this maximum value as the target matching probability value of the field description information matching the target data type. For example, as shown in Table 3, for the field description information of "equipment number", the maximum value among its matching values with each data type is 0.83. Therefore, the data type "code" corresponding to 0.83 can be used as the target data type of "equipment number", and 0.83 can be used as the target matching probability value of "equipment number".
[0067] Step 205: Perform word segmentation on the field description information to obtain the sub-texts corresponding to the field description information, and perform word nature tagging on the word segmentation result.
[0068] For the field description information, tools such as HanLP tool or LTP tool can be used for processing to divide the text content in each field description information into multiple sub-texts and tag the word natures of each sub-text.
[0069] Table 4 shows the word segmentation and word nature tagging results for the above field description information. For example, for "equipment in-use status", it can be divided into four sub-texts: "equipment", "in", "use", and "status". Among them, "equipment", "in", "use", and "status" are respectively tagged as "noun", "preposition", "preposition", and "noun".
[0070]
[0071]
[0072] Table 4
[0073] Step 206: Determine the sub-text weights of each sub-text.
[0074] As shown in Table 5, in the predefined correspondence between data types, word natures, and weights, different word natures under the same data type correspond to different weights, and the same word nature can also be set with different weights in different data types.
[0075]
[0076] Table 5
[0077] For example, the target data type corresponding to the field description information of "equipment number" is code. Among them, the word nature corresponding to "equipment" is noun, and its sub-text weight is 0.46. While the target data type corresponding to the field description information of "equipment in-use status" is code, and the word nature corresponding to "in" is preposition, and its sub-text weight is 0.32.
[0078] Step 207: Determine the word vector of each subtext.
[0079] Each sub-text is input into a word vector model that has been pre-trained based on a large amount of Chinese or industry-specific corpus, and the word vector corresponding to each sub-text is output by the word vector model.
[0080] Step 208: Determine the revised weight of each subtext in the field description information.
[0081] After determining the sub-text weight corresponding to the target data type, the target matching probability value can be multiplied by the sub-text weight corresponding to each sub-text to obtain the corrected weight. It should be noted that this application does not limit the order of the above steps 207 and 208.
[0082] For example, if the target data type for the "Device Number" field description is code, and its target match probability is 0.83, then the corrected weight for "Device" is 0.83*0.46, and the corrected weight for "Code" is 0.83*0.46. If the target data type for the "Device In Use Status" field description is code, and its target match probability is 0.79, then for this field description, the corrected weight for "Device" is 0.79*0.36, the corrected weight for "In" is 0.79*0.32, the corrected weight for "Used" is 0.79*0.32, and the corrected weight for "Status" is 0.79*0.36.
[0083] Step 209: Calculate and obtain the word vector of the field description information.
[0084] For each field description information, the corrected weight determined above is first normalized to obtain the normalized weight corresponding to each sub-text, and then the word vectors of each sub-text determined above are weighted and summed based on the normalized weight, so that the word vector of the field description information can be calculated.
[0085] Corresponding to the above method embodiment, this specification also provides an embodiment of a device.
[0086] Figure 3 This is a schematic diagram of a structure of an electronic device for generating a word vector according to an exemplary embodiment of the present application. Figure 3At the hardware level, the electronic device includes a processor 302, an internal bus 304, a network interface 306, a memory 308, and a non-volatile memory 310. Of course, it may also include hardware required for other services. The processor 302 reads the corresponding computer program from the non-volatile memory 310 into the memory 308 and then runs it. Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0087] Figure 4 1 is a block diagram of a device for generating a word vector according to an exemplary embodiment of the present application. Figure 4 The apparatus includes an acquiring unit 402, a first determining unit 404, a second determining unit 406, and a first calculating unit 408, wherein:
[0088] The acquiring unit 402 is configured to acquire field description information and field enumeration value corresponding to the target field.
[0089] Optionally, obtaining the field description information and field enumeration value corresponding to the target field includes: performing data exploration on the target data table to obtain a data exploration table corresponding to the target data table, the data exploration table containing field description information and field enumeration values corresponding to each field in the target data table; determining any field in the target data table as a target field, and obtaining the field description information and field enumeration value corresponding to the target field in the data exploration table.
[0090] The first determining unit 404 is configured to determine a target data type and a corresponding target matching probability value corresponding to the data feature that best matches the field enumeration value according to a predefined correspondence between data types and data features.
[0091] Optionally, the method of determining the target data type and the corresponding target matching probability value corresponding to the data feature that best matches the field enumeration value based on the correspondence between predefined data types and data features includes: inputting the field enumeration value into a pre-trained multi-classification model to obtain the matching probability values of each data feature that matches the field enumeration value; wherein the multi-classification model is constructed based on the correspondence between predefined data types and data features; determining the maximum value among each matching probability value, and determining the data type corresponding to the maximum value as the target data type, and determining the maximum value as the target matching probability value.
[0092] Optionally, determining the target data type and the corresponding target matching probability value corresponding to the data feature that best matches the field enumeration value based on the correspondence between predefined data types and data features includes: obtaining the original word vector of the field description information; inputting the original word vector and the field enumeration value into a pre-trained multi-classification model to obtain the matching probability values of each data feature that matches the field enumeration value and the original word vector; wherein the pre-trained multi-classification model is constructed based on the correspondence between predefined data types and data features and the correspondence between predefined word vectors and data features; determining the maximum value among each matching probability value, and determining the data type corresponding to the maximum value as the target data type, and determining the maximum value as the target matching probability value.
[0093] Optionally, there are multiple target fields, and obtaining the original word vector of the field description information includes: sorting each field description information according to the field length; grouping each field description information according to the sorting result to divide a preset number of adjacent field description information into the same text group; inputting each text group into a pre-trained original word vector model to obtain the original word vector of each field description information in each text group.
[0094] The second determining unit 406 is configured to segment the field description information to obtain at least one subtext, and determine the subtext weight of each subtext according to the part of speech of each subtext, the target data type and the predefined correspondence between data type, part of speech and weight.
[0095] The first calculation unit 408 is configured to calculate the word vector of each subtext according to the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information.
[0096] Optionally, the word vector of each sub-text is calculated based on the sub-text weight of each sub-text and the target matching probability value to obtain the word vector of the field description information, including: multiplying the sub-text weight of each sub-text by the target matching probability value to obtain the corrected weight of each sub-text; normalizing the corrected weight of each sub-text to obtain the normalized weight of each sub-text; and weighted summing the word vectors of each sub-text according to the normalized weight of each sub-text to obtain the word vector of the field description information.
[0097] Optionally, the above device further includes:
[0098] The second calculation unit 410 is configured to calculate the semantic similarity between the word vector of the field description information and the word vector of each predefined standard data element.
[0099] The third determining unit 412 is configured to determine the standard data element having the highest semantic similarity with the word vector of the field description information as the target data element associated with the field description information.
[0100] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0101] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0102] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory including instructions. The instructions may be executed by a processor of a word vector generation device to implement any of the methods described in the above embodiments. For example, the method may include:
[0103] Obtain field description information and field enumeration value corresponding to the target field; determine the target data type and the corresponding target matching probability value corresponding to the data feature that best matches the field enumeration value based on the predefined correspondence between the data type and the data feature; segment the field description information to obtain at least one subtext, and determine the subtext weight of each subtext based on the part of speech of each subtext, the target data type, and the predefined correspondence between the data type, part of speech and weight; calculate the word vector of each subtext based on the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information.
[0104] The non-temporary computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc., and this application does not limit this.
[0105] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for generating a word vector, characterized in that: The method comprises: Get the field description information and field enumeration value corresponding to the target field; According to the correspondence between the predefined data types and data features, determine the target data type corresponding to the data feature that best matches the field enumeration value and the corresponding target matching probability value; Segmenting the field description information to obtain at least one subtext, and determining a subtext weight of each subtext according to the part of speech of each subtext, the target data type, and a predefined correspondence between the data type, part of speech, and weight; Calculating the word vector of each subtext according to the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information; The step of calculating the word vector of each subtext according to the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information includes: Multiplying the subtext weight of each subtext by the target matching probability value respectively to obtain a modified weight of each subtext; Normalizing the corrected weights of each sub-text to obtain a normalized weight of each sub-text; The word vectors of each sub-text are weighted and summed according to the normalized weights of each sub-text to obtain the word vector of the field description information.
2. The method according to claim 1, characterized in that: The obtaining of field description information and field enumeration values corresponding to the target field includes: Performing data exploration on the target data table to obtain a data exploration table corresponding to the target data table, wherein the data exploration table includes field description information and field enumeration values corresponding to each field in the target data table; Any field in the target data table is determined as a target field, and field description information and field enumeration value corresponding to the target field in the data exploration table are obtained.
3. The method according to claim 1, characterized in that: The determining, according to the correspondence between the predefined data types and the data features, the target data type corresponding to the data feature that best matches the field enumeration value and the corresponding target matching probability value comprises: Input the field enumeration value into a pre-trained multi-classification model to obtain a matching probability value of each data feature matching the field enumeration value; wherein the multi-classification model is constructed according to a pre-defined correspondence between data types and data features; The maximum value among the matching probability values is determined, and the data type corresponding to the maximum value is determined as the target data type, and the maximum value is determined as the target matching probability value.
4. The method according to claim 1, characterized in that: The determining, according to the correspondence between the predefined data types and the data features, the target data type corresponding to the data feature that best matches the field enumeration value and the corresponding target matching probability value comprises: Obtaining the original word vector of the field description information; Inputting the original word vector and the field enumeration value into a pre-trained multi-classification model to obtain a matching probability value of each data feature that matches the field enumeration value and the original word vector; wherein the pre-trained multi-classification model is constructed according to a pre-defined correspondence between data types and data features and a pre-defined correspondence between word vectors and data features; The maximum value among the matching probability values is determined, and the data type corresponding to the maximum value is determined as the target data type, and the maximum value is determined as the target matching probability value.
5. The method according to claim 4, characterized in that: There are multiple target fields, and obtaining the original word vector of the field description information includes: Sort the description information of each field according to the length of the field; Grouping each field description information according to the sorting result to divide a preset number of adjacent field description information into the same text group; Each text group is input into the pre-trained original word vector model to obtain the original word vector of each field description information in each text group.
6. The method according to claim 1, characterized in that: The method further comprises: Calculating the semantic similarity between the word vector of the field description information and the word vector of each predefined standard data element; The standard data element having the highest semantic similarity with the word vector of the field description information is determined as the target data element associated with the field description information.
7. A device for generating a word vector, characterized in that: The device comprises: An acquisition unit, used to acquire field description information and field enumeration value corresponding to a target field; A first determining unit, configured to determine a target data type corresponding to a data feature that best matches the field enumeration value and a corresponding target matching probability value according to a predefined correspondence between the data type and the data feature; A second determination unit is used to segment the field description information to obtain at least one subtext, and determine a subtext weight of each subtext according to the part of speech of each subtext, the target data type, and a predefined correspondence between the data type, the part of speech, and the weight; A first calculation unit, configured to calculate a word vector of each subtext according to a subtext weight of each subtext and the target matching probability value, so as to obtain a word vector of the field description information; The step of calculating the word vector of each subtext according to the subtext weight of each subtext and the target matching probability value to obtain the word vector of the field description information includes: Multiplying the subtext weight of each subtext by the target matching probability value respectively to obtain a modified weight of each subtext; Normalizing the corrected weights of each sub-text to obtain a normalized weight of each sub-text; The word vectors of each sub-text are weighted and summed according to the normalized weights of each sub-text to obtain the word vector of the field description information.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 6 by running the executable instructions.
9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Short-text content classification method and system
CN108595440A
Word vector generation method and device, computer equipment and storage medium
CN110888984A