Text vector generation method and related device
By using text vector model and Bert layer in text processing equipment, using prior vector assistance to automatically learn word segmentation, the limitations of relying on dictionaries in the prior art are solved, and the accuracy of text sequence conversion is improved.
Patent Information
- Application Number
- CN202210290851.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-23
AI Technical Summary
When processing language texts without vocabulary delimiters, the prior art relies on the dictionary for word segmentation, resulting in the effect of word segmentation being affected by dictionary selection and has limitations.
By using the text vector model, including the Bert layer, in the text processing device, we obtain the prior vector of the text sequence, and input it into the Bert layer with word vectors, position vectors, and segment vectors, to automatically learn word segmentation.
Without relying on dictionary, word segmentation is performed through a priori vector auxiliary text vector model, which improves the accuracy of text sequence conversion and obtains more accurate text vectors.
Smart Images

Figure CN114611511B_ABST
Abstract
Claims
1. A method for generating text vectors, characterized in that, applied to a text processing device, the text processing device is configured with a text vector model, the text vector model includes a Bert layer, and the method includes: Obtain a text sequence; according to the text sequence, construct a plurality of square matrices to be initialized, wherein the rows and columns of each square matrix correspond one-to-one with the text sequence; Select different texts from the text sequence for each of the plurality of square matrices as the starting text of the square matrix, wherein the starting text of each square matrix and the text corresponding to each row of the square matrix form a plurality of text pairs; For each square matrix, initialize each row of the square matrix respectively according to the text fragments intercepted from the text sequence by the text pairs of each row of the square matrix, wherein if the text fragments intercepted from the text sequence by the text pairs of each row can form a vocabulary, initialize the positions corresponding to the text fragments in that row to a first preset value, and initialize the positions in that row that do not correspond to the text fragments to a second preset value; Take all the initialized square matrices as the word matrix of the text sequence; Convert the word matrix into a prior vector, wherein the prior vector has the same vector dimension as the word vector, position vector and segment vector of the text sequence; Generate the word vector, position vector and segment vector of the text sequence according to the convention of the Bert layer for input vectors; Input the word vector, position vector, segment vector and the prior vector of the text sequence into the Bert layer to obtain the text vector of the text sequence.
2. The method for generating text vectors according to claim 1, characterized in that, the text vector model further includes a feature extraction layer, and the converting the word matrix into the prior vector includes: Input the word matrix into the feature extraction layer to obtain the prior vector.
3. The method for generating text vectors according to claim 1, characterized in that, the inputting the word vector, position vector, segment vector and the prior vector of the text sequence into the Bert layer to obtain the text vector of the text sequence includes: Fuse the word vector, position vector, segment vector and the prior vector of the text sequence to obtain a fused vector; Input the fused vector into the Bert layer to obtain the text vector of the text sequence.
4. The method for generating text vectors according to claim 3, characterized in that, the fusing the word vector, position vector, segment vector and the prior vector of the text sequence to obtain a fused vector includes: Fuse the word vector, position vector, segment vector and the prior vector of the text sequence in a summing manner to obtain the fused vector.
5. A text vector generation device, characterized in that, applied to a text processing device, the text processing device is configured with a text vector model, the text vector model includes a Bert layer, and the text vector generation device includes: A prior information module, configured to obtain a text sequence; and construct a plurality of square matrices to be initialized according to the text sequence, wherein the rows and columns of each square matrix respectively correspond to the text sequence one by one; Select different texts from the text sequence for each of the plurality of square matrices as the starting text of the square matrix, wherein the starting text of each square matrix and the text corresponding to each row of the square matrix form a plurality of text pairs; For each square matrix, initialize each row of the square matrix respectively according to the text segment intercepted from the text sequence by the text pair of each row of the square matrix. If the text segments intercepted from the text sequence by the text pairs of each row can form a vocabulary, initialize the positions corresponding to the text segments in that row to a first preset value, and initialize the positions in that row not corresponding to the text segments to a second preset value; Use all the initialized square matrices as the word matrix of the text sequence; Convert the word matrix into a prior vector, wherein the prior vector has the same vector dimension as the word vector, position vector and segment vector of the text sequence; A vector generation module, configured to generate the word vector, position vector and segment vector of the text sequence according to the convention of the Bert layer on the input vector; The vector generation module is further configured to input the word vector, position vector, segment vector and the prior vector of the text sequence into the Bert layer to obtain the text vector of the text sequence.
6. A text processing device, characterized in that, the text processing device includes a processor and a memory, the memory stores a computer program, and when the computer program is executed by the processor, the text vector generation method according to any one of claims 1-4 is implemented.
7. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the text vector generation method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Chinese grammar debugging method and system based on multivariate text features
CN112183094A