Text vector generation method and related device

By using text vector model and Bert layer in text processing equipment, using prior vector assistance to automatically learn word segmentation, the limitations of relying on dictionaries in the prior art are solved, and the accuracy of text sequence conversion is improved.

CN114611511BActive Publication Date: 2025-05-30SHANGHAI ZHENGDA XIMALAYA NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210290851.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-05-30
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

When processing language texts without vocabulary delimiters, the prior art relies on the dictionary for word segmentation, resulting in the effect of word segmentation being affected by dictionary selection and has limitations.

Method used

By using the text vector model, including the Bert layer, in the text processing device, we obtain the prior vector of the text sequence, and input it into the Bert layer with word vectors, position vectors, and segment vectors, to automatically learn word segmentation.

Benefits of technology

Without relying on dictionary, word segmentation is performed through a priori vector auxiliary text vector model, which improves the accuracy of text sequence conversion and obtains more accurate text vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611511B_ABST
    Figure CN114611511B_ABST
Patent Text Reader

Abstract

In the text vector generation method, model training method, and related devices provided by this application, for the obtained text sequence, the text processing device inputs the prior vector of the text sequence, together with the word vectors, position vectors, and segment vectors of the text sequence, into the Bert layer of the text vector model, enabling the text vector model to use the prior vector of the text sequence as a reference to obtain possible lexical knowledge in the text sequence for converting the text sequence into a text vector. Since the prior vector carries prior information of the words in the text sequence, it is possible to assist the text vector model in converting the text sequence through this prior information without relying on a dictionary for word segmentation, and obtain a more accurate text vector of the text sequence.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for generating text vectors, characterized in that, applied to a text processing device, the text processing device is configured with a text vector model, the text vector model includes a Bert layer, and the method includes: Obtain a text sequence; according to the text sequence, construct a plurality of square matrices to be initialized, wherein the rows and columns of each square matrix correspond one-to-one with the text sequence; Select different texts from the text sequence for each of the plurality of square matrices as the starting text of the square matrix, wherein the starting text of each square matrix and the text corresponding to each row of the square matrix form a plurality of text pairs; For each square matrix, initialize each row of the square matrix respectively according to the text fragments intercepted from the text sequence by the text pairs of each row of the square matrix, wherein if the text fragments intercepted from the text sequence by the text pairs of each row can form a vocabulary, initialize the positions corresponding to the text fragments in that row to a first preset value, and initialize the positions in that row that do not correspond to the text fragments to a second preset value; Take all the initialized square matrices as the word matrix of the text sequence; Convert the word matrix into a prior vector, wherein the prior vector has the same vector dimension as the word vector, position vector and segment vector of the text sequence; Generate the word vector, position vector and segment vector of the text sequence according to the convention of the Bert layer for input vectors; Input the word vector, position vector, segment vector and the prior vector of the text sequence into the Bert layer to obtain the text vector of the text sequence.

2. The method for generating text vectors according to claim 1, characterized in that, the text vector model further includes a feature extraction layer, and the converting the word matrix into the prior vector includes: Input the word matrix into the feature extraction layer to obtain the prior vector.

3. The method for generating text vectors according to claim 1, characterized in that, the inputting the word vector, position vector, segment vector and the prior vector of the text sequence into the Bert layer to obtain the text vector of the text sequence includes: Fuse the word vector, position vector, segment vector and the prior vector of the text sequence to obtain a fused vector; Input the fused vector into the Bert layer to obtain the text vector of the text sequence.

4. The method for generating text vectors according to claim 3, characterized in that, the fusing the word vector, position vector, segment vector and the prior vector of the text sequence to obtain a fused vector includes: Fuse the word vector, position vector, segment vector and the prior vector of the text sequence in a summing manner to obtain the fused vector.

5. A text vector generation device, characterized in that, applied to a text processing device, the text processing device is configured with a text vector model, the text vector model includes a Bert layer, and the text vector generation device includes: A prior information module, configured to obtain a text sequence; and construct a plurality of square matrices to be initialized according to the text sequence, wherein the rows and columns of each square matrix respectively correspond to the text sequence one by one; Select different texts from the text sequence for each of the plurality of square matrices as the starting text of the square matrix, wherein the starting text of each square matrix and the text corresponding to each row of the square matrix form a plurality of text pairs; For each square matrix, initialize each row of the square matrix respectively according to the text segment intercepted from the text sequence by the text pair of each row of the square matrix. If the text segments intercepted from the text sequence by the text pairs of each row can form a vocabulary, initialize the positions corresponding to the text segments in that row to a first preset value, and initialize the positions in that row not corresponding to the text segments to a second preset value; Use all the initialized square matrices as the word matrix of the text sequence; Convert the word matrix into a prior vector, wherein the prior vector has the same vector dimension as the word vector, position vector and segment vector of the text sequence; A vector generation module, configured to generate the word vector, position vector and segment vector of the text sequence according to the convention of the Bert layer on the input vector; The vector generation module is further configured to input the word vector, position vector, segment vector and the prior vector of the text sequence into the Bert layer to obtain the text vector of the text sequence.

6. A text processing device, characterized in that, the text processing device includes a processor and a memory, the memory stores a computer program, and when the computer program is executed by the processor, the text vector generation method according to any one of claims 1-4 is implemented.

7. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the text vector generation method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Chinese grammar debugging method and system based on multivariate text features

    CN112183094A