Video script classification method based on word meaning and part-of-speech base model

Through the method based on the word-sentence part-of-speech base model, the multi-faceted features of the video script text are extracted and encoded, which solves the problem of limited classification accuracy in the prior art, and realizes the precise classification and recommendation possibility evaluation of video scripts.

CN120086377APending Publication Date: 2025-06-03WUHAN CHIDARUI ADVERTISING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076421.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate the multi-faceted high-dimensional features of video script text, resulting in limited classification accuracy and difficulty in accurately recommending the recommendation possibility of video scripts.

Method used

The method based on the word-sentence part-of-speech base model is adopted to extract the context semantic features of the text through the ERNIE model, and the StructBERT model extracts part-of-speech vectors and part-of-speech statistical features, and encodes these features using an autoencoder. Finally, a full connection layer model based on MLP is built for training to realize the precise classification of video scripts.

Benefits of technology

Through the combination of multi-feature encoder and neural network classification model, text-level embedding vectors and part-of-speech vectors can be automatically encoded, and efficient low-dimensional representations can be extracted, so as to achieve accurate classification and recommendation possibility evaluation of video scripts, which are suitable for multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086377A_ABST
    Figure CN120086377A_ABST
Patent Text Reader

Abstract

The invention provides a video script classification method based on a word meaning and part-of-speech base model, and relates to the field of short videos and artificial intelligence, and the method comprises the steps: extracting context semantic features of an input text of an audio and video script through an ERNIE model, and obtaining a text-level embedded vector; extracting part-of-speech vectors and part-of-speech statistical features by using a StrauctBERT model; encoding the text-level embedded vector, the part-of-speech vector and the part-of-speech statistical feature through a preset number of auto-encoders to obtain a feature tensor; training the constructed full connection layer model through the feature tensor to obtain an optimal video script classification evaluation model; obtaining a feature tensor of an audio and video script text to be classified; and inputting the to-be-classified feature tensor into the optimal video script classification evaluation model, and outputting a text classification and evaluation result of the to-be-classified audio and video script text. The input text of the video script can be effectively classified and evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of short videos and artificial intelligence, and particularly to a method for classifying video scripts based on a word sense and part-of-speech base model. Background Art

[0002] Text classification and evaluation based on a video word sense and part-of-speech base model are of great significance in multiple fields, especially in content recommendation, user personalized experience, and data analysis. By classifying and evaluating text, it can help authors quickly determine the applicable fields of the text, enable a more accurate understanding and marking of the content to be presented in the video, and even indicate the direction of modification. It also avoids overly single or redundant information in the text, can better match the target users, improve the user experience, increase content diversity, and enhance the platform's recommendation probability, which is very important for improving the dissemination and recommendation degree of the video. The classification and evaluation of video scripts are problems of feature learning and classification of high-dimensional data. In the fields of big data, short videos, and artificial intelligence, feature learning and classification of high-dimensional data are an important research direction. Existing technologies usually use a single model to encode features, making it difficult to fully utilize the multi-modal features of the data or improve the classification accuracy. In addition, traditional classification models are difficult to effectively fuse various features in high-dimensional complex data, resulting in limited classification performance. Feature extraction large models such as the "ERNIE" series are natural language processing models based on deep learning, widely used in feature extraction and generation tasks. Multi-Feature Encoder technology (Multi-Feature Encoder) is a deep learning technology that can extract information from multiple different feature sources and fuse them. Its goal is to capture complex patterns and information by effectively encoding different types of features (such as text, images, time series, audio, etc.). Neural Network Classification Models perform classification tasks by simulating the structure and function of the human nervous system.

[0003] However, for video script text, before the video is completed, how to fuse multi-faceted high-dimensional features and comprehensively describe the recommendation probability of the script is a difficult problem. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for classifying video scripts based on a word sense and part-of-speech base model to solve the problem that the prior art cannot accurately classify based on the high-dimensional features of video script text.

[0005] The above object of this application is achieved through the following technical solutions: S1: Obtain the script text of the audio - video and perform pre - processing to obtain a script dataset; S2: Through the ERNIE model, extract the context semantic features of the input text of the script dataset to obtain text - level embedding vectors; S3: Use the StructBERT model to extract the part - of - speech vectors and part - of - speech statistical features of individual texts in the script dataset; S4: Through a preset number of auto - encoders, encode the text - level embedding vectors, part - of - speech vectors, and part - of - speech statistical features to obtain feature tensors; S5: Construct a fully - connected layer model based on MLP; through the feature tensors, train the fully - connected layer model to obtain an optimal video script classification evaluation model; S6: Obtain the audio - video script text to be classified, repeat the operations in steps S1 to S4 to obtain the feature tensors to be classified; input the feature tensors to be classified into the optimal video script classification evaluation model, and output the text classification and evaluation results of the audio - video script text to be classified.

[0006] Optionally, step S1 includes: Pre - processing includes: cleaning, deduplication, stop - word removal, word segmentation, and word embedding.

[0007] Optionally, step S2 includes: Through the tokenizer, tokenize the input text of the script dataset and convert it into a tensor form; Input the input text in tensor form into the ERNIE model, extract the context embedding vectors of fixed length of the input text, obtain high - dimensional vectors and transfer them to the GPU to get the context - based semantic representation of the input text, and output text - level embedding vectors.

[0008] Optionally, step S3 includes: Use 37 - dimensional vectors to represent the part - of - speech vectors and part - of - speech statistical features.

[0009] Optionally, step S4 includes: The input data includes: text - level embedding vectors, part - of - speech vectors, and part - of - speech statistical features; S41: Through 4 auto - encoders, learn the low - dimensional representation in the input data; through the low - dimensional representation, reconstruct the input data; S42: Through the mean - square error (MSE), calculate the reconstruction loss of the input data; through the reconstruction loss, adjust the model parameters of the 4 auto - encoders; S43: Perform a feature splicing operation on the reconstructed input data to obtain feature tensors; Feature tensors , are represented as follows:

[0010] Among them, is the feature expression of single data, that is, the feature expression of the account source; is the basic characteristic of the source; represents the result obtained by struct_encoded; represents the result obtained by ernie_encoded.

[0011] Optionally, step S41 includes: Construct a first autoencoder network structure through the autoencoder t_e_encoder and the autoencoder c_e_encoder; Construct a second autoencoder network structure through the autoencoder t_s_encoder and the autoencoder c_s_encoder; Learn the low-dimensional representation in the text-level embedding vector through the first autoencoder network structure; reconstruct the text-level embedding vector through the low-dimensional representation in the text-level embedding vector; Learn the low-dimensional representation in the part-of-speech vector and the part-of-speech statistical features through the second autoencoder network structure; reconstruct the part-of-speech vector and the part-of-speech statistical features through the low-dimensional representation in the part-of-speech vector and the part-of-speech statistical features.

[0012] Optionally, step S41 further includes: The autoencoder includes: an encoder and a decoder; the encoder is connected to the decoder; The encoder of the first autoencoder network structure includes: three fully connected layers, a ReLU activation function, and a LayerNorm unit; The decoder of the first autoencoder network structure includes: three fully connected layers and a ReLU activation function; The encoder of the second autoencoder network structure includes: three fully connected layers, a ReLU activation function, and a LayerNorm unit; The decoder of the second autoencoder network structure includes: three fully connected layers and a ReLU activation function.

[0013] An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes a video script classification method based on a word-meaning and part-of-speech base model.

[0014] A computer-readable storage medium stores instructions. When the instructions are executed, a video script classification method based on a word-meaning and part-of-speech base model is executed.

[0015] The beneficial effects brought by the technical solution provided by this application are as follows: By using a multi-feature encoder and a neural network classification model, it is possible to automatically encode the text-level embedding vectors, part-of-speech vectors, and part-of-speech statistical features in the input data, transform the script text and source information into a comprehensive feature vector description that integrates basic features, semantic, and structural features, extract its efficient low-dimensional representation, and achieve accurate classification through the classification model, which is applicable to multiple fields such as finance, medical care, and social network analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The present application will be further described below in conjunction with the drawings and embodiments. In the drawings: Figure 1 is a schematic diagram of the algorithm in the embodiment of this application; Figure 2 is a structural diagram of the autoencoder in the embodiment of this application; Figure 3 is a schematic diagram of the structure of the electronic device in the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In order to have a clearer understanding of the technical features, objectives, and effects of this application, the specific embodiments of this application will now be described in detail with reference to the drawings.

[0018] The embodiment of this application provides a video script classification method based on a word meaning and part-of-speech base model.

[0019] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the algorithm of a video script classification method based on a word meaning and part-of-speech base model in the embodiment of this application, including: S1: Obtain the script text of the audio-visual and perform preprocessing to obtain a script data set; As an embodiment, the script text data includes basic features of the account source such as the number of likes and the number of fans, as well as the theme, script content, etc.

[0020] S2: Through the ERNIE model, extract the context semantic features of the input text of the script data set to obtain text-level embedding vectors; S3: Use the StructBERT model to extract the part-of-speech vectors and part-of-speech statistical features of a single text in the script data set; S4: Through a preset number of autoencoders, encode the text-level embedding vectors, part-of-speech vectors, and part-of-speech statistical features to obtain a feature tensor; S5: Construct a fully connected layer model based on MLP; through the feature tensor, train the fully connected layer model to obtain an optimal video script classification evaluation model; As an example, the unified text feature dataset (feature tensor) is input into the fully connected layer, that is, classification calculation is performed by a classifier based on a multi-layer fully connected neural network (MLP), and text classification and evaluation are performed according to the text features. The unified feature vector is sent into the fully connected layer model based on the neural network to realize multi-class decomposition of the spliced feature vector, combined with the evaluation index of the text, so as to realize the accurate classification of the script text through the output of the fully connected layer.

[0021] As an example, a multi-layer fully connected neural network (MLP) is used here as a classifier, and its structure includes: Input layer: Receives the spliced 82-dimensional feature vector. Hidden layer: Contains two fully connected layers, uses the ReLU activation function and the Dropout layer to reduce overfitting. Output layer: Outputs multi-class classification results, and uses the Softmax activation function to calculate the classification probability. The classification model performs forward propagation, the input features pass through the classification network, and finally the classification probability is output.

[0022] As an example, during the training process of the fully connected layer model based on MLP (Multilayer Perceptron), NLLLoss (negative log-likelihood loss) is used as the loss function, which is suitable for multi-class classification tasks. The optimizer uses AdamW, and weight decay is introduced to improve the generalization ability of the model. The evaluation method is to calculate the number of matches between the prediction result and the true label through the count_identical_rows function for evaluation.

[0023] S6: Obtain the audio-visual script text to be classified, repeat the operations in steps S1 to S4 to obtain the feature tensor to be classified; input the feature tensor to be classified into the optimal video script classification evaluation model, and output the text classification and evaluation results of the audio-visual script text to be classified.

[0024] As an example, the formed feature tensor after connection enters a simple fully connected layer to map the low-dimensional feature vector to the predefined class label for text classification and evaluation. And use a fully connected layer model (such as BERT, etc.) to complete the parameter selection of the corresponding machine learning model for the input text; Verify whether the trained corresponding machine learning model meets the selected index threshold on the validation set of the feature tensor, determine the optimal video script classification evaluation model for multi-factor mapping, and then use the optimal video script classification evaluation model to predict the evaluation index value of the newly input text.

[0025] Step S1 includes: Preprocessing includes: cleaning, deduplication, stop word removal, word segmentation, and word embedding.

[0026] Step S2 includes: The input text of the script dataset is tokenized by a tokenizer and converted into a tensor form; The input text in tensor form is input into the ERNIE model to extract the context embedding vectors of a fixed length of the input text, obtain high-dimensional vectors and transfer them to the GPU, obtain the semantic representation of the input text based on the context, and output the text-level embedding vectors.

[0027] As an embodiment, the text-level embedding vectors are generally n×768 (a matrix of N words × d-dimensional vectors), and the text is converted into a tensor form that the model can accept (ensuring that the length does not exceed 512).

[0028] Step S3 includes: A 37-dimensional vector is used to represent the part-of-speech vector and part-of-speech statistical features.

[0029] As an embodiment, the present invention selects a 37-dimensional vector to represent features. The part-of-speech vector and part-of-speech statistical features include 37 types of parts of speech. Based on the Structured Pre-training Model (StructBERT), a part-of-speech type dictionary is designed to map different token types in the text to unique indices, obtain the type of each word or token, a total of 37-dimensional features are selected, and a statistical feature vector expression of these types is generated. Furthermore, the preprocessed text data is used with this model to extract and generate the part-of-speech vector and part-of-speech statistical features of each text, that is, a vector with the length of the number of token types.

[0030] Step S4 includes: The input data includes: text-level embedding vectors, and part-of-speech vectors and part-of-speech statistical features; S41: Through 4 autoencoders, learn the low-dimensional representation in the input data; through the low-dimensional representation, reconstruct the input data; S42: Calculate the reconstruction loss of the input data through the mean squared error MSE; adjust the model parameters of the 4 autoencoders through the reconstruction loss; S43: Perform a feature splicing operation on the reconstructed input data to obtain a feature tensor; Feature tensor , is expressed as follows:

[0031] Among them, is the feature expression of a single data, that is, the account source feature expression; is the basic characteristic of the source; represents the result obtained by struct_encoded; represents the result obtained by ernie_encoded.

[0032] As an example, through four autoencoders (such as Figure 2 shown), during the feature learning process, useful low-dimensional representations can be automatically learned from the input data, effectively compressing the original features. During the data reconstruction process, the model reconstructs the original input based on the learned low-dimensional representations. The quality of the reconstruction is evaluated using the mean squared error (MSE). After completing the feature extraction, feature transformation, and feature concatenation operations, all the features are concatenated into a large tensor, and these feature tensors are passed to a neural network model for classification and evaluation, finally outputting an evaluation result based on the probability distribution of the recommended features.

[0033] Step S41 includes: Construct a first autoencoder network structure through the autoencoder t_e_encoder and the autoencoder c_e_encoder; Construct a second autoencoder network structure through the autoencoder t_s_encoder and the autoencoder c_s_encoder; Through the first autoencoder network structure, learn the low-dimensional representation in the text-level embedding vector; through the low-dimensional representation in the text-level embedding vector, reconstruct the text-level embedding vector; Through the second autoencoder network structure, learn the low-dimensional representation in the part-of-speech vector and the part-of-speech statistical features; through the low-dimensional representation in the part-of-speech vector and the part-of-speech statistical features, reconstruct the part-of-speech vector and the part-of-speech statistical features.

[0034] As an example, the different four autoencoders can respectively achieve the extraction of high- and low-dimensional semantic features of the text (text-level embedding vector) and the extraction of part-of-speech features of the text (part-of-speech vector and part-of-speech statistical features).

[0035] Step S41 further includes: The autoencoder includes: an encoder and a decoder; the encoder is connected to the decoder; The encoder of the first autoencoder network structure includes: three fully connected layers, a ReLU activation function, and a LayerNorm unit; The decoder of the first autoencoder network structure includes: three fully connected layers and a ReLU activation function; As an embodiment, the first autoencoder network structure includes two parts: an encoder and a decoder. The encoder (Encoder) compresses the input 768-dimensional matrix data into 20-dimensional and 50-dimensional representations. This process is completed through three fully-connected layers (Linear) and an activation function (ReLU). Finally, normalization is performed through LayerNorm. The decoder (Decoder): Through three fully-connected layers and the ReLU activation function, the 20-dimensional and 50-dimensional representations are reconstructed back into a 768-dimensional output. By creating an Autoencoder model, the model parameters are continuously adjusted through a loss function and an optimizer.

[0036] The encoder of the second autoencoder network structure includes: three fully-connected layers, a ReLU activation function, and a LayerNorm unit; The decoder of the second autoencoder network structure includes: three fully-connected layers and a ReLU activation function.

[0037] As an embodiment, the encoder of the second autoencoder network structure includes two parts: an encoder and a decoder. The encoder (Encoder) compresses the input 37-class common part-of-speech feature-dimensional data into 4-dimensional and 7-dimensional representations (the 4-dimensional and 7-dimensional features of the title and the text), achieving structured feature dimensionality reduction. This process is completed through three fully-connected layers (Linear) and an activation function (ReLU). Finally, normalization is performed through LayerNorm. The decoder (Decoder): Through three fully-connected layers and the ReLU activation function, the 4-dimensional and 7-dimensional representations are reconstructed back into an output of 37 classes of part-of-speech and statistical features. By creating an Autoencoder model, the model parameters are continuously adjusted through a loss function and an optimizer.

[0038] This application also discloses an electronic device. Referring to Figure 3 , Figure 3 is a schematic structural diagram of an electronic device disclosed in an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.

[0039] Among them, the communication bus 502 is used to realize the connection and communication between these components.

[0040] Among them, the user interface 503 may include a display screen. Optionally, the user interface 503 may further include a standard wired interface and a wireless interface.

[0041] Among them, the network interface 504 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0042] The present application also discloses a computer-readable storage medium storing a plurality of instructions adapted to be loaded by a processor to execute the above-described method for classifying video scripts based on a word-sense and part-of-speech base model.

[0043] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure.

[0044] The present application aims to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not described in the present disclosure. The description and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A video script classification method based on a word meaning part of speech base model, characterized in that: The method comprises the following steps: S1: Obtain the script text of the audio and video and preprocess it to obtain the script dataset; S2: Through the ERNIE model, the contextual semantic features of the input text of the script dataset are extracted to obtain the text-level embedding vector; S3: Use the StructBERT model to extract the part-of-speech vector and part-of-speech statistical features of a single text in the script dataset; S4: Encode the text-level embedding vector, part-of-speech vector, and part-of-speech statistical features through a preset number of autoencoders to obtain a feature tensor; S5: Build a fully connected layer model based on MLP; train the fully connected layer model through feature tensors to obtain the optimal video script classification and evaluation model; S6: Obtain the audio and video script text to be classified, repeat the operations from step S1 to step S4, and obtain the feature tensor to be classified; input the feature tensor to be classified into the optimal video script classification and evaluation model, and output the text classification and evaluation results of the audio and video script text to be classified.

2. A video script classification method based on a word meaning part of speech base model as claimed in claim 1, characterized in that: Step S1 includes: Preprocessing includes: cleaning, deduplication, removal of stop words and word segmentation, and word embedding.

3. A video script classification method based on a word meaning part of speech base model as claimed in claim 1, characterized in that: Step S2 includes: The input text of the script dataset is tokenized through the tokenizer and converted into a tensor form; The input text in tensor form is input into the ERNIE model, the fixed-length context embedding vector of the input text is extracted, the high-dimensional vector is obtained and transferred to the GPU, the context-based semantic representation of the input text is obtained, and the text-level embedding vector is output.

4. A video script classification method based on a word meaning part of speech base model as claimed in claim 1, characterized in that: Step S3 includes: A 37-dimensional vector is used to represent the part-of-speech vector and part-of-speech statistical features.

5. A video script classification method based on a word meaning part of speech base model as claimed in claim 1, characterized in that: Step S4 includes: The input data includes: text-level embedding vectors, part-of-speech vectors, and part-of-speech statistical features; S41: Learn low-dimensional representations of input data through 4 autoencoders; reconstruct input data through low-dimensional representations; S42: Calculate the reconstruction loss of the input data through the mean square error MSE; adjust the model parameters of the four autoencoders through the reconstruction loss; S43: performing feature concatenation operation on the reconstructed input data to obtain a feature tensor; Feature Tensor , which is expressed as follows: in, It is the feature expression of a single data, that is, the feature expression of the account source; Basic characteristics of the source; Indicates the result obtained by struct_encoded; Indicates the result obtained by ernie_encoded.

6. A video script classification method based on a word meaning part of speech base model as claimed in claim 5, characterized in that: Step S41 includes: Through the autoencoder t_e_encoder and the autoencoder c_e_encoder, a first autoencoder network structure is constructed; Through the autoencoder t_s_encoder and the autoencoder c_s_encoder, a second autoencoder network structure is constructed; Through the first autoencoder network structure, a low-dimensional representation in the text-level embedding vector is learned; through the low-dimensional representation in the text-level embedding vector, the text-level embedding vector is reconstructed; Through the second autoencoder network structure, low-dimensional representations in part-of-speech vectors and part-of-speech statistical features are learned; through the low-dimensional representations in part-of-speech vectors and part-of-speech statistical features, part-of-speech vectors and part-of-speech statistical features are reconstructed.

7. A video script classification method based on a word meaning part of speech base model as claimed in claim 6, characterized in that: Step S41 also includes: The autoencoder includes: an encoder and a decoder; the encoder is connected to the decoder; The encoder of the first autoencoder network structure includes: three fully connected layers, ReLU activation function and LayerNorm unit; The decoder of the first autoencoder network structure includes: three fully connected layers and ReLU activation function; The encoder of the second autoencoder network structure includes: three fully connected layers, ReLU activation function and LayerNorm unit; The decoder of the second autoencoder network structure includes: three fully connected layers and a ReLU activation function.

8. An electronic device, characterized in that: The electronic device comprises a processor (501), a memory (505), a user interface (503) and a network interface (504), wherein the memory (505) is used to store instructions, the user interface (503) and the network interface (504) are used to communicate with other devices, and the processor (501) is used to execute the instructions stored in the memory (505) so that the electronic device executes the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed by a computer, the method according to any one of claims 1 to 7 is executed.