A data processing method and device for video annotation
By constructing multi-dimensional, time-series labeling information, and using the target text processing model to analyze and integrate video comments and emotional information, the problem of singularity and subjectivity of video emotional labeling in the existing technology is solved, and more accurate and efficient emotional labeling is achieved, providing richer data support for the field of emotional computing.
Patent Information
- Application Number
- CN202410918448.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-10
AI Technical Summary
The prior art has problems such as single labeling information, fewer labeling personnel, time-consuming and laborious and subjective in the emotional labeling of video-induced materials, which makes it difficult to accurately represent changes in emotional intensity, limiting the development of the field of emotional research.
By obtaining video comment information and video emotional information, the target text processing model is used to perform feature analysis and feature fusion, and multi-dimensional and time-series annotation information is constructed to improve the accuracy and efficiency of emotional type annotation of video materials.
It realizes more accurate and efficient labeling of the emotions types of video materials, provides richer and more accurate emotional data support, and opens up a new perspective for research in the field of emotional computing.
Smart Images

Figure CN118887581B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data processing method and device for video annotation. Background Art
[0002] In the field of emotional computing, the emotion labeling of induced materials is of great significance. Accurate labeling of emotional materials helps to build a benchmark data set for emotional computing and provide a unified and reliable standard for the performance evaluation and comparison of algorithms. At present, there are two main methods for the emotional labeling of video induced materials: one is to invite labelers to watch the video and use professional questionnaires in the field of psychology to subjectively evaluate the emotions of the video; the other is to label and classify emotions by analyzing visual elements such as color and brightness in the video, or by text parsing information such as voice and sentences. However, the above-mentioned emotion labeling methods face the disadvantages of single labeling information, few labelers, time-consuming and laborious labeling, and high subjectivity. These problems make it difficult to characterize the changes in emotion intensity, thereby limiting the development of the field of emotion research. Therefore, a data processing method and device for video labeling are provided to improve the accuracy and labeling efficiency of characterizing the emotion type of video materials by constructing multi-dimensional and temporal labeling information, open up new perspectives for research in the field of emotional computing, and provide more abundant and accurate emotional data support. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a data processing method and device for video annotation, which is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, open up new perspectives for research in the field of emotional computing, and provide richer and more accurate emotion data support.
[0004] In order to solve the above technical problems, a first aspect of an embodiment of the present invention discloses a data processing method for video annotation, the method comprising:
[0005] Acquire video information to be processed; the video information to be processed includes video comment information and video emotion information;
[0006] Using a target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain target comment sentiment information;
[0007] Feature fusion processing is performed on the target comment emotion information and the video emotion information in the to-be-processed video information to obtain target video annotation information.
[0008] A second aspect of an embodiment of the present invention discloses a data processing device for video annotation, the device comprising:
[0009] An acquisition module is used to acquire video information to be processed; the video information to be processed includes video comment information and video emotion information;
[0010] A first processing module is used to perform feature analysis on the video comment information in the video information to be processed using a target text processing model to obtain target comment sentiment information;
[0011] The second processing module is used to perform feature fusion processing on the target comment emotion information and the video emotion information in the to-be-processed video information to obtain target video annotation information.
[0012] A third aspect of the present invention discloses another data processing device for video annotation, the device comprising:
[0013] A memory storing executable program code;
[0014] a processor coupled to the memory;
[0015] The processor calls the executable program code stored in the memory to execute part or all of the steps in the data processing method for video annotation disclosed in the first aspect of the embodiment of the present invention.
[0016] The fourth aspect of the present invention discloses a computer-readable storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in the data processing method for video annotation disclosed in the first aspect of an embodiment of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 is a schematic diagram of a scenario of a data processing system for video annotation provided by an embodiment of the present invention;
[0019] Figure 2 It is a flowchart of a data processing method for video annotation disclosed in an embodiment of the present invention;
[0020] Figure 3 is a structural schematic diagram of a data processing device for video annotation disclosed in an embodiment of the present invention;
[0021] Figure 4 is a structural schematic diagram of another data processing device for video annotation disclosed in an embodiment of the present invention;
[0022] Figure 5 It is a structural schematic diagram of a target processing text model disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or equipment.
[0025] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0026] In this application, the word "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described in this application as "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments. The following description is given to enable any technician in the field to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present application.
[0027] It should be noted that since the method of the embodiment of the present application is executed in a computer device, the processing objects of each computer device exist in the form of data or information. For example, time is actually time information. It can be understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data for processing by the computer device. The details will not be repeated here.
[0028] The embodiments of the present application provide a data processing method, apparatus, computer device, and computer-readable storage medium for video annotation, which are described in detail below.
[0029] See also Figure 1 , Figure 1 Schematic diagram of a scenario of a data processing system for video annotation provided in an embodiment of the present application. The data processing system for video annotation may include a computer device 100, in which a data processing device for video annotation is integrated, such as Figure 1 Computer equipment in.
[0030] In the embodiment of the present application, the computer device 100 is mainly used to obtain video information to be processed; the video information to be processed includes video comment information and video emotion information;
[0031] The target text processing model is used to perform feature analysis on the video comment information in the video information to be processed, and the target comment sentiment information is obtained;
[0032] The target comment sentiment information and the video sentiment information in the video information to be processed are subjected to feature fusion processing to obtain the target video annotation information.
[0033] It can improve the accuracy and annotation efficiency of characterizing the emotional types of video materials by constructing multi-dimensional and temporal annotation information, open up new perspectives for research in the field of emotional computing, and provide richer and more accurate emotional data support.
[0034] In the embodiment of the present application, the computer device 100 may be an independent server, or a server network or server cluster composed of servers. For example, the computer device 100 described in the embodiment of the present application includes but is not limited to a computer, a network host, a single network server, a plurality of network server sets or a cloud server composed of a plurality of servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.
[0035] It is understandable that the computer device 100 used in the embodiments of the present application may be a device including both receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. Such a device may include: a cellular or other communication device having a single-line display or a multi-line display or a cellular or other communication device without a multi-line display. The specific computer device 100 may be a desktop terminal or a mobile terminal, and the computer device 100 may also be one of a mobile phone, a tablet computer, a laptop computer, etc.
[0036] Those skilled in the art will understand that Figure 1 The application environment shown in the figure is only one application scenario of the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1 More or less computer equipment as shown in Figure 1 Only one computer device is shown in the figure. It can be understood that the data processing system for video annotation can also include one or more other services, which are not specifically limited here.
[0037] In addition, if Figure 1 As shown, the data processing system for video annotation may also include a memory 200 for storing data, such as image data, location information, etc.
[0038] It should be noted that Figure 1 The scenario diagram of the data processing system for video annotation shown is merely an example. The data processing system and scenario for video annotation described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided in the embodiment of the present application. A person of ordinary skill in the art can appreciate that with the evolution of the data processing system for video annotation and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is equally applicable to similar technical problems.
[0039] The present invention discloses a data processing method and device for video annotation, which is beneficial to improve the accuracy and annotation efficiency of the emotion type characterization of video materials by constructing multi-dimensional and temporal annotation information, opening up a new perspective for the research in the field of emotional computing and providing more abundant and accurate emotion data support. The following are detailed descriptions.
[0040] Embodiment 1
[0041] See also Figure 2 , Figure 2 is a flow chart of a data processing method for video annotation disclosed in an embodiment of the present invention. Figure 2The data processing method for video annotation described is applied to a management system, such as a local server or a cloud server for management, and the embodiments of the present invention do not limit this. Figure 2 As shown, the data processing method for video annotation may include the following operations:
[0042] 101. Obtain video information to be processed.
[0043] The video information to be processed includes video comment information and video emotion information.
[0044] 102. Use the target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain the target comment sentiment information.
[0045] 103. Perform feature fusion processing on the target comment sentiment information and the video sentiment information in the to-be-processed video information to obtain the target video annotation information.
[0046] It should be noted that the above-mentioned target text processing model is a model that conducts in-depth learning of a large amount of text data in a pre-training stage to grasp the deep semantics and grammatical structure of the language. It can simultaneously consider the contextual information of the word and generate a context-related representation of each word, thereby capturing the keywords in the text or sentences on the basis of understanding the complex meanings of the keywords in different contexts, and the embodiments of the present invention are not limited to this.
[0047] It should be noted that the "sequential distribution" referred to in this application means the distribution and arrangement according to the time sequence of video playback, that is, the sequence of each data information is the time sequence corresponding to the video frames corresponding to the data information, and the embodiment of the present invention is not limited to this.
[0048] It should be noted that the above-mentioned video comment information represents the barrage data containing comment text and timestamp information, which is not limited in the embodiment of the present invention. Furthermore, when the original barrage data containing comment text and timestamp information is obtained, the non-text content such as HTML tags, emoticons, etc. in the comments and barrage should be removed first to ensure the purity of the data. Subsequently, deduplication processing is performed to eliminate the impact of duplicate data on subsequent analysis. At the same time, a reasonable filling strategy or elimination is performed for missing values to ensure the integrity of the data. For the barrage data, its time information is extracted for time series analysis. Finally, the two types of data are merged into video comment information, which is not limited in the embodiment of the present invention.
[0049] It should be noted that the above video emotional information represents facial expression features, which is not limited in the embodiment of the present invention. Further, the above video emotional information is obtained in the following manner: using short video materials, designing an emotion induction experiment, and inviting subjects to participate in the experiment, by inducing different emotional states of the subjects, and using a high frame rate camera to record the facial movements of the subjects, analyzing the temporal changes in the facial expressions of the subjects, and outputting the emotional intensity vector v2 representing the temporal changes in the emotional intensity, that is, obtaining the video emotional information, which is not limited in the embodiment of the present invention. The above analysis of the temporal changes in the facial expressions of the subjects and outputting the emotional intensity vector v2 representing the temporal changes in the emotional intensity is to identify the collected facial movements of the subjects through an open source facial recognition and expression analysis library (such as OpenFace), and extract Action Units (AUs). At this time, OpenFace will generate a CSV file containing timestamps and AUs data. Then, a preset machine learning model or neural network model (such as RNN, etc.) is used to read the CSV file, extract expression features, and obtain video emotional information, which is not limited in the embodiment of the present invention.
[0050] It should be noted that the above-mentioned feature fusion processing of the target comment emotion information and the video emotion information in the to-be-processed video information is to first generate a time-domain-varying emotion intensity vector v1 for the target comment type emotion in the target comment emotion information in combination with the timestamp (video frame time), and then weightedly fuse the emotion intensity vector v1 and the emotion intensity vector v2 (video emotion information) to output the target video annotation information. Further, v1 and v2 are fused and the probability score of each time window is output, and the formula is:
[0051] s emo (t i )=σ(N[v1(t i )])+N[v2(t i )])
[0052] Among them, s emo (t i ) indicates t i The emotional intensity of the time period (video frame time); σ is the sigmoid function; N is the max-min normalization operation; v1(t i ) and v2(t i ) are respectively at t i The vector value of the time period (the vector values are arranged by time period, and the time period is the order of the elements corresponding to the vector value). Then all s emo (t i ) The emotion intensity is mapped back to the time window, and the label of the temporal emotion intensity is output, that is, the target video annotation information is obtained, which is not limited in the embodiment of the present invention.
[0053] It should be noted that the data processing method for video annotation in this application analyzes the temporal changes of the emotional semantics of bullet comments, and combines the temporal changes of the facial expressions of the subjects in the short video emotion induction experiment, and weightedly fuses the two modalities to achieve adaptive annotation of multi-label information such as the time distribution of the emotion type and emotion intensity induced by the short video. This automated annotation method can further save the time and cost of manual annotation, and at the same time provide an important research basis for video-based emotion analysis.
[0054] It can be seen that implementing the data processing method for video annotation described in the embodiment of the present invention is beneficial to improving the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0055] In an optional embodiment, if Figure 5 As shown, the target text processing model includes a first feature processing module 401, a second feature processing module 402, a third feature processing module 403, a fourth feature processing module 404, a fifth feature processing module 405, a sixth feature processing module 406, a seventh feature processing module 407, an eighth feature processing module 408, a linear module 409 and an activation module 410; wherein,
[0056] The input end of the first feature processing module is configured as the model input of the target text processing model, the output end of the first feature processing module is connected to the input end of the second feature processing module; the output end of the second feature processing module is connected to the input end of the third feature processing module; the output end of the third feature processing module is connected to the input end of the fourth feature processing module; the output end of the fourth feature processing module is connected to the input end of the fifth feature processing module; the output end of the fifth feature processing module is connected to the input end of the sixth feature processing module; the output end of the sixth feature processing module is connected to the input end of the seventh feature processing module; the input end of the seventh feature processing module is connected to the input end of the eighth feature processing module; the output end of the eighth feature processing module is connected to the input end of the linear module; the output end of the linear module is connected to the input end of the activation module; the output end of the activation module is configured as the model output of the target text processing model.
[0057] It should be noted that the model structures of the above-mentioned first feature processing module, second feature processing module, third feature processing module, fourth feature processing module, fifth feature processing module, sixth feature processing module, seventh feature processing module, and eighth feature processing module are consistent, and the embodiments of the present invention do not limit this. Furthermore, the semantic feature information of words and sentences can be deeply extracted by processing the above-mentioned first feature processing module, second feature processing module, third feature processing module, fourth feature processing module, fifth feature processing module, sixth feature processing module, seventh feature processing module, and eighth feature processing module in sequence, thereby realizing the deep capture of contextual relationships and identifying the words and sentences in which they play a main role, and the embodiments of the present invention do not limit this.
[0058] It should be noted that the above linear module is constructed based on the fully connected layer, which is composed of a series of weight matrices and bias vectors, and is used to perform linear transformation on the input features, which is not limited in the embodiment of the present invention. Furthermore, the linear module provides powerful feature processing for the target text processing model through simple mathematical operations to cooperate with subsequent activation operations and increase the nonlinear expression ability of the target text processing model, which is not limited in the embodiment of the present invention.
[0059] It should be noted that the above activation module is constructed based on the softmax activation function, which is not limited in the embodiment of the present invention.
[0060] It can be seen that implementing the data processing method for video annotation described in the embodiment of the present invention is beneficial to improving the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0061] In another optional embodiment, the first feature processing module includes a first multi-head attention unit 4011, a second multi-head attention unit 4017, a third multi-head attention unit 40110, a first fusion unit 4012, a second fusion unit 4015, a third fusion unit 4018, a fourth fusion unit 40111, a fifth fusion unit 40114, a first normalization unit 4013, a second normalization unit 4016, a third normalization unit 4019, a fourth normalization unit 40112, a fifth normalization unit 40115, a first network unit 4014 and a second network unit 40113; wherein,
[0062] The input end of the first multi-head attention unit and the input end of the first fusion unit are both configured as the input end of the first feature processing module, the output end of the first multi-head attention unit is connected to the input end of the first fusion unit; the output end of the first fusion unit is connected to the input end of the first normalization unit; the output end of the first normalization unit is respectively connected to the input end of the first network unit and the input end of the second fusion unit; the output end of the first network unit is connected to the input end of the second fusion unit; the output end of the second fusion unit is connected to the input end of the second normalization unit; the output end of the second normalization unit is respectively connected to the input end of the second multi-head attention unit, the input end of the third multi-head attention unit and the input end of the third fusion unit; the second multi-head attention unit The output end of the network is connected to the input end of the third fusion unit; the output end of the third fusion unit is connected to the input end of the third normalization unit; the output end of the third normalization unit is respectively connected to the input end of the third multi-head attention unit and the input end of the fourth fusion unit; the output end of the third multi-head attention unit is connected to the input end of the fourth fusion unit; the output end of the fourth fusion unit is connected to the input end of the fourth normalization unit; the output end of the fourth normalization unit is respectively connected to the input end of the second network unit and the input end of the fifth fusion unit; the output end of the second network unit is connected to the input end of the fifth fusion unit; the output end of the fifth fusion unit is connected to the input end of the fifth normalization unit; the output end of the fifth normalization unit is configured as the output end of the first feature processing module.
[0063] It should be noted that the above-mentioned first fusion unit, second fusion unit, third fusion unit, fourth fusion unit, and fifth fusion unit are constructed based on the add operation, which is not limited in the embodiment of the present invention.
[0064] It should be noted that the above-mentioned first normalization unit, second normalization unit, third normalization unit, fourth normalization unit, and fifth normalization unit are constructed based on the Layer Normalization layer, and the embodiment of the present invention does not limit this.
[0065] It should be noted that the first network unit and the second network unit are constructed based on a feedforward neural network, which is not limited in the embodiment of the present invention. Further, the size of the first network unit and the second network unit is 4×768, which is not limited in the embodiment of the present invention.
[0066] It should be noted that the above-mentioned first multi-head attention unit, second multi-head attention unit, and third multi-head attention unit are constructed based on a multi-head attention mechanism (Multi-Head Attention), which is not limited in the embodiments of the present invention.
[0067] It should be noted that the output of the second normalization unit mentioned above is input to the third multi-head attention unit respectively as query (Query, Q) (i.e., query vector) and key (Key, K) (i.e., key vector). Furthermore, the query vector represents the currently processed input element, which is used to compare with the key vector to determine the attention score. In multi-head attention, the query is copied to each head, but the query vector of each head may be transformed through different linear layers to capture different information. Furthermore, key vectors are similar to query vectors, but they represent other elements in the input sequence. The dot product of the query vector and the key vector is used to calculate the attention scores, which determine the weight of the value vector in the final output.
[0068] It should be noted that the output of the third normalization unit is input to the third multi-head attention unit as value (Value, V) (i.e., value vector). Furthermore, the value vector contains information about the input elements, which will be weighted according to the attention scores calculated by the query and the key. In multi-head attention, the value vector is also transformed through different linear layers to capture information in different representation subspaces.
[0069] It should be noted that the output data from the second normalization unit must first undergo position encoding processing before being input into the second multi-head attention unit, which is not limited in this embodiment of the present invention.
[0070] It can be seen that implementing the data processing method for video annotation described in the embodiment of the present invention is beneficial to improving the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0071] In yet another optional embodiment, the target text processing model is used to perform feature analysis on the video comment information in the video information to be processed to obtain the target comment sentiment information, including:
[0072] The target text processing model is used to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information; the video comment information includes a plurality of comment data information distributed in sequence; the target comment word information includes a plurality of preferred comment word information distributed in sequence; the preferred comment word information includes a plurality of preferred comment words;
[0073] The target comment word information is matched and converted to obtain the target comment sentiment information.
[0074] It should be noted that the above-mentioned matching and conversion processing of the target comment word information is to select keywords related to emotions from the preferred comment words selected in the previous step. Match the candidate keywords with the emotion words in the emotion dictionary (such as DUTIR3 emotion dictionary, etc.) to find the corresponding emotion polarity (type emotion, such as happy, fearful, sad, etc.). According to the matched emotion words, emotion polarity voting is performed, that is, the frequency of occurrence of different types of emotions corresponding to the emotion keywords in the document is counted, and the overall emotion polarity of the document is determined according to the voting results. The emotion polarity output here is the emotion classification result (i.e., the target comment type emotion), thereby obtaining the target comment type emotion in the target comment emotion information, which is not limited in the embodiments of the present invention.
[0075] It can be seen that implementing the data processing method for video annotation described in the embodiment of the present invention is beneficial to improving the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0076] In another optional embodiment, the target text processing model is used to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information, including:
[0077] For any comment data information in the video comment information in the video information to be processed, based on the comment data information, determine the first comment word information corresponding to the comment data information; the first comment word information includes N first comment words;
[0078] Using the target text processing model to perform feature analysis on the first comment word information to obtain comment word feature information; the comment word feature information includes N word feature vector information; each word feature vector information corresponds to a first comment word;
[0079] Using the target text processing model to perform feature analysis on the comment data information to obtain text feature information;
[0080] Based on the comment word feature information and the text feature information, the preferred comment word information corresponding to the comment data information is determined.
[0081] It should be noted that the above-mentioned word feature vector information and text feature information are both represented in the form of vectors, which is not limited in the embodiment of the present invention.
[0082] It should be noted that before using the target text processing model to perform feature analysis on the comment data information and the first comment word information, it is necessary to first perform encoding embedding processing on the comment data information and the first comment word information, that is, to first perform word embedding, segment embedding and position embedding. Furthermore, word embedding is to map words or subwords (such as WordPiece) to a vector space of fixed dimension, segment embedding is to assign a unique identifier (usually a learned vector) to distinguish the roles of different sequences, and position embedding is to map the position index to a vector space so that the target text processing model can understand complex language structures and perform keyword screening, which is not limited in the embodiments of the present invention.
[0083] It can be seen that implementing the data processing method for video annotation described in the embodiment of the present invention is beneficial to improving the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0084] In an optional embodiment, the step of determining the first comment word information corresponding to the comment data information based on the comment data information includes:
[0085] The comment data information is segmented to obtain comment word information; the comment word information includes M video comment words distributed in sequence;
[0086] The comment word information is evaluated and analyzed to obtain first analysis value information; the first analysis value information includes M first analysis values distributed in sequence; each first analysis value corresponds to a video comment word in the same distribution order;
[0087] Performing semantic analysis on the comment word information to obtain second analysis value information; the second analysis value information includes M second analysis values distributed in sequence; each second analysis value corresponds to the video comment word in the same distribution order;
[0088] Performing information entropy analysis and calculation processing on the comment word information to obtain third analysis value information; the third analysis value information includes M third analysis values distributed in sequence; each third analysis value corresponds to the video comment words in the same distribution order;
[0089] Normalizing the first analysis value information, the second analysis value information, and the third analysis value information to obtain first normalized value information, second normalized value information, and third normalized value information;
[0090] Acquire evaluation coefficient information; the evaluation coefficient information includes a first evaluation coefficient, a second evaluation coefficient and a third evaluation coefficient;
[0091] The first evaluation coefficient, the second evaluation coefficient and the third evaluation coefficient in the evaluation coefficient information are used to perform weighted sum processing with the first normalized value information, the second normalized value information and the third normalized value information in sequence to obtain the comment word evaluation value information; the comment word evaluation value information includes M comment word evaluation values distributed in sequence;
[0092] Sorting the evaluation values of the comment words in the evaluation value information of the comment words in descending order to obtain an evaluation value sequence;
[0093] The video comment words corresponding to the top N comment word evaluation values in the evaluation value sequence are taken as the first comment words.
[0094] It should be noted that, based on the comment data information, the first comment word information corresponding to the comment data information is determined by using a Chinese word segmentation tool (such as Jieba), and then keyword mining is performed by combining word frequency, importance, contextual information and neighbor diversity to obtain the first comment word in the video comment information that better reflects its emotional relevance, which is not limited to the embodiments of the present invention.
[0095] It should be noted that the first analysis value information obtained by evaluating and analyzing the comment word information is to evaluate the importance of the video comment word in the video comment information, which can be implemented based on the Term Frequency-Inverse Document Frequency algorithm, and is not limited in the embodiment of the present invention.
[0096] It should be noted that the above-mentioned semantic analysis processing of the comment word information to obtain the second analysis value information is to extract key video comment words by utilizing the co-occurrence information (semantics) between words within the video comment information, which can be implemented by a graph-based sorting algorithm for keyword extraction and document summarization, and the embodiments of the present invention are not limited to this.
[0097] It should be noted that the above-mentioned information entropy analysis and calculation processing of the comment word information obtains the third analysis value information to calculate the information entropy of the left neighbor and the right neighbor of each word, which can be achieved through left and right information entropy (Left and Right Information Entropy), which is not limited in the embodiment of the present invention.
[0098] It should be noted that the above-mentioned normalization processing of the first analysis value information, the second analysis value information and the third analysis value information is to divide the monomer value in each of the above-mentioned first analysis value information, the second analysis value information and the third analysis value information by the sum of all monomer values in each information to achieve separate normalization processing of each information, and the embodiment of the present invention is not limited to this.
[0099] It should be noted that the first evaluation coefficient, the second evaluation coefficient and the third evaluation coefficient may be set by the user or may be a default value given by the system, which is not limited in the embodiment of the present invention. Further, the first evaluation coefficient, the second evaluation coefficient and the third evaluation coefficient are between [0, 1], which is not limited in the embodiment of the present invention.
[0100] It can be seen that implementing the data processing method for video annotation described in the embodiment of the present invention is beneficial to improving the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0101] In another optional embodiment, based on the comment word feature information and the text feature information, determining the preferred comment word information corresponding to the comment data information includes:
[0102] For any word feature vector information in the comment word feature information, the feature matching model is used to calculate and process the word feature vector information and the text feature information to obtain the feature matching value corresponding to the word feature vector information;
[0103] Among them, the feature matching model is:
[0104]
[0105] Among them, TP is the feature matching value; X is the word feature vector information, Y is the text feature information; s1 and s2 are the first matching coefficient and the second matching coefficient respectively;
[0106] Sort all feature matching values in descending order to obtain a feature matching value sequence;
[0107] The first comment word corresponding to the word feature vector information corresponding to the first L feature matching values in the feature matching value sequence is determined as the preferred comment word in the preferred comment word information.
[0108] It should be noted that the first matching coefficient and the second matching coefficient may be set by the user or by a default value given by the system, which is not limited in the embodiment of the present invention. Further, the first matching coefficient and the second matching coefficient are between [0, 1], which is not limited in the embodiment of the present invention.
[0109] It should be noted that, the above-mentioned determination of the preferred comment word information corresponding to the comment data information based on the comment word feature information and the text feature information is to embed the document embedding (text feature information, expressed in the form of a vector) and the keyword candidate embedding (word feature vector information, expressed in the form of a vector) obtained by processing the target text processing model into the vector space, and then calculate the feature matching association between the document embedding (text feature information) and the keyword embedding (word feature vector information), and sort the feature matching values corresponding to the first comment word according to the degree of association of the feature matching association (the order of feature matching values from large to small), and select the top L most relevant first comment words, which is not limited to the embodiments of the present invention.
[0110] It should be noted that the above L is a positive integer not less than 1, and the embodiment of the present invention does not limit this.
[0111] It can be seen that implementing the data processing method for video annotation described in the embodiment of the present invention is beneficial to improving the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0112] Embodiment 2
[0113] See also Figure 3 , Figure 3 is a structural diagram of a data processing device for video annotation disclosed in an embodiment of the present invention. Figure 3 The described device can be applied to a management system, such as a local server or a cloud server for management, etc., which is not limited in the embodiments of the present invention. Figure 3 As shown, the device may include:
[0114] The acquisition module 201 is used to acquire the video information to be processed; the video information to be processed includes video comment information and video emotion information;
[0115] The first processing module 202 is used to perform feature analysis on the video comment information in the to-be-processed video information using the target text processing model to obtain target comment sentiment information;
[0116] The second processing module 203 is used to perform feature fusion processing on the target comment emotion information and the video emotion information in the video information to be processed to obtain the target video annotation information.
[0117] It can be seen that implementation Figure 3The described data processing device for video annotation is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0118] In another optional embodiment, Figure 3 As shown, the target text processing model includes a first feature processing module, a second feature processing module, a third feature processing module, a fourth feature processing module, a fifth feature processing module, a sixth feature processing module, a seventh feature processing module, an eighth feature processing module, a linear module and an activation module; wherein,
[0119] The input end of the first feature processing module is configured as the model input of the target text processing model, the output end of the first feature processing module is connected to the input end of the second feature processing module; the output end of the second feature processing module is connected to the input end of the third feature processing module; the output end of the third feature processing module is connected to the input end of the fourth feature processing module; the output end of the fourth feature processing module is connected to the input end of the fifth feature processing module; the output end of the fifth feature processing module is connected to the input end of the sixth feature processing module; the output end of the sixth feature processing module is connected to the input end of the seventh feature processing module; the input end of the seventh feature processing module is connected to the input end of the eighth feature processing module; the output end of the eighth feature processing module is connected to the input end of the linear module; the output end of the linear module is connected to the input end of the activation module; the output end of the activation module is configured as the model output of the target text processing model.
[0120] It can be seen that implementation Figure 3 The described data processing device for video annotation is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0121] In yet another optional embodiment, Figure 3 As shown, the first feature processing module includes a first multi-head attention unit, a second multi-head attention unit, a third multi-head attention unit, a first fusion unit, a second fusion unit, a third fusion unit, a fourth fusion unit, a fifth fusion unit, a first normalization unit, a second normalization unit, a third normalization unit, a fourth normalization unit, a fifth normalization unit, a first network unit, and a second network unit; wherein,
[0122] The input end of the first multi-head attention unit and the input end of the first fusion unit are both configured as the input end of the first feature processing module, the output end of the first multi-head attention unit is connected to the input end of the first fusion unit; the output end of the first fusion unit is connected to the input end of the first normalization unit; the output end of the first normalization unit is respectively connected to the input end of the first network unit and the input end of the second fusion unit; the output end of the first network unit is connected to the input end of the second fusion unit; the output end of the second fusion unit is connected to the input end of the second normalization unit; the output end of the second normalization unit is respectively connected to the input end of the second multi-head attention unit, the input end of the third multi-head attention unit and the input end of the third fusion unit; the second multi-head attention unit The output end of the network is connected to the input end of the third fusion unit; the output end of the third fusion unit is connected to the input end of the third normalization unit; the output end of the third normalization unit is respectively connected to the input end of the third multi-head attention unit and the input end of the fourth fusion unit; the output end of the third multi-head attention unit is connected to the input end of the fourth fusion unit; the output end of the fourth fusion unit is connected to the input end of the fourth normalization unit; the output end of the fourth normalization unit is respectively connected to the input end of the second network unit and the input end of the fifth fusion unit; the output end of the second network unit is connected to the input end of the fifth fusion unit; the output end of the fifth fusion unit is connected to the input end of the fifth normalization unit; the output end of the fifth normalization unit is configured as the output end of the first feature processing module.
[0123] It can be seen that the implementation Figure 3 The described data processing device for video annotation is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0124] In yet another optional embodiment, Figure 3 As shown, the first processing module 202 uses the target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain the target comment sentiment information, including:
[0125] The target text processing model is used to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information; the video comment information includes a plurality of comment data information distributed in sequence; the target comment word information includes a plurality of preferred comment word information distributed in sequence; the preferred comment word information includes a plurality of preferred comment words;
[0126] The target comment word information is matched and converted to obtain the target comment sentiment information.
[0127] It can be seen that the implementation Figure 3The described data processing device for video annotation is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0128] In yet another optional embodiment, Figure 3 As shown, the first processing module 202 uses the target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information, including:
[0129] For any comment data information in the video comment information in the video information to be processed;
[0130] Based on the comment data information, determining first comment word information corresponding to the comment data information; the first comment word information includes N first comment words;
[0131] Using the target text processing model to perform feature analysis on the first comment word information to obtain comment word feature information; the comment word feature information includes N word feature vector information; each word feature vector information corresponds to a first comment word;
[0132] Using the target text processing model to perform feature analysis on the comment data information to obtain text feature information;
[0133] Based on the comment word feature information and the text feature information, the preferred comment word information corresponding to the comment data information is determined.
[0134] It can be seen that the implementation Figure 3 The described data processing device for video annotation is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0135] In yet another optional embodiment, Figure 3 As shown, the first processing module 202 determines the first comment word information corresponding to the comment data information based on the comment data information, including:
[0136] The comment data information is segmented to obtain comment word information; the comment word information includes M video comment words distributed in sequence;
[0137] The comment word information is evaluated and analyzed to obtain first analysis value information; the first analysis value information includes M first analysis values distributed in sequence; each first analysis value corresponds to a video comment word in the same distribution order;
[0138] Performing semantic analysis on the comment word information to obtain second analysis value information; the second analysis value information includes M second analysis values distributed in sequence; each second analysis value corresponds to the video comment word in the same distribution order;
[0139] Performing information entropy analysis and calculation processing on the comment word information to obtain third analysis value information; the third analysis value information includes M third analysis values distributed in sequence; each third analysis value corresponds to the video comment words in the same distribution order;
[0140] Normalizing the first analysis value information, the second analysis value information, and the third analysis value information to obtain first normalized value information, second normalized value information, and third normalized value information;
[0141] Acquire evaluation coefficient information; the evaluation coefficient information includes a first evaluation coefficient, a second evaluation coefficient and a third evaluation coefficient;
[0142] The first evaluation coefficient, the second evaluation coefficient and the third evaluation coefficient in the evaluation coefficient information are used to perform weighted sum processing with the first normalized value information, the second normalized value information and the third normalized value information in sequence to obtain the comment word evaluation value information; the comment word evaluation value information includes M comment word evaluation values distributed in sequence;
[0143] Sorting the evaluation values of the comment words in the evaluation value information of the comment words in descending order to obtain an evaluation value sequence;
[0144] The video comment words corresponding to the top N comment word evaluation values in the evaluation value sequence are taken as the first comment words.
[0145] It can be seen that implementation Figure 3 The described data processing device for video annotation is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0146] In yet another optional embodiment, Figure 3 As shown, the first processing module 202 determines the preferred comment word information corresponding to the comment data information based on the comment word feature information and the text feature information, including:
[0147] For any word feature vector information in the comment word feature information, the feature matching model is used to calculate and process the word feature vector information and the text feature information to obtain the feature matching value corresponding to the word feature vector information;
[0148] Among them, the feature matching model is:
[0149]
[0150] Among them, TP is the feature matching value; X is the word feature vector information, Y is the text feature information; s1 and s2 are the first matching coefficient and the second matching coefficient respectively;
[0151] Sort all feature matching values in descending order to obtain a feature matching value sequence;
[0152] The first comment word corresponding to the word feature vector information corresponding to the first L feature matching values in the feature matching value sequence is determined as the preferred comment word in the preferred comment word information.
[0153] It can be seen that the implementation Figure 3 The described data processing device for video annotation is beneficial to improve the accuracy and annotation efficiency of characterizing the emotion types of video materials by constructing multi-dimensional and temporal annotation information, opening up new perspectives for research in the field of emotional computing and providing richer and more accurate emotion data support.
[0154] Embodiment 3
[0155] See also Figure 4 , Figure 4 is a structural diagram of another data processing device for video annotation disclosed in an embodiment of the present invention. Figure 4 The described device can be applied to a management system, such as a local server or a cloud server for management, etc., which is not limited in the embodiments of the present invention. Figure 4 As shown, the device may include:
[0156] A memory 301 storing executable program codes;
[0157] a processor 302 coupled to the memory 301;
[0158] The processor 302 calls the executable program code stored in the memory 301 to execute the steps in the data processing method for video annotation described in the first embodiment.
[0159] Embodiment 4
[0160] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the data processing method for video annotation described in the first embodiment.
[0161] Embodiment 5
[0162] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps of the data processing method for video annotation described in the first embodiment.
[0163] The device embodiments described above are only illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, i.e., they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.
[0164] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution can be essentially or partly contributed to the prior art in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0165] Finally, it should be noted that the data processing method and device for video annotation disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, which are only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method for video annotation, characterized in that: The method comprises: Obtaining video information to be processed; the video information to be processed includes video comment information and video emotion information; the video emotion information represents facial expression features; the video emotion information is obtained in the following manner: using short video materials, designing an emotion induction experiment, and inviting subjects to participate in the experiment, inducing different emotional states of the subjects, and using high-frame rate camera equipment to record the facial movements of the subjects, analyzing the temporal changes of the facial expressions of the subjects, and outputting an emotion intensity vector v2 representing the temporal changes of the emotion intensity, so as to obtain the video emotion information; Using a target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain target comment sentiment information; Performing feature fusion processing on the target comment emotion information and the video emotion information in the to-be-processed video information to obtain target video annotation information; The step of using the target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain target comment sentiment information includes: The target text processing model is used to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information; the video comment information includes a plurality of comment data information distributed in sequence; the target comment word information includes a plurality of preferred comment word information distributed in sequence; the preferred comment word information includes a plurality of preferred comment words; Performing matching conversion processing on the target comment word information to obtain the target comment sentiment information; The step of using the target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information includes: For any of the comment data information in the video comment information in the to-be-processed video information, based on the comment data information, determine first comment word information corresponding to the comment data information; the first comment word information includes N first comment words; Using the target text processing model to perform feature analysis on the first comment word information to obtain comment word feature information; the comment word feature information includes N word feature vector information; each word feature vector information corresponds to one of the first comment words; Using the target text processing model to perform feature analysis on the comment data information to obtain text feature information; Based on the comment word feature information and the text feature information, the preferred comment word information corresponding to the comment data information is determined.
2. The data processing method for video annotation according to claim 1, characterized in that: The target text processing model includes a first feature processing module, a second feature processing module, a third feature processing module, a fourth feature processing module, a fifth feature processing module, a sixth feature processing module, a seventh feature processing module, an eighth feature processing module, a linear module and an activation module; wherein, The input end of the first feature processing module is configured as the model input of the target text processing model, the output end of the first feature processing module is connected to the input end of the second feature processing module; the output end of the second feature processing module is connected to the input end of the third feature processing module; the output end of the third feature processing module is connected to the input end of the fourth feature processing module; the output end of the fourth feature processing module is connected to the input end of the fifth feature processing module; the output end of the fifth feature processing module is connected to the input end of the sixth feature processing module; the output end of the sixth feature processing module is connected to the input end of the seventh feature processing module; the input end of the seventh feature processing module is connected to the input end of the eighth feature processing module; the output end of the eighth feature processing module is connected to the input end of the linear module; the output end of the linear module is connected to the input end of the activation module; the output end of the activation module is configured as the model output of the target text processing model.
3. The data processing method for video annotation according to claim 2, characterized in that: The first feature processing module includes a first multi-head attention unit, a second multi-head attention unit, a third multi-head attention unit, a first fusion unit, a second fusion unit, a third fusion unit, a fourth fusion unit, a fifth fusion unit, a first normalization unit, a second normalization unit, a third normalization unit, a fourth normalization unit, a fifth normalization unit, a first network unit, and a second network unit; wherein, The input end of the first multi-head attention unit and the input end of the first fusion unit are both configured as the input end of the first feature processing module, the output end of the first multi-head attention unit is connected to the input end of the first fusion unit; the output end of the first fusion unit is connected to the input end of the first normalization unit; the output end of the first normalization unit is respectively connected to the input end of the first network unit and the input end of the second fusion unit; the output end of the first network unit is connected to the input end of the second fusion unit; the output end of the second fusion unit is connected to the input end of the second normalization unit; the output end of the second normalization unit is respectively connected to the input end of the second multi-head attention unit, the input end of the third multi-head attention unit and the input end of the third fusion unit; The output end is connected to the input end of the third fusion unit; the output end of the third fusion unit is connected to the input end of the third normalization unit; the output end of the third normalization unit is respectively connected to the input end of the third multi-head attention unit and the input end of the fourth fusion unit; the output end of the third multi-head attention unit is connected to the input end of the fourth fusion unit; the output end of the fourth fusion unit is connected to the input end of the fourth normalization unit; the output end of the fourth normalization unit is respectively connected to the input end of the second network unit and the input end of the fifth fusion unit; the output end of the second network unit is connected to the input end of the fifth fusion unit; the output end of the fifth fusion unit is connected to the input end of the fifth normalization unit; the output end of the fifth normalization unit is configured with the output end of the first feature processing module.
4. The data processing method for video annotation according to claim 1, characterized in that: The step of determining first comment word information corresponding to the comment data information based on the comment data information includes: Performing word segmentation processing on the comment data information to obtain comment word information; the comment word information includes M video comment words distributed in sequence; The comment word information is evaluated and analyzed to obtain first analysis value information; the first analysis value information includes M first analysis values distributed in sequence; each of the first analysis values corresponds to the video comment words in the same distribution order; Performing semantic analysis on the comment word information to obtain second analysis value information; the second analysis value information includes M second analysis values distributed in sequence; each of the second analysis values corresponds to the video comment words in the same distribution order; Performing information entropy analysis and calculation processing on the comment word information to obtain third analysis value information; the third analysis value information includes M third analysis values distributed in sequence; each of the third analysis values corresponds to the video comment words in the same distribution order; Normalizing the first analysis value information, the second analysis value information, and the third analysis value information to obtain first normalized value information, second normalized value information, and third normalized value information; Acquire evaluation coefficient information; the evaluation coefficient information includes a first evaluation coefficient, a second evaluation coefficient and a third evaluation coefficient; The first evaluation coefficient, the second evaluation coefficient and the third evaluation coefficient in the evaluation coefficient information are used to perform weighted sum processing with the first normalized value information, the second normalized value information and the third normalized value information in sequence to obtain comment word evaluation value information; the comment word evaluation value information includes M comment word evaluation values distributed in sequence; Sorting the evaluation values of the comment words in the evaluation value information of the comment words in descending order to obtain an evaluation value sequence; The video comment words corresponding to the top N comment word evaluation values in the evaluation value sequence are used as the first comment words.
5. The data processing method for video annotation according to claim 1, characterized in that: The determining, based on the comment word feature information and the text feature information, the preferred comment word information corresponding to the comment data information includes: For any word feature vector information in the comment word feature information, use a feature matching model to calculate and process the word feature vector information and the text feature information to obtain a feature matching value corresponding to the word feature vector information; Wherein, the feature matching model is: Wherein, TP is the feature matching value; X is the word feature vector information, Y is the text feature information; s1 and s2 are the first matching coefficient and the second matching coefficient respectively; Sorting all the feature matching values in descending order of the feature matching values to obtain a feature matching value sequence; The first comment word corresponding to the word feature vector information corresponding to the first L feature matching values in the feature matching value sequence is determined as the preferred comment word in the preferred comment word information.
6. A data processing device for video annotation, characterized in that: The device comprises: An acquisition module is used to acquire video information to be processed; the video information to be processed includes video comment information and video emotion information; the video emotion information represents facial expression features; the video emotion information is acquired in the following manner: using short video materials, designing an emotion induction experiment, and inviting subjects to participate in the experiment, by inducing different emotional states of the subjects, and using high-frame rate camera equipment to record the facial movements of the subjects, analyzing the temporal changes of the facial expressions of the subjects, and outputting an emotion intensity vector v2 representing the temporal changes of the emotion intensity, that is, obtaining the video emotion information; A first processing module is used to perform feature analysis on the video comment information in the video information to be processed using a target text processing model to obtain target comment sentiment information; The second processing module is used to perform feature fusion processing on the target comment emotion information and the video emotion information in the to-be-processed video information to obtain target video annotation information; The step of using the target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain target comment sentiment information includes: The target text processing model is used to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information; the video comment information includes a plurality of comment data information distributed in sequence; the target comment word information includes a plurality of preferred comment word information distributed in sequence; the preferred comment word information includes a plurality of preferred comment words; Performing matching conversion processing on the target comment word information to obtain the target comment sentiment information; The step of using the target text processing model to perform feature analysis on the video comment information in the video information to be processed to obtain target comment word information includes: For any of the comment data information in the video comment information in the to-be-processed video information, based on the comment data information, determine first comment word information corresponding to the comment data information; the first comment word information includes N first comment words; Using the target text processing model to perform feature analysis on the first comment word information to obtain comment word feature information; the comment word feature information includes N word feature vector information; each word feature vector information corresponds to one of the first comment words; Using the target text processing model to perform feature analysis on the comment data information to obtain text feature information; Based on the comment word feature information and the text feature information, the preferred comment word information corresponding to the comment data information is determined.
7. A data processing device for video annotation, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the data processing method for video annotation according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the data processing method for video annotation according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video classification method and apparatus
CN105868686A
Emotional tendency information obtaining method and device
CN110516249A