Content recognition method, device, computer equipment and storage medium
By correlating and adjusting text features based on associated data, the method enhances content recognition accuracy by focusing on important text features, addressing the low accuracy issues in existing methods.
Patent Information
- Application Number
- CN202110325997.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-26
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-03-26
AI Technical Summary
The content recognition method in the prior art has a problem of low accuracy.
By determining the text and association data of the target content, feature extraction and association calculation are performed, text features are adjusted based on feature correlation and attention intensity to improve recognition accuracy.
By adaptive adjustment of the associated data and text features, we will increase the attention to important features and improve the accuracy of content recognition.
Smart Images

Figure CN113723166B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a content recognition method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of natural language processing technologies and artificial intelligence technologies, content recognition is required in many cases, such as video recognition. When recognizing content, the content can be recognized based on an artificial intelligence model, and the required information can be obtained from the content. For example, text can be recognized to obtain the required content entities from the text.
[0003] Currently, the method for recognizing content has a situation where the information of the content cannot be accurately recognized, resulting in a low accuracy of content recognition. Summary of the Invention
[0004] Based on this, it is necessary to provide a content recognition method, apparatus, computer device, and storage medium that can improve the accuracy of content recognition for the above technical problems.
[0005] A content recognition method, the method includes: determining target content to be recognized, obtaining target text in the target content and text association data associated with the target text; extracting features from the target text to obtain text extraction features; extracting features from the text association data to obtain association extraction features; performing an association calculation on the association extraction features and the text extraction features, and obtaining a feature attention intensity corresponding to the text extraction features based on the calculated feature association degree, where the feature association degree is positively correlated with the feature attention intensity; adjusting the text extraction features based on the feature attention intensity to obtain adjusted text features; and performing recognition based on the adjusted text features to obtain a content recognition result corresponding to the target content.
[0006] A content recognition device, the device comprising: a target content determination module for determining a target content to be recognized, obtaining target text in the target content and text association data associated with the target text; a feature extraction module for extracting features from the target text to obtain text extraction features; extracting features from the text association data to obtain association extraction features; a feature attention intensity obtaining module for performing an association calculation on the association extraction features and the text extraction features, and obtaining a feature attention intensity corresponding to the text extraction features based on a calculated feature association degree, where the feature association degree is positively correlated with the feature attention intensity; an adjusted text feature obtaining module for adjusting the text extraction features based on the feature attention intensity to obtain adjusted text features; and a content recognition result obtaining module for performing recognition based on the adjusted text features to obtain a content recognition result corresponding to the target content.
[0007] In some embodiments, the content recognition result obtaining module includes: a first fused text feature obtaining unit for fusing the adjusted text features and the text extraction features to obtain fused text features; and a first content recognition result obtaining unit for performing recognition based on the fused text features to obtain a content recognition result corresponding to the target content.
[0008] In some embodiments, the first fused text feature obtaining unit is further configured to encode the text extraction features to obtain first encoded features, encode the adjusted text features to obtain second encoded features; fuse the first encoded features and the second encoded features to obtain fused encoded features; obtain an adjusted feature weight corresponding to the adjusted text features based on the fused encoded features; and fuse the adjusted text features and the text extraction features based on the adjusted feature weight to obtain fused text features.
[0009] In some embodiments, the first encoded features are obtained by encoding through a first encoder in a trained content recognition model, the second encoded features are obtained by encoding through a second encoder in the content recognition model, and the first fused text feature obtaining unit is further configured to input the fused encoded features into a target activation layer in the content recognition model for activation processing to obtain a target activation value, and use the target activation value as the adjusted feature weight corresponding to the adjusted text features, where the activation layer is a shared activation layer of the first encoder and the second encoder.
[0010] In some embodiments, the first fused text feature obtaining unit is further configured to obtain the text feature weight corresponding to the text extraction feature based on the adjusted feature weight; perform a product calculation on the adjusted feature weight and the adjusted text feature to obtain the calculated adjusted text feature; perform a product calculation on the text feature weight and the text extraction feature to obtain the calculated text extraction feature; and add the calculated adjusted text feature and the calculated text extraction feature to obtain the fused text feature.
[0011] In some embodiments, the target content is a target video; the target content determination module includes: a target text obtaining unit, configured to obtain the text corresponding to the target time in the target video to obtain the target text; and a text associated data obtaining unit, configured to obtain the video-related data corresponding to the target time in the target video, and use the video-related data as the text associated data associated with the target text, where the video-related data includes at least one of video frames or audio frames.
[0012] In some embodiments, the adjusted text feature includes a first adjusted text feature adjusted according to the video frame and a second adjusted text feature adjusted according to the audio frame; the content recognition result obtaining module includes: a second fused text feature obtaining unit, configured to fuse the first adjusted text feature, the second adjusted text feature, and the text extraction feature to obtain the fused text feature; and a second content recognition result obtaining unit, configured to perform recognition based on the fused text feature to obtain the content recognition result corresponding to the target content.
[0013] In some embodiments, the adjusted text feature obtaining module includes: a feature value product obtaining unit, configured to multiply the feature attention intensity by each feature value of the text extraction feature to obtain the feature value product; and an adjusted text feature obtaining unit, configured to arrange the feature value product according to the positions of the feature values in the text extraction feature, and use the arranged feature value sequence as the adjusted text feature.
[0014] In some embodiments, the text extraction feature is the feature corresponding to the word segmentation in the target text; each adjusted text feature forms a feature sequence according to the order of the word segmentation in the target text; the content recognition result obtaining module includes: a position relationship obtaining unit, configured to obtain the position relationship of each word segmentation relative to the named entity based on the feature sequence; and a third content recognition result obtaining unit, configured to obtain the target named entity from the target text based on each position relationship, and use the target named entity as the content recognition result corresponding to the target content.
[0015] In some embodiments, the third content recognition result obtaining unit is further configured to obtain the word segment at the starting position of the named entity as the starting word of the named entity; use, among the backward word segments corresponding to the starting word of the named entity, the word segments whose position relationship is inside the named entity as the constituent words of the named entity; and combine the starting word of the named entity and the constituent words of the named entity to obtain the target named entity.
[0016] In some embodiments, the position relationship obtaining unit is further configured to obtain, based on the feature sequence, the position relationship of each word segment relative to the named entity and the entity type corresponding to the word segment; the third content recognition result obtaining unit is further configured to use, among the backward word segments corresponding to the starting word of the named entity, the word segments whose position relationship is inside the named entity and whose entity type is the same as that of the starting word of the named entity as the constituent words of the named entity.
[0017] In some embodiments, the feature attention intensity obtaining module includes: a product operation value obtaining unit configured to perform a product operation on the association feature value in the association extraction feature and the text feature value at the corresponding position in the text extraction feature to obtain a product operation value; and a feature attention intensity obtaining unit configured to perform statistics on the product operation value to obtain the feature association degree between the association extraction feature and the text extraction feature, and use the feature association degree as the feature attention intensity corresponding to the text extraction feature.
[0018] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above content recognition method are implemented.
[0019] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the above content recognition method are implemented.
[0020] In some embodiments, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the steps in the above method embodiments.
[0021] The above content recognition method, device, computer equipment and storage medium determine the target content to be recognized, obtain the target text in the target content and the text association data associated with the target text, extract features from the target text to obtain text extraction features, extract features from the text association data to obtain associated extraction features, perform association calculation on the associated extraction features and the text extraction features, obtain the feature attention intensity corresponding to the text extraction features based on the calculated feature association degree, the feature association degree is positively correlated with the feature attention intensity, adjust the text extraction features based on the feature attention intensity to obtain adjusted text features, and perform recognition based on the adjusted text features to obtain the content recognition result corresponding to the target content. Since the feature association degree is positively correlated with the feature attention intensity, the greater the association relationship between the target text and the text association data, the greater the feature attention intensity, and the more attention is paid to the text extraction features. Therefore, when performing recognition based on the adjusted text features, the greater the association relationship between the target text and the text association data, the greater the influence degree of the text association data on the recognition result, so that the text features can be adjusted adaptively according to the relationship between the text association data and the target text, making the important text features more concerned during content recognition and improving the accuracy of content recognition. Description of the Drawings
[0022] Figure 1 It is an application environment diagram of the content recognition method in some embodiments;
[0023] Figure 2 It is a flowchart of the content recognition method in some embodiments;
[0024] Figure 3 It is a schematic diagram of video recognition using the content recognition method in some embodiments;
[0025] Figure 4 It is a framework diagram of the content recognition model in some embodiments;
[0026] Figure 5 It is a framework diagram of the content recognition model in some embodiments;
[0027] Figure 6 It is a schematic diagram of entity recognition using the entity recognition network in some embodiments;
[0028] Figure 7 It is a framework diagram of the content recognition network in some embodiments;
[0029] Figure 8 It is a schematic diagram of entity recognition using the entity recognition model in some embodiments;
[0030] Figure 9 It is a structural block diagram of the content recognition device in some embodiments;
[0031] Figure 10 Internal structure diagrams of computer devices in some embodiments;
[0032] Figure 11 Internal structure diagrams of computer devices in some embodiments. Specific implementation manners
[0033] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0034] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.
[0035] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0036] Computer Vision Technology (CV) Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking and measurement on targets, and further performing graphic processing to make the computer process into images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0037] The key technologies of speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction. Among them, speech has become one of the most promising human-computer interaction methods in the future.
[0038] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.
[0039] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0040] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0041] The solution provided in the embodiments of this application involves technologies such as computer vision technology, speech technology, natural language processing, and machine learning in artificial intelligence, which will be specifically described through the following embodiments:
[0042] The content recognition method provided in this application can be applied to an Figure 1 application environment as shown. This application environment includes a terminal 102 and a server 104. Among them, the terminal 102 and the server 104 communicate through a network.
[0043] Specifically, the server 104 can, in response to a content recognition request, obtain the target content to be recognized. The target content to be recognized can be carried in the content recognition request or the content obtained according to the content identifier carried in the content recognition request. The server 104 can obtain the target text in the target content and the text association data associated with the target text, extract features from the target text to obtain text extraction features, extract features from the text association data to obtain associated extraction features, perform an association calculation on the associated extraction features and the text extraction features, obtain the feature attention intensity corresponding to the text extraction features based on the calculated feature association degree, where the feature association degree is positively correlated with the feature attention intensity, adjust the text extraction features based on the feature attention intensity to obtain adjusted text features, perform recognition based on the adjusted text features to obtain the content recognition result corresponding to the target content, and the server 104 can store the content recognition result and the target content in an associated manner. For example, the content recognition result can be used as a label for the target content. Among them, the content recognition request can be triggered by the server 104 or sent to the server 104 by other devices such as a terminal.
[0044] Among them, a client can be installed on the terminal 102. For example, at least one of a video client, a browser client, an instant messaging client, or an education client can be installed. The terminal 102 can send a content search request to the server 104 through the client in response to a content search operation triggered by a user. The content search request can carry search information. The server 104 can match the search information with the content recognition result. When the search information matches the content recognition result, the content corresponding to the content recognition result is sent to the terminal 102, and the terminal 102 can display the content returned by the server 104 in the client.
[0045] Among them, the terminal 102 can be, but is not limited to, a laptop computer, a smart phone, a smart TV, a desktop computer, a tablet computer, and a portable wearable device. The server 104 can be implemented by an independent server or a server cluster or cloud server composed of multiple servers. It can be understood that the above application scenario is only an example and does not constitute a limitation on the content recognition method provided by the embodiments of the present application. The method provided by the embodiments of the present application can also be applied in other application scenarios. For example, the content recognition method provided by the present application can be executed by the terminal 102 or the server 104, or can be executed jointly by the terminal 102 and the server 104. The terminal 102 can upload the recognized content recognition result to the server 104, and the server 104 can store the target content and the content recognition result in an associated manner.
[0046] In some embodiments, as Figure 2 shown, a content recognition method is provided, and this method is applied to Figure 1Taking the server 104 in [it] as an example, the following steps are included:
[0047] S202. Determine the target content to be recognized, and obtain the target text in the target content and the text association data associated with the target text.
[0048] Among them, the content can be any one of video, audio or text. The content includes text data, and can also include at least one of image data or audio data. The audio data can be, for example, voice data. When the content is video, the text data in the content can include at least one of subtitles, bullet screens, comments or titles in the video. The image data in the content can be video frames in the video, and the audio data in the content can be dubbing or music and other audio data in the video. When the content is audio data, the text data in the content can be the text data corresponding to the audio data. For example, when the content is a song, the text data in the content can be the lyrics corresponding to the song, and the audio data in the content can be audio frames. Audio frames are obtained by frame-dividing the audio. Frame-dividing means dividing the audio into multiple small segments, and each small segment is one frame.
[0049] The target content refers to the content to be recognized, and can be at least one of the content to be identity-recognized or the content to be scene-recognized. Identity recognition refers to recognizing the identity of the person appearing in the target content. For example, the identity of the person can be determined by recognizing the person information appearing in the target content. The person information can include at least one of the name of the person or the face of the person. Scene recognition refers to recognizing the scene to which the target content belongs. For example, the scene can be determined by recognizing the location appearing in the target content. The target text refers to the text data in the target content, and can include the text data at any moment in the target content. For example, when the target content is video, the target text can include at least one of subtitles, bullet screens, comments or titles shown at any moment or time period in the video. When the target content is a song, the target text data can be the lyrics corresponding to the song.
[0050] The text association data refers to the data in the target content that has an association relationship with the target text. For example, it can include at least one of the target image data or the target audio data in the target content that has an association relationship with the target text. The target image data is the image data in the target content that has an association relationship with the target text, and the target audio data is the audio data in the target content that has an association relationship with the target text. The target image data can include one or more images. Multiple means at least two. The target audio data can include one or more audio frames. Multiple segments means at least two segments.
[0051] The association relationship may include an association relationship in terms of time. For example, the text association data may include the data that appears within the time when the target text appears in the target content, or may include the data that appears within the time when the time interval between the appearance of the target text in the target content is less than the time interval threshold. For example, when the target content is a target video and the target text is the subtitle text of the video, the text association data may be the video frames and voices that match the subtitle. For example, the target text and the corresponding text association data may be the data that describes the same video scene. For example, when the target text is the subtitle that appears at the target time in the target video, the text association data may include the data that appears at the target time in the target video. For example, it may include at least one of the video frames, bullet comments, or audio frames that appear at the target time in the target video, or may include the data that appears within the time when the time interval between the target time in the target video is less than the time interval threshold. The time interval threshold may be preset or may be set as needed. Among them, the target video may be any video, which may be a directly shot video or a video clip intercepted from the shot video. The target video may be any type of video, including but not limited to at least one of advertising videos, TV drama videos, or news videos. The target video may also be a video to be pushed to the user. The video frames that appear at the target time in the target video may include one or more frames, and the audio frames that appear at the target time in the target video may include one or more frames. Multiple frames refer to at least two frames.
[0052] The association relationship may also include a semantic association relationship. For example, the text association data may include the data that semantically matches the target text in the target content. The data that semantically matches the target text may include the data that is semantically consistent with the target text, or may include the data whose semantic difference from the semantic of the target text is less than the semantic difference threshold. The semantic difference threshold may be preset or may be set as needed.
[0053] Specifically, the server may obtain the content to be subjected to entity recognition, such as a video to be subjected to entity recognition, use the content recognition method provided by this application to recognize the content to be subjected to entity recognition, obtain the recognized entity words, construct a knowledge graph based on the recognized entity words, or may use the recognized entity words as the labels corresponding to the target content. When it is necessary to push the target content, the user who matches the target content may be determined according to the labels corresponding to the target content, and the target content may be pushed to the terminal of the matching user.
[0054] Among them, an entity refers to a thing with a specific meaning, which can include at least one of place names, organization names, or proper nouns, etc. The target text may include one or more entities, and an entity word is a word representing an entity. For example, assuming the target text is "Monkeys like to eat bananas", the entities included in the target text are "monkeys" and "bananas", "monkey" is an entity word, and "banana" is an entity word. A knowledge graph is a graph-based data structure that includes nodes (points) and edges. Each node represents an entity, and each edge is a relationship between entities.
[0055] Entity recognition can also be called entity word recognition or named entity recognition (NER). Entity word recognition is an important research direction in the field of natural language processing (NLP). There are many methods for entity word recognition, such as methods based on dictionaries and rules, machine learning methods including hidden Markov model (HMM), maximum entropy Markov model (MEMM), conditional random fields (CRF), etc., deep learning models including recurrent neural networks (RNN) and long short-term memory networks (LSTM), and recognition methods combining LSTM and CRF. Among them, natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language.
[0056] In some embodiments, the first terminal may send a content push request to the server. The content push request may carry a content identifier corresponding to the content to be pushed. The content identifier is used to uniquely identify the content. The content to be pushed may be, for example, a video to be pushed. The server may, in response to the content push request, obtain the content to be pushed corresponding to the content identifier carried in the content push request as the target content to be recognized. For example, the first terminal may display a content push interface. The content push interface may display a push content acquisition area and a content push trigger control. The push content acquisition area is used to receive content information corresponding to the content to be pushed. The content information includes one or more content identifiers. "Multiple" means at least two. The content push trigger control is used to trigger the first terminal to send a content push request to the server. When the first terminal obtains a trigger operation on the content push trigger control, it obtains the content information received in the push content acquisition area and sends a content push request carrying the content information to the server. The server may obtain the content corresponding to each content identifier included in the content information respectively as each target content to be recognized. The server may use the content recognition method provided in this application to recognize each content to be recognized, determine the users respectively matching each target content according to the recognition result, and push the target content to the terminals of the matching users. For example, the recognition result may be matched with the user's user image description. When the match is successful, the target content is pushed to the terminal of the user.
[0057] For example, when the content is a video, the content recognition method may also be referred to as a video recognition method, the content push request may also be referred to as a video push request, and the content push interface may be, for example, Figure 3 the video push interface 300 in Figure 3 the area 302 in Figure 3 The "OK" control 304 in. When the terminal obtains a click operation on the "OK" control 304, it may send a video push request to the server. The server recognizes videos A and B according to the video recognition method, determines user 1 matching video A and user B matching video B, pushes video A to the terminal of user 1, and pushes video B to the terminal of user 2.
[0058] S204. Extract features from the target text to obtain text extraction features; extract features from the text-associated data to obtain associated extraction features.
[0059] Among them, the text extraction feature is the feature obtained by performing feature extraction on the target text. The text extraction feature can be the feature obtained by further performing feature extraction on the target word vectors of the target word segmentation corresponding to the target text. The target word segmentation is obtained by segmenting the target text, and the segmentation granularity can be set as needed. For example, it can be segmented by character, word, or sentence as the unit to obtain the segmented text blocks, and each text block is used as a word segmentation. When segmented by character as the unit, one character corresponds to one text block, that is, one character is one word segmentation. For example, when the target text is "I'm so thirsty", when segmented by character as the unit, the obtained word segmentations are respectively "I", "so", and "thirsty". The target word vector is the vector representation form of the target word segmentation. The target word segmentations obtained by segmenting the target text can be one or more, and "more" means at least two.
[0060] The associated extraction feature is the feature obtained by performing feature extraction on the text associated data. When the text associated data is the target image data, the associated extraction feature can be the target image feature obtained by performing image feature extraction on the target image data. When the text associated data is the target audio data, the associated extraction feature can be the target audio feature obtained by performing audio feature extraction on the target audio data. The target image feature is the image feature extracted by performing image feature extraction on the target image data. The target audio feature is the audio feature extracted by performing audio feature extraction on the target audio feature. The text extraction feature and the associated extraction feature can be of the same dimension, for example, they can be vectors or matrices of the same dimension.
[0061] Specifically, the server can input the target text into the text feature extraction network in the trained content recognition model, use the text feature extraction network to perform feature extraction on the target text to obtain the text extraction feature, input the text associated data into the associated feature extraction network in the trained content recognition model, and use the associated extraction network to perform feature extraction on the text associated data to obtain the associated extraction feature. The trained content recognition model is used to recognize the content to obtain the content recognition result, for example, used to recognize at least one of the entity words included in the subtitle of the video or the scene of the video. There can be multiple associated feature extraction networks in the trained content recognition model. For example, the associated feature extraction network can include at least one of the image feature extraction network or the audio feature extraction network. The image feature extraction network is used to extract the features of the image, and the audio feature extraction network is used to extract the features of the audio. When the text associated data is the target image data, the text associated data can be input into the image feature extraction network, and the image feature extracted by the image feature extraction network is used as the associated extraction feature. When the text associated data is the target audio data, the text associated data can be input into the audio feature extraction network, and the audio feature extracted by the audio feature extraction network is used as the associated extraction feature.
[0062] Among them, the text feature extraction network, the image feature extraction network, and the audio feature extraction network can be neural networks based on artificial intelligence. For example, they can be Convolutional Neural Networks (CNNs). Of course, they can also be other types of neural networks. The text feature extraction network can be, for example, a Transformer network or a Bidirectional Encoder Representations from Transformers (BERT) network based on Transformer. The image feature extraction network can be, for example, a Residual Network (ResNet). The audio feature extraction network can be, for example, a VGG (Visual Geometry Group) convolutional network. VGG represents the Visual Geometry Group at the University of Oxford. For example, the server can perform scale transformation on the target image to obtain the scaled image, input the scaled image data into the Residual Network for image feature extraction, pool the features output by the feature map extraction layer in the Residual Network. For example, pool them into a fixed-size n*n size, and use the pooled features as the associated extraction features. n is a positive number greater than or equal to 1.
[0063] In some embodiments, the steps of extracting text features from the target text include: segmenting the target text to obtain target word segments, performing vector transformation on the target word segments to obtain target word vectors corresponding to the target word segments, and using the target word vectors as text extraction features.
[0064] In some embodiments, the server can input the target text into a Transformer model based on attention. The Transformer model, as the encoder of text features, can encode the target text to obtain the encoded features in the form of embedding representations of each character in the target text, and can use the encoded features corresponding to each character as text extraction features.
[0065] In some embodiments, the server may perform a spectrum calculation on the target audio data to obtain a spectrogram corresponding to the target audio data, extract features from the spectrogram corresponding to the target audio data, and use the extracted features as associated extraction features. For example, the server may perform a sound spectrum calculation on the spectrogram corresponding to the target audio data to obtain sound spectrum information corresponding to the target audio data, and extract features from the sound spectrum information of the target audio data to obtain associated extraction features. For example, the server may use a hann (Hanning window) time window to perform a Fourier transform on the target audio data to obtain a spectrogram corresponding to the target audio data, calculate the spectrogram through a mel (Mel) filter to obtain sound spectrum information corresponding to the target audio data, use a VGG convolutional network to extract features from the sound spectrum information, and use the audio features obtained by feature extraction as associated extraction features.
[0066] S206. Perform an association calculation on the associated extraction features and the text extraction features, and obtain the feature attention intensity corresponding to the text extraction features based on the calculated feature association degree. The feature association degree is positively correlated with the feature attention intensity.
[0067] Among them, the feature association degree is the result obtained by performing an association calculation on the associated extraction features and the text extraction features. The larger the association degree, the stronger the association relationship. The feature association degree is positively correlated with the feature attention intensity. Different target word segmentations corresponding to the text extraction features may result in different feature attention intensities corresponding to the text extraction features. For example, the feature association degree may be used as the feature attention intensity, or a linear operation or a non-linear operation may be performed on the feature association degree, and the result of the operation may be used as the feature attention intensity. The linear operation includes at least one of addition operation or multiplication operation, and the non-linear operation includes at least one of exponential operation or logarithmic operation. The positive correlation relationship means that: under the condition that other conditions remain unchanged, the change directions of two variables are the same. When one variable changes from large to small, the other variable also changes from large to small. It can be understood that the positive correlation relationship here means that the change directions are the same, but it does not require that when one variable changes a little, the other variable must also change. For example, it can be set that when variable a is from 10 to 20, variable b is 100, and when variable a is from 20 to 30, variable b is 120. In this way, the change directions of a and b are both that when a becomes larger, b also becomes larger. However, within the range of a from 10 to 20, b may not change.
[0068] The feature correlation degree may include at least one of an image correlation degree or an audio correlation degree. The image correlation degree refers to the result obtained by performing a correlation calculation on the target image features and the text extraction features. The audio correlation degree refers to the result obtained by performing a correlation calculation on the target audio features and the text extraction features. The feature attention intensity may include at least one of an image attention intensity or an audio attention intensity. The image attention intensity is obtained based on the image correlation degree, and the image attention intensity has a positive correlation with the image correlation degree. The audio attention intensity is obtained based on the audio correlation degree, and the audio attention intensity has a positive correlation with the audio correlation degree. The feature attention intensity is used to reflect the intensity of attention to the feature. The greater the feature attention intensity, the more attention needs to be paid to this feature when performing content recognition.
[0069] The associated extraction features may include a plurality of ordered associated feature values, and the text extraction features may include a plurality of ordered text feature values. The text feature value refers to the feature value included in the text extraction features, and the associated feature value refers to the feature value included in the associated extraction features. The associated extraction features and the text extraction features may be of the same dimension. For example, they may be vectors or matrices of the same dimension. That is to say, the number of associated feature values included in the associated extraction features may be the same as the number of text feature values included in the text extraction features. For example, assume that the text extraction feature is vector A = [a1, a2, a3], and the associated extraction feature is vector B = [b1, b2, b3]. Among them, vector A includes 3 elements, namely a1, a2, and a3, and each element in vector A is a text feature value. Similarly, vector B includes 3 elements, namely b1, b2, and b3, and each element in vector B is an associated feature value.
[0070] Specifically, the correlation calculation may be at least one of a product operation or a summation operation. When the correlation calculation is a product operation, the associated feature values in the associated extraction features may be multiplied by the text feature values at the corresponding positions in the text extraction features to obtain product operation values. Statistical operations are performed on each product operation value. For example, summation operations or mean operations are performed on each product operation value to obtain a statistical operation result. The feature correlation degree is obtained based on the statistical operation result. For example, the statistical operation result may be used as the feature correlation degree, or the statistical operation result may be normalized, and the result of the normalization process may be used as the feature correlation degree. When the correlation calculation is a summation operation, the associated feature values in the associated extraction features may be added to the text feature values at the corresponding positions in the text extraction features to obtain summation operation values. Statistical operations are performed on each summation operation value. For example, summation operations or mean operations may be performed on each summation operation value to obtain a statistical operation result.
[0071] In some embodiments, there are multiple target word segments obtained by segmenting the target text. The server can obtain text extraction features respectively obtained according to each target word segment, form a matrix with each text extraction feature, and use the formed matrix as a text extraction feature matrix. Each column in the text extraction feature matrix is a text extraction feature. The server can perform a multiplication operation on the associated extraction feature and the text extraction feature matrix to obtain a total multiplication operation result, and determine the feature correlation degree corresponding to each text extraction feature based on the total multiplication operation result. Among them, the step of performing a multiplication operation on the associated extraction feature and the text extraction feature matrix to obtain a total multiplication operation result may include: performing a multiplication operation on each text extraction feature in the text extraction feature matrix and the associated extraction feature respectively to obtain a sub-multiplication operation result corresponding to each text extraction feature, and using each sub-multiplication operation result as the total multiplication operation result. Among them, the step of performing a multiplication operation on each text extraction feature in the text extraction feature matrix and the associated extraction feature respectively to obtain a sub-multiplication operation result corresponding to each text extraction feature may include: performing a multiplication operation on the text feature value in the text extraction feature and the associated feature value at the corresponding position in the associated extraction feature to obtain a sub-multiplication operation result corresponding to the text extraction feature. The step of determining the feature correlation degree corresponding to each text extraction feature based on the total multiplication operation result may include: performing a normalization process on each sub-multiplication operation result in the total multiplication operation result to obtain each normalized sub-multiplication operation result, and using the normalized sub-multiplication operation result as the feature correlation degree corresponding to the text extraction feature.
[0072] In some embodiments, when the text association data is target image data and there are multiple target image data, the server can form a matrix with the target image features corresponding to each target image data, and use the formed matrix as an image feature matrix. Each column in the image feature matrix is a target image feature. The server can perform a matrix multiplication operation on the transposed matrix of the target image feature matrix and the text extraction feature matrix to obtain a first product matrix, perform a normalization process on each matrix value in the first product matrix to obtain a normalized first product matrix, and determine the image correlation degree corresponding to each text extraction feature based on the normalized first product matrix. The normalized first product matrix includes the image correlation degree corresponding to each text extraction feature.
[0073] For example, assume the target text is "I'm so thirsty". Segment the target text by character as a unit to obtain 3 target word segments, namely "I", "so", and "thirsty". One word segment is one character, the dimension of the target word vector corresponding to the target word segment is 2, and the target word vector corresponding to "I" is A = (a1, a2) T , the target word vector corresponding to "so" is B = (b1, b2) T, the target word vector corresponding to "thirsty" is C = (c1, c2) T , taking each target word vector as a text extraction feature, then the text extraction feature matrix feature text can be expressed by formula (1). Suppose there are 3 target image data, and these 3 target image data can be the same or different. For example, they are 3 images, and the target image features corresponding to each target image data are R = (r1, r2) T , M = (m1, m2) T , N = (n1, n2) T , R, M, and N can be the same or different, then the target image feature matrix feature image can be expressed by formula (2). Then the first product matrix L1 can be expressed by formula (3).
[0074]
[0075] L1 = [feature image T [feature text (3)
[0076] In some embodiments, the steps of normalizing each matrix value in the first product matrix to obtain the normalized first product matrix include: determining a scaling factor, dividing each matrix value in the first product matrix by the scaling factor respectively to obtain the scaling values corresponding to the matrix values, normalizing the scaling values, and taking the matrix composed of the scaling values as the normalized first product matrix. Among them, the scaling factor can be preset or set as needed. For example, the scaling factor can be determined according to the dimension of the text extraction feature. For example, the scaling factor can be positively correlated with the dimension of the text extraction feature. For example, the square root of the dimension of the text extraction feature can be calculated to obtain the scaling factor. For example, the square root of the dimension of the text extraction feature can be processed, and the ratio of the result after the square root processing to the first value is used as the scaling factor. The first value can be preset. The method used for normalization processing can be any function that can convert the input data into a number between 0 and 1. For example, the function softmax can be used for normalization processing. For example, the normalized first product matrix L2 can be calculated using formula (4).
[0077] Among them, d is the dimension of the text extraction feature, and m is the first value.
[0078]
[0079] Similarly, when the text-associated data is the target audio data and there are multiple pieces of target audio data, the server may form a target audio feature matrix by combining the target audio features corresponding to each piece of target audio data. Each column in the target audio feature matrix is a target audio feature. The server may perform a matrix multiplication operation on the transposed matrix of the target audio feature matrix and the text extraction feature matrix to obtain a second product matrix, perform a normalization process on each matrix value in the second product matrix to obtain a normalized second product matrix, determine the audio association degrees corresponding to each text extraction feature based on the normalized second product matrix, and the normalized second product matrix includes the audio association degrees corresponding to each text extraction feature.
[0080] In some embodiments, performing an association calculation on the association extraction feature and the text extraction feature, and obtaining the feature attention intensity corresponding to the text extraction feature based on the calculated feature association degree includes: calculating a feature similarity between the association extraction feature and the text extraction feature to obtain a feature similarity, using the feature similarity as the feature association degree, and obtaining the feature attention intensity corresponding to the text extraction feature based on the feature association degree. For example, the cosine similarity calculation formula may be used to calculate the similarity between the association extraction feature and the text extraction feature, and the calculated cosine similarity is used as the feature similarity.
[0081] S208. Adjust the text extraction feature based on the feature attention intensity to obtain an adjusted text feature.
[0082] Among them, the adjusted text feature is the feature obtained by adjusting the text extraction feature based on the feature attention intensity, and the adjusted text feature may include at least one of a first adjusted text feature or a second adjusted text feature. The first adjusted text feature refers to the feature obtained by adjusting the text extraction feature based on the image attention intensity. The second adjusted text feature refers to the feature obtained by adjusting the text extraction feature based on the audio attention intensity.
[0083] Specifically, the server can adjust each text feature value in the text extraction features by using the feature attention intensity to obtain adjusted text features. For example, a linear operation can be performed on the text feature value and the feature attention intensity to obtain a text feature value after the linear operation, and the adjusted text features are obtained based on each text feature value after the linear operation. Among them, the linear operation can include at least one of addition operation or multiplication operation. For example, the server can perform a multiplication operation on the feature attention intensity and each feature value in the text extraction features respectively to obtain each feature value product, sort the feature value products according to the positions of the feature values in the text extraction features to obtain a feature value sequence, and use the feature value sequence as the adjusted text features. The position of the text feature value in the text extraction features is the same as the position of the feature value product calculated from this text feature value in the feature value sequence. For example, assuming that the text extraction feature is a vector [a1, a2, a3], then a1, a2, and a3 are the feature values in the text extraction features. When the feature attention intensity is c, the feature value sequence is a vector [a1*c, a2*c, a3*c], a1*c, a2*c, and a3*c are the feature value products, and the position of a1*c in the feature value sequence [a1*c, a2*c, a3*c] is the same as the position of a1 in the text extraction feature [a1, a2, a3].
[0084] In some embodiments, the server can adjust the text extraction feature matrix by using the normalized first product matrix to obtain a first adjusted text extraction feature matrix. The normalized first product matrix includes the image correlation degrees corresponding to each text extraction feature respectively, and the first adjusted text extraction feature matrix can include the first adjusted text features corresponding to each text extraction feature respectively. For example, the server can perform a matrix multiplication operation on the normalized first product matrix and the transpose matrix of the text extraction feature matrix, and use the transpose matrix of the obtained matrix as the first adjusted text extraction feature matrix. For example, the first adjusted text extraction feature matrix feature can be calculated by using formula (5). fusion1 , where feature fusion1 represents the first adjusted text extraction feature matrix, [feature fusion1 T represents the transpose matrix of feature fusion1 . Similarly, the server can perform a matrix multiplication operation on the normalized second product matrix and the transpose matrix of the text extraction feature to obtain a second adjusted text extraction feature matrix. The normalized second product matrix includes the audio correlation degrees corresponding to each text extraction feature respectively, and the second adjusted text extraction feature matrix can include the second adjusted text features corresponding to each text extraction feature respectively. For example, the second adjusted text extraction feature matrix feature can be calculated by using formula (6), where feature fusion2 , where, feature audio is the target audio feature matrix. [feature audio T represents the transpose matrix corresponding to the target audio feature matrix.
[0085]
[0086] S210. Based on the adjusted text features, perform recognition to obtain the content recognition result corresponding to the target content.
[0087] Among them, the content recognition result is the result obtained by performing recognition based on the adjusted text features. The content recognition result can be determined according to the content recognition network used during recognition. Different content recognition networks may result in the same or different content recognition results. The content recognition network may include at least one of a scene recognition network or an entity recognition network. The scene recognition network is used to recognize scenes, and the entity recognition network is used to recognize entities. When the content recognition network is a scene recognition network, the content recognition model can also be called a scene recognition model. When the content recognition network is an entity recognition network, the content recognition model can also be called an entity recognition model or an entity word recognition model.
[0088] Specifically, the server can input the adjusted text features into the content recognition network of the trained content recognition model, and use the content recognition model to perform recognition on the adjusted text features to obtain the content recognition result corresponding to the target content. For example, when the text extraction feature is the feature corresponding to the target word segmentation in the target text, the adjusted text features corresponding to each target word segmentation can be sorted in the order of the target word segmentation in the target text, and the sorted sequence can be used as the feature sequence. The server can perform recognition based on the feature sequence to obtain the content recognition result corresponding to the target content. For example, the feature sequence can be input into the content recognition network in the content recognition model to obtain the content recognition result. For example, when the content recognition network is an entity recognition network, the entity words included in the target content can be recognized.
[0089] Such as Figure 4 As shown, a content recognition model 400 is presented. The content recognition model 400 includes a text feature extraction network, a correlation feature extraction network, an attention intensity calculation module, a feature adjustment module, and a content recognition network. Among them, the attention intensity calculation module is used to perform a correlation calculation on the correlation extraction features and the text extraction features to obtain the feature attention intensity. The feature adjustment module is used to adjust the text extraction features based on the feature attention intensity to obtain adjusted text features, and input the adjusted text features into the content recognition network to obtain the content recognition result corresponding to the target content. Each network and module in the content recognition model 400 can be obtained through joint training. The server obtains the target text and text correlation data from the target, inputs the target text into the text feature extraction network to obtain text extraction features, inputs the associated text data into the correlation feature extraction network to obtain correlation extraction features, inputs the text extraction features and the correlation extraction features into the attention intensity calculation module to obtain the feature attention intensity, inputs the feature attention intensity and the text extraction features into the feature adjustment module to obtain adjusted text features, and inputs the adjusted text features into the content recognition network to obtain the content recognition result.
[0090] In some embodiments, the server can also fuse the adjusted text features and the text extraction features to obtain fused text features. For example, statistical operations can be performed on the adjusted text features and the text extraction features, such as weighted calculation or mean calculation, to obtain fused text features. For example, the server can determine the adjusted feature weights corresponding to the adjusted text features, and fuse the adjusted text features and the text extraction features based on the adjusted feature weights to obtain fused text features. The server can perform recognition based on the fused text features to obtain the content recognition result corresponding to the target content.
[0091] In some embodiments, the adjusted text features include first adjusted text features and second adjusted text features. The server can fuse the first adjusted text features, the second adjusted text features, and the text extraction features, such as weighted calculation or mean calculation, to obtain fused text features. For example, the adjusted feature weights can include first feature weights corresponding to the first adjusted text features and second feature weights corresponding to the second adjusted text features. The server can fuse the first adjusted text features and the text extraction features based on the first feature weights to obtain first fused features, fuse the second adjusted text features and the text extraction features based on the second feature weights to obtain second fused features, perform statistical operations on the first fused features and the second fused features, and use the result of the statistical operation as the fused text features. For example, the feature values at the corresponding positions in the first fused features and the second fused features are added together to obtain respective sums, and the sums are sorted according to the positions of the feature values in the first fused features or the second fused features, and the sorted sequence is used as the fused text features.
[0092] In the above content recognition method, the target content to be recognized is determined, the target text in the target content and the text association data associated with the target text are obtained, feature extraction is performed on the target text to obtain text extraction features, feature extraction is performed on the text association data to obtain association extraction features, association calculation is performed on the association extraction features and the text extraction features, and based on the calculated feature association degree, the feature attention intensity corresponding to the text extraction features is obtained. The feature association degree is positively correlated with the feature attention intensity. The text extraction features are adjusted based on the feature attention intensity to obtain adjusted text features, and recognition is performed based on the adjusted text features to obtain the content recognition result corresponding to the target content. Since the feature association degree is positively correlated with the feature attention intensity, the greater the association relationship between the target text and the text association data, the greater the feature attention intensity and the more attention is paid to the text extraction features. Therefore, when recognition is performed based on the adjusted text features, the greater the association relationship between the target text and the text association data, the greater the influence degree of the text association data on the recognition result. Thus, the text features can be adjusted adaptively according to the relationship between the text association data and the target text, so that more attention is paid to important text features during content recognition, improving the accuracy of content recognition.
[0093] In some embodiments, recognition is performed based on the adjusted text features, and obtaining the content recognition result corresponding to the target content includes: fusing the adjusted text features and the text extraction features to obtain fused text features; and performing recognition based on the fused text features to obtain the content recognition result corresponding to the target content.
[0094] Among them, the fused text features are the features obtained by fusing the adjusted text features and the text extraction features. The dimensions of the fused text features, the adjusted text features, and the text extraction features can be the same. For example, they can be vectors or matrices of the same dimension.
[0095] Specifically, the server can perform statistical operations on the adjusted text features and the text extraction features, such as mean operation or summation operation, and use the result of the statistical operation as the fused text features. For example, the server can encode the text extraction features to obtain the encoded features corresponding to the text extraction features as the first encoded features, encode the adjusted text features to obtain the encoded features corresponding to the adjusted text features as the second encoded features, and perform statistical operations on the first encoded features and the second encoded features, such as mean operation or summation operation, and use the result of the operation as the fused text features.
[0096] In some embodiments, the server can input the fused text features into the content recognition network of the trained content recognition model, and use the content recognition network to recognize the fused text features to obtain the content recognition result corresponding to the target content.
[0097] In this embodiment, the adjusted text features and the text extraction features are fused to obtain fused text features, and recognition is performed based on the fused text features to obtain a content recognition result corresponding to the target content, which can improve the accuracy of content recognition.
[0098] In some embodiments, fusing the adjusted text features and the text extraction features to obtain fused text features includes: encoding the text extraction features to obtain first encoded features, encoding the adjusted text features to obtain second encoded features; fusing the first encoded features and the second encoded features to obtain fused encoded features; obtaining adjusted feature weights corresponding to the adjusted text features based on the fused encoded features; and fusing the adjusted text features and the text extraction features based on the adjusted feature weights to obtain fused text features.
[0099] Among them, the first encoded features are the features obtained by encoding the text extraction features. The second encoded features are the features obtained by encoding the adjusted text features. The fused encoded features are the features obtained by fusing the first encoded features and the second encoded features. The adjusted feature weights are obtained based on the fused encoded features.
[0100] Specifically, the content recognition model may further include a first encoder, a second encoder, and a feature fusion module. The feature fusion module is used to fuse the first encoded features and the second encoded features to obtain fused encoded features. The server may input the text extraction features into the first encoder in the trained content recognition model for encoding to obtain first encoded features, input the adjusted text features into the second encoder in the trained content recognition model for encoding to obtain second encoded features, and fuse the first encoded features and the second encoded features. For example, the first encoded features and the second encoded features may be input into the feature fusion module to obtain fused encoded features. Among them, the first encoder and the second encoder may be neural networks based on artificial intelligence, and each network and module in the content recognition model may be jointly trained. For example, the first encoder and the second encoder are jointly trained.
[0101] In some embodiments, the server may perform statistical operations on the first encoded features and the second encoded features to obtain encoded statistical features. For example, the first encoded features and the second encoded features may be added, and the result of the addition operation may be used as the fused encoded features, or the mean operation may be performed on the first encoded features and the second encoded features, and the calculated mean may be used as the fused encoded features. The server may determine the fused encoded features based on the encoded statistical features. For example, the encoded statistical features may be used as the fused encoded features.
[0102] In some embodiments, the server may normalize the fused encoded features and use the result of the normalization as the adjusted feature weight corresponding to the adjusted text features. For example, the trained content recognition model may include an activation layer, and the activation layer may convert the data into data between 0 and 1. The steps of normalizing the fused encoded features and using the result of the normalization as the adjusted feature weight may include: inputting the fused encoded features into the activation layer of the content recognition model for activation processing, and using the result of the activation processing as the adjusted feature weight corresponding to the adjusted text features.
[0103] In some embodiments, the server may perform a multiplication calculation on the adjusted feature weight and the adjusted text features to obtain the calculated adjusted text features, and perform a statistical operation on the calculated adjusted text features and the text extraction features, such as an addition operation or a mean operation, and use the result of the statistical operation as the fused text features.
[0104] In some embodiments, the server may determine the text feature weight corresponding to the text extraction features. For example, a preset weight may be obtained and used as the text feature weight. The preset weight may be a weight set in advance according to needs. The text feature weight may also be determined based on the adjusted feature weight. For example, the adjusted feature weight and the text feature weight may be negatively correlated, and the sum of the adjusted feature weight and the text feature weight may be a preset value. The preset value may be set in advance according to needs, for example, it may be 1. For example, the result obtained by subtracting the text feature weight from the preset value may be used as the text feature weight. For example, when the adjusted feature weight is 0.3, the text feature weight may be 0.7. Here, the preset value is greater than the text feature weight and greater than the adjusted feature weight. The negatively correlated relationship means that: under other unchanged conditions, the change directions of two variables are opposite. When one variable changes from large to small, the other variable changes from small to large. It can be understood that the negatively correlated relationship here refers to the opposite change directions, but it does not require that when one variable changes a little, the other variable must also change.
[0105] In some embodiments, the first encoder may include at least one of a first text encoder or a second text encoder, and the second encoder may include at least one of an image encoder or an audio encoder. The first encoded feature may include at least one of a first text feature or a second text feature. The first text feature is a feature obtained by encoding text extraction features using the first text encoder, and the second text feature is a feature obtained by encoding text extraction features using the second text encoder. The second encoded feature may include at least one of an image encoded feature or an audio encoded feature. The image encoded feature is a feature obtained by encoding the first adjusted text features using the image encoder, and the audio encoded feature is a feature obtained by encoding the second adjusted text features using the audio encoded feature. The fused encoded feature may include at least one of a text-image encoded feature or a text-audio encoded feature. The text-image encoded feature is a feature obtained by fusing the first text encoded feature and the image encoded feature. The text-audio encoded feature is a feature obtained by fusing the second text encoded feature and the audio encoded feature. For example, when the adjusted text feature is the first adjusted text feature, the server may input the text extraction features into the first text encoder for encoding to obtain the first text feature, input the first adjusted text features into the image encoder for encoding to obtain the image encoded feature, and fuse the first text feature and the image encoded feature to obtain the text-image encoded feature. When the adjusted text feature is the second adjusted text feature, the server may input the text extraction features into the second text encoder for encoding to obtain the second text feature, input the second adjusted text features into the audio encoder for encoding to obtain the audio encoded feature, and fuse the second text feature and the audio encoded feature to obtain the text-audio encoded feature. The text-image encoded feature and the text-audio encoded feature may be used as the fused encoded feature. Among them, the first text encoder and the second text encoder may be the same encoder or different encoders, and the image encoder and the audio encoder may be the same encoder or different encoders.
[0106] In this embodiment, the text extraction features are encoded to obtain the first encoded feature, the adjusted text features are encoded to obtain the second encoded feature, the first encoded feature and the second encoded feature are fused to obtain the fused encoded feature, the adjusted feature weight corresponding to the adjusted text features is obtained based on the fused encoded feature, and the adjusted text features and the text extraction features are fused based on the adjusted feature weight to obtain the fused text features. Thus, the fused text features can reflect both the text extraction features and the adjusted text features, improving the expression ability of the fused text features. When recognition is performed based on the adjusted text features, the recognition accuracy can be improved.
[0107] In some embodiments, the first encoded feature is encoded by a first encoder in a trained content recognition model, and the second encoded feature is encoded by a second encoder in the content recognition model. Obtaining the adjusted feature weights corresponding to the adjusted text features based on the fused encoded features includes: inputting the fused encoded features into a target activation layer in the content recognition model for activation processing to obtain a target activation value, and using the target activation value as the adjusted feature weights corresponding to the adjusted text features. The activation layer is a shared activation layer of the first encoder and the second encoder.
[0108] Among them, the activation layer is used to convert data into data between 0 and 1, which can be implemented by an activation function. The activation function includes but is not limited to at least one of the Sigmoid function, the tanh function, or the Relu function. The target activation layer is the activation layer in the trained content recognition model and is a shared activation layer of the first encoder and the second encoder, that is, the target activation layer set can receive the output data of the first encoder and can also receive the output data of the second encoder. The target activation value is the result obtained by activating the fused encoded features using the target activation layer. The target activation value and the dimension of the fused encoded features can be the same. For example, they can be vectors or matrices of the same dimension. As Figure 5 shown, a content recognition module 500 is shown. The content recognition module 500 includes an associated feature extraction network, an attention intensity calculation module, a text feature extraction network, a feature adjustment module, a first encoder, a second encoder, a feature fusion module, a target activation layer, a fused text feature generation module, and a content recognition network. The feature fusion module is used to fuse the first encoded feature and the second encoded feature to obtain a fused encoded feature. The fused text feature generation is used to fuse the adjusted text features and the text extraction features based on the adjusted feature weights to obtain fused text features.
[0109] Specifically, the target activation layer can include at least one of a first activation layer shared by the first text encoder and the image encoder and a second activation layer shared by the second text encoder and the audio encoder. The target activation value can include at least one of a first activation value obtained by activating the text image encoded features or a second activation value obtained by activating the text audio encoded features. When the fused encoded feature is the text image encoded feature, the server can input the text image encoded features into the first activation layer for activation to obtain a first activation value, and use the first activation value as the first feature weights corresponding to the first adjusted text features; when the fused encoded feature is the text audio encoded feature, the server can input the text audio encoded features into the second activation layer for activation to obtain a second activation value, and use the second activation value as the second feature weights corresponding to the second adjusted text features, and use the first feature weights and the second feature weights as the adjusted feature weights.
[0110] In some embodiments, when the adjusted text feature is the first adjusted text feature and there are multiple first adjusted text features, the server may perform matrix fusion on the first adjusted text feature matrix and the text extraction feature matrix. For example, the text extraction feature matrix may be input into the first text encoder for encoding to obtain the first matrix encoded feature, and the first adjusted text feature matrix may be input into the image encoder for encoding to obtain the second matrix encoded feature. Statistical operations are performed on the first matrix encoded feature and the second matrix encoded feature to obtain the first matrix feature statistical result. The first matrix feature statistical result is normalized. For example, the first matrix feature statistical result may be input into the first activation layer for activation to obtain the normalized first matrix feature statistical result. The normalized first matrix feature statistical result may include the first feature weights corresponding to each of the first adjusted text features. For example, the normalized first matrix feature statistical result gate1 may be calculated using formula (7). Where gate1 represents the normalized first matrix feature statistical result, sigmoid is the activation function, W1 T is the model parameter of the first text encoder, is the model parameter of the image encoder.
[0111]
[0112] In some embodiments, when the adjusted text feature is the second adjusted text feature and there are multiple second adjusted text features, the server may perform matrix fusion on the second adjusted text feature matrix and the text extraction feature matrix. For example, the text extraction feature matrix may be input into the second text encoder for encoding to obtain the third matrix encoded feature, and the second adjusted text feature matrix may be input into the audio encoder for encoding to obtain the fourth matrix encoded feature. Statistical operations are performed on the third matrix encoded feature and the fourth matrix encoded feature to obtain the second matrix feature statistical result. The second matrix feature statistical result is normalized. For example, the first matrix feature statistical result may be input into the second activation layer for activation to obtain the normalized second matrix feature statistical result. The normalized second matrix feature statistical result may include the second feature weights corresponding to each of the second adjusted text features. For example, the normalized second matrix feature statistical result gate2 may be calculated using formula (8). Where gate2 represents the normalized second matrix feature statistical result, is the model parameter of the second text encoder, is the model parameter of the audio encoder.
[0113]
[0114] In this embodiment, the fused coding features are input into the target activation layer of the content recognition model for activation processing to obtain target activation values, and the target activation values are used as the adjusted feature weights corresponding to the adjusted text features, so that the adjusted feature weights are normalized values, improving the rationality of the adjusted feature weights.
[0115] In some embodiments, fusing the adjusted text features and the text extraction features based on the adjusted feature weights to obtain fused text features includes: obtaining text feature weights corresponding to the text extraction features based on the adjusted feature weights; calculating the product of the adjusted feature weights and the adjusted text features to obtain the calculated adjusted text features; calculating the product of the text feature weights and the text extraction features to obtain the calculated text extraction features; adding the calculated adjusted text features and the calculated text extraction features to obtain fused text features.
[0116] Among them, the text feature weights can be determined according to the adjusted feature weights, and the text feature weights can be negatively correlated with the adjusted feature weights. For example, the result obtained by subtracting the text feature weights from a preset value can be used as the text feature weights. The preset value is greater than the text feature weights and greater than the adjusted feature weights.
[0117] Specifically, the server can use the result of multiplying the adjusted feature weights and the adjusted text features as the calculated adjusted text features, use the result of multiplying the text feature weights and the text extraction features as the calculated text extraction features, and use the result obtained by adding the calculated adjusted text features and the calculated text extraction features as the fused text features.
[0118] In some embodiments, adjusting the feature weights includes a first feature weight and a second feature weight. The text feature weights may include a first text weight obtained based on the first feature weight and a second text weight obtained based on the second feature weight. The first text weight is negatively correlated with the first feature weight. The second text weight is negatively correlated with the second feature weight. The server may use the first feature weight, the second feature weight, the first text weight, and the second text weight to perform weighted calculations on the first adjusted text feature, the second adjusted text feature, and the text extraction feature, and use the result of the weighted calculation as the fused text feature. For example, the server may use the first feature weight and the first text weight to perform weighted calculations on the first adjusted text feature and the text extraction feature to obtain a first weighted value, use the second feature weight and the second text weight to perform weighted calculations on the second adjusted text feature and the text extraction feature to obtain a second weighted value, and use the sum of the first weighted value and the second weighted value as the fused text feature. Specifically, the server may calculate the product of the first text weight and the text extraction feature to obtain a first product value, calculate the product of the first feature weight and the first adjusted text feature to obtain a second product value, calculate the product of the second text weight and the text extraction feature to obtain a third product value, calculate the product of the second feature weight and the second adjusted text feature to obtain a fourth product value, add the first product value, the second product value, the third product value, and the fourth product value, and use the added result as the fused text feature.
[0119] In some embodiments, the server may use the normalized first matrix feature statistical result and the normalized second matrix feature statistical result to perform weighted calculations on the first adjusted text feature matrix and the second adjusted text feature matrix to obtain a fused text feature matrix, and the fused text feature matrix may include the fused text features corresponding to the respective text extraction features. For example, the fused text feature matrix output may be calculated using formula (9). Where output refers to the fused text feature matrix.
[0120] output = feature fusion1 ·gate1+(1 - gate1)feature text +feature fusion2 ·gate2+(1 - gate2)feature text (9)
[0121] In this embodiment, the adjusted feature weights are multiplied by the adjusted text features to obtain the calculated adjusted text features. The text feature weights are multiplied by the text extraction features to obtain the calculated text extraction features. The calculated adjusted text features and the calculated text extraction features are added together to obtain the fused text features. Since the text feature weights are obtained based on the adjusted feature weights, the accuracy of the text feature weights is improved, thereby improving the accuracy of the fused text features.
[0122] In some embodiments, the target content is a target video; obtaining the target text in the target content and the text association data associated with the target text includes: obtaining the text corresponding to the target time in the target video to obtain the target text; obtaining the video-related data corresponding to the target time in the target video, and using the video-related data as the text association data associated with the target text. The video-related data includes at least one of video frames or audio frames.
[0123] Among them, a video frame is the smallest unit that makes up a video. A video is composed of multiple images, and an image in the video is called a frame, or can also be called a video frame. The target video can be any video, which can be a directly shot video or a video clip intercepted from a shot video. The target video can be any type of video, including but not limited to at least one of advertising videos, TV drama videos, or news videos. The target video can also be a video to be pushed to users. The target time can be any time point or time period from the start time point to the end time point of the target video. The video-related data refers to any data displayed or played at the target time in the target video, which can include at least one of the video frames displayed at the target time in the target video or the audio frames played at the target time. The video frames displayed at the target time can include one or more frames, and the audio frames played at the target time can include one or more frames.
[0124] Specifically, the server can obtain the text displayed at the target time in the target video as the target text, such as at least one of the subtitles, bullet screens, or comments displayed at the target time, as the target text. The server can obtain at least one of the video frames displayed at the target time in the target video or the audio frames played at the target time as the video-related data.
[0125] In this embodiment, obtaining the text corresponding to the target time in the target video to obtain the target text, and obtaining the video-related data corresponding to the target time in the target video, and using the video-related data as the text association data associated with the target text. Since the video-related data includes at least one of video frames or audio frames, text data and image data or audio data other than text data are obtained, so that the video can be recognized by combining image data or audio data on the basis of the text data, which is conducive to improving the recognition accuracy.
[0126] In some embodiments, adjusting text features includes a first adjusted text feature obtained by adjusting according to a video frame and a second adjusted text feature obtained by adjusting according to an audio frame; based on the adjusted text features for recognition, obtaining a content recognition result corresponding to the target content includes: fusing the first adjusted text feature, the second adjusted text feature, and the text extraction feature to obtain a fused text feature; based on the fused text feature for recognition, obtaining a content recognition result corresponding to the target content.
[0127] Specifically, the server can obtain a video frame from text-associated data, extract features from the obtained video frame to obtain target image features, obtain a first adjusted text feature based on the target image features, obtain an audio frame from the text-associated data, extract features from the obtained audio frame to obtain target audio features, and obtain a second adjusted text feature based on the target audio features.
[0128] In some embodiments, the server can perform a weighted calculation on the first adjusted text feature, the second adjusted text feature, and the text extraction feature, and use the result of the weighted calculation as the fused text feature. For example, the server can perform a multiplication calculation on the first text weight and the text extraction feature to obtain a first product value, perform a multiplication calculation on the first feature weight and the first adjusted text feature to obtain a second product value, perform a multiplication calculation on the second text weight and the text extraction feature to obtain a third product value, perform a multiplication calculation on the second feature weight and the second adjusted text feature to obtain a fourth product value, add the first product value, the second product value, the third product value, and the fourth product value, and use the added result as the fused text feature.
[0129] In this embodiment, the first adjusted text feature, the second adjusted text feature, and the text extraction feature are fused to obtain a fused text feature, so that the fused text feature is obtained based on these three features: the first adjusted text feature, the second adjusted text feature, and the text extraction feature, thereby improving the feature richness of the fused text feature. Therefore, when recognition is performed based on the fused text feature, the recognition accuracy can be improved.
[0130] In some embodiments, adjusting the text extraction feature based on the feature attention intensity to obtain an adjusted text feature includes: multiplying the feature attention intensity by each feature value of the text extraction feature to obtain a feature value product; arranging the feature value products according to the positions of the feature values in the text extraction feature, and using the arranged feature value sequence as the adjusted text feature.
[0131] Among them, the product of feature values refers to the result obtained by multiplying the text feature value by the feature attention intensity. The feature value sequence is obtained by arranging the products of feature values calculated from the text feature values according to the positions of the text feature values in the text extraction features. That is, the position of the text feature value in the text extraction features is the same as the position of the product of the feature value calculated from this text feature value in the feature value sequence.
[0132] In this embodiment, the feature attention intensity is multiplied by each feature value of the text extraction features to obtain the product of feature values. Thus, the product of feature values can reflect the attention degree of the text associated data to the text feature values. The products of feature values are arranged according to the positions of the feature values in the text extraction features, and the arranged feature value sequence is used as the adjusted text features. Thus, the adjusted text features can reflect the attention degree of the text associated data to the text extraction features.
[0133] In some embodiments, the text extraction features are the features corresponding to the word segments in the target text; each adjusted text feature forms a feature sequence according to the order of the word segments in the target text; the content recognition result corresponding to the target content obtained by recognition based on the adjusted text features includes: obtaining the position relationship of each word segment relative to the named entity based on the feature sequence; obtaining the target named entity from the target text based on each position relationship, and taking the target named entity as the content recognition result corresponding to the target content.
[0134] Among them, the feature sequence is a sequence obtained by sorting the adjusted text features corresponding to the target word segments according to the order of the target word segments in the target text. The target word segment refers to the word segment in the target text. A named entity refers to an entity identified by a name, which may include at least one of a person name, a place name, or an organization name. For example, the named entity may be "Zhang San", "Area A", or "Organization B".
[0135] The position relationship relative to the named entity may include at least one of the named entity position or the non-named entity position. The named entity position refers to the position where the named entity is located, which may include at least one of the start position of the named entity, the end position of the named entity, or the middle position of the named entity. The middle position of the named entity may include each position between the start position and the end position of the named entity. The non-named entity position refers to the position where the word segment other than the named entity is located.
[0136] Specifically, the server may determine the position relationship of each target word segment relative to the named entity based on the feature sequence, obtain the position relationship corresponding to each target word segment respectively, and from each position relationship, obtain the target word segment corresponding to the position relationship belonging to the named entity position as the entity word segment, and obtain the target named entity based on each entity word segment.
[0137] In some embodiments, the trained content recognition model may include an entity recognition network. The server may input the feature sequence into the entity recognition network and use the entity recognition network to perform position recognition on each adjusted text feature in the feature sequence. For example, the entity recognition network may determine the probability that the target word segment corresponding to the adjusted text feature is in the named entity position based on the adjusted text feature, obtain the named entity probability, and determine the position relationship of the target word segment with the named entity probability greater than the named entity probability threshold as the named entity position. The named entity probability threshold can be set as needed. The entity recognition network may also determine the probability that the target word segment corresponding to the adjusted text feature is in the starting position of the named entity based on the adjusted text feature, obtain the starting probability, and determine the position relationship of the target word segment with the starting probability greater than the starting probability threshold as the starting position of the named entity. The starting probability threshold can be set as needed. The entity recognition network may also determine the probability that the target word segment corresponding to the adjusted text feature is in the ending position of the named entity based on the adjusted text feature, obtain the ending probability, and determine the position relationship of the target word segment with the ending probability greater than the ending probability threshold as the ending position of the named entity. The ending probability threshold can be set as needed.
[0138] In this embodiment, the position relationship of each word segment relative to the named entity is obtained based on the feature sequence, and the target named entity is obtained from the target text based on each position relationship. The target named entity is used as the content recognition result corresponding to the target content. Thus, content recognition can be performed based on the feature sequence formed by the adjusted text features, improving the accuracy of content recognition.
[0139] In some embodiments, obtaining the target named entity from the target text based on each position relationship includes: obtaining the word segment with the position relationship as the starting position of the named entity as the starting word of the named entity; taking the word segment with the position relationship as being inside the named entity among the backward word segments corresponding to the starting word of the named entity as the constituent words of the named entity; and combining the starting word of the named entity and the constituent words of the named entity to obtain the target named entity.
[0140] Among them, the starting word of a named entity refers to the segmented word at the starting position of the named entity, and the backward segmented words corresponding to the starting word of the named entity refer to the segmented words sorted after the starting word of the named entity in the target text. The component words of a named entity refer to the segmented words inside the named entity in the target text. The inside of the named entity includes the ending position and the middle position of the named entity. The ending position and the middle position of the named entity word can be the same position. For example, when the segmented word is a single character, assuming the target text is "Zhang San likes flowers", since the named entity is "Zhang San" and consists of two characters, and "Zhang" is at the starting position of the named entity, the starting word of the named entity is "Zhang". The backward segmented words corresponding to the starting word of the named entity include "San", "Xi", "Huan", and "Hua". Since "San" is inside the named entity, the component word of the named entity is "San". The target named entity is the entity included in the target text and is obtained by combining the starting word of the named entity and the corresponding component words of the named entity. The target text may include one or more target named entities. "Multiple" means at least two. For example, assuming the target text is "Zhang San likes Li Si", then the target text includes 2 target named entities, namely "Zhang San" and "Li Si".
[0141] Specifically, the server can, based on the positional relationships corresponding to each target segmented word, obtain the segmented word with the positional relationship of the starting position of the named entity from the target text as the starting word of the named entity. In the order from front to back, it sequentially obtains one backward segmented word from each of the backward segmented words of the starting word of the named entity as the current backward segmented word. When the positional relationship of the current backward segmented word is inside the named entity, the current backward segmented word is used as the component word of the named entity corresponding to the starting word of the named entity. When the positional relationship of the current backward segmented word is outside the named entity, it stops obtaining backward segmented words from the backward segmented words of the starting word of the named entity. According to the positions of the starting word of the named entity and the component words of the named entity in the target text, it sorts the starting word of the named entity and the component words of the named entity from front to back to obtain the target named entity. For example, since the position of "Zhang" is before "San", the sorted result is "Zhang San", that is, "Zhang San" is the target named entity.
[0142] In this embodiment, the segmented word with the positional relationship of the starting position of the named entity is obtained as the starting word of the named entity. Among the backward segmented words corresponding to the starting word of the named entity, the segmented words with the positional relationship inside the named entity are used as the component words of the named entity. The starting word of the named entity and the component words of the named entity are combined to obtain the target named entity, so that entity recognition can be performed based on the feature sequence formed by adjusting the text features, improving the accuracy of entity recognition.
[0143] In some embodiments, obtaining the positional relationship of each word segment relative to a named entity based on the feature sequence includes: obtaining the positional relationship of each word segment relative to the named entity and the entity type corresponding to the word segment based on the feature sequence; taking, as the constituent words of the named entity, the word segments whose positional relationship is inside the named entity among the backward word segments corresponding to the starting word of the named entity includes: taking, as the constituent words of the named entity, the word segments whose positional relationship is inside the named entity and whose entity type is the same as the type of the starting word of the named entity among the backward word segments corresponding to the starting word of the named entity.
[0144] Among them, the entity type refers to the type of the named entity, including at least one of the types of person names, organization names, or place names. The starting word of the named entity and the constituent words of the named entity can each correspond to an entity type.
[0145] Specifically, the server can identify the entity type of each feature in the feature sequence, determine the entity type corresponding to each feature in the feature sequence, and sequentially obtain a backward word segment from each backward word segment of the starting word of the named entity in the order from front to back as the current backward word segment. When the positional relationship of the current backward word segment is inside the named entity and the entity type is the same as the entity type of the starting word of the named entity, the current backward word segment is taken as the constituent word of the named entity corresponding to the starting word of the named entity. When the positional relationship of the current backward word segment is outside the named entity or the entity type is different from the entity type of the starting word of the named entity, stop obtaining backward word segments from each backward word segment of the starting word of the named entity.
[0146] In some embodiments, the text extraction feature is the feature corresponding to the target word segment in the target text; the fused text features corresponding to each target word segment form a fused feature sequence in the order of the target word segments in the target text; performing recognition based on the adjusted text feature to obtain the content recognition result corresponding to the target content includes: obtaining the positional relationship of each word segment relative to the named entity based on the fused feature sequence; obtaining the target named entity from the target text based on each positional relationship, and taking the target named entity as the content recognition result corresponding to the target content.
[0147] In some embodiments, the fused feature sequence can be input into an entity recognition network, and the entity recognition network performs entity word recognition on each fused text feature in the fused feature sequence. The entity recognition network can be, for example, Figure 6 the CRF network in Figure 6Among them, the target text is "Zhang Xiaohua loves to laugh", the fusion feature sequence is [h1, h2, h3, h4, h5], h1 is the fusion text feature corresponding to the word segmentation "Zhang", h2 is the fusion text feature corresponding to the word segmentation "Xiao", h3 is the fusion text feature corresponding to the word segmentation "Hua", h4 is the fusion text feature corresponding to the word segmentation "loves", and h5 is the fusion text feature corresponding to the word segmentation "laugh". The fusion feature sequence is input into the CRF network for entity recognition. The CRF network can score the word segmentations in the target text based on each feature in the fusion feature sequence to obtain the scores corresponding to each word segmentation. The softmax can be used to normalize the scores of the word segmentations to obtain the probability distribution corresponding to the word segmentations. The CRF network is used to identify the position of the person name in "Zhang Xiaohua loves to laugh". The CRF network can use the "BIO" annotation method to annotate each target word segmentation in "Zhang Xiaohua loves to laugh" to obtain the annotations corresponding to each fusion text feature. Among them, B is the abbreviation of begin, indicating the start of the entity word, I is the abbreviation of inside, indicating the inside of the entity word, and O is the abbreviation of outside, indicating the outside of the entity word. As shown in the figure, the annotation of "Zhang Xiaohua loves to laugh" is "B-PER, I-PER, I-PER, O, O", where "PER" indicates that the entity word type is a person name. From "B-PER, I-PER, I-PER, O, O", it can be determined that "Zhang Xiaohua" in "Zhang Xiaohua loves to laugh" is the target named entity.
[0148] In this embodiment, based on the feature sequence, the position relationship of each word segmentation relative to the named entity and the entity type corresponding to the word segmentation are obtained. Among the backward word segmentations corresponding to the starting word of the named entity, the word segmentations with the position relationship being inside the named entity and the entity type being the same as that of the starting word of the named entity are used as the constituent words of the named entity, which improves the accuracy of entity recognition.
[0149] In some embodiments, the correlation calculation is performed on the associated extraction feature and the text extraction feature, and the feature attention intensity corresponding to the text extraction feature is obtained based on the calculated feature correlation degree, including: performing a multiplication operation on the associated feature value in the associated extraction feature and the text feature value at the corresponding position in the text extraction feature to obtain a multiplication operation value; performing statistics on the multiplication operation value to obtain the feature correlation degree between the associated extraction feature and the text extraction feature, and using the feature correlation degree as the feature attention intensity corresponding to the text extraction feature.
[0150] Specifically, the associated extraction features can be vectors or matrices of the same dimension as the text extraction features. The server can obtain the associated feature value at the target sorting position from the associated extraction features as the first target feature value, and obtain the text feature value at the target position from the text extraction features as the second target feature value. Then, the second target feature value has a position correspondence with the second target feature value. The server can perform a multiplication operation on the first target feature value and the second target feature value to obtain a multiplication operation value calculated from the text feature value and the associated feature value at the target position. The target position can be any position in the associated extraction features or the text extraction features. For example, when the associated extraction features are vectors, the target position can be any sorting position, such as the first position.
[0151] In some embodiments, the server can perform statistics on each multiplication operation value to obtain a product statistic value, and perform normalization processing on the product statistic value, and use the result of the normalization processing as the feature correlation degree.
[0152] In this embodiment, the associated feature value in the associated extraction features is multiplied by the text feature value at the corresponding position in the text extraction features to obtain a multiplication operation value. Statistical operations are performed on each multiplication operation value to obtain the feature correlation degree between the associated extraction features and the text extraction features. The feature correlation degree is used as the feature attention intensity corresponding to the text extraction features. Thus, the feature attention intensity can accurately reflect the association relationship between the text association data and the target text. Therefore, when adjusting the text extraction features based on the feature attention intensity, the accuracy of the adjustment can be improved.
[0153] In some embodiments, a content recognition method is provided, including the following steps:
[0154] 1. Determine the target video to be recognized, and obtain the target text in the target video, as well as the target image data and target audio data associated with the target text.
[0155] 2. Extract features from the target text to obtain text extraction features, extract features from the target image data to obtain target image features, and extract features from the target audio data to obtain target audio features.
[0156] Specifically, as Figure 7 shown, a trained entity recognition network 700 is shown. The server can use the text feature extraction network in the trained entity recognition model to extract features from the target text to obtain text extraction features. Similarly, the server can use the image feature extraction network to extract features from the target image data to obtain target image features, and use the audio feature extraction network to extract features from the target audio data to obtain target audio features.
[0157] 3. Perform an association calculation on the target image features and the text extraction features to obtain an image association degree, and use the image association degree as the image attention intensity. Perform an association calculation on the target audio features and the text extraction features to obtain an audio association degree, and use the audio association degree as the audio attention intensity.
[0158] Specifically, as Figure 7 shown, an image attention intensity calculation module can be used to perform an association calculation on the target image features and the text extraction features to obtain the image attention intensity. An audio attention intensity calculation module can be used to perform an association calculation on the target audio features and the text extraction features to obtain the audio attention intensity. The image attention intensity calculation module includes a product operation unit and a normalization processing unit. The image attention intensity calculation module can perform a product operation on the target image features and the text extraction features through the product operation unit, and input the result of the operation into the normalization operation unit for normalization processing to obtain the image attention intensity. The process of the audio attention intensity calculation module calculating the audio attention intensity can refer to the image attention intensity calculation module and will not be elaborated here.
[0159] 4. Adjust the text extraction features based on the image attention intensity to obtain the first adjusted text features, and adjust the text extraction features based on the audio attention intensity to obtain the second adjusted text features.
[0160] Specifically, as Figure 7 shown, the image attention intensity and the text extraction features can be input into the first feature adjustment module. The first feature adjustment module can multiply the image attention intensity by each feature value of the text extraction features, and arrange the obtained values according to the positions of the feature values in the text extraction features to obtain the first adjusted text features. Similarly, the second adjusted text features can be obtained using the second feature adjustment module.
[0161] 5. Determine the first feature weight corresponding to the first adjusted text features, and determine the second feature weight corresponding to the second adjusted text features.
[0162] Specifically, as Figure 7As shown, the server can input the first adjusted text feature into an image encoder for encoding to obtain an image encoding feature, input the text extraction feature into a first text encoder for encoding to obtain a first text feature, and input the first text feature and the image encoding feature into a first feature fusion module to obtain a text-image encoding feature. The server can input the second adjusted text feature into an audio encoder for encoding to obtain an audio encoding feature, input the text extraction feature into a second text encoder for encoding to obtain a second text feature, input the second text feature and the audio encoding feature into a second feature fusion module to obtain a text-audio encoding feature, input the text-audio encoding feature into a second activation layer for activation to obtain a second feature weight corresponding to the second adjusted text feature.
[0163] 6. Fusion of the first adjusted text feature and the text extraction feature is performed based on the first feature weight to obtain a first fusion feature, fusion of the second adjusted text feature and the text extraction feature is performed based on the second feature weight to obtain a second fusion feature, a statistical operation is performed on the first fusion feature and the second fusion feature, and the result of the statistical operation is used as the fusion text feature.
[0164] Specifically, as Figure 7 shown, the server can input the first feature weight, the first adjusted text feature, and the text extraction feature into a first fusion text feature generation module to obtain a first fusion feature, and input the second feature weight, the second adjusted text feature, and the text extraction feature into a second fusion text feature generation module to obtain a second fusion feature.
[0165] 7. Named entity recognition is performed based on the fusion text feature to obtain a target named entity corresponding to the target content, and the target named entity is used as the content recognition result corresponding to the target content.
[0166] For example, as Figure 8 shown, the target video is a video of "Zhang Xiaohua", the target text is the subtitle "Zhang Xiaohua loves to laugh" in the video of "Zhang Xiaohua", the target image data is the image associated with the subtitle "Zhang Xiaohua loves to laugh" in the video of "Zhang Xiaohua" in terms of time, that is, the image including "Zhang Xiaohua", and the target audio data is the audio associated with the subtitle "Zhang Xiaohua loves to laugh" in the video of "Zhang Xiaohua" in terms of time, that is, the audio including "Zhang Xiaohua". Inputting the subtitle "Zhang Xiaohua loves to laugh", the image including "Zhang Xiaohua", and the audio including "Zhang Xiaohua" into an entity recognition model can determine the entity word "Zhang Xiaohua".
[0167] When performing entity recognition, the above content recognition method utilizes not only the text information in the video, such as the title, subtitles, or description information in the video, but also the audio features and image features of the video, and fuses various modal features, enabling more accurate and effective extraction of video information, enhancing the recognition effect of entity word recognition, for example, improving the accuracy and efficiency of entity word recognition. It can improve the accuracy and recall rate on the test data set. Among them, one modality can be a data type, for example, text, audio, and image are each a modality, and multi-modal includes at least two modalities. Modal features can be, for example, any one of text features, audio features, or image features. Multi-modal features include at least two modal features. The entity word recognition model (i.e., the entity recognition model) provided in this application can effectively extract video information.
[0168] This application also provides an application scenario. Applying the above content recognition method, it can perform entity recognition on the text in the video. Specifically, the application of this content recognition method in this application scenario is as follows:
[0169] Receive a video label generation request for a target video. In response to the video label generation request, use the content recognition method provided in this application to perform entity word recognition on the target video, obtain the recognized entity words, and use each recognized entity word as the video label corresponding to the target video.
[0170] The content recognition method provided in this application, when applied to video recognition, can save the time for obtaining video information and improve the efficiency of understanding video information.
[0171] This application also provides an application scenario. Applying the above content recognition method, it can perform entity recognition on the text in the video. Specifically, the application of this content recognition method in this application scenario is as follows:
[0172] Receive a video recommendation request corresponding to a target user, obtain candidate videos, use the content recognition method provided in this application to perform entity word recognition on the candidate videos, use the recognized entity words as the video labels corresponding to the candidate videos, obtain the user information corresponding to the target user, and when it is determined that the video label matches the user information, for example, when the video label matches the user's user image description, push the candidate video to the terminal corresponding to the target user.
[0173] The content recognition method provided in this application, when applied to video recommendation, can provide high-quality features for the video recommendation algorithm and optimize the video recommendation effect.
[0174] It should be understood that although Figure 2-8The steps in the flowchart are displayed in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-8 at least a part of the steps in Figure 2-8 may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0175] In some embodiments, as Figure 9 shown, a content recognition device is provided. This device can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the device includes: a target content determination module 902, a feature extraction module 904, a feature attention intensity obtaining module 906, an adjusted text feature obtaining module 908, and a content recognition result obtaining module 910, where: The target content determination module 902 is configured to determine the target content to be recognized, and obtain the target text in the target content and the text association data associated with the target text. The feature extraction module 904 is configured to extract features from the target text to obtain text extraction features; extract features from the text association data to obtain association extraction features. The feature attention intensity obtaining module 906 is configured to perform an association calculation on the association extraction features and the text extraction features, and obtain the feature attention intensity corresponding to the text extraction features based on the calculated feature association degree. The feature association degree is positively correlated with the feature attention intensity. The adjusted text feature obtaining module 908 is configured to adjust the text extraction features based on the feature attention intensity to obtain adjusted text features. The content recognition result obtaining module 910 is configured to perform recognition based on the adjusted text features to obtain the content recognition result corresponding to the target content.
[0176] In the above content recognition device, the target content to be recognized is determined, the target text in the target content and the text association data associated with the target text are obtained, feature extraction is performed on the target text to obtain text extraction features, feature extraction is performed on the text association data to obtain association extraction features, association calculation is performed on the association extraction features and the text extraction features, and based on the calculated feature association degree, the feature attention intensity corresponding to the text extraction features is obtained. The feature association degree is positively correlated with the feature attention intensity. The text extraction features are adjusted based on the feature attention intensity to obtain adjusted text features, and recognition is performed based on the adjusted text features to obtain the content recognition result corresponding to the target content. Since the feature association degree can reflect the magnitude of the association relationship between the target text and the text association data, the greater the feature association degree, the greater the association relationship between the target text and the text association data, and the smaller the feature association degree, the smaller the association relationship between the target text and the text association data. And since the feature association degree is positively correlated with the feature attention intensity, the greater the association relationship between the target text and the text association data, the greater the feature attention intensity, and the greater the degree of adjustment of the text extraction features. The smaller the association relationship between the target text and the text association data, the smaller the feature attention intensity, and the smaller the degree of adjustment of the text extraction features. Therefore, when recognizing based on the adjusted text features, the greater the association relationship between the target text and the text association data, the greater the influence degree of the text association data on the recognition result, and the smaller the association relationship between the target text and the text association data, the smaller the influence degree of the text association data on the recognition result. Thus, the features used for recognition can be adaptively adjusted according to the relationship between the text association data and the target text, improving the accuracy of the features used for recognition and the accuracy of content recognition.
[0177] In some embodiments, the content recognition result obtaining module 910 includes: a first fused text feature obtaining unit, configured to fuse the adjusted text features and the text extraction features to obtain fused text features. A first content recognition result obtaining unit, configured to perform recognition based on the fused text features to obtain the content recognition result corresponding to the target content.
[0178] In this embodiment, the adjusted text features and the text extraction features are fused to obtain fused text features, and recognition is performed based on the fused text features to obtain the content recognition result corresponding to the target content, which can improve the accuracy of content recognition.
[0179] In some embodiments, the first fused text feature obtaining unit is further configured to encode the text extraction features to obtain first encoded features, and encode the adjusted text features to obtain second encoded features; fuse the first encoded features and the second encoded features to obtain fused encoded features; obtain adjusted feature weights corresponding to the adjusted text features based on the fused encoded features; and fuse the adjusted text features and the text extraction features based on the adjusted feature weights to obtain fused text features.
[0180] In this embodiment, the text extraction features are encoded to obtain first encoded features, the adjusted text features are encoded to obtain second encoded features, the first encoded features and the second encoded features are fused to obtain fused encoded features, adjusted feature weights corresponding to the adjusted text features are obtained based on the fused encoded features, and the adjusted text features and the text extraction features are fused based on the adjusted feature weights to obtain fused text features. Thus, the fused text features can reflect both the text extraction features and the adjusted text features, improving the expression ability of the fused text features. When recognition is performed based on the adjusted text features, the recognition accuracy can be improved.
[0181] In some embodiments, the first encoded features are obtained by encoding through a first encoder in a pre-trained content recognition model, the second encoded features are obtained by encoding through a second encoder in the content recognition model, and the first fused text feature obtaining unit is further configured to input the fused encoded features into a target activation layer in the content recognition model for activation processing to obtain activation values, and use the activation values as the adjusted feature weights corresponding to the adjusted text features, where the activation layer is a shared activation layer of the first encoder and the second encoder.
[0182] In this embodiment, the fused encoded features are input into a target activation layer in the content recognition model for activation processing to obtain target activation values, and the target activation values are used as the adjusted feature weights corresponding to the adjusted text features, making the adjusted feature weights be normalized values and improving the rationality of the adjusted feature weights.
[0183] In some embodiments, the first fused text feature obtaining unit is further configured to obtain text feature weights corresponding to the text extraction features based on the adjusted feature weights; perform a multiplication calculation on the adjusted feature weights and the adjusted text features to obtain calculated adjusted text features; perform a multiplication calculation on the text feature weights and the text extraction features to obtain calculated text extraction features; and add the calculated adjusted text features and the calculated text extraction features to obtain fused text features.
[0184] In this embodiment, the adjusted feature weight and the adjusted text feature are multiplied to obtain the calculated adjusted text feature. The text feature weight and the text extraction feature are multiplied to obtain the calculated text extraction feature. The calculated adjusted text feature and the calculated text extraction feature are added together to obtain the fused text feature. Since the text feature weight is obtained based on the adjusted feature weight, the accuracy of the text feature weight is improved, thereby improving the accuracy of the fused text feature.
[0185] In some embodiments, the target content is a target video; the target content determination module 902 includes: a target text obtaining unit, configured to obtain the text corresponding to the target time in the target video to obtain the target text; and a text association data obtaining unit, configured to obtain the video-related data corresponding to the target time in the target video, and use the video-related data as the text association data associated with the target text, where the video-related data includes at least one of video frames or audio frames.
[0186] In this embodiment, the text corresponding to the target time in the target video is obtained to obtain the target text, and the video-related data corresponding to the target time in the target video is obtained, and the video-related data is used as the text association data associated with the target text. Since the video-related data includes at least one of video frames or audio frames, text data and image data or audio data other than the text data are obtained, so that the video can be recognized by combining the image data or audio data on the basis of the text data, which is beneficial to improving the recognition accuracy.
[0187] In some embodiments, the adjusted text feature includes a first adjusted text feature adjusted according to video frames and a second adjusted text feature adjusted according to audio frames; the content recognition result obtaining module 910 includes: a second fused text feature obtaining unit, configured to fuse the first adjusted text feature, the second adjusted text feature, and the text extraction feature to obtain a fused text feature; and a second content recognition result obtaining unit, configured to perform recognition based on the fused text feature to obtain the content recognition result corresponding to the target content.
[0188] In this embodiment, the first adjusted text feature, the second adjusted text feature, and the text extraction feature are fused to obtain the fused text feature, so that the fused text feature is obtained based on these three features: the first adjusted text feature, the second adjusted text feature, and the text extraction feature, thereby improving the feature richness of the fused text feature. Therefore, when recognition is performed based on the fused text feature, the recognition accuracy can be improved.
[0189] In some embodiments, the module 908 for adjusting text features includes: an eigenvalue product obtaining unit, configured to multiply the feature attention intensity by each eigenvalue of the text extraction features to obtain an eigenvalue product; and an adjusted text feature obtaining unit, configured to arrange the eigenvalue product according to the positions of the eigenvalues in the text extraction features, and use the obtained eigenvalue sequence as the adjusted text features.
[0190] In this embodiment, the feature attention intensity is multiplied by each eigenvalue of the text extraction features to obtain an eigenvalue product, so that the eigenvalue product can reflect the attention degree of the text associated data to the text feature values. The eigenvalue product is sorted according to the sorting of the eigenvalues in the text extraction features, and the obtained eigenvalue sequence is used as the adjusted text features, so that the adjusted text features can reflect the attention degree of the text associated data to the text extraction features.
[0191] In some embodiments, the text extraction features are the features corresponding to the word segments in the target text; each adjusted text feature forms a feature sequence according to the order of the word segments in the target text; the module 910 for obtaining the content recognition result includes: a position relationship obtaining unit, configured to obtain the position relationship of each word segment relative to the named entity based on the feature sequence; and a third content recognition result obtaining unit, configured to obtain the target named entity from the target text based on each position relationship, and use the target named entity as the content recognition result corresponding to the target content.
[0192] In this embodiment, the position relationship of each word segment relative to the named entity is obtained based on the feature sequence, the target named entity is obtained from the target text based on each position relationship, and the target named entity is used as the content recognition result corresponding to the target content, so that content recognition can be performed based on the feature sequence formed by the adjusted text features, and the accuracy of content recognition is improved.
[0193] In some embodiments, the third content recognition result obtaining unit is further configured to obtain the word segment whose position relationship is the starting position of the named entity as the starting word of the named entity; use the word segment whose position relationship is inside the named entity among the backward word segments corresponding to the starting word of the named entity as the constituent words of the named entity; and combine the starting word of the named entity and the constituent words of the named entity to obtain the target named entity.
[0194] In this embodiment, the word segment whose position relationship is the starting position of the named entity is obtained as the starting word of the named entity, the word segment whose position relationship is inside the named entity among the backward word segments corresponding to the starting word of the named entity is used as the constituent words of the named entity, and the starting word of the named entity and the constituent words of the named entity are combined to obtain the target named entity, so that entity recognition can be performed based on the feature sequence formed by the adjusted text features, and the accuracy of entity recognition is improved.
[0195] In some embodiments, the position relationship obtaining unit is further configured to obtain the position relationship of each word segment relative to the named entity and the entity type corresponding to the word segment based on the feature sequence; the third content recognition result obtaining unit is further configured to use, as the named entity component words, the word segments in the backward word segments corresponding to the start word of the named entity, where the position relationship is inside the named entity and the entity type is the same as that of the start word of the named entity.
[0196] In this embodiment, the position relationship of each word segment relative to the named entity and the entity type corresponding to the word segment are obtained based on the feature sequence. Among the backward word segments corresponding to the start word of the named entity, the word segments with a position relationship inside the named entity and the same entity type as the start word of the named entity are used as the named entity component words, which improves the accuracy of entity recognition.
[0197] In some embodiments, the feature attention intensity obtaining module 906 includes: a product operation value obtaining unit configured to perform a product operation on the association feature value in the association extraction feature and the text feature value at the corresponding position in the text extraction feature to obtain a product operation value; a feature attention intensity obtaining unit configured to perform statistics on the product operation value to obtain the feature correlation degree between the association extraction feature and the text extraction feature, and use the feature correlation degree as the feature attention intensity corresponding to the text extraction feature.
[0198] In this embodiment, a product operation is performed on the association feature value in the association extraction feature and the text feature value at the corresponding position in the text extraction feature to obtain a product operation value. Statistical operations are performed on each product operation value to obtain the feature correlation degree between the association extraction feature and the text extraction feature, and the feature correlation degree is used as the feature attention intensity corresponding to the text extraction feature. Thus, the feature attention intensity can accurately reflect the association relationship between the text association data and the target text. Therefore, when adjusting the text extraction feature based on the feature attention intensity, the adjustment accuracy can be improved.
[0199] For the specific limitations of the content recognition device, reference can be made to the limitations on the content recognition method in the foregoing text, which will not be elaborated here. Each module in the above content recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0200] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store content recognition data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a content recognition method.
[0201] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 11 shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a content recognition method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0202] Those skilled in the art can understand that Figure 10 and Figure 11 the structures shown in the figure are only block diagrams of some structures related to the solution of the present application, and do not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0203] In some embodiments, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0204] In some embodiments, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0205] In some embodiments, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.
[0206] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0207] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0208] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A content recognition method, characterized in that, The method includes: Determine the target content to be recognized, and obtain the target text in the target content and text association data associated with the target text; Extract features from the target text to obtain text extraction features; extract features from the text association data to obtain association extraction features; Perform an association calculation on the association extraction features and the text extraction features, and obtain the feature attention intensity corresponding to the text extraction features based on the calculated feature association degree, where the feature association degree is positively correlated with the feature attention intensity; Adjust the text extraction features based on the feature attention intensity to obtain adjusted text features; Encode the text extraction features to obtain first encoded features, and encode the adjusted text features to obtain second encoded features; Fuse the first encoded features and the second encoded features to obtain fused encoded features; Obtain the adjusted feature weight corresponding to the adjusted text features based on the fused encoded features; Fuse the adjusted text features and the text extraction features based on the adjusted feature weight to obtain fused text features; Perform recognition based on the fused text features to obtain a content recognition result corresponding to the target content.
2. The method according to claim 1, wherein The first encoded features are encoded by a first encoder in a trained content recognition model, and the second encoded features are encoded by a second encoder in the content recognition model. The obtaining the adjusted feature weight corresponding to the adjusted text features based on the fused encoded features includes: Input the fused encoded features into a target activation layer in the content recognition model for activation processing to obtain a target activation value, and use the target activation value as the adjusted feature weight corresponding to the adjusted text features. The activation layer is a shared activation layer of the first encoder and the second encoder.
3. The method according to claim 1, wherein The fusing the adjusted text features and the text extraction features based on the adjusted feature weight to obtain fused text features includes: Obtain the text feature weight corresponding to the text extraction features based on the adjusted feature weight; Perform a multiplication calculation on the adjusted feature weight and the adjusted text features to obtain the calculated adjusted text features; Perform a multiplication calculation on the text feature weight and the text extraction features to obtain the calculated text extraction features; Add the calculated adjusted text features and the calculated text extraction features to obtain fused text features.
4. The method according to claim 1, wherein The target content is a target video. The obtaining the target text in the target content and text association data associated with the target text includes: Obtain the text corresponding to the target time in the target video to obtain the target text; Obtain the video-related data corresponding to the target time in the target video, and use the video-related data as the text association data associated with the target text. The video-related data includes at least one of video frames or audio frames.
5. The method according to claim 4, wherein The adjusted text features include a first adjusted text feature adjusted according to the video frames and a second adjusted text feature adjusted according to the audio frames.
6. The method according to claim 1, wherein Adjusting the text extraction features based on the feature attention intensity to obtain adjusted text features includes: Multiplying the feature attention intensity by each feature value of the text extraction features to obtain a feature value product; Arranging the feature value products according to the positions of the feature values in the text extraction features, and using the arranged feature value sequence as the adjusted text features.
7. The method according to claim 1, characterized in that, The text extraction features are the features corresponding to the word segments in the target text; each adjusted text feature forms a feature sequence in the order of the word segments in the target text; the method further includes: Obtaining the position relationship of each word segment relative to the named entity based on the feature sequence; Obtaining the target named entity from the target text based on each position relationship, and using the target named entity as the content recognition result corresponding to the target content.
8. The method according to claim 7, wherein The obtaining the target named entity from the target text based on each position relationship includes: Obtaining the word segment with the position relationship as the starting position of the named entity as the starting word of the named entity; Regarding the word segments with the position relationship as being inside the named entity among the backward word segments corresponding to the starting word of the named entity as the constituent words of the named entity; Combining the starting word of the named entity with the constituent words of the named entity to obtain the target named entity.
9. The method according to claim 8, characterized in that, The obtaining the position relationship of each word segment relative to the named entity based on the feature sequence includes: Based on the feature sequence, obtaining the position relationship of each word segment relative to the named entity and the entity type corresponding to the word segment; The regarding the word segments with the position relationship as being inside the named entity among the backward word segments corresponding to the starting word of the named entity as the constituent words of the named entity includes: Regarding the word segments with the position relationship as being inside the named entity and having the same entity type as the starting word of the named entity among the backward word segments corresponding to the starting word of the named entity as the constituent words of the named entity.
10. The method according to claim 1, characterized in that, The performing an association calculation on the associated extraction features and the text extraction features, and obtaining the feature attention intensity corresponding to the text extraction features based on the calculated feature association degree includes: Performing a product operation on the associated feature values in the associated extraction features and the text feature values at the corresponding positions in the text extraction features to obtain a product operation value; Statistically analyzing the product operation value to obtain the feature association degree between the associated extraction features and the text extraction features, and using the feature association degree as the feature attention intensity corresponding to the text extraction features.
11. A content recognition device, characterized in that, The device includes: A target content determination module, configured to determine the target content to be recognized, obtain the target text in the target content, and text association data associated with the target text; A feature extraction module, configured to perform feature extraction on the target text to obtain text extraction features; perform feature extraction on the text association data to obtain associated extraction features; A feature attention intensity obtaining module, configured to perform an association calculation on the associated extraction features and the text extraction features, and obtain the feature attention intensity corresponding to the text extraction features based on the calculated feature association degree, where the feature association degree has a positive correlation with the feature attention intensity; An adjusted text feature obtaining module, configured to adjust the text extraction features based on the feature attention intensity to obtain adjusted text features; An encoding module, configured to encode the text extraction features to obtain first encoded features, and encode the adjusted text features to obtain second encoded features; A fused encoding module, configured to fuse the first encoded features and the second encoded features to obtain fused encoded features; An adjusted feature weight obtaining module, configured to obtain adjusted feature weights corresponding to the adjusted text features based on the fused encoded features; A fused text feature obtaining module, configured to fuse the adjusted text features and the text extraction features based on the adjusted feature weights to obtain fused text features; A content recognition result obtaining module, configured to perform recognition based on the fused text features to obtain a content recognition result corresponding to the target content.
12. The content recognition device according to claim 11, wherein The first encoded features are obtained by encoding through a first encoder in a pre-trained content recognition model, the second encoded features are obtained by encoding through a second encoder in the content recognition model, and the adjusted feature weight obtaining module is further configured to input the fused encoded features into a target activation layer in the content recognition model for activation processing to obtain a target activation value, and use the target activation value as the adjusted feature weights corresponding to the adjusted text features, and the activation layer is a shared activation layer of the first encoder and the second encoder.
13. The content recognition device according to claim 11, characterized in that, The fused text feature obtaining module is further configured to obtain text feature weights corresponding to the text extraction features based on the adjusted feature weights; perform a product calculation on the adjusted feature weights and each feature value of the adjusted text features to obtain a calculated adjusted text feature; perform a product calculation on the text feature weights and the text extraction features to obtain a calculated text extraction feature; and add the calculated adjusted text feature and the calculated text extraction feature to obtain fused text features.
14. The content recognition device according to claim 11, characterized in that The target content is a target video; The target content determination module includes: A target text obtaining unit, configured to obtain text corresponding to a target time in the target video to obtain a target text; A text associated data obtaining unit, configured to obtain video-related data corresponding to the target time in the target video, and use the video-related data as text associated data associated with the target text, where the video-related data includes at least one of video frames or audio frames.
15. The content recognition device according to claim 14, wherein The adjusted text features include first adjusted text features adjusted according to the video frames and second adjusted text features adjusted according to the audio frames.
16. The content recognition device according to claim 11, characterized in that, The adjusted text feature obtaining module is further configured to multiply the feature attention intensity by each feature value of the text extraction features to obtain a feature value product; arrange the feature value product according to the positions of the feature values in the text extraction features, and use the arranged feature value sequence as the adjusted text features.
17. The content recognition device according to claim 11, characterized in that, The text extraction feature is the feature corresponding to the word segmentation in the target text; each adjusted text feature forms a feature sequence in the order of the word segmentation in the target text; the content recognition result obtaining module is further configured to obtain the position relationship of each word segmentation relative to the named entity based on the feature sequence; and obtain the target named entity from the target text based on each position relationship, and use the target named entity as the content recognition result corresponding to the target content.
18. The content recognition device according to claim 17, wherein The content recognition result obtaining module is further configured to obtain the word segmentation with the position relationship being the starting position of the named entity as the starting word of the named entity; use the word segmentation with the position relationship being inside the named entity among the backward word segmentations corresponding to the starting word of the named entity as the constituent words of the named entity; and combine the starting word of the named entity and the constituent words of the named entity to obtain the target named entity.
19. The content recognition device according to claim 18, characterized in that, The content recognition result obtaining module is further configured to obtain the position relationship of each word segmentation relative to the named entity and the entity type corresponding to the word segmentation based on the feature sequence; and use the word segmentation with the position relationship being inside the named entity and the entity type being the same as that of the starting word of the named entity among the backward word segmentations corresponding to the starting word of the named entity as the constituent words of the named entity.
20. The content recognition device according to claim 11, characterized in that, The feature attention intensity obtaining module is further configured to perform a multiplication operation on the correlation feature value in the correlation extraction feature and the text feature value at the corresponding position in the text extraction feature to obtain a multiplication operation value; perform statistics on the multiplication operation value to obtain the feature correlation degree between the correlation extraction feature and the text extraction feature, and use the feature correlation degree as the feature attention intensity corresponding to the text extraction feature.
21. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 10 is implemented.
22. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 10 is implemented.
23. A computer program product comprising computer instructions, characterized in that, When the computer instruction is executed by the processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Text processing method and device, model training method and device, computer equipment and storage medium
CN112084331A
Speaking person separation method and apparatus based on recurrent neural network and acoustic features
WO2020258661A1