Text content recognition method, device, equipment, storage medium and program product
By combining video type features with text semantic features to perform text content recognition, the problem of misrecognition in the existing technology is solved and the accuracy of recognition is improved.
Patent Information
- Application Number
- CN202111506620.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-12-10
AI Technical Summary
In the prior art, when identifying illegal text based on the semantic features of video-related text, it is easy to cause misrecognition and the accuracy is low.
Text content recognition is performed by combining text semantic features with video type features. By extracting video type features and fusing them with text semantic features, specific types of text can be recognized.
The probability of misidentification of video-associated text is reduced, and the accuracy of text content recognition is improved.
Smart Images

Figure CN116263788B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a method, apparatus, device, storage medium, and program product for recognizing text content. Background Art
[0002] Currently, many video playback platforms support the ability for users to comment on videos, such as sending barrage comments. The text of video comments needs to be filtered.
[0003] Related technologies often perform text content recognition on video-related text by directly extracting its semantic features, thereby identifying, for example, illegal content based on these features. However, the same comment content may have different meanings in different contexts. Recognition based solely on the semantic features of the comment content can easily lead to misidentification, resulting in low accuracy in identifying illegal content. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, device, storage medium, and program product for identifying text content, which can reduce the probability of misidentification of video-associated text and improve the accuracy of text content recognition.
[0005] The technical solution is as follows:
[0006] In one aspect, an embodiment of the present application provides a method for identifying text content, the method comprising:
[0007] Acquire a target text and a target video type, wherein the target text is text associated with a target video, and the target video type is a video type of the target video;
[0008] Extracting semantic features of the target text to obtain text semantic features of the target text;
[0009] Extracting type features of the target video type to obtain video type features of the target video;
[0010] The target text is subjected to text content recognition based on the text semantic features and the video type features to obtain a text content recognition result, wherein the text content recognition result is used to indicate the relationship between the target text and the target video and a specific type of text.
[0011] On the other hand, an embodiment of the present application provides a device for identifying text content, the device comprising:
[0012] An acquisition module, configured to acquire a target text and a target video type, wherein the target text is text associated with a target video, and the target video type is a video type of the target video;
[0013] A first feature extraction module is used to extract semantic features of the target text to obtain text semantic features of the target text;
[0014] A second feature extraction module is used to extract type features of the target video type to obtain video type features of the target video;
[0015] The first recognition module is used to perform text content recognition on the target text based on the text semantic features and the video type features to obtain a text content recognition result, wherein the text content recognition result is used to indicate the relationship between the target text and the target video and a specific type of text.
[0016] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text content recognition method as described in the above aspects.
[0017] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the text content recognition method as described in the above aspects.
[0018] In another aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text content recognition method provided in the above aspects.
[0019] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:
[0020] In an embodiment of the present application, when performing text content recognition on text associated with a target video, the computer device first extracts the text semantic features of the text, obtains the video type of the target video, performs feature extraction on the video type to obtain video type features, and then identifies the relationship between the target text and a specific type of text based on the text semantic features and the video type features. That is, in an embodiment of the present application, by introducing video type features when performing text content recognition, specific text types can be recognized for different video types. Compared with the solution in the related art that performs text content recognition based only on text semantic features, the probability of misidentification of video-associated text can be reduced, and the accuracy of text content recognition can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 A schematic diagram showing the principle of the text content recognition method provided by an embodiment of the present application is shown;
[0023] Figure 2 A schematic diagram showing an implementation environment provided by an exemplary embodiment of the present application is shown;
[0024] Figure 3 A flowchart of a method for identifying text content provided by an exemplary embodiment of the present application is shown;
[0025] Figure 4 A flowchart of a method for identifying text content provided by another exemplary embodiment of the present application is shown;
[0026] Figure 5 A schematic diagram illustrating an implementation of a text content recognition process provided by an exemplary embodiment of the present application is shown;
[0027] Figure 6 A flowchart of a classification network training method provided by an exemplary embodiment of the present application is shown;
[0028] Figure 7 This is a structural block diagram of a text content recognition device provided by an exemplary embodiment of the present application;
[0029] Figure 8 A schematic structural diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0030] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0031] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0032] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0033] Natural language processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology generally includes technologies such as text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, etc. The text content recognition method provided in the embodiment of the present application is an application in semantic understanding, which improves the accuracy of text content recognition by introducing video type features in the text content recognition process.
[0034] In the related art, when performing text content recognition on text associated with a video, such as barrage, to determine whether the text is of a specific text type, text semantic features are usually directly extracted for recognition. For example, in the process of identifying whether a text is of an illegal text type, recognition is based on text semantic features, and the same text may have different meanings in different videos. For example, when "This pig is so funny" is an associated text with a variety show video, it may be identified as an illegal text because it contains abusive words, and when it is an associated text with a pet video, it may be a normal text, not an illegal text. If the solution in the related art is used, only text semantic features are used for recognition, it is impossible to recognize different video types, which may cause misidentification of the corresponding type of text. Therefore, in an embodiment of the present application, when performing text content recognition, text semantic features and video type features are used together for recognition, thereby reducing the probability of misidentification of video-associated text and improving the accuracy of text content recognition.
[0035] like Figure 1 , which shows a schematic diagram of the principle of the text content recognition method provided by an embodiment of the present application. When performing text content recognition on a target text, semantic feature extraction 102 can be performed on the target text 101 to obtain text semantic features 103, and type feature extraction 105 can be performed on the target video type 104 to obtain video type features 106. Based on the text semantic features 103 and the video type features 106, text content recognition 107 is performed to obtain a text content recognition result.
[0036] In the embodiment of the present application, since video type features are introduced in the text content recognition process, text content recognition can be performed on the associated texts of different video types in a targeted manner, thereby reducing the probability of associated texts being misrecognized and improving the accuracy of text content recognition.
[0037] The method provided in the embodiment of the present application can be applied to any scenario where text content recognition of video-related text is required. The following schematically illustrates the application scenario of the text content recognition method provided in the embodiment of the present application.
[0038] 1. Applicable to video playback scenarios
[0039] When applied to a video playback scenario, the method provided by the embodiment of the present application can be applied to the background server of the video playback platform. When a user posts text associated with the played video, the background server obtains the associated text and performs semantic feature extraction on the associated text to obtain text semantic features. The background server also obtains the video type of the played video and performs type feature extraction on the video type to obtain video type features. Based on the text semantic features and video type features, the text content of the associated text is identified to determine whether it is a specific type of text, thereby screening out specific types of text from the text associated with the played video, such as high-quality text related to the video or identifying illegal text.
[0040] 2. Applicable to live viewing scenarios
[0041] When applied to live broadcast viewing scenarios, the method provided by the embodiments of the present application can be applied to the backend server of the live broadcast viewing platform. When a user posts text related to the live broadcast content, the backend server can obtain the related text and extract the text semantic features, obtain the live broadcast content type, and extract the video type features. Based on the text semantic features and video type features, the related text is identified and determined to be a specific type of text, thereby filtering out the specific type of text related to the live broadcast.
[0042] The above is only a schematic illustration of the application scenario. The method provided in the embodiment of the present application can also be applied to other scenarios requiring text content recognition. The embodiment of the present application does not limit the actual application scenario.
[0043] Figure 2 A schematic diagram of an implementation environment provided by an exemplary embodiment of the present application is shown. This embodiment is described using the text content recognition method applied to a video playback scenario as an example. The implementation environment includes a terminal 210 and a server 220. Data communication is performed between the terminal 210 and the server 220 via a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network.
[0044] Terminal 210 is an electronic device with a video playback function. This electronic device can be a mobile terminal such as a smartphone, tablet computer, or laptop computer, or a desktop computer, projector computer, or other terminal, which is not limited in this embodiment of the present application. The terminal can provide the video playback function through, for example, a movie and TV series playback client, a short video playback client, or other client, such as a web page in a browser, which is not limited in this embodiment of the present application.
[0045] Server 220 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In the embodiment of the present application, server 220 is a backend server that provides a video playback client in terminal 210 and can perform text content recognition on text posted by users in connection with the played video.
[0046] like Figure 2 As shown, when the terminal 210 receives the target text 221 sent by the user, the terminal 210 sends the target text 221 to the server 220, and the server 220 extracts semantic features of the target text 221 to obtain text semantic features 223, and obtains the target video type 222 of the played video, and extracts type features of the target video type 222 to obtain video type features 224, thereby performing text content recognition 225 based on the text semantic features 223 and the video type features 224 to obtain text content recognition results 226, and feeds back the text content recognition results 226 to the terminal 210.
[0047] In another possible implementation, the text content recognition process may also be performed by the terminal 210. The server 220 trains the model used for text content recognition and sends the trained model to the terminal 210, which then performs text content recognition locally without the need for the server 220. Alternatively, the model used for text content recognition may also be trained on the terminal 210 side, and the terminal 210 performs text content recognition. This embodiment of the present application is not limited to this.
[0048] For the convenience of description, the following embodiments are described by taking the text content recognition method executed by a computer device as an example.
[0049] Please refer to Figure 3 , which shows a flowchart of a text content recognition method provided by an exemplary embodiment of the present application. This embodiment uses the method applied to a computer device as an example to illustrate, and the method includes the following steps.
[0050] Step 301: Acquire target text and target video type, where the target text is text associated with the target video, and the target video type is the video type of the target video.
[0051] The target text may be a comment text on the target video, such as a barrage message sent by a user during video viewing, or a comment message sent by a user in a comment area corresponding to the video.
[0052] Optionally, videos can be categorized by content, such as food, animals, fashion, entertainment, sports, etc. Alternatively, videos can be categorized by format, such as TV series, movies, variety shows, animation, live broadcasts, short videos, documentaries, etc.
[0053] In a possible implementation, after the computer device receives the text sent by the user and associated with the target video, it can obtain the text information therein to obtain the target text.
[0054] When the computer device acquires the target text, it also acquires the video type of the target video associated with the target text. Different videos have different type tags. After acquiring the target text, the computer device also acquires the video type tag corresponding to the target video, thereby determining the target video type. Optionally, the target video type includes at least one type corresponding to the target video. For example, if the target video is an entertainment-related variety show, the target video type may be "Entertainment" or "Variety Show."
[0055] Step 302: extract semantic features from the target text to obtain text semantic features of the target text.
[0056] After acquiring the target text, the computer device performs semantic feature extraction on the target text. In one possible embodiment, when performing semantic feature extraction, the word vector sequence of the target text is first obtained, wherein the word vector sequence can be obtained by converting the target text using the word vector (wordtovector, word2vec) model. Afterwards, the computer device uses the bidirectional long short-term memory (Bi-LSTM) model to mine the semantic feature information between words based on the word vector sequence to obtain the text semantic features.
[0057] It should be noted that the above is only an exemplary description of the text semantic feature extraction process. In addition to using the word2vec model and the Bi-LSTM model, other models can also be used for semantic feature extraction, and the embodiments of the present application do not limit this.
[0058] Step 303: extract type features of the target video type to obtain video type features of the target video.
[0059] The computer device then performs type feature extraction on the acquired target video type to obtain a video type feature corresponding to the target video. The type feature extraction process refers to the process of converting the target video type into a feature vector. Through type feature extraction, the types contained in the target video type are converted into a feature vector to obtain a video type feature, thereby performing text content recognition based on the video type feature.
[0060] Step 304 : performing text content recognition on the target text based on the text semantic features and the video type features to obtain a text content recognition result. The text content recognition result is used to indicate the relationship between the target text, the target video, and the specific type of text.
[0061] In one possible implementation, a computer device uses text semantic features and video type features to perform text content recognition on a target text. Due to the introduction of video type features, when performing text content recognition, recognition can also be performed based on the video type, thereby identifying the relationship between the target text associated with the target video and a specific type of text, for example, identifying whether the target text belongs to a specific type of text.
[0062] Optionally, the characteristic type text may be high-quality text, illegal text, irrelevant text (low-quality text), etc. In different scenarios, different types of text may need to be screened.
[0063] Among them, high-quality text refers to commendatory text that is highly relevant to the target video. In one possible case, the text associated with the video may contain a lot of text that is unrelated to the current video. When it is necessary to screen out commendatory text that is highly relevant to the video, the feature type text can be determined as high-quality text. When the feature type text is high-quality text, the text semantic features and video type features can be used to identify the text content and identify high-quality commentary content for the target video. When it is necessary to screen out text that is unrelated to the target video, the feature type text can be determined as low-quality text, and the text semantic features and video type features can be used to identify the text content and identify text that is unrelated to the target video.
[0064] In another possible scenario, the text associated with the video may contain illegal content, such as sensitive or derogatory content, and the illegal text needs to be filtered. In this case, the feature type text can be identified as illegal text. If the feature type text is illegal text, the text semantic features and video type features can be used to identify the text content and identify the illegal comments about the video.
[0065] Illustratively, when identifying whether the target text is illegal text, if the target text is "This pig is so funny" and the video type is a pet type, when performing text content recognition based on text semantic features and video type features, the correlation between "pig" and pets can be identified, and it can be identified as normal text rather than illegal text, reducing the probability of misidentification.
[0066] In summary, in the embodiment of the present application, when performing text content recognition on the text associated with the target video, the computer device first extracts the text semantic features of the text, obtains the video type of the target video, performs feature extraction on the video type to obtain the video type features, and then identifies the relationship between the target text and the specific type of text based on the text semantic features and the video type features. That is, in the embodiment of the present application, the video type features are introduced when performing text content recognition, and specific types of text can be recognized for different video types. Compared with the scheme of performing text content recognition based only on text semantic features in the related art, the probability of misrecognition of video-associated text can be reduced, and the accuracy of text content recognition can be improved.
[0067] In an embodiment of the present application, text content recognition is performed jointly using text semantic features and video type features. During the recognition process, the text semantic features and video type features need to be fused, so that text content recognition is performed based on the fused features. This will be explained below with an exemplary embodiment.
[0068] Please refer to Figure 4 , which shows a flow chart of a text content recognition method provided by another exemplary embodiment of the present application. This embodiment uses the method applied to a computer device as an example to illustrate, and the method includes the following steps.
[0069] Step 401: Obtain target text and target video type.
[0070] The implementation of this step can refer to the above-mentioned step 301, and will not be described in detail in this embodiment.
[0071] Step 402 : extracting semantic features of the target text to obtain text semantic features of the target text and extracting type features of the target video type to obtain video type features of the target video.
[0072] When extracting semantic features, the computer first segments the target text into a word sequence w = {w_1, w_2, ..., w_n}. This word sequence is then fed into the word2vec model for word embedding, resulting in a word embedding sequence wv = {wv_1, wv_2, ..., wv_n}. The computer then feeds this word embedding sequence into a Bi-LSTM model to obtain the text's semantic features, namely, a semantic feature matrix v = {rv_1, rv_2, ..., rv_n}.
[0073] The process of extracting the type features of the target video type is the process of converting the target video type into a feature vector. Optionally, feature encoding is performed on the content recognition target video type to obtain the video type features of the content recognition target video.
[0074] In one possible embodiment, the feature encoding process for the target video type may be a quantization process, which includes: setting the numerical values at the positions corresponding to the target video type in the target sequence to the first numerical value, and setting the numerical values at other positions in the target sequence to the second numerical value, to obtain a type feature sequence of the target video, wherein the numerical values at different positions in the target sequence correspond to different video types.
[0075] Optionally, the target sequence is used to represent the video type, and the values at different positions in the sequence correspond to different video types. For example, when the video types include sports, variety shows, TV series, food, entertainment, and animals, each video type can be numbered, and the video type corresponding to each position in the target sequence can be determined based on the number. For example, sports corresponds to number 1, variety shows correspond to number 2, TV series correspond to number 3, food corresponds to number 4, entertainment corresponds to number 5, and animals correspond to number 6. Then, the first position in the target sequence corresponds to the sports type, the second position corresponds to the variety show type, the third position corresponds to the TV series type, the fourth position corresponds to the food type, the fifth position corresponds to the entertainment type, and the sixth position corresponds to the animal type.
[0076] The first value is used to indicate that the target video belongs to the video type corresponding to the position, while the second value is used to indicate that the target video does not belong to the video type corresponding to the position. For example, when the target video type is determined to be variety show and entertainment, the values at the second and fifth positions in the target sequence can be set to the first value "1", and the values at the first, third, fourth, and sixth positions in the target sequence can be set to the second value "0", resulting in a video type feature, i.e., a type feature sequence c = {0, 1, 0, 0, 1, 0}.
[0077] In another possible implementation, different video types correspond to different n-dimensional vectors. When encoding the target video type, the n-dimensional vectors corresponding to each type contained in the target video type can be fused to obtain a feature vector corresponding to the target video type, wherein the fusion processing may include vector splicing, weighted averaging processing, etc.
[0078] For example, the word2vec model can be used to convert different genres into corresponding n-dimensional vectors: sports corresponds to the first vector, variety shows to the second vector, TV dramas to the third vector, food to the fourth vector, entertainment to the fifth vector, and animals to the sixth vector. If the target video genre includes variety shows and entertainment, the second and fifth vectors can be fused to obtain the feature vector corresponding to the target video genre.
[0079] Alternatively, other feature encoding methods may be used to encode the target video type and convert the target video type into a feature vector, which is not limited in this embodiment.
[0080] Step 403 : obtaining a semantic fusion feature based on the text semantic feature and the video type feature. The semantic fusion feature is used to indicate the meaning of the target text when the target text is the associated text of the video corresponding to the target video type.
[0081] In one possible implementation, a computer device performs feature fusion on text semantic features and video type features to obtain fused semantic fusion features, wherein the semantic fusion features refer to the meaning when the target text is the video-associated text corresponding to the target video type. For example, the target text "This pig is so funny" has a complimenting meaning for a pet-type video.
[0082] Optionally, the process of performing feature fusion to obtain semantic fusion features may include the following steps:
[0083] Step 403a: determining a correlation coefficient based on the text semantic features and the video type features. The correlation coefficient is used to indicate the correlation between the target text and the target video type.
[0084] To consider the impact of video type features on text semantic features, in this embodiment, the computer device determines the correlation between the target text and the target video type by determining the correlation coefficient between the two features. Determining the correlation coefficient may include the following steps:
[0085] Step 1: Perform full connection processing on the text semantic features through the first fully connected layer to obtain the text semantic feature vector.
[0086] In one possible implementation, a fully connected layer is first used to transform the text semantic features and video type features into vectors of the same dimension. Optionally, a first fully connected layer is used to fully connect the extracted text semantic features. Specifically, the semantic feature matrix v = {rv_1, rv_2, …, rv_n} is input into the first fully connected layer for mapping, resulting in a text semantic feature vector V of dimension K.
[0087] Step 2: Perform full connection processing on the video type features through the second fully connected layer to obtain a video type feature vector. The video type feature vector has the same vector dimension as the text semantic feature vector.
[0088] Correspondingly, the video type feature is input into the second fully connected layer for mapping to obtain a video type feature vector C with a dimension of K, thereby obtaining a text semantic feature vector V and a video type feature vector C with a vector dimension of K.
[0089] Step 3: Multiply the text semantic feature vector and the video type feature vector to obtain the correlation coefficient.
[0090] Optionally, the text semantic feature vector V is vector-multiplied with the video type feature vector C to obtain a correlation coefficient. That is, the correlation between the text semantic feature vector and the video type feature vector of the same dimension is determined by using a dot product method.
[0091] Step 403b: Based on the correlation coefficient, the text semantic features are weighted to obtain semantic fusion features.
[0092] After the computer device determines the correlation coefficient, it multiplies the correlation coefficient with the text semantic feature vector to achieve weighting of the text semantic features, thereby obtaining a semantic fusion feature, that is, a semantic fusion feature vector.
[0093] Step 404 : performing text content recognition on the target text based on the semantic fusion feature and the video type feature to obtain a text content recognition result.
[0094] In a possible implementation, a computer device uses semantic fusion features and video type features to perform text content recognition on a target text, which may include the following steps:
[0095] Step 404a: concatenate the semantic fusion features with the video type features to obtain target text features.
[0096] Optionally, when performing feature concatenation, the semantic fusion feature vector is concatenated with the video type feature vector to obtain a target text feature vector, wherein the target text feature vector is the text feature vector after fusion of the video type.
[0097] In another possible implementation, in order to increase the complexity of the model and improve the model learning ability, after obtaining the video type feature vector, it can be input into the fourth fully connected layer, and it can be vector transformed again to obtain the transformed video type feature vector, and the transformed video type feature vector can be concat spliced with the semantic fusion feature vector to obtain the target text feature vector.
[0098] In step 404b, the target text features are input into a classifier to obtain the probability that the target text belongs to a specific type of text. The classifier includes a third fully connected layer and a Softmax layer.
[0099] After obtaining the target text features, the computer device uses a classifier to identify the text content. The classifier includes a third fully connected layer and a softmax layer. That is, the computer device inputs the target text feature vector into the third fully connected layer, performs a nonlinear transformation, and outputs the result. The nonlinear transformation is as follows:
[0100] Y=f(Wx+b)
[0101] Among them, f is the activation function, W is the weight parameter, and b is the bias constant.
[0102] After the nonlinear transformation output is performed through the third fully connected layer, the output result is input into the Softmax layer to convert it into the probability that the target text is a violation text, as follows:
[0103]
[0104] Among them, z j That is the output result of the third fully connected layer, that is, the output result of the jth node, z k Output result for the kth node, where K is the number of nodes in the third fully connected layer, i.e. the number of classifications.
[0105] Optionally, when the output probability is greater than a preset threshold, the target text may be determined to be a specific type of text.
[0106] In one possible implementation, the text content recognition process can be as follows: Figure 5 As shown, the target text 501 is segmented by word to obtain a word sequence w = {w_1, w_2, ..., w_n}, and the word sequence w is input into the word2vec model 502 for word vector conversion to obtain a word vector sequence wv = {wv_1, wv_2, ..., wv_n}, and then the word vector sequence is input into the Bi-LSTM model 503 to obtain a semantic feature matrix v = {rv_1, rv_2, ..., rv_n}, and then the semantic feature matrix v is input into the first fully connected layer 504 for mapping to obtain a text semantic feature vector V.
[0107] At the same time, the target video type 505 is quantized to obtain video type features, and the video type features are input into the second fully connected layer 506 for mapping to obtain a video type feature vector C. Afterwards, the computer device multiplies the text semantic feature vector V with the video type feature vector C to obtain a correlation coefficient a, and multiplies the correlation coefficient a with the text semantic feature vector V to obtain a semantic fusion feature vector vec, and inputs the video type feature vector C into the fourth fully connected layer 507 to obtain a transformed video type feature vector type-c, and concats the transformed video type feature vector type-c with the semantic fusion feature vector vec to obtain a target text feature vector, and outputs the target text feature vector to input into the third fully connected layer 508 for nonlinear transformation. After the transformation, the output result is input into the Softmax layer to obtain the probability that the final target text is a specific type of text, that is, the text content recognition result.
[0108] Step 405 , in response to the text content recognition result indicating that the target bullet comment belongs to a specific type of text, an associated screen of the target bullet comment is determined, and the associated screen is selected according to the sending time of the target bullet comment.
[0109] In one possible case, the target text is the target barrage. Since the same video contains different video screens, different video screens correspond to different screen contents. When the user sends the target barrage during the playback of the target video, the target barrage may be a comment only on the content in the current video playback screen. If the target barrage is identified as a specific type of text based on the target video type and the target text features, since the target barrage is a comment text only for the current screen, there may be a case of misidentification. For example, when the specific type of text is an illegal text, if the video type is variety show or entertainment, and the target barrage is "This pig is so funny", after the text content is identified based on the text semantic features and the video type features, the recognition result indicates that it is an illegal text. If the target barrage contains "pig" in the screen content when it is sent, it may cause misidentification. Therefore, in order to further ensure the recognition accuracy, when the text content recognition result indicates that the target barrage is a specific type of text, the associated screen of the target barrage is first obtained.
[0110] In one possible implementation, an associated screen may be selected based on the sending time of the target barrage, wherein the sending time is the time when the computer device receives the barrage edited by the user. Optionally, the playback time of the corresponding video may be determined based on the sending time, thereby obtaining the screen corresponding to the video playback time and the screen at the time adjacent to the playback time. For example, the video screen 5 seconds before the video playback time corresponding to the barrage sending time and the video screen 5 seconds after the video playback time may be selected. Schematically, if the corresponding video playback time is 12:15 of the target video, the video screen from 12:10 to 12:20 of the target video may be selected as the associated screen of the target barrage.
[0111] Step 406: Perform content recognition on the associated pictures to obtain picture content features.
[0112] In one possible implementation, when obtaining a screen associated with the target bullet comment, the content of the associated screen may be identified to obtain screen content features. Optionally, the screen content features may include object information in the associated screen, such as information about people, animals, or plants, scene information, color information, and texture information, etc., which is not limited in this embodiment.
[0113] Step 407: Determine the degree of relationship between the target bullet comment and the specific type based on the screen content features and the text semantic features.
[0114] After obtaining the image content features, the computer device further performs text content recognition based on the image content features and the text semantic features of the target text, thereby determining the degree of relationship between the target barrage and the specific type, that is, determining whether the target barrage is a specific type of text. In one possible embodiment, text content recognition based on image content features and text semantic features may include the following steps:
[0115] Step 407a: Determine the correlation between the picture content feature and the text semantic feature.
[0116] The computer device first determines the degree of correlation between the image content features and the text semantic features. The correlation refers to the degree of association between the associated image and the target text. If the correlation is high, the target text is determined to be related to the associated image. For example, if the target text is "This pig is so funny" and the associated image contains a pig, the correlation is determined to be high.
[0117] Step 407b: In response to the correlation degree being higher than the correlation threshold, determining that the target bullet comment is a non-specific type of text.
[0118] Optionally, when determining the degree of relationship between the target barrage and a specific type, the specific type may be a violation type or a low-quality type, that is, the specific type of text is a violation text or a low-quality text. After determining the correlation between the screen content and the target text, the correlation is compared with the correlation threshold, where the correlation threshold is a preset threshold, for example, 70%. When the correlation is higher than 70%, it is determined that the target text is related to the associated screen, thereby determining that the target barrage is not a specific type of text, for example, non-violation text or non-low-quality text.
[0119] Schematically, when the feature type text is illegal text, if the video type is variety show or entertainment, and the target barrage is "This pig is so funny", the target barrage is identified as illegal text based on the text semantic features and the video type features, and the associated picture of the target barrage is obtained, wherein the associated picture contains "pig", and it is determined that the correlation degree is higher than the correlation threshold, thereby determining that the target barrage is pseudo-illegal text.
[0120] Step 407c: In response to the correlation degree being lower than the correlation threshold, determining that the target bullet comment is a specific type of text.
[0121] When the correlation degree is lower than the correlation threshold, it is determined that the recognition result based on the text semantic features and the video type features is correct, that is, the target barrage is determined to be a specific type of text.
[0122] In the embodiments of this application, after determining the correlation coefficient between text semantic features and video type features, the text semantic features are weighted using the correlation coefficient to obtain semantic fusion features guided by video type information. Simultaneously, the semantic fusion features are concatenated with the video type features, and the concatenated features are used to perform text content recognition, thereby determining the probability that the target text is of a specific type. By combining the video type with text content recognition, the probability of misidentification is reduced, thereby improving the accuracy of text content recognition.
[0123] In this embodiment, in the process of identifying illegal text, after the target text is identified as a specific type of text, a video image associated with the target text is further obtained, and the correlation between the video image content and the text is used again to perform secondary recognition, thereby further improving the accuracy of identifying whether the target text belongs to a specific type of text and improving the recognition accuracy.
[0124] Optionally, in the process of performing text content recognition on the target text based on text semantic features and video type features, the computer device inputs the text semantic features and video type features into the classification network to obtain text content recognition results. The classification network is trained based on sample text, sample video type and sample text label. The sample text is the text associated with the sample video, the sample video type is the video type of the sample video, and the sample text label is used to indicate the relationship between the sample text and a specific type of text.
[0125] In one possible approach, before performing text content recognition on the target text, the classification network is pre-trained using multiple sets of training samples, where each set of training samples includes sample text, sample video type, and sample text labels. Optionally, the sample text can be text previously posted by users and associated with the sample video. The sample text labels can be manually annotated labels.
[0126] The classification network may include a first fully connected layer, a second fully connected layer, a third fully connected layer, and a fourth fully connected layer. During the training process using multiple sets of training samples, the parameters in the fully connected layers are trained.
[0127] Please refer to Figure 6 , which shows a flow chart of a classification network training method provided by an exemplary embodiment of the present application. This embodiment uses the method applied to a computer device as an example to illustrate, and the method includes the following steps.
[0128] Step 601 : extracting semantic features from the sample text to obtain sample text semantic features of the sample text, and extracting type features from the sample video type to obtain sample video type features of the sample video.
[0129] In a possible implementation, when training a classification network using training samples, semantic features are first extracted from the sample text to obtain sample text semantic features of the sample text, and type features are extracted from the sample video type to obtain sample video type features of the sample video.
[0130] The process of semantic feature extraction and type feature extraction can refer to the implementation in the above step 402, and will not be repeated in this embodiment.
[0131] Step 602: Input the sample text semantic features and the sample video type features into the classification network to obtain the predicted text content recognition result.
[0132] After obtaining the sample text semantic features and sample video type features, the sample text semantic features and sample video type features are input into the classification network to perform text content recognition.
[0133] The computer device inputs the sample text semantic features and sample video type features into the first fully connected layer and the second fully connected layer respectively for mapping, obtaining a sample text semantic feature vector and a sample video type feature vector of the same dimension, multiplying the two vectors to obtain a correlation coefficient, and using the correlation coefficient to weight the sample text semantic feature vector to obtain a sample semantic fusion feature vector. Subsequently, to improve the learning ability of the classification network, the sample video type feature vector may also be input into the fourth fully connected layer to obtain a transformed sample video type feature vector, which is finally concatenated with the sample semantic fusion feature vector and input into a classifier consisting of the third fully connected layer and the Softmax layer for classification, obtaining a predicted text content recognition result. The predicted text content recognition result is the relationship between the sample text and the specific type of text predicted by the classification network.
[0134] Step 603: Training the classification network based on the predicted text content recognition result and the sample text label.
[0135] The computer device can use the error between the predicted text content recognition result and the sample text label to train the classification network. Optionally, the classification network can be trained using a back propagation algorithm.
[0136] Once the classification network is trained, it can be used for text content recognition. When performing text content recognition, semantic features are extracted from the target text to obtain text semantic features, and type features are extracted from the target video type to obtain video type features. These features are then fed into the trained classification network to determine whether the target text is of a specific type. This allows for specific text type recognition for different video types, improving text content recognition accuracy.
[0137] Figure 7 This is a structural block diagram of a text content recognition device provided by an exemplary embodiment of the present application. Figure 7 As shown, the device includes:
[0138] An acquisition module 701 is configured to acquire a target text and a target video type, wherein the target text is text associated with a target video, and the target video type is a video type of the target video;
[0139] A first feature extraction module 702 is used to extract semantic features of the target text to obtain text semantic features of the target text;
[0140] The second feature extraction module 703 is used to extract the type feature of the target video type to obtain the video type feature of the target video;
[0141] The first recognition module 704 is used to perform text content recognition on the target text based on the text semantic features and the video type features to obtain a text content recognition result, which is used to indicate the relationship between the target text and the target video and a specific type of text.
[0142] Optionally, the first identification module 704 includes:
[0143] a feature fusion unit, configured to obtain a semantic fusion feature based on the text semantic feature and the video type feature, wherein the semantic fusion feature is used to indicate the meaning of the target text when the target text is an associated text of a video corresponding to the target video type;
[0144] The recognition unit is configured to perform the text content recognition on the target text based on the semantic fusion feature and the video type feature to obtain the text content recognition result.
[0145] Optionally, the feature fusion unit is further configured to:
[0146] Determining a correlation coefficient based on the text semantic feature and the video type feature, wherein the correlation coefficient is used to indicate the correlation between the target text and the target video type;
[0147] Based on the correlation coefficient, the text semantic features are weighted to obtain the semantic fusion features.
[0148] Optionally, the feature fusion unit is further configured to:
[0149] Performing full connection processing on the text semantic features through a first fully connected layer to obtain a text semantic feature vector;
[0150] Performing the fully connected processing on the video type feature through a second fully connected layer to obtain a video type feature vector, wherein the video type feature vector has the same vector dimension as the text semantic feature vector;
[0151] The text semantic feature vector is multiplied by the video type feature vector to obtain the correlation coefficient.
[0152] Optionally, the identification unit is further configured to:
[0153] Perform feature splicing on the semantic fusion feature and the video type feature to obtain the target text feature;
[0154] The target text features are input into a classifier to obtain the probability that the target text belongs to the specific type of text, and the classifier includes a third fully connected layer and a Softmax layer.
[0155] Optionally, the second feature extraction module 703 is further configured to:
[0156] Feature encoding is performed on the target video type to obtain video type features of the target video.
[0157] Optionally, the first identification module 704 is further configured to:
[0158] The text semantic features and the video type features are input into a classification network to obtain the text content recognition result. The classification network is trained based on sample text, sample video type and sample text label. The sample text is text associated with a sample video, the sample video type is the video type of the sample video, and the sample text label is used to indicate whether the sample text is related to the specific type of text.
[0159] Optionally, the device further includes:
[0160] a third feature extraction module, configured to extract the semantic features of the sample text to obtain the sample text semantic features of the sample text and to extract the type features of the sample video type to obtain the sample video type features of the sample video;
[0161] A second recognition module is configured to input the sample text semantic features and the sample video type features into the classification network to obtain a predicted text content recognition result;
[0162] A training module is used to train the classification network based on the predicted text content recognition result and the sample text label.
[0163] Optionally, the target text is a target barrage.
[0164] Optionally, the device further includes:
[0165] a picture determination module, configured to determine, in response to the text content recognition result indicating that the target bullet comment belongs to the specific type of text, an associated picture of the target bullet comment, wherein the associated picture is selected based on the time when the target bullet comment appears in the target video;
[0166] A third recognition module is used to perform content recognition on the associated screen to obtain screen content features;
[0167] A determination module is used to determine the degree of relationship between the target bullet comment and a specific type based on the screen content features and the text semantic features.
[0168] Optionally, the determining module includes:
[0169] A first determining unit, configured to determine a correlation between the picture content feature and the text semantic feature;
[0170] a second determining unit, configured to determine, in response to the correlation degree being higher than a correlation threshold, that the target bullet comment is a non-specific type of text;
[0171] The third determining unit is configured to determine that the target bullet comment is the specific type of text in response to the correlation degree being lower than the correlation threshold.
[0172] In summary, in the embodiment of the present application, when performing text content recognition on the text associated with the target video, the computer device first extracts the text semantic features of the text, obtains the video type of the target video, performs feature extraction on the video type to obtain the video type features, and then identifies the relationship between the target text and the specific type of text based on the text semantic features and the video type features. That is, in the embodiment of the present application, the video type features are introduced when performing text content recognition, and specific types of text can be recognized for different video types. Compared with the scheme of performing text content recognition based only on text semantic features in the related art, the probability of misrecognition of video-associated text can be reduced, and the accuracy of text content recognition can be improved.
[0173] It should be noted that the apparatus provided in the above embodiments is merely exemplified by the division of the above functional modules. In actual applications, the above functions can be distributed among different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.
[0174] Please refer to Figure 8 , which shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Specifically, the computer device 800 includes a central processing unit (CPU) 801, a system memory 804 including a random access memory 802 and a read-only memory 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The computer device 800 also includes a basic input / output system (I / O system) 806 that facilitates information transmission between various components within the computer, and a mass storage device 807 for storing an operating system 813, application programs 814, and other program modules 815.
[0175] The basic input / output system 806 includes a display 808 for displaying information and an input device 809 such as a mouse and keyboard for user input. The display 808 and the input device 809 are both connected to the central processing unit 801 via an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include an input / output controller 810 for receiving and processing input from a variety of other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides output to a display screen, printer, or other types of output devices.
[0176] The mass storage device 807 is connected to the central processing unit 801 via a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable media provide non-volatile storage for the computer device 800. In other words, the mass storage device 807 may include a computer-readable medium (not shown) such as a hard disk or drive.
[0177] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, tape cassette, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above-mentioned ones. The above-mentioned system memory 804 and mass storage device 807 can be collectively referred to as memory.
[0178] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 801. The one or more programs contain instructions for implementing the above-mentioned methods. The central processing unit 801 executes the one or more programs to implement the methods provided by the above-mentioned various method embodiments.
[0179] According to various embodiments of the present application, the computer device 800 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 800 may be connected to a network 812 via a network interface unit 811 connected to the system bus 805, or the network interface unit 811 may be used to connect to other types of networks or remote computer systems (not shown).
[0180] The memory also includes one or more programs, which are stored in the memory and include steps executed by a computer device in the method provided in the embodiment of the present application.
[0181] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the text content recognition method described in any of the above embodiments.
[0182] The present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text content recognition method provided in the above aspects.
[0183] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program. The program can be stored in a computer-readable storage medium, which can be the computer-readable storage medium contained in the memory in the above embodiments; or it can be a separate computer-readable storage medium that is not installed in the terminal. The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the text content recognition method described in any of the above method embodiments.
[0184] Optionally, the computer-readable storage medium may include: ROM, RAM, solid-state drives (SSDs), or optical disks. Among them, RAM may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are for descriptive purposes only and do not represent the advantages or disadvantages of the embodiments.
[0185] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0186] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for identifying text content, characterized in that: The method comprises: Acquire a target text and a target video type, wherein the target text is text associated with a target video, and the target video type is a video type of the target video; Extracting semantic features of the target text to obtain text semantic features of the target text; Extracting type features of the target video type to obtain video type features of the target video; Based on the text semantic features and the video type features, performing text content recognition on the target text to obtain a text content recognition result, wherein the text content recognition result is used to indicate the relationship between the target text and the target video and a specific type of text; In response to the text content recognition result indicating that the target text belongs to the specific type of text, determining an associated screen for the target text, wherein the associated screen is selected according to the sending time of the target text; Performing content recognition on the associated pictures to obtain picture content features; Based on the picture content features and the text semantic features, the degree of relationship between the target text and the specific type of text is determined.
2. The method according to claim 1, characterized in that The performing text content recognition on the target text based on the text semantic features and the video type features to obtain a text content recognition result includes: Obtaining a semantic fusion feature based on the text semantic feature and the video type feature, wherein the semantic fusion feature is used to indicate the meaning when the target text is an associated text of a video corresponding to the target video type; Based on the semantic fusion feature and the video type feature, the text content recognition is performed on the target text to obtain the text content recognition result.
3. The method according to claim 2, characterized in that The obtaining of semantic fusion features based on the text semantic features and the video type features includes: Determining a correlation coefficient based on the text semantic feature and the video type feature, wherein the correlation coefficient is used to indicate the correlation between the target text and the target video type; Based on the correlation coefficient, the text semantic features are weighted to obtain the semantic fusion features.
4. The method according to claim 3, characterized in that The determining of the correlation coefficient based on the text semantic feature and the video type feature includes: Performing full connection processing on the text semantic features through a first fully connected layer to obtain a text semantic feature vector; Performing the fully connected processing on the video type feature through a second fully connected layer to obtain a video type feature vector, where the vector dimension of the video type feature vector is the same as the vector dimension of the text semantic feature vector; The text semantic feature vector is multiplied by the video type feature vector to obtain the correlation coefficient.
5. The method according to claim 2, characterized in that The performing the text content recognition on the target text based on the semantic fusion feature and the video type feature to obtain the text content recognition result includes: Perform feature splicing on the semantic fusion feature and the video type feature to obtain the target text feature; The target text features are input into a classifier to obtain the probability that the target text belongs to the specific type of text, and the classifier includes a third fully connected layer and a Softmax layer.
6. The method according to any one of claims 1 to 5, characterized in that: The extracting type features of the target video type to obtain the video type features of the target video includes: Feature encoding is performed on the target video type to obtain video type features of the target video.
7. The method according to any one of claims 1 to 5, characterized in that: The performing text content recognition on the target text based on the text semantic features and the video type features to obtain a text content recognition result includes: The text semantic features and the video type features are input into a classification network to obtain the text content recognition result. The classification network is trained based on sample text, sample video type and sample text label. The sample text is text associated with a sample video, the sample video type is the video type of the sample video, and the sample text label is used to indicate the relationship between the sample text and the specific type of text.
8. The method according to claim 7, characterized in that The method further comprises: Performing the semantic feature extraction on the sample text to obtain a sample text semantic feature of the sample text, and performing the type feature extraction on the sample video type to obtain a sample video type feature of the sample video; Inputting the sample text semantic features and the sample video type features into the classification network to obtain a predicted text content recognition result; The classification network is trained based on the predicted text content recognition result and the sample text label.
9. The method according to any one of claims 1 to 5, characterized in that: The target text is a target bullet comment; and determining the associated screen of the target text includes: Determine the associated screen of the target barrage, where the associated screen is selected according to the sending time of the target barrage.
10. The method according to claim 9, characterized in that The determining, based on the picture content features and the text semantic features, the degree of relationship between the target text and the specific type of text includes: Determining the degree of correlation between the picture content feature and the text semantic feature; In response to the correlation degree being higher than a correlation threshold, determining that the target bullet comment is a non-specific type of text; In response to the correlation degree being lower than the correlation threshold, determining that the target bullet comment is the specific type of text.
11. A text content recognition device, characterized in that: The device comprises: An acquisition module, configured to acquire a target text and a target video type, wherein the target text is text associated with a target video, and the target video type is a video type of the target video; A first feature extraction module is used to extract semantic features of the target text to obtain text semantic features of the target text; A second feature extraction module is used to extract type features of the target video type to obtain video type features of the target video; A first recognition module is configured to perform text content recognition on the target text based on the text semantic features and the video type features to obtain a text content recognition result, wherein the text content recognition result is used to indicate a relationship between the target text and the target video and a specific type of text; a picture determination module, configured to determine a picture associated with the target text in response to the text content recognition result indicating that the target text belongs to the specific type of text, wherein the picture associated with the target text is selected according to the time when the target text is sent; A third recognition module is used to perform content recognition on the associated screen to obtain screen content features; A determination module is used to determine the degree of relationship between the target text and the specific type of text based on the picture content features and the text semantic features.
12. The device according to claim 11, characterized in that The first identification module includes: a feature fusion unit, configured to obtain a semantic fusion feature based on the text semantic feature and the video type feature, wherein the semantic fusion feature is used to indicate the meaning when the target text is an associated text of a video corresponding to the target video type; The recognition unit is configured to perform the text content recognition on the target text based on the semantic fusion feature and the video type feature to obtain the text content recognition result.
13. The device according to claim 12, characterized in that The feature fusion unit is further used for: Determining a correlation coefficient based on the text semantic feature and the video type feature, wherein the correlation coefficient is used to indicate the correlation between the target text and the target video type; Based on the correlation coefficient, the text semantic features are weighted to obtain the semantic fusion features.
14. The device according to claim 13, characterized in that The feature fusion unit is further used for: Performing full connection processing on the text semantic features through a first fully connected layer to obtain a text semantic feature vector; Performing the fully connected processing on the video type feature through a second fully connected layer to obtain a video type feature vector, where the vector dimension of the video type feature vector is the same as the vector dimension of the text semantic feature vector; The text semantic feature vector is multiplied by the video type feature vector to obtain the correlation coefficient.
15. The device according to claim 12, characterized in that The identification unit is further configured to: Perform feature splicing on the semantic fusion feature and the video type feature to obtain the target text feature; The target text features are input into a classifier to obtain the probability that the target text belongs to the specific type of text, and the classifier includes a third fully connected layer and a Softmax layer.
16. The device according to any one of claims 11 to 15, characterized in that The second feature extraction module is further configured to: Feature encoding is performed on the target video type to obtain video type features of the target video.
17. The device according to any one of claims 11 to 15, characterized in that The first identification module is further configured to: The text semantic features and the video type features are input into a classification network to obtain the text content recognition result. The classification network is trained based on sample text, sample video type and sample text label. The sample text is text associated with a sample video, the sample video type is the video type of the sample video, and the sample text label is used to indicate the relationship between the sample text and the specific type of text.
18. The device according to claim 17, characterized in that The device further comprises: a third feature extraction module, configured to extract the semantic features of the sample text to obtain the sample text semantic features of the sample text, and to extract the type features of the sample video type to obtain the sample video type features of the sample video; A second recognition module is configured to input the sample text semantic features and the sample video type features into the classification network to obtain a predicted text content recognition result; A training module is used to train the classification network based on the predicted text content recognition result and the sample text label.
19. The device according to any one of claims 11 to 15, characterized in that The target text is a target bullet screen; the picture determination module is further used to: Determine the associated screen of the target barrage, where the associated screen is selected according to the sending time of the target barrage.
20. The device according to claim 19, characterized in that The determining module includes: A first determining unit, configured to determine a correlation between the picture content feature and the text semantic feature; a second determining unit, configured to determine, in response to the correlation degree being higher than a correlation threshold, that the target bullet comment is a non-specific type of text; The third determining unit is configured to determine that the target bullet comment is the specific type of text in response to the correlation degree being lower than the correlation threshold.
21. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the text content recognition method according to any one of claims 1 to 10.
22. A computer-readable storage medium, characterized in that The readable storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the text content recognition method according to any one of claims 1 to 10.
23. A computer program product, characterized in that The computer program product includes computer instructions, and the computer instructions are executed by a processor to implement the text content recognition method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Bullet screen information interception method and device, storage medium and equipment
CN110674247A
Data processing method and device, equipment and storage medium
CN113705563A