Content classification method and device, storage medium and computer equipment
By constructing a cross-modal map based on images and text, combining object recognition and text encoding technology, the problem that the existing technology cannot identify hidden inferior content is solved, and the accurate identification of satirical content is achieved.
Patent Information
- Application Number
- CN202311577436.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art cannot accurately and effectively identify inferior content, especially those that are more concealed in the form presented through metaphor or satirical techniques.
By obtaining the object recognition results of the image in the content to be classified and encoding it based on the object information, an image feature matrix is obtained; at the same time, the text is encoded to obtain the word participle feature matrix. Then, the edge weights of the node pair are determined based on the image characteristics and word participle characteristics, a cross-modal map is constructed, and cross-modal feature extraction and classification are performed.
Effectively identify content to be classified with ironic attributes with interactive characteristics, and improve the accuracy of recognition of inferior content.
Smart Images

Figure CN120030379A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a content classification method, apparatus, storage medium and computer equipment. Background Art
[0002] Content review is an important task that Internet platforms cannot avoid. The stock of video content and graphic content on the Internet has grown exponentially with the development of information technology, but the quality of content varies. Content review can ensure the safety and health of network content. Due to the large amount of new graphic content and video content added daily on various Internet platforms, it is not realistic to rely entirely on manual review.
[0003] Existing technologies mainly use recognition algorithms to automatically identify content. However, there are many low-quality contents in a hidden form. For example, a common method is to modify or package the sentences in the article or title by using metaphors or irony, combined with the content of the article or video, to achieve the purpose of innuendo. The current recognition algorithm cannot accurately and effectively identify low-quality content. Summary of the invention
[0004] The embodiments of the present application provide a content classification method, apparatus, storage medium and computer equipment to solve the problem in the related art that low-quality content cannot be accurately and effectively identified.
[0005] On the one hand, an embodiment of the present application provides a content classification method, the method comprising: obtaining an object recognition result of an image in content to be classified, the object recognition result comprising object information of each object recognized from the image, the object information comprising at least a pixel area where the object is located; encoding based on the object information in the object recognition result to obtain a first encoding matrix; the first encoding matrix comprising image features of each object; encoding text in the content to be classified to obtain a second encoding matrix, the second encoding matrix comprising segmentation features of each word in the text; determining an edge weight of each node pair according to the image features of each object and the segmentation features of each word; the node pair is formed by two nodes representing different segmentation features, or by a node representing a segmentation feature and a node representing an image feature; constructing a cross-modal graph for the content to be classified based on the edge weights, segmentation features and image features of the node pairs; extracting and classifying cross-modal features based on the cross-modal graph to obtain a classification result of the content to be classified.
[0006] On the other hand, an embodiment of the present application also provides a content classification device, which includes: an object recognition module, which is used to obtain an object recognition result of an image in the content to be classified, the object recognition result includes object information of each object recognized from the image, and the object information includes at least a pixel area where the object is located; a first encoding module, which is used to encode based on the object information in the object recognition result to obtain a first encoding matrix; the first encoding matrix includes image features of each object; a second encoding module, which is used to encode the text in the content to be classified to obtain a second encoding matrix, and the second encoding matrix includes segmentation features of each word in the text; a weight determination module, which is used to determine the edge weight of each node pair according to the image features of each object and the segmentation features of each word; the node pair is formed by two nodes representing different segmentation features, or by a node representing a segmentation feature and a node representing an image feature; a graph construction module, which is used to construct a cross-modal graph for the content to be classified based on the edge weights, segmentation features and image features of the node pairs; a content classification module, which is used to extract and classify cross-modal features based on the cross-modal graph to obtain a classification result of the content to be classified.
[0007] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the above-mentioned content classification method is executed when the computer program is executed by a processor.
[0008] On the other hand, an embodiment of the present application further provides a computer device, which includes a processor and a memory, wherein the memory stores a computer program, and the computer program executes the above-mentioned content classification method when called by the processor.
[0009] On the other hand, an embodiment of the present application also provides a computer program product, which includes a computer program stored in a storage medium; a processor of a computer device reads the computer program from the storage medium, and the processor executes the computer program, so that the computer device executes the steps in the above-mentioned content classification method.
[0010] The content classification method provided by the present application can obtain the object recognition result of the image in the content to be classified, the object recognition result includes the object information of each object recognized from the image, and the object information at least includes the pixel area where the object is located, and then, based on the object information in the object recognition result, encoding is performed to obtain a first encoding matrix, the first encoding matrix includes the image features of each object, and the text in the content to be classified is encoded to obtain a second encoding matrix, the second encoding matrix includes the word segmentation features of each word in the text, and further, according to the image features of each object and the word segmentation features of each word, the edge weight of each node pair is determined, and then, based on the edge weights, word segmentation features and image features of the node pairs, a cross-modal graph for the content to be classified is constructed, and cross-modal features are extracted and classified based on the cross-modal graph to obtain the classification result of the content to be classified.
[0011] Therefore, the cross-modal graph for the content to be classified constructed based on the edge weights, word segmentation features and image features of node pairs can cross-modally connect the image features about the image and the word segmentation features about the text in the content to be classified in the form of a graph, so that cross-modal feature extraction can be performed on the content to be classified based on the cross-modal graph, and cross-modal interaction information that effectively represents the image modality and the text modality can be obtained, thereby effectively identifying the content to be classified with sarcastic attributes with interactive characteristics, thereby improving the accuracy of identifying inferior content. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0013] Figure 1 A schematic diagram of a system architecture provided in an embodiment of the present application is shown.
[0014] Figure 2 An application scenario diagram of a content classification method provided in an embodiment of the present application is shown.
[0015] Figure 3 A flow chart of a content classification method provided in an embodiment of the present application is shown.
[0016] Figure 4 The figure shows an architecture diagram of a target generation network provided in an embodiment of the present application.
[0017] Figure 5 A schematic diagram of image recognition provided by an embodiment of the present application is shown.
[0018] Figure 6 A schematic diagram of a dependency tree provided in an embodiment of the present application is shown.
[0019] Figure 7 A flow chart of another content classification method provided in an embodiment of the present application is shown.
[0020] Figure 8 A network training framework diagram provided in an embodiment of the present application is shown.
[0021] Fig. 9 It is a module block diagram of a content classification device provided in an embodiment of the present application.
[0022] Fig.10 It is a module block diagram of a computer device provided in an embodiment of the present application.
[0023] Fig.11 It is a module block diagram of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0025] In some processes described in the specification, claims and the above drawings, multiple steps appearing in a specific order are included, but it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, the descriptions such as "first" and "second" in this document are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0026] The specification refers to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0027] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0028] It should be noted that in the specific implementation of this application, the images and texts of the content to be classified and other related data, when applied to the specific products or technologies of the embodiments of this application, need to obtain user permission or consent, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and carry out subsequent data use and processing within the scope of authorization of the laws, regulations and personal information subjects.
[0029] The content classification method proposed in this application involves artificial intelligence (AI) technology, which is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0030] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called the big model or the basic model. After fine-tuning, it can be widely used in downstream tasks in various major directions of artificial intelligence. Among them, the pre-trained network is also called the big network or the basic network. After fine-tuning, it can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0031] Among them, deep learning (DL) is a sub-problem of machine learning, and its main purpose is to automatically learn effective feature representations from data. Through multi-layer feature conversion, the original data is transformed into a higher-level, more abstract representation. These learned representations can replace manually designed features, thereby avoiding "feature engineering". The abstract representation is further input into the prediction function to obtain the final result. For example, in an embodiment of the present application, the text in the classified content is encoded.
[0032] Computer Vision Technology (CV) Computer vision is a science that studies how to make machines "see". To put it more specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing so that the computer processes the images into images that are more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained networks in the visual field such as Swin-Transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning (Fine Tune).
[0033] Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition, etc., and also includes common biometric recognition technologies such as face recognition, fingerprint recognition, palm print recognition, etc. For example, in an embodiment of the present application, an object recognition result of an image in the content to be classified is obtained, and the object recognition result includes object information of each object recognized from the image based on image recognition.
[0034] The rapid development of Internet information has led to an increasing amount of video content and graphic content on the Internet, but the quality of various contents varies. In order to ensure the safety and health of online content, Internet platforms need to review various contents. At present, there are many hidden forms of low-quality content (that is, illegal or illegal content) on online platforms. The hidden means include but are not limited to: homophone replacement, synonym replacement, abbreviation, rewriting, expansion, homophony, etc. For example, by combining the visual information in the article or video (cover image, illustration, video frame, etc.), the harmful text or title is rewritten using irony.
[0035] Existing technologies mainly focus on processing single-mode texts, including keyword matching-based methods, which use keyword filtering to identify text content containing corresponding metaphors and sarcastic words as low-quality. Methods based on various text similarity calculations, which calculate the similarity between the text to be identified and the sarcastic text, and identify the corresponding text as sarcastic when the similarity is greater than a threshold. Machine learning / deep learning-based methods use large-scale labeled data to train machine learning or deep learning networks to identify sarcastic texts.
[0036] Among them, the methods based on keyword matching and the methods based on various text similarity calculations can only rely on the existing semantic rules of the vocabulary and the cases of sarcastic texts. It is difficult to capture the changes in the meaning of words and cannot be updated in real time. In addition, this rule-based recognition method performs poorly in the accuracy of verbal abuse recognition. In addition, the methods based on various text similarity calculations require the use of text similarity matching algorithms, which are limited by the accuracy of various similarity algorithms. The methods based on machine learning / deep learning cannot avoid the need to use a large amount of labeled data and computing resources for network training.
[0037] Since this kind of ironic content uses a lot of knowledge, language background and cross-modal visual information, it is highly confusing and easily leads to recognition failure. In particular, as a large number of new words and hot topics continue to emerge on the Internet platform, the expression of irony is becoming more and more diverse. The existing technology cannot effectively identify the ironic content of interactive pictures and texts, especially for cases of ironic content with strong correlation between pictures and texts. The recognition network cannot provide timely and effective feedback. In order to solve the above problems, the inventor has proposed a content classification method provided in the embodiment of this application after research.
[0038] The following is an introduction to the system architecture and application scenarios of the content classification method involved in this application.
[0039] like Figure 1 As shown, the content classification method provided in the embodiment of the present application can be applied in the system 100, and the data acquisition device 110 is used to obtain training data. For the content classification method of the embodiment of the present application, the training data can be a training sample used for network training, and the training sample includes sample content and a label of the sample content. Among them, the training sample can be obtained after data preprocessing based on the collected original graphic data and video data. After acquiring the training data, the data acquisition device 110 can store the training data in the database 120, and the training device 130 can train the target network 101 based on the training data maintained in the database 120.
[0040] Specifically, the training device 130 can train the preset neural network based on the input training data until the preset neural network meets the preset conditions to obtain the trained target network 101. Among them, the preset conditions can be: the total loss value of the target loss is less than the preset value, the total loss value of the target loss is within the threshold range, or the number of training times reaches the preset number of times, etc. The target network 101 can be used to implement the content classification method in the embodiment of the present application. The target network 101 in the embodiment of the present application can be a deep neural network network, for example, a content classification network composed of a text encoder BERT, a visual encoder VisionTransformer, and a graph convolution (Graph convolution Network, GCN), etc., which is not limited here.
[0041] In actual application scenarios, the training data maintained in the database 120 may not necessarily all come from the data acquisition device 110, but may also be received from other devices. For example, the execution device 140 may also serve as a data acquisition terminal, and use the acquired data as new training data and store it in the database 120. In addition, the training device 130 may not necessarily train the preset neural network based entirely on the training data maintained by the database 120, but may also train the preset neural network based on the training data obtained from the cloud or other devices. For example, the execution device 140 may use the collected graphic data as training data. The above description should not be used as a limitation on the embodiments of the present application.
[0042] The target network 101 obtained by training the training device 130 can be applied to different systems or devices, such as Figure 1 The execution device 140 shown. The training device 130 and the execution device 140 can be a server or a terminal, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), blockchains, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, etc.
[0043] During the process of the processing module 141 of the execution device 140 performing processing related to calculations and the like, the execution device 140 can call data, programs, etc. in the data storage system 150 for corresponding calculation processing, and store data and instructions such as the processing results obtained from the calculation processing into the data storage system 150. The training device 130 can generate corresponding target networks 101 based on different training data for different targets or different tasks, and the corresponding target networks 101 can be used to complete the training tasks of the corresponding application clustering network and the application clustering tasks performed using the application clustering network.
[0044] It should be noted that Figure 1 This is only a schematic diagram of the architecture of a system provided by the embodiments of the present application. The architecture and application scenarios of the system described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. For example, Figure 1 the data storage system 150 in is an external memory relative to the execution device 140. In other cases, the data storage system 150 can also be placed in the execution device 140.
[0045] Please refer to Figure 2 , Figure 2 which shows an application scenario diagram of a content classification method. As Figure 2 shown, the content classification method provided by the embodiments of the present application can also be applied in a Client-Server (C / S) architecture system. In the application scenario shown in Figure 2 , the content classification service provider provides a server side, and the server side can include a training server 210 and a running server 230 in the cloud. The server side can communicate with the first terminal 220, the second terminal 240, and the third terminal 260 through the network respectively to provide content classification services for the users of each terminal. Among them, the first terminal 220 can communicate with the server side through the network router 250.
[0046] The training server 210 can perform network training on a preset neural network for content classification based on training data to obtain a content classification network. The running server 230 can deploy the content classification network so as to, in response to a plurality of contents to be classified sent by the terminal, obtain classification results of the plurality of contents to be classified based on the content classification network, and then send the classification results to the corresponding terminal.
[0047] For example, the second terminal 240 installs a content sharing client 241, on which the user can send the social content edited by the user, which can be composed of images and text. When the user finishes editing the social content and clicks "Publish", the content sharing client 241 can send the social content to the operation server 230, and then the operation server 230 can classify the social content based on the content classification network to determine whether it is low-quality content. Furthermore, the content sharing client 241 can receive the content classification result sent by the operation server 230. If the classification result of the social content is low-quality content, the content sharing client 241 can end the publishing process of the social content and prompt the user to modify the content.
[0048] Among them, the operation server 230 may include at least one processor 231, a memory 232, and at least one communication interface 234. The various components in the operation server 230 are coupled together through a bus system 233. It is understandable that the bus system 233 is used to realize the connection and communication between these components. The memory 232 can store data to support various operations, and examples of these data include programs, modules and data structures or subsets or supersets thereof. The operating system 2321 in the memory 232 includes system programs for processing various basic system services and performing hardware-related tasks. For example, the framework layer, the core library layer, the driver layer, etc. are used to implement various basic businesses and process hardware-based tasks.
[0049] The image generation device provided in the embodiment of the present application can be implemented in a software manner. Figure 2 The content classification device 2322 stored in the memory 232 is shown, which can be software in the form of a program and a plug-in, etc., and includes the following modules: an object recognition module 310, a first encoding module 320, a second encoding module 330, a weight determination module 340, a graph construction module 350, and a content classification module 360. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0050] It should be noted that Figure 2It is only a schematic diagram of an application scenario provided by an embodiment of the present application. The application scenario and system framework described in the embodiment of the present application are only to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided by the embodiment of the present application. For example, the first terminal 220, the second terminal 240 or the third terminal 260 can refer to one of the multiple terminals, and this embodiment is only illustrated by the first terminal 220, the second terminal 240 or the third terminal 260. In addition to the application scenario in which the platform classifies the content to be sent by the user, the present application can also be used to classify the content that has been published in the platform, such as graphic content, live content, etc., which is not limited here. In addition, the bus system 233 includes a power bus, a control bus and a status signal bus in addition to the data bus, which is not limited here. The target generation network can also be deployed directly on the terminal. It is known to those of ordinary skill in the art that with the evolution of the application scenario or system architecture, the technical solution provided in the embodiment of the present invention is also applicable to similar technical problems.
[0051] See also Figure 3 , Figure 3 The following is a flow chart of a content classification method provided by an embodiment of the present application. In this embodiment, the content classification method can be executed by a computer device (the computer device can be a server, an edge computing device, or other device with certain processing capabilities), and the computer device has at least storage, computing, and communication functions. Figure 3 The process shown combines Figure 4 The network training diagram shown in FIG. Figure 3 As shown, the content classification method may specifically include the following steps:
[0052] Step S110: Obtaining an object recognition result of an image in the content to be classified, wherein the object recognition result includes object information of each object recognized from the image, and the object information at least includes a pixel area where the object is located.
[0053] Considering that the existing technology cannot effectively classify the content of interactive satirical content, this application proposes a method for content classification based on a cross-modal graph between images and texts, so as to capture the interactive information between different modalities and improve the accuracy of content classification. To this end, a series of objects need to be identified from the image of the content to be classified.
[0054] Among them, the content to be classified refers to the content consisting of pictures and texts that need to be classified. The content to be classified can be graphic content, or it can be composed of video pictures or live pictures, and text extracted from the pictures or based on the corresponding voice of the pictures. Objects refer to target objects of various patterns or types identified by image recognition algorithms, including plants, animals or other entities, such as trees, clouds, houses, etc. Object information may include pixel areas and object attribute descriptions. The pixel area refers to the area where the object is located in the image of the content to be classified. The object attribute description is information used to describe the attributes of the object, including the object name and attribute description words.
[0055] As an implementation method, a computer device can obtain content to be classified. For example, a graphic content composed of images and texts posted by a user on a social platform is obtained as content to be classified. The video to be detected is subjected to frame extraction processing to obtain a video frame, and optical character recognition (OCR) is used to perform text recognition on the video frame to obtain the text corresponding to the video frame, and then the video frame and its text are used as content to be classified. During the live broadcast process, the live screen in the live stream is obtained, and speech recognition (Automatic Speech Recognition, ASR) is performed on the live audio to generate corresponding text content, or the subtitles in the live stream are directly obtained as text content, and then the live screen and the text content are used as content to be classified. Furthermore, based on the image recognition algorithm, image recognition is performed on the image of the content to be classified, and an object recognition result can be obtained, wherein the object recognition result includes the object information of each object recognized from the image.
[0056] See also Figure 5 , Figure 5 A schematic diagram of image recognition is shown. Figure 5 As shown, a deep image recognition network based on a convolutional neural network can be used to perform image recognition on an image of content to be classified. Specifically, the visual area of each object in the image can be determined through image recognition, and an object attribute description corresponding to each visual area can be generated. The object attribute description includes an object name and an attribute description word, for example, Figure 5 The visual area A of the object A is identified and the corresponding object attribute description A of the object A is generated. Further, the visual area of each object can be formatted uniformly to obtain the corresponding pixel area for image feature extraction. For example, the visual area of each object identified Perform uniform scaling to obtain the corresponding pixel area Optionally, L h =L w =224,L h and Lw Represents the height and width of the pixel area respectively.
[0057] Step S120: Encoding is performed based on the object information in the object recognition result to obtain a first encoding matrix; the first encoding matrix includes image features of each object.
[0058] As an implementation method, the computer device can encode the pixel area of the object information in the object recognition result through an image encoder, that is, extract the image features, and obtain a first encoding matrix. Specifically, the pixel area of each object can be segmented to obtain a representation sequence corresponding to each pixel area, the representation sequence is a data format that can be encoded, and each representation sequence is encoded based on the image encoder to obtain a first encoding matrix, which includes the image features of each object represented based on the feature vector.
[0059] For example, each pixel region I is converted into its corresponding representation sequence after segmentation operation Where r = p × p, p represents the unit area in the pixel area, and r represents the number of unit areas obtained after the pixel area is segmented, that is, the length of the representation sequence. After encoding by the image encoder, the first encoding matrix Z = {z 1 ,z 2 ,…,z n}, where n>0&n∈N * , z is the feature vector representing the image features.
[0060] Step S130: Encode the text in the content to be classified to obtain a second encoding matrix, where the second encoding matrix includes the segmentation features of each segmentation in the text.
[0061] As an implementation method, the text in the content to be classified can be encoded by a text encoder, that is, word segmentation features can be extracted to obtain a second encoding matrix. Specifically, the text can be segmented to obtain a word segmentation sequence S corresponding to the text = {w 1 ,w 2 ,…,w m}, where m is the length of the word sequence, that is, the number of words in the text. Furthermore, each word w in the word sequence S can be mapped into a low-dimensional text vector through a pre-trained language model. For example, each word in the word sequence S is mapped through a pre-trained BERT-Base model to obtain a word vector matrix Among them, X T T in the text is a transposition symbol. [CLS] and [SEP] represent the sequence start symbol and word boundary symbol of the word sequence S, respectively. [CLS] and [SEP] do not participate in text encoding.
[0062] Furthermore, the text encoder can be used based on the word vector matrix X T Encode to obtain the second encoding matrix. For example, the word vector matrix X is encoded by a bidirectional long short-term memory network (Bidirectional Long Short-Term Memory, Bi-LSTM) T Encode and obtain the second encoding matrix T = {t 1 ,t 2 ,…,t m}, where m>0&m∈N * , t is the feature vector representing the word segmentation features.
[0063] Step S140: Determine the edge weight of each node pair according to the image features of each object and the segmentation features of each segmentation; a node pair is formed by two nodes representing different segmentation features, or by a node representing a segmentation feature and a node representing an image feature.
[0064] Since the existing technology cannot obtain the deep interactive information between images and texts from interactive content, it is impossible to accurately classify the content. To this end, this application proposes to construct a cross-modal graph between images and texts of the content to classify the content, so as to mine the interactive information between the two modes of images and texts based on the cross-modal graph and improve the accuracy of content classification. The cross-modal graph is an undirected graph composed of different nodes. The cross-modal graph G =<V,E> .
[0065] Where V is a non-empty set, used to represent the node set {v 1 ,v 2 ,…,v n+m},v i represents the i-th node (vertex) in the cross-modal graph G, node v i It can be a node representing a word segmentation feature or a node representing an image feature. E is a non-empty set, used to represent the edge set {e 1 ,e 2 ,…,e k},e i represents an edge of a node pair in the cross-modal graph G, and an edge of a node pair e i It is formed by two nodes representing different word segmentation features, or by a node representing word segmentation features and a node representing image features. The edge e of each node pair i There are corresponding edge weights.
[0066] Since each node in the cross-modal graph can aggregate information from its neighboring nodes through the edges between nodes, the key to constructing a cross-modal graph lies in the setting of the edge weights in the graph. Considering that there is related semantic information between the segmentations in the text of the content to be classified, the edge weights between two nodes representing segmentation features are determined. And in order to mine the interactive information between the two modalities of picture and text, the edge weights between the nodes representing segmentation features and the nodes representing image features are determined. In some embodiments, step S140 may specifically include steps S141 and S142:
[0067] Step S141: for a first node pair formed by two nodes representing different word segmentation features, determine an edge weight of the first node pair according to a dependency relationship between two word segmentations involved in the first node pair.
[0068] As an implementation method, a computer device can form a first node pair based on two nodes representing different word segmentation features, and determine whether there is a dependency relationship between the word segmentations corresponding to the two word segmentation features in the first node pair based on a dependency tree corresponding to the text in the content to be classified. If there is a dependency relationship, the edge weight of the first node pair is 1. Otherwise, if there is no dependency relationship, the edge weight of the first node pair is 0.
[0069] Please participate Figure 6 , Figure 6 A schematic diagram of a dependency tree is shown. Figure 6 As shown, a computer device can obtain text of content to be classified, for example, a sentence: "What a wonderful weather!", and can perform dependency parsing on the sentence. Specifically, each word after word segmentation is input into a dependency parser. The dependency parser (such as spaCy) can create a dependency tree for the input word segmentation and output the dependency relationship between the word segments in the form of typed dependency relationship connection. Figure 6 The dependency relationships between "weather" and "really", "one", "wonderful" and "!" are shown in .
[0070] Since ironic information in textual modality is often expressed through multiple words, as Figure 6 Therefore, the grammatical perception relationship between word segments on the sentence dependency tree is represented in the cross-modal graph through the edge weights of the node pairs, which can promote the learning of contextual dependencies. Thus, the accuracy of the feature representation of the classified content is improved.
[0071] Step S142: For a second node pair formed by a node representing a segmentation feature and a node representing an image feature, determine the edge weight of the second node pair based on the feature similarity between the target segmentation involved in the second node pair and the object name of the target object involved, the sentiment weight of the target segmentation, and the sentiment weight of the attribute description word of the target object.
[0072] As an implementation, the computer device may form a node pair, namely, a second node pair, with the node representing the word segmentation feature and the node representing the image feature, and then calculate the edge weight of the second node pair. Specifically, step S142 may include (1) to (5).
[0073] (1) Based on the segmentation features of the target segmentation and the semantic features of the object name of the target object, the feature similarity between the target segmentation and the object name of the target object is calculated.
[0074] As an implementation method, the computer device can obtain all node pairs consisting of nodes representing word segmentation features and nodes representing image features. Further, the feature similarity of each node pair is calculated. Specifically, the node pair whose feature similarity is to be calculated, that is, the target node pair, is obtained, and the target object corresponding to the node representing the image feature in the target node pair is determined, and the semantic features of the object name of the target object are obtained. Then, based on the word segmentation features of the target word segmentation and the semantic features of the object name of the target object, the feature similarity Sim(w i ,o j ), where w i To represent the target word, o j is the object name, Sim(·) represents an algorithm or tool for calculating the similarity between two elements, and Sim(·) can calculate the feature similarity between the target segmentation and the object name of the target object based on the segmentation features of the target segmentation and the semantic features of the object name of the target object. Optionally, Sim(·) can be NLTK, etc., which is not limited here.
[0075] (2) Based on the segmentation features of the target segmentation, the sentiment weight of the target segmentation is obtained.
[0076] (3) Based on the semantic features of the attribute descriptors of the target object, the sentiment weight of the attribute descriptors of the target object is obtained.
[0077] Considering that the satirical content in the content to be classified all has strong emotional conflicts, this application proposes to accurately identify the satirical content by capturing the emotional conflict features in the content. To this end, an adjustment factor can be determined based on the emotional weight of the target segmentation and the emotional weight of the attribute description word of the target object, and the feature similarity between the target segmentation and the object name of the target object is weighted based on the adjustment factor, so that the interactive information between the image modality and the text modality with emotional conflict features can be mined on the cross-modal graph.
[0078] As an implementation method, the computer device can determine the sentiment weight of the target word segment and the sentiment weight of the attribute description word of the target object based on the sentiment dictionary. For example, the segmentation feature of the target word segment is input into the sentiment dictionary to obtain the sentiment weight ω(w i ), input the semantic features of the attribute description words of the target object into the sentiment dictionary to obtain the sentiment weight ω(a j ). Among them, ω(w i ),ω(a j )∈[―1,1], if the target word or the attribute description of the target object is not included in the sentiment dictionary, then ω(w i ) or ω(a j ) has a value of 0. Optionally, the sentiment dictionary may be SenticNet, which is not limited here.
[0079] (4) According to the sentiment weight of the target segmentation word and the sentiment weight of the attribute description word of the target object, an adjustment factor is determined. The adjustment factor is used to characterize the inconsistency of the sentiment tendency between the target segmentation word and the target object.
[0080] As an implementation method, the computer device may multiply the sentiment weight of the target segmentation word with the sentiment weight of the attribute description word of the target object to obtain an intermediate adjustment factor, and determine a first adjustment coefficient based on the intermediate adjustment factor, and the first adjustment coefficient is negatively correlated with the intermediate adjustment factor. Further, the computer device may subtract the sentiment weight of the target segmentation word from the sentiment weight of the attribute description word of the target object, and use the absolute value of the subtraction result as the second adjustment coefficient, and multiply the first adjustment coefficient by the second adjustment coefficient to obtain the adjustment factor. The calculation formula is as follows:
[0081]
[0082] Among them, ξ i,j represents the adjustment factor, ω(w i )ω(a j ) represents the intermediate adjustment factor, represents the first adjustment coefficient, |ω(w i )―ω(a j)| represents the second adjustment coefficient. γ is a hyperparameter used to control the value of the adjustment factor, that is, to adjust the deviation caused by the opposite emotional relationship. The value of γ can be dynamically adjusted manually based on experimental experience and actual application requirements. Specifically, if ω(w i ) and ω(a j ) has the opposite emotional polarity, the value of the first adjustment coefficient will increase. Otherwise, it will decrease. The second adjustment coefficient means that the greater the emotional weight, the higher the confidence of the increase or decrease of the first adjustment coefficient.
[0083] (5) Determine the edge weight of the second node pair based on the feature similarity and the adjustment factor.
[0084] As an implementation, the computer device may determine the edge weight of the second node pair based on the product of the feature similarity and the adjustment factor. The calculation is as follows:
[0085] e i,j =Sim(w i ,o j )×ξ i,j +1 (2)
[0086] Among them, e i,j Represents the edge weight of the second node pair.
[0087] Step S150: construct a cross-modal graph for the content to be classified based on the edge weights, word segmentation features and image features of the node pairs.
[0088] As an implementation method, the computer device may represent the adjacent relationship between nodes in the cross-modal graph of the content to be classified based on the edge weights of the node pairs, the word segmentation features, and the adjacency matrix determined by the image features. Specifically, the adjacency matrix corresponding to the cross-modal graph is Can be defined as:
[0089]
[0090] Among them, A i,j represents the matrix element in the adjacency matrix A, that is, the edge weight of the corresponding node pair. Since the cross-modal graph is an undirected graph, in the adjacency matrix corresponding to the cross-modal graph, A i,j =A j,i , and each node is a self-loop, that is, A i,i = 1. Then, the adjacency matrix and the nodes represented by different image features and word segmentation features in the cross-modal graph are combined to represent the cross-modal graph.
[0091] Step S160: extract and classify cross-modal features based on the cross-modal graph to obtain classification results for the content to be classified.
[0092] This application proposes to extract and identify key clues of irony in the content to be classified by aggregating the relevance of nodes in the cross-modal graph. To this end, the relevance of nodes in the cross-modal graph can be aggregated to obtain the feature representation of each node on the cross-modal graph to classify the content to be classified.
[0093] As an implementation method, a computer device can extract and classify cross-modal features based on the adjacency matrix of a cross-modal graph through a modal fusion network to obtain a classification result of the content to be classified. Among them, the modal fusion network may include a representation subnetwork and a classification subnetwork. Specifically, the adjacency matrix of the cross-modal graph is input into the representation subnetwork for cross-modal feature extraction to obtain cross-modal features. Further, the classification subnetwork is used to classify the content based on the cross-modal features to obtain the classification result of the content to be classified. Optionally, the representation subnetwork can be a graph convolutional neural network with a joint attention mechanism, and the classification subnetwork can be a fully connected layer, which is not limited here.
[0094] This embodiment can obtain object recognition results of images in content to be classified, wherein the object recognition results include object information of each object recognized from the image, and the object information at least includes the pixel area where the object is located. Then, encoding is performed based on the object information in the object recognition results to obtain a first encoding matrix, which includes image features of each object, and the text in the content to be classified is encoded to obtain a second encoding matrix, which includes word segmentation features of each word in the text. Further, according to the image features of each object and the word segmentation features of each word, the edge weight of each node pair is determined. Then, based on the edge weights, word segmentation features and image features of the node pairs, a cross-modal graph for the content to be classified is constructed, and cross-modal features are extracted and classified based on the cross-modal graph to obtain a classification result of the content to be classified.
[0095] In this way, the cross-modal graph for the content to be classified constructed based on the edge weights, word segmentation features and image features of node pairs can cross-modally connect the image features about the image and the word segmentation features about the text in the content to be classified in the form of a graph, so that cross-modal feature extraction can be performed on the content to be classified based on the cross-modal graph, and cross-modal interaction information that effectively represents the image modality and the text modality can be obtained, thereby effectively identifying the content to be classified with sarcastic attributes with interactive characteristics, thereby improving the accuracy of identifying inferior content.
[0096] See also Figure 7 , Figure 7 The following is a flow chart of a content classification method provided by an embodiment of the present application. In this embodiment, the method may also include: Figure 7The content classification method can be executed by a computer device, which has at least storage, calculation and communication functions. Figure 7 The process shown combines Figure 8 The network training framework diagram shown in the figure is described in detail. Figure 7 As shown, the content classification method may specifically include the following steps:
[0097] Step S210: Acquire training samples, where the training samples include sample content and labels of the sample content, and the sample content includes sample images and sample texts.
[0098] In an embodiment of the present application, the classification method for the content to be classified belongs to a two-classification scheme, that is, the classification result finally outputted is used to indicate whether the content to be classified is low-quality content, and the classification result can be represented in the form of a label, for example, label "1" indicates that the content to be classified is low-quality content, and label "0" indicates that the content to be classified is non-low-quality content. Correspondingly, network training can be performed based on minimizing the difference between the predicted sample recognition result and the true label (Ground-truth). To this end, the training sample may include sample content and a label of the sample content, and the sample content includes a sample image and a sample text.
[0099] As an implementation, the computer device may obtain a training set from a database, wherein the training set includes a number of training samples that can meet the number of training samples used for network training. For example, the training samples may be normalized sample content obtained through data preprocessing, such as collected graphic content, video screenshots, and live streaming.
[0100] For example, the training set s = {(x 1 ,y 1 ),(x 2 ,y 2 ),…,(x n ,y n )}. The training set s includes n sampled training samples, each of which includes the corresponding sample content and the label of the sample content. For example, the i-th training sample includes the corresponding sample image x i and the label y of the sample content i , where n>0&n∈N * .
[0101] Step S220: Perform word segmentation processing on the sample text to obtain a sample word segmentation sequence.
[0102] As an implementation method, the computer device may obtain a sample text in the sample content, and perform tokenization on the sample text based on word granularity to obtain a sample token sequence. For example, the sample text "What a wonderful weather!" is obtained, and tokenization (Tokenization) is performed on the sample text based on word granularity to obtain a corresponding sample token sequence [CLS] What a wonderful weather! [SEP], which includes five tokens: "really", "one", "wonderful", "weather", "!".
[0103] Step S230: performing object recognition on the sample image to obtain a sample object recognition result.
[0104] The sample object recognition result includes sample object information of each sample object recognized from the sample image, and the sample object information at least includes the pixel area where the sample object is located.
[0105] Exemplarily, the computer device may perform object recognition on the sample image through an image recognition algorithm (e.g., YOLO) to obtain a sample object recognition result, such as Figure 8 The sample object recognition results shown include object information of five sample objects, each of which includes a pixel region where the sample object is located. Optionally, the pixel region can be formatted uniformly to have the same format size. For example, the sample object information corresponding to the first sample object includes a first pixel region.
[0106] Step S240: The image encoder performs encoding based on the sample object recognition result to obtain a first sample encoding matrix; the first sample encoding matrix includes image features of each sample object.
[0107] In an embodiment of the present application, the image encoder may include a first linear projection network, a second linear projection network, and a feature extraction network. Specifically, step S240 may also include steps S241 to S243:
[0108] Step S241: Mapping the pixel regions where each object in the sample object recognition result is located based on the first linear projection network to obtain a sample region feature matrix corresponding to the pixel region.
[0109] As an implementation method, the computer device can divide the pixel area where each object in the sample object recognition result is located into regions to obtain a preset number of sub-regions, and map each sub-region based on the first linear projection network to obtain a regional feature vector corresponding to each sub-region.
[0110] Furthermore, the computer device can obtain a region position matrix and a region mark, wherein the region position matrix is used to characterize the relative position information of each sub-region in the sample region, and the region mark is used to characterize the image information of the sample region, and then determine the sample region feature matrix based on the region feature vector, region position matrix and region mark corresponding to each sub-region.
[0111] For example, each pixel region I is converted into its corresponding representation sequence after segmentation operation Wherein, it indicates that the sequence includes a preset number r of sub-regions p j . Based on the first linear projection network, each sub-region is mapped, that is, the sub-region p j Mapped into a d I Vector of dimension z j =p j Emb, where Emb represents the embedding space.
[0112] For each pixel region I, a [class] tag can be added in advance, that is, a region tag, which is used to represent the overall image information of the pixel region, so as to avoid the network focusing too much on the local information of the sub-region and ignoring the overall image information of the pixel region. And the region position matrix Add to the sequence representing the sequence to retain the previous position information of each sub-region in the pixel area, so that the sample area feature matrix Z corresponding to each pixel area can be obtained i =[z [class] ; z 1 ; z 2 ;…;z r ]+F pos .
[0113] Step S242: encoding each sample region feature matrix based on the feature extraction network to obtain a corresponding sample region intermediate matrix.
[0114] For example, the feature extraction network may be a ViT encoder, which is not limited here. The computer device may input each sample region feature matrix into the ViT encoder for encoding processing to obtain the corresponding sample region intermediate matrix. The calculation process is as follows:
[0115] H i =ViT(Z i ),h i =H i,[class] (3)
[0116] Among them, h iis the intermediate matrix of the sample region corresponding to the i-th pixel region. As shown in (3), the representation vector labeled [class] is used as the feature of the pixel region. The image feature representation of the entire sample image, that is, the intermediate matrix of the sample region X I ={h 1 ,h 2 ,…,h n}, where n is the number of sample objects identified from the sample image.
[0117] Step S243: mapping the sample region intermediate matrix based on the second linear projection network to obtain a first sample encoding matrix corresponding to each sample region.
[0118] The vector dimension of the image feature of each sample object in the first sample encoding matrix is the same as the vector dimension of the word segmentation feature of each word in the second sample encoding matrix. For example, the computer device can input the sample region intermediate matrix into the second linear projection network to obtain the first sample encoding matrix. The calculation process is as follows:
[0119] Z={z 1 ,z 2 ,…,z n}=X I W V (4)
[0120] Where Z represents the first sample encoding matrix, z i is the image feature of the i-th sample object, is the parameter matrix of the second linear projection network.
[0121] Step S250: The text encoder encodes the sample word segmentation sequence to obtain a second sample encoding matrix, where the second sample encoding matrix includes the word segmentation features of each word segment in the sample text.
[0122] For example, the text encoder may be a bidirectional long short-term memory network (Bi-LSTM), and the computer device may use a pre-trained BERT-Base model to segment the word sequence S = {w 1 ,w 2 ,…,w m} to map each word segment in the word vector matrix Among them, X T T is the transposition symbol, [CLS] and [SEP] represent the sequence start symbol and word boundary symbol of the word segmentation sequence S respectively, [CLS] and [SEP] do not participate in text encoding, where m>0&m∈N * , m is the length of the word segmentation sequence, that is, the number of word segments in the text.
[0123] Furthermore, the text encoder can be a bidirectional long short-term memory network (Bi-LSTM), based on the Bi-LSTM word vector matrix X T Encode and obtain the second encoding matrix. The calculation process is as follows:
[0124] T={t 1 ,t 2 ,…,t m}=Bi﹣LSTM(X T ) (5)
[0125] Where T represents the second encoding matrix, t i , Represents the hidden state vector of the i-th time step of Bi-LSTM, that is, the word segmentation feature corresponding to the i-th word segmentation. h Represents the dimension of the hidden state vector.
[0126] Step S260: Determine the weight of each node pair for the sample content according to the image features of each sample object and the segmentation features of each segmentation in the sample text.
[0127] Step S270: construct a sample cross-modal graph for the sample content based on the weights of each node pair for the sample content, the segmentation features of each segmentation in the sample text, and the image features of the sample object.
[0128] Specifically, steps S260 to S270 may refer to the contents of step S140 and step S150 in the aforementioned embodiment, which will not be repeated here.
[0129] Step S280: The modal fusion network extracts and classifies cross-modal features based on the sample cross-modal graph to obtain a sample recognition result of the sample content.
[0130] In some embodiments, the modality fusion network may include a representation subnetwork and a classification subnetwork. The representation subnetwork may aggregate the correlations of nodes in the cross-modality graph to extract ironic information in the content to be classified, so that the classification based on the classification subnetwork is more accurate. Specifically, step S280 may include steps S281 and S282:
[0131] Step S281: Through the representation sub-network, based on the sample cross-modal graph, obtain the sample graph feature representation vector.
[0132] As an implementation method, the computer device can input the sample adjacency matrix and the sample merged feature vector into the representation subnetwork for representation learning to obtain the sample graph feature vector, and based on the sample graph feature vector and the sample merged feature vector, calculate the attention weight vector corresponding to the sample graph feature vector, and further determine the sample feature representation vector based on the sample graph feature vector and the attention weight vector. Among them, the attention weight vector is used to retrieve the explicit connection existing in the cross-modal graph, and makes the final sample feature representation vector more accurate in representing the sarcastic information. The modal fusion network can also effectively extract cross-modal features based on the adjacency matrix of the cross-modal graph and the hidden representation of its neighborhood, thereby reducing the consumption of computing resources.
[0133] Step S282: Input the sample feature representation vector into the classification subnetwork for content classification to obtain a sample recognition result corresponding to the sample content.
[0134] As an implementation, the classification subnetwork may be a fully connected layer (Fully Connected Layer, FCL). The computer device may input the sample feature representation vector into the fully connected layer to capture the probability distribution in the irony decision space, that is, the sample recognition result.
[0135] Exemplarily, the representation subnetwork in the modal fusion network can be a graph convolutional neural network (GCN), TransE, TransH, TransR, TransD, etc., which are not limited here. The classification subnetwork can be a fully connected layer with a Softmax activation function. For example, the neighbor matrix A of the cross-modal graph and the sample merge feature vector Q = {q 1 ,q,…,q n+m}={t 1 ,…,t n ,v 1 ,…,v m} is input into GCN, and GCN outputs the sample graph feature vector Q′={q′ 1 ,q′ 2 ,…,q′ n+m}, the calculation process is as follows:
[0136]
[0137] in, is the normalized symmetric adjacency matrix, D is the degree matrix of A, and D ii =∑ j A i,j . G l―1 W represents the hidden layer representation vector of the previous layer output in GCN. l , b lRepresents the weight parameter of the current layer in GCN, The input of the GCN input layer is the concatenated feature vector of each word feature and each image feature. The calculation process of the attention weight vector is as follows:
[0138]
[0139] Among them, C represents the node set corresponding to the cross-modal edges in the cross-modal graph, that is, the node set composed of the nodes in the second node pair. Represents the word segmentation feature set {t 1 ,…,t n Finally, the calculation process of the sample feature representation vector is as follows:
[0140]
[0141] Furthermore, the sample feature representation vector f is input into the fully connected layer with the Softmax activation function to obtain the sample recognition result. The calculation process is as follows:
[0142]
[0143] Among them, y represents the sample recognition result, d p W represents the dimension of the sample recognition result, which is the same as the dimension of the label corresponding to the sample content. o , b o represents the weight parameter of the fully connected layer, and
[0144] Step S290: Based on the target loss determined by the sample recognition result and the label of the sample content, iteratively update the weight parameters of the image encoder, text encoder and modality fusion network until the training end condition is reached.
[0145] As an implementation, the computer device may use a standard gradient descent algorithm to minimize the cross entropy loss between the sample recognition result and the label of the sample content to perform network training. The calculation process of the cross entropy loss is as follows:
[0146]
[0147] Wherein, N represents the number of sample contents in the training set. Θ represents all trainable parameters, including the weight parameters of the representation subnetwork and the classification subnetwork in the first linear projection network, the second linear projection network and the feature extraction network in the image encoder, the text encoder, and the modality fusion network. After reaching the training end condition, the classified content can be classified based on the trained image encoder, text encoder, and modality fusion network. The specific process can refer to the content of the aforementioned embodiment, which will not be described in detail here.
[0148] The training end conditions may include: the target loss is less than a preset value, the target loss is within a preset threshold range, or the number of training times reaches a preset number of times, etc. Optionally, an optimizer may be used to optimize the target loss, and the learning rate, batch size, and epoch number may be set based on experimental experience.
[0149] In some embodiments, the existing content classification methods (Solution 1, Solution 2 and Solution 3) and the content classification method of the present application can be tested for actual effects. Specifically, the accuracy, precision, recall and F1 value of various methods can be tested on the text content + cover image data set. The test results are shown in the following table:
[0150]
[0151] As shown in the table above, compared with Solution 1, Solution 2 and Solution 3, the test results of this application are significantly improved. For example, the Acc is increased by 20%, 11% and 7% respectively, proving the effectiveness of the method of this application. The reason why the method of this application can achieve such a big improvement can be mainly due to two aspects:
[0152] (1) The prior art only uses the data and features of the single modality of text to identify sarcastic content in text, without considering that some low-quality content of the text-image interaction type is based on the content of the illustration and utilizes the interaction between visual information and text information to achieve a sarcastic effect. This application can mine the interactive information between multiple modalities by automatically constructing a cross-modal graph of text and images, and can perform better in the task of identifying low-quality multimodal content.
[0153] (2) Due to the rapid update and iteration of Internet information, new words and hot memes continue to appear, and the existing technology is very dependent on thesaurus or training corpus, which makes it difficult for the existing technology to be updated synchronously with the network environment. This application introduces new vocabulary knowledge in the stage of constructing a cross-modal graph, so that new words and hot memes can be checked and supplemented in a targeted manner without introducing a large amount of training corpus, avoiding the risk of missing harmful content.
[0154] See also Fig. 9 , which shows a structural block diagram of a content classification device 300 provided in an embodiment of the present application. The content classification device 300 may include: an object recognition module 310, used to obtain an object recognition result of an image in the content to be classified, the object recognition result includes object information of each object recognized from the image, and the object information at least includes the pixel area where the object is located; a first encoding module 320, used to encode based on the object information in the object recognition result to obtain a first encoding matrix; the first encoding matrix includes the image features of each object; a second encoding module 330, used to encode the text in the content to be classified to obtain a second encoding matrix, the second encoding matrix includes the The segmentation features of each segmentation in the text; a weight determination module 340, used to determine the edge weight of each node pair according to the image features of each object and the segmentation features of each segmentation; the node pair is formed by two nodes representing different segmentation features, or by a node representing a segmentation feature and a node representing an image feature; a graph construction module 350, used to construct a cross-modal graph for the content to be classified based on the edge weights of the node pairs, the segmentation features and the image features; a content classification module 360, used to extract and classify cross-modal features based on the cross-modal graph to obtain a classification result for the content to be classified.
[0155] In some embodiments, the object information also includes an object attribute description; the object attribute description includes an object name and attribute descriptors; the image features include semantic features of the object name and semantic features of the attribute descriptors; the weight determination module 340 may include: a first determination unit, for determining, for a first node pair formed by two nodes representing different segmentation features, an edge weight of the first node pair according to a dependency relationship between two segmentations involved in the first node pair; a second determination unit, for determining, for a second node pair formed by a node representing segmentation features and a node representing image features, an edge weight of the second node pair according to a feature similarity between a target segmentation involved in the second node pair and an object name of a target object involved, the sentiment weight of the target segmentation, and the sentiment weight of the attribute descriptor of the target object.
[0156] In some embodiments, the second determination unit may further include: a similarity subunit, used to calculate the feature similarity between the target participle and the object name of the target object based on the participle features of the target participle and the semantic features of the object name of the target object; a participle weight subunit, used to obtain the sentiment weight of the target participle based on the participle features of the target participle; a sentiment weight subunit, used to obtain the sentiment weight of the attribute descriptor of the target object based on the semantic features of the attribute descriptor of the target object; an adjustment factor subunit, used to determine the adjustment factor based on the sentiment weight of the target participle and the sentiment weight of the attribute descriptor of the target object, the adjustment factor being used to characterize the inconsistency of the sentiment tendency between the target participle and the target object; an edge weight subunit, used to determine the edge weight of the second node pair based on the feature similarity and the adjustment factor.
[0157] In some embodiments, the adjustment factor subunit can be specifically used to: multiply the sentiment weight of the target participle by the sentiment weight of the attribute description word of the target object to obtain an intermediate adjustment factor; determine a first adjustment coefficient based on the intermediate adjustment factor, and the first adjustment coefficient is negatively correlated with the intermediate adjustment factor; subtract the sentiment weight of the target participle from the sentiment weight of the attribute description word of the target object, and use the absolute value of the subtraction result as the second adjustment coefficient; multiply the first adjustment coefficient by the second adjustment coefficient to obtain the adjustment factor.
[0158] In some embodiments, encoding is performed based on the object information in the object recognition result by an image encoder; the text in the content to be classified is encoded by a text encoder; cross-modal feature extraction and classification are performed based on the cross-modal graph by a modal fusion network; the content classification device 300 may also include: a sample acquisition module, used to obtain training samples, the training samples include sample content and labels of the sample content, and the sample content includes sample images and sample text; a word segmentation processing module, used to perform word segmentation processing on the sample text to obtain a sample word segmentation sequence; an object recognition module, used to perform object recognition on the sample image to obtain a sample object recognition result; the sample object recognition result includes sample object information of each sample object recognized from the sample image, and the sample object information at least includes the pixel area where the sample object is located; an image encoding module, used to encode by the image encoder based on the sample object recognition result to obtain a first sample encoding matrix; the first sample encoding matrix includes The image features of each sample object; a word segmentation encoding module, which is used to encode the sample word segmentation sequence by the text encoder to obtain a second sample encoding matrix, wherein the second sample encoding matrix includes the word segmentation features of each word in the sample text; a weight calculation module, which is used to determine the edge weight of each node pair for the sample content according to the image features of each sample object and the word segmentation features of each word in the sample text; a graph generation module, which is used to construct a sample cross-modal graph for the sample content based on the edge weight of each node pair for the sample content, the word segmentation features of each word in the sample text and the image features of the sample object; a result prediction module, which is used to extract and classify cross-modal features based on the sample cross-modal graph by the modal fusion network to obtain a sample recognition result of the sample content; a network training module, which is used to iteratively update the weight parameters of the image encoder, the text encoder and the modal fusion network based on the sample recognition result and the target loss determined by the label of the sample content until the training end condition is reached.
[0159] In some embodiments, the image encoder includes a first linear projection network, a second linear projection network and a feature extraction network; the image encoding module may include: a first mapping unit, used to map the pixel area where each object in the sample object recognition result is located based on the first linear projection network to obtain a sample area feature matrix corresponding to the pixel area; a region encoding unit, used to encode each sample area feature matrix based on the feature extraction network to obtain a corresponding sample area intermediate matrix; a second mapping unit, used to map the sample area intermediate matrix based on the second linear projection network to obtain a first sample encoding matrix corresponding to each sample area; the vector dimension of the image features of each sample object in the first sample encoding matrix is the same as the vector dimension of the segmentation features of each segmentation in the second sample encoding matrix.
[0160] In some embodiments, the first mapping unit can be specifically used to: perform regional segmentation on the pixel area where each object in the sample object recognition result is located to obtain a preset number of sub-regions; perform mapping processing on each sub-region based on the first linear projection network to obtain a regional feature vector corresponding to each sub-region; obtain a regional position matrix and a regional mark; the regional position matrix is used to characterize the relative position information of each sub-region in the sample region in the sample region; the regional mark is used to characterize the image information of the sample region; based on the regional feature vector corresponding to each sub-region, the regional position matrix and the regional mark, determine the sample region feature matrix.
[0161] In some embodiments, the result prediction module may include: a feature representation unit, used to obtain a sample graph feature representation vector based on the sample cross-modal graph through the representation sub-network; a content classification unit, used to input the sample feature representation vector into the classification sub-network for content classification to obtain a sample recognition result corresponding to the sample content.
[0162] In some embodiments, the feature representation unit can be specifically used to: input the sample adjacency matrix and the sample merged feature vector into the representation subnetwork for representation learning to obtain a sample graph feature vector; based on the sample graph feature vector and the sample merged feature vector, calculate the attention weight vector corresponding to the sample graph feature vector; based on the sample graph feature vector and the attention weight vector, determine the sample feature representation vector.
[0163] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0164] In several embodiments provided in the present application, the coupling between modules may be electrical, mechanical or other forms of coupling.
[0165] In addition, each functional module in each embodiment of the present application can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or software functional modules.
[0166] The solution provided by the present application can obtain object recognition results of images in content to be classified, and the object recognition results include object information of each object recognized from the image, and the object information includes at least the pixel area where the object is located. Then, encoding is performed based on the object information in the object recognition results to obtain a first encoding matrix, and the first encoding matrix includes image features of each object, and the text in the content to be classified is encoded to obtain a second encoding matrix, and the second encoding matrix includes word segmentation features of each word in the text. Further, according to the image features of each object and the word segmentation features of each word, the edge weight of each node pair is determined, and then, based on the edge weights, word segmentation features and image features of the node pairs, a cross-modal graph for the content to be classified is constructed, and cross-modal features are extracted and classified based on the cross-modal graph to obtain a classification result of the content to be classified. The cross-modal graph for the content to be classified, which is constructed based on the edge weights, word segmentation features and image features of node pairs, can cross-modally connect the image features about the image and the word segmentation features about the text in the content to be classified in the form of a graph, so that cross-modal feature extraction can be performed on the content to be classified based on the cross-modal graph, and cross-modal interaction information between the image modality and the text modality can be obtained, thereby effectively identifying the content to be classified with sarcastic attributes with interactive characteristics, and improving the accuracy of identifying inferior content.
[0167] like Fig.10 As shown, the embodiment of the present application also provides a computer device 400, which includes a processor 410, a memory 420, a power supply 430 and an input unit 440. The memory 420 stores a computer program. When the computer program is called by the processor 410, the various method steps provided in the above embodiment can be implemented. Those skilled in the art can understand that the structure of the computer device shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0168] The processor 410 may include one or more processing cores. The processor 410 uses various interfaces and lines to connect various parts of the entire battery management system, and by running or executing instructions, programs, instruction sets or program sets stored in the memory 420, calls the data stored in the memory 420, executes various functions of the battery management system and processes data, and executes various functions and processes data of the computer device, thereby controlling the computer device as a whole. Optionally, the processor 410 can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 410 can integrate one or more combinations of a central processing unit 410 (CPU), an image processor 410 (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 410, but may be implemented separately through a communication chip.
[0169] The memory 420 may include a random access memory 420 (Random Access Memory, RAM), and may also include a read-only memory 420 (Read-Only Memory). The memory 420 may be used to store instructions, programs, instruction sets or program sets. The memory 420 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data (such as a phone book and audio and video data) created by the computer device during use. Accordingly, the memory 420 may also include a memory controller to provide the processor 410 with access to the memory 420.
[0170] The power supply 430 can be logically connected to the processor 410 through a power management system, so as to manage charging, discharging, power consumption and other functions through the power management system. The power supply 430 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.
[0171] The input unit 440 may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0172] Although not shown, the computer device 400 may also include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 410 in the computer device will load the executable files corresponding to the processes of one or more computer programs into the memory 420 according to the following instructions, and the processor 410 will run the data stored in the memory 420, such as the phone book and audio and video data, so as to implement the various method steps provided in the aforementioned embodiments.
[0173] like Fig.11 As shown, the embodiment of the present application further provides a computer-readable storage medium 500, in which a computer program 510 is stored. The computer program 510 can be called by a processor to execute various method steps provided in the embodiment of the present application.
[0174] The computer readable storage medium may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Optionally, the computer readable storage medium includes a non-volatile computer readable storage medium (Non-Transitory Computer-Readable Storage Medium). The computer readable storage medium 500 has storage space for a computer program that executes any of the method steps in the above embodiments. These computer programs can be read from or written to one or more computer program products. The computer program can be compressed in an appropriate form.
[0175] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program, the computer program being stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes various method steps provided in the above embodiments.
[0176] The above are only preferred embodiments of the present application, and are not intended to limit the present application in any form. Although the present application has been disclosed as above with preferred embodiments, it is not intended to limit the present application. Any technical personnel in the field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present application. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. A content classification method, It is characterized in that The method comprises: Obtaining an object recognition result of an image in the content to be classified, wherein the object recognition result includes object information of each object recognized from the image, and the object information includes at least a pixel area where the object is located; Encoding is performed based on the object information in the object recognition result to obtain a first encoding matrix; the first encoding matrix includes image features of each object; Encoding the text in the content to be classified to obtain a second encoding matrix, wherein the second encoding matrix includes the segmentation features of each segmentation in the text; Determining the edge weight of each node pair according to the image feature of each object and the segmentation feature of each segmentation; the node pair is formed by two nodes representing different segmentation features, or by a node representing a segmentation feature and a node representing an image feature; Based on the edge weights of the node pairs, the word segmentation features, and the image features, construct a cross-modal graph for the content to be classified; Cross-modal feature extraction and classification are performed based on the cross-modal graph to obtain a classification result of the content to be classified.
2. The method according to claim 1, It is characterized in that The object information also includes object attribute description; the object attribute description includes object name and attribute description words; the image features include semantic features of the object name and semantic features of the attribute description words; The step of determining the edge weight of each node pair according to the image features of each object and the segmentation features of each segmentation word includes: For a first node pair formed by two nodes representing different segmentation features, determining an edge weight of the first node pair according to a dependency relationship between two segmentations involved in the first node pair; For a second node pair formed by a node representing a segmentation feature and a node representing an image feature, the edge weight of the second node pair is determined based on the feature similarity between the target segmentation involved in the second node pair and the object name of the target object involved, the sentiment weight of the target segmentation, and the sentiment weight of the attribute description word of the target object.
3. The method according to claim 2, It is characterized in that The step of determining the edge weight of the second node pair according to the feature similarity between the target word involved in the second node pair and the object name of the target object involved, the sentiment weight of the target word, and the sentiment weight of the attribute description word of the target object includes: Calculating the feature similarity between the target segmentation and the object name of the target object according to the segmentation feature of the target segmentation and the semantic feature of the object name of the target object; Based on the segmentation features of the target segmentation, obtaining the sentiment weight of the target segmentation; Based on the semantic features of the attribute description words of the target object, obtaining the sentiment weight of the attribute description words of the target object; Determining an adjustment factor according to the sentiment weight of the target participle and the sentiment weight of the attribute description word of the target object, wherein the adjustment factor is used to characterize the inconsistency of the sentiment tendency between the target participle and the target object; An edge weight of the second node pair is determined according to the feature similarity and the adjustment factor.
4. The method according to claim 3, It is characterized in that The step of determining the adjustment factor according to the sentiment weight of the target word segment and the sentiment weight of the attribute description word of the target object includes: Multiplying the sentiment weight of the target word segment by the sentiment weight of the attribute description word of the target object to obtain an intermediate adjustment factor; Determine a first adjustment coefficient according to the intermediate adjustment factor, wherein the first adjustment coefficient is negatively correlated with the intermediate adjustment factor; Subtracting the sentiment weight of the target word from the sentiment weight of the attribute description word of the target object, and taking the absolute value of the subtraction result as the second adjustment coefficient; The first adjustment coefficient is multiplied by the second adjustment coefficient to obtain the adjustment factor.
5. The method according to claim 1, It is characterized in that Encoding the object information in the object recognition result by using an image encoder; encoding the text in the content to be classified by using a text encoder; Extracting and classifying cross-modal features based on the cross-modal graph through a modality fusion network; the method further includes: Acquire a training sample, wherein the training sample includes sample content and a label of the sample content, and the sample content includes a sample image and a sample text; Performing word segmentation processing on the sample text to obtain a sample word segmentation sequence; Performing object recognition on the sample image to obtain a sample object recognition result; the sample object recognition result includes sample object information of each sample object recognized from the sample image, and the sample object information at least includes a pixel area where the sample object is located; The image encoder performs encoding based on the sample object recognition result to obtain a first sample encoding matrix; the first sample encoding matrix includes image features of each sample object; The text encoder encodes the sample word segmentation sequence to obtain a second sample encoding matrix, wherein the second sample encoding matrix includes word segmentation features of each word in the sample text; Determining edge weights of each node pair for the sample content according to image features of each sample object and segmentation features of each segmentation in the sample text; Constructing a sample cross-modal graph for the sample content based on the edge weights of each node pair for the sample content, the word segmentation features of each word in the sample text, and the image features of the sample object; The modal fusion network extracts and classifies cross-modal features based on the sample cross-modal graph to obtain a sample recognition result of the sample content; Based on the target loss determined by the sample recognition result and the label of the sample content, the weight parameters of the image encoder, text encoder and modality fusion network are iteratively updated until the training end condition is reached.
6. The method according to claim 5, It is characterized in that The image encoder includes a first linear projection network, a second linear projection network and a feature extraction network; The encoding by the image encoder based on the sample object recognition result to obtain a first sample encoding matrix includes: Based on the first linear projection network, mapping processing is performed on the pixel area where each object in the sample object recognition result is located to obtain a sample area feature matrix corresponding to the pixel area; Based on the feature extraction network, encoding processing is performed on each sample region feature matrix to obtain a corresponding sample region intermediate matrix; The sample area intermediate matrix is mapped based on the second linear projection network to obtain a first sample encoding matrix corresponding to each sample area; the vector dimension of the image features of each sample object in the first sample encoding matrix is the same as the vector dimension of the segmentation features of each word in the second sample encoding matrix.
7. The method according to claim 6, It is characterized in that The mapping process is performed on the pixel regions where each object in the sample object recognition result is located based on the first linear projection network to obtain a sample region feature matrix corresponding to the pixel region, including: Performing region segmentation on the pixel region where each object in the sample object recognition result is located to obtain a preset number of sub-regions; Performing mapping processing on each sub-region based on the first linear projection network to obtain a regional feature vector corresponding to each sub-region; Acquire a region position matrix and a region mark; the region position matrix is used to characterize the relative position information of each sub-region in the sample region; the region mark is used to characterize the image information of the sample region; A sample region feature matrix is determined based on the region feature vector corresponding to each sub-region, the region position matrix and the region mark.
8. The method according to claim 5, It is characterized in that The modality fusion network includes a representation subnetwork and a classification subnetwork. The modality fusion network extracts and classifies cross-modality features based on the sample cross-modality graph to obtain a sample recognition result of the sample content, including: Obtaining a sample graph feature representation vector based on the sample cross-modal graph through the representation subnetwork; The sample feature representation vector is input into the classification subnetwork to perform content classification, and a sample recognition result corresponding to the sample content is obtained.
9. The method according to claim 8, It is characterized in that The passing of the representation sub-network to obtain a sample graph feature representation vector based on the sample cross-modal graph includes: Inputting the sample adjacency matrix and the sample merged feature vector into the representation subnetwork for representation learning to obtain a sample graph feature vector; Based on the sample atlas feature vector and the sample merged feature vector, calculating an attention weight vector corresponding to the sample atlas feature vector; Based on the sample graph feature vector and the attention weight vector, a sample feature representation vector is determined.
10. A content classification device, It is characterized in that The device comprises: An object recognition module, used to obtain an object recognition result of an image in the content to be classified, wherein the object recognition result includes object information of each object recognized from the image, and the object information includes at least a pixel area where the object is located; A first encoding module, configured to perform encoding based on the object information in the object recognition result to obtain a first encoding matrix; the first encoding matrix includes image features of each object; A second encoding module, used for encoding the text in the content to be classified to obtain a second encoding matrix, wherein the second encoding matrix includes the segmentation features of each segmentation in the text; A weight determination module, used to determine the edge weight of each node pair according to the image features of each object and the segmentation features of each segmentation; the node pair is formed by two nodes representing different segmentation features, or by a node representing a segmentation feature and a node representing an image feature; A graph construction module, used to construct a cross-modal graph for the content to be classified based on the edge weights of node pairs, the word segmentation features and the image features; The content classification module is used to extract and classify cross-modal features based on the cross-modal graph to obtain a classification result of the content to be classified.
11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
12. A computer device, It is characterized in that include: Memory; A processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product or a computer program, It is characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.