Methods, devices, equipment and media for evaluating article content

By fusing the image and text information of the article content, a multi-dimensional evaluation result is generated, which solves the problem of unreasonable evaluation in the existing technology and improves the accuracy and robustness of the evaluation.

CN114818691BActive Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110125965.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-29
Publication Date
2025-09-16
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

Existing technologies can only evaluate article content through text information and are unable to make reasonable evaluations, especially when images and text are combined.

Method used

Extract image and text information from the article content, integrate multiple evaluation results, including image and text evaluation, text evaluation, objective prior evaluation and layout evaluation, and perform weighted averaging through the image and text prior high-quality evaluation model and attention mechanism to generate multi-dimensional evaluation results.

Benefits of technology

The generated evaluation results are more in line with the user's actual reading experience, reduce erroneous evaluations, and improve the robustness of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114818691B_ABST
    Figure CN114818691B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and medium for evaluating article content, relating to the field of artificial intelligence. The method comprises: extracting image and text information from the article content; obtaining multiple evaluation results for the article content based on the image and text information, wherein the multiple evaluation results include at least two of the following: image and text evaluation results, text evaluation results, objective a priori evaluation results, and typesetting evaluation results; and fusing the multiple evaluation results to obtain a multi-dimensional evaluation result for the article content. This application enables a relatively comprehensive evaluation of article content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a method, apparatus, device, and medium for evaluating article content. Background Art

[0002] Article content evaluation involves computer equipment assessing the quality of the text and graphics within an article, determining the proportion of high-quality and low-quality content, and judging the quality of the article based on this proportion. This allows merchants to recommend appropriate articles to users, ensuring a better reading experience.

[0003] The relevant technology is that computer equipment evaluates the content of an article from the perspective of text information through supervised learning or unsupervised learning. For example, the article content is evaluated from statistical dimensions such as the number of words in the article and the diversity of vocabulary, and the proportion of high-quality text information and low-quality text information in the article content is obtained. The evaluation results of the relevant text information are obtained through the proportion, and the obtained evaluation results are used as the evaluation results of the entire article content.

[0004] However, related technologies can only evaluate the content of an article through text information, and in some cases cannot make a reasonable evaluation of the content of the article. Summary of the Invention

[0005] This application provides a method, device, equipment, and storage medium for evaluating article content, which can evaluate article content from multiple dimensions and obtain reasonable article evaluation results. The technical solution is as follows:

[0006] According to one aspect of the present application, a method for evaluating article content is provided, the method comprising:

[0007] Extracting image information and text information from the article content;

[0008] Obtaining multiple evaluation results of the article content based on the image information and the text information, the multiple evaluation results including at least two of a graphic evaluation result, a text evaluation result, an objective priori evaluation result, and a typesetting evaluation result;

[0009] The multiple evaluation results are integrated to obtain a multi-dimensional evaluation result of the article content.

[0010] According to another aspect of the present application, a device for evaluating article content is provided, the device comprising:

[0011] An extraction module, used to extract image information and text information from the article content;

[0012] an evaluation module, configured to obtain a plurality of evaluation results of the article content based on the image information and the text information, the plurality of evaluation results comprising at least two of a graphic evaluation result, a text evaluation result, an objective priori evaluation result, and a typesetting evaluation result;

[0013] The evaluation fusion module is used to fuse the multiple evaluation results to obtain a multi-dimensional evaluation result of the article content.

[0014] In an optional design of the present application, the extraction module is also used to extract multimodal feature information from the image information and the text information.

[0015] The evaluation module is also used to input the multimodal feature information into the image-text prior high-quality evaluation model to obtain the multiple evaluation results of the article content. The image-text prior high-quality evaluation model is used to obtain the multiple multi-dimensional evaluation results based on the multimodal feature information.

[0016] In an optional design of the present application, the extraction module is further used to extract image feature vectors and text feature vectors from the multimodal feature information, where the text feature vector is a feature vector corresponding to all or part of the text in the article content.

[0017] The evaluation module is further configured to input the image feature vector and the text feature vector into the image-text multimodal sub-network to obtain the image-text evaluation result of the article content.

[0018] In an optional design of the present application, the evaluation module is also used to generate a text feature representation of fused image information and an image feature representation of fused text information through the image-text multimodal sub-network; and to fuse the text feature representation and the image feature representation through the image-text multimodal sub-network to generate the image-text evaluation result of the article content.

[0019] In an optional design of the present application, the extraction module is also used to extract objective prior features from the multimodal feature information, and the objective prior features include at least one of statistical features, linguistic features, image quality features, and account features.

[0020] The evaluation module is further configured to input the objective priori features into the objective priori feature sub-network to obtain an objective priori evaluation result of the article content.

[0021] In an optional design of the present application, the extraction module is also used to extract article word vectors from the multimodal feature information.

[0022] The evaluation module is further used to input the article word vector into the text sub-network to obtain the text evaluation result of the article content.

[0023] In an optional design of the present application, the extraction module is further used to extract image feature vectors and text feature vectors from the multimodal feature information, where the text feature vector is a feature vector corresponding to all or part of the text in the article content.

[0024] The evaluation module is further configured to input the image feature vector and the text feature vector into the typesetting subnetwork to obtain the typesetting evaluation result of the article content.

[0025] In an optional design of the present application, the evaluation fusion module is also used to assign corresponding weight values ​​to the multiple evaluation results through the image-text prior high-quality evaluation model and the attention mechanism; and to perform weighted averaging on the multiple evaluation results through the image-text prior high-quality evaluation model and the weight values ​​to obtain the multi-dimensional evaluation results of the article content.

[0026] In an optional design of the present application, the device further includes: a training module.

[0027] The training module is used to obtain a picture and text training set, which includes sample articles and real picture and text evaluation results corresponding to the sample articles; extract sample image information and sample text information from the sample articles; extract sample image feature vectors and sample text feature vectors based on the sample image information and the sample text information; input the sample image feature vectors and the sample text feature vectors into the picture and text multimodal subnetwork to obtain predicted picture and text evaluation results; and train the picture and text multimodal subnetwork based on the error loss between the predicted picture and text evaluation results and the real picture and text evaluation results.

[0028] In an optional design of the present application, the training module is also used to obtain an objective prior training set, which includes sample articles and real objective prior evaluation results corresponding to the sample articles; extract sample image information and sample text information in the sample articles; obtain sample objective prior features of the sample articles based on the sample image information and the sample text information; input the sample objective prior features into the objective prior feature sub-network to obtain predicted objective prior evaluation results; and train the objective prior feature sub-network based on the error loss between the predicted objective prior evaluation results and the real objective prior evaluation results.

[0029] In an optional design of the present application, the training module is also used to obtain a text training set, which includes sample articles and real text evaluation results corresponding to the sample articles; extract sample text information from the sample articles; extract sample article word vectors based on the sample text information; input the sample article word vectors into the text sub-network to obtain predicted text evaluation results; and train the text sub-network based on the error loss between the predicted text evaluation results and the real text evaluation results.

[0030] In an optional design of the present application, the training module is also used to obtain a typesetting training set, which includes sample articles and real typesetting evaluation results corresponding to the sample articles; extract sample image information and sample text information in the sample articles; extract sample image feature vectors and sample text feature vectors based on the sample image information and the sample text information; input the sample image vector and the sample text feature vector into the typesetting sub-network to obtain a predicted typesetting evaluation result; and train the typesetting sub-network based on the error loss between the predicted typesetting evaluation result and the real typesetting evaluation result.

[0031] According to another aspect of the present application, a computer device is provided, comprising: a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the method for evaluating the content of an article as described above.

[0032] According to another aspect of the present application, a computer storage medium is provided, in which at least one program code is stored. The program code is loaded and executed by a processor to implement the article content evaluation method as described above.

[0033] According to another aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the article content evaluation method described above.

[0034] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:

[0035] By integrating multiple evaluation results from multiple dimensions of article content, we can evaluate the article content from multiple perspectives, making the final evaluation result more closely aligned with the user's actual reading experience. Furthermore, since the final evaluation result integrates evaluations from multiple dimensions, it can effectively reduce erroneous evaluation results and improve the robustness of the entire solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0037] Figure 1 It is a structural diagram of the VistaNet model provided by an exemplary embodiment of the present application;

[0038] Figure 2 is a schematic diagram of the structure of a computer system provided by an exemplary embodiment of the present application;

[0039] Figure 3 is a flowchart of a method for evaluating article content provided by an exemplary embodiment of the present application;

[0040] Figure 4 is a flowchart of a method for evaluating article content provided by an exemplary embodiment of the present application;

[0041] Figure 5 This is an exemplary structural diagram of a picture-text prior quality evaluation model provided by an exemplary embodiment of the present application;

[0042] Figure 6 This is an exemplary structural diagram of a graphic-text multimodal sub-network provided by an exemplary embodiment of the present application;

[0043] Figure 7 is an exemplary structural diagram of an objective prior feature subnetwork provided by an exemplary embodiment of the present application;

[0044] Figure 8 is an exemplary structural diagram of a text sub-network provided by an exemplary embodiment of the present application;

[0045] Figure 9 is an exemplary structural diagram of a typesetting sub-network provided by an exemplary embodiment of the present application;

[0046] Figure 10 This is an exemplary complete structural diagram of a picture-text prior quality evaluation model provided by an exemplary embodiment of the present application;

[0047] Figure 11 This is a flowchart of a method for training a multimodal sub-network of images and texts provided by an exemplary embodiment of the present application;

[0048] Figure 12 1 is a flow chart of an objective prior feature sub-network training method provided by an exemplary embodiment of the present application;

[0049] Figure 13 1 is a flowchart of a text sub-network training method provided by an exemplary embodiment of the present application;

[0050] Figure 14 1 is a flow chart of a typesetting sub-network training method provided by an exemplary embodiment of the present application;

[0051] Figure 15 is an exemplary business architecture diagram provided by an exemplary embodiment of the present application;

[0052] Figure 16 This is a schematic diagram of the structure of an article content evaluation device provided by an exemplary embodiment of the present application;

[0053] Figure 17 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0055] First, the nouns involved in the embodiments of this application are introduced:

[0056] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0057] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0058] Computer Vision (CV): Computer vision is the science of making machines "see." Specifically, it refers to machine vision, where cameras and computers replace the human eye in identifying, tracking, and measuring targets. Further image processing is performed to transform computer-generated images into images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0059] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs.

[0060] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0061] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0062] Image and text prior quality: Image and text prior quality constructs a rational evaluation system for article quality based on the article content itself, helping recommenders better understand and apply the image and text content released by the content center. To comprehensively evaluate article quality, we modeled the multimodality of image and text, account number, article layout experience, and article linguistic atomic features (such as the lexical diversity used in the article, whether the article uses diverse syntax such as metaphors and parallelism, and whether the article quotes ancient poetry). Ultimately, we constructed an integrated method for identifying article prior quality.

[0063] Multimodal image and text: refers to Multimodal Machine Learning (MMML), specifically the ability to process and understand multi-source modal information through machine learning methods. The main research direction is multimodal learning between semantics, images, and videos. Multimodal learning can be divided into the following five research directions: multimodal representation learning, modal transformation, alignment, multimodal fusion, and collaborative learning. Unimodal representation learning is responsible for representing information as numerical vectors that can be processed by computers or further abstracted into higher-level feature vectors. Multimodal representation learning refers to learning better feature representations by leveraging the complementarity between multiple modalities and eliminating redundancy between modalities.

[0064] Linguistics: is the scientific study of human language, involving the analysis of language form, language meaning and context. Linguistics has an important thematic division between the study of language structure (grammar) and the study of meaning (semantics and pragmatics). Grammar includes morphology (the formation and composition of words), syntax (the rules that determine how words are combined into phrases or sentences), and phonetics (the study of sound systems and abstract sound units). In order to comprehensively evaluate the quality of articles, this application implements a comprehensive assessment of high-quality articles from various feature dimensions such as the overall semantic information of the article, the semantic relationship between sentences in the article, the lexical diversity of the article, the diversity of rhetoric used in the article (such as metaphors, parallelism, etc.), and the citation of ancient poems in the article.

[0065] Ensemble model: Ensemble learning is a machine learning paradigm. In ensemble learning, we train multiple models to solve the same problem and combine them to achieve better results. A single learner is prone to either underfitting or overfitting. To obtain a learner with good generalization performance, we can train multiple individual learners and, through a certain combination strategy, ultimately form a strong learner. This method of combining multiple individual learners is called ensemble learning. Ensemble learning builds multiple models on the data and integrates the modeling results of all models. The ensemble algorithm considers the modeling results of multiple evaluators and aggregates them to obtain a comprehensive result, thereby achieving better regression or classification performance than a single model.

[0066] BERT (Bidirectional Encoder Representations from Transformers) model: BERT is a new language representation model. BERT is designed to pre-train deep bidirectional representations based on both left and right context at all layers. As a result, pre-trained BERT representations can be fine-tuned with just one additional output layer to create state-of-the-art models for many tasks, such as question answering and language reasoning, without requiring extensive modifications to task-specific architectures. BERT is conceptually simple but its experimental results are powerful. It has refreshed the current best results on 11 Natural Language Processing (NLP) tasks, including raising the GLUE (General Language Understanding Evaluation) benchmark to 80.4% (7.6% absolute improvement), improving the accuracy of MultiNLI (a public natural language dataset) to 86.7% (5.6% absolute improvement), and improving the score of the SQuADv1.1 (a dataset) question-answering test to 93.2 points (1.5 points absolute improvement) - 2.0 points higher than human performance.

[0067] HAN (Hierarchical Attention Network) model: The HAN (Hierarchical Attention Network) model has good classification accuracy in long text classification tasks. The overall structure of the model is as follows: The input word vector sequence is w 2x After passing through the word-level Bi-GRU (Bi-Gated Recurrent Unit, gated recurrent unit model), each word will have a corresponding Bi-GRU output latent vector h, and then through u wThe dot product of the vector h with each h vector yields the attention weight. The h sequence is then weighted summed according to the attention weights to produce the sentence summary vector s2. Each sentence is then processed through the same Bi-GRU structure with attention added to produce the final document feature vector v. This v vector is then passed through a fully connected layer and a classifier to produce the final text classification result. In summary, the HAN model structure closely aligns with the human understanding process, from words to sentences to passages. It not only addresses the problem of Text Convolutional Neural Networks (Text CNNs) losing text structure information, but also offers strong interpretability.

[0068] The attention mechanism is a problem-solving method that mimics human attention and can quickly filter out high-value information from a large amount of information. It is commonly used in encoder-decoder models. The attention mechanism helps the model assign different weights to each part of the input, extracting more critical and important information, enabling the model to make more accurate judgments without increasing the model's computational and storage overhead. For example, when the encoder + decoder model is used for translation, the input sentence and the output sentence often correspond to one or several input words and one or several output words. If each word in the sentence is given the same weight, this is unreasonable. Therefore, different weight values ​​are given to different words to distinguish the important parts of the sentence. Suppose the input sentence is "Today, Mingruns", and the output sentence is "Today, Xiaoming runs". The words "today", "Xiaoming" and "running" can be extracted from the translated sentence. Obviously, in the translated sentence, the three words have different importance. Among them, the importance of the word "today" is not as high as that of the words "Xiaoming" and "running". Therefore, the weight value of "today" can be set to 0.2, and the weight values ​​of the words "Xiaoming" and "running" can be set to 0.4 to increase the importance of the words "Xiaoming" and "running".

[0069] VistaNet model: It uses the Attention mechanism to fuse image and text information, cleverly solving the problem of inconsistent vector spaces between different modal data and enhancing the model's ability to analyze sentiment in comments. The VistaNet model is divided into three layers from bottom to top: word encoder + attention layer 11 (Word Encoder + Attention layer), sentence encoder + attention layer 12 (Sentence Encoder + Attention layer) and text encoder + attention layer 13 (Document Encoder + Attention layer). The following is an introduction to the structure of each layer, please refer to Figure 1:

[0070] 1. Word encoder + attention layer 11;

[0071] The input data of this layer is the article word vector w of each word in each sentence in the article content (the maximum single number is T). Optionally, the article word vector can be obtained through a pre-trained neural network model. After inputting the word vector w, the hidden state of each RNN in two directions is obtained through a bidirectional recurrent neural network (RNN). The two hidden states obtained are spliced ​​as the output of the time step (timestep). The attention mechanism is then used to calculate the importance weight α of each time step. After normalizing the importance weight α, the output of all time steps is weighted and summed to obtain the vector representation si of the sentence. The specific calculation formula is as follows:

[0072] u i,t =U T tanh(W w h i,t +b w );

[0073]

[0074] s i =∑ t α i,t h i,t ;

[0075] Among them, u i,t represents the weight of the tth word in the ith sentence. U is a randomly initialized value. tanh() represents the hyperbolic tangent function. W w It is the word embedding matrix corresponding to the article word vector w. i,t Represents the hidden state of the tth word in the ith sentence. w is a constant. i,t Represents the normalized weight of the tth word in the ith sentence. exp() represents the exponential function with the natural logarithm e as the base. s i The vector representation of the sentence is the article sentence vector.

[0076] 2. Sentence encoder + attention layer 12;

[0077] This layer inputs the sentence vector s of each sentence in the article content i (up to L sentences), the output is the text representation d for the jth image j, which is the text feature vector of the jth image. The input article sentence vector passes through the bidirectional RNN to obtain the hidden states of each RNN in two directions, and the hidden state hi of each sentence is obtained by splicing. On the other hand, the image feature vector mj of the jth image in the article content is extracted. Exemplary, a method for obtaining the image feature vector is as follows: image a in the article content is j (1≤j≤M, M is the number of images in the article content), is input into the CNN network, the image features are extracted, and the features are input into the fully connected layer, weights are assigned to each feature, and then the fully connected layer outputs the image feature vector mj. The attention mechanism is implemented on mj to obtain the importance weight β corresponding to each hi. Based on the obtained β, the hi for the jth image is weighted averaged to obtain the text representation dj for the jth image. The specific calculation formula is as follows:

[0078] p j =tanh(W p m j +b p );

[0079] q i =tanh(W q h i +b q );

[0080] v j,i =V T (p j ⊙q i +q i );

[0081]

[0082] d j =∑ i β j,i h i ;

[0083] Among them, p j Indicates the contribution of the j-th image to the weight value. p represents the embedding matrix of the jth image. b p Is a constant. i W represents the contribution of the i-th sentence to the weight value. q represents the embedding matrix of the i-th sentence. b q is a constant. V is a randomly initialized value. j,i Indicates the weight value of the i-th sentence corresponding to the j-th image. ⊙ represents the XOR operation. β j,i Represents the normalized weight value of the i-th sentence corresponding to the j-th image. i Represents the i-th hidden state. mj Represents the image feature vector of the j-th image.

[0084] 3. Text encoder + attention layer 13;

[0085] The final layer inputs multiple text feature vectors dj generated for different images. Using the attention mechanism, the corresponding weights are calculated and then weighted to obtain a text feature vector representing the entire article content. The article content is evaluated based on the resulting text feature vector. The specific calculation formula is as follows:

[0086] k j =K T tanh(W d d j +b d );

[0087]

[0088] d=∑ j γ j d j ;

[0089] Among them, k j Represents the weight value corresponding to the jth image. K is a randomly initialized value. W d is the embedding matrix corresponding to the text feature vector d. d is a constant. j is the normalized weight value corresponding to the jth image. d represents the evaluation result of the entire article content.

[0090] Figure 2 FIG2 is a block diagram of a computer system according to an exemplary embodiment of the present application. The computer system 200 includes a terminal 220 and a server 240 .

[0091] Terminal 220 has an application installed that is related to article content evaluation. This application can be a small program within an app (application), a dedicated application, or a web client. Users can receive article content evaluation results on terminal 220, or they can send article content to server 240, which will generate a corresponding evaluation result and return the evaluation result to terminal 220 or another terminal. Terminal 220 is at least one of a smartphone, an in-vehicle computer, a tablet computer, an e-book reader, an MP3 player, an MP4 player, a laptop computer, and a desktop computer.

[0092] The terminal 220 is connected to the server 240 via a wireless network or a wired network.

[0093] Server 240 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 240 is used to provide background services for applications that support article content evaluation. Optionally, server 240 undertakes the main computing work and terminal 220 undertakes the secondary computing work; or, server 240 undertakes the secondary computing work and terminal 220 undertakes the main computing work; or, server 240 and terminal 220 both adopt a distributed computing architecture for collaborative computing.

[0094] Figure 3 The flow chart of the method for evaluating the content of an article provided by an exemplary embodiment of the present application is shown. Figure 2 The terminal 220 or the server 240 shown in the figure performs the method, which includes the following steps:

[0095] Step 302: Extract image information and text information from the article content.

[0096] There are many ways to obtain article content, such as downloading article content from the Internet, receiving article content sent by other terminals, obtaining article content from local storage, or obtaining article content input in real time. This application does not impose any restrictions on this.

[0097] Image information refers to the images in the article content. For example, if the article content contains image A and image B, then image A and image B are image information.

[0098] Text information refers to the text in the article content. For example, if the article content contains text C and text D, then text C and text D are text information.

[0099] Step 304: obtaining multiple evaluation results of the article content based on the image information and text information, wherein the multiple evaluation results include at least two of the following: image and text evaluation results, text evaluation results, objective priori evaluation results, and typesetting evaluation results.

[0100] The evaluation results can be expressed in various forms. Optionally, the evaluation results are the identification and determination of high-quality and low-quality content in the article content. Optionally, the evaluation results are the scoring of the article content.

[0101] The image-text evaluation results are used to indicate the relevance between images and text within an article. For example, if the text in an article is "I went swimming yesterday," and the corresponding image is an image of a plum blossom, the relevance between the image and text is low, and we can conclude that the text and image are unrelated.

[0102] The text evaluation results are used to evaluate the quality of the text in the article content. The text evaluation results can identify high-quality text in the article content. Here, high-quality text refers to text that carries rich information and can be used by users to obtain information through reading.

[0103] Objective prior evaluation results refer to the evaluation results that can be obtained from the article content or related information without considering the article content. For example, if the article content is published by Account A, and Account A has published many high-quality articles, the objective prior evaluation here will be higher based on Account A's historical performance.

[0104] The layout evaluation results are used to assess the layout of the article content. For example, if Text A is on page 5 of the article, while the corresponding Image A is on page 12, the layout of Text A and Image A is unreasonable, making it very inconvenient for users to read the article. Therefore, the layout evaluation results will be poor.

[0105] In this embodiment, a method of obtaining multiple evaluation results through a neural network model is used as an example to illustrate:

[0106] 1. Extract multimodal feature information from image information and text information;

[0107] Multimodal feature information refers to multi-dimensional information that can be extracted from image information and text information, including but not limited to at least one of text-level information, image-level information, information corresponding to the combination of text and image, information corresponding to the account where the article content is published, and typesetting information. Exemplarily, multimodal feature information includes at least one of image feature vectors, article word vectors, article sentence vectors, text feature vectors (feature vectors corresponding to all or part of the text content in the article content), statistical features, linguistic features, image quality features, and account features. Here

[0108] 2. Input the multimodal feature information into the image-text prior quality evaluation model to obtain the multiple evaluation results of the article content. The image-text prior quality evaluation model is used to obtain multiple multi-dimensional evaluation results based on the multimodal feature information.

[0109] Step 306: Fuse multiple evaluation results to obtain a multi-dimensional evaluation result of the article content.

[0110] The multi-dimensional evaluation results are used to evaluate the content of the article from multiple dimensions. In this embodiment, the multi-dimensionality here refers to at least two dimensions among the above-mentioned graphic evaluation results, text evaluation results, objective prior evaluation results, and typesetting evaluation results.

[0111] Optionally, multiple evaluation results are fused through a neural network.

[0112] Optionally, corresponding weights are assigned to multiple evaluation results, and corresponding multi-dimensional evaluation results are calculated using the weights.

[0113] In summary, this embodiment integrates multiple evaluation results from multiple dimensions of article content to provide a corresponding evaluation of the article content from multiple perspectives, making the final evaluation result more closely aligned with the user's actual reading experience. Furthermore, since the final evaluation result integrates evaluations from multiple dimensions, it can effectively reduce erroneous evaluation results and improve the robustness of the entire solution.

[0114] Figure 4 The flow chart of the method for evaluating the content of an article provided by an exemplary embodiment of the present application is shown. Figure 2 The terminal 220 or the server 240 shown in the figure performs the method, which includes the following steps:

[0115] Step 401: extract image information and text information from the article content.

[0116] Step 402: Extract multimodal feature information from image information and text information.

[0117] In this embodiment, the multimodal feature information is processed by a priori quality evaluation model of images and texts. Figure 5 The following diagram illustrates an exemplary structural diagram of a priori quality evaluation model for images and texts, provided by an exemplary embodiment of the present application. The priori quality evaluation model for images and texts includes, but is not limited to, four subnetworks and an attention fusion layer 55, namely, an image-text modality subnetwork 51, an objective prior feature subnetwork 52, a text subnetwork 53, and a typesetting subnetwork 54. The priori quality evaluation model for images and texts is composed of at least two of the four subnetworks. The inputs of the four subnetworks are corresponding multimodal feature information, and the outputs are their respective evaluation results.

[0118] The attention fusion layer 55 receives multiple evaluation results as input and outputs a multi-dimensional evaluation result 56. The attention fusion layer 55 is used to assign corresponding weights to the input multiple evaluation results and perform corresponding weighted calculations to obtain the multi-dimensional evaluation result 56.

[0119] Step 403: extracting image feature vectors and text feature vectors from the multimodal feature information. The text feature vector is a feature vector corresponding to all or part of the text in the article content.

[0120] The extracted image feature vector can be obtained by directly extracting it or by processing the original data.

[0121] The text feature vector may be directly extracted or processed after extracting the original data. For example, the text feature vector is obtained by directly extracting the feature vectors of words. For example, the text feature vector is obtained by extracting the feature vectors of words and then processing the feature vectors of the words to obtain the feature vectors of the sentence.

[0122] The image feature vector is used to represent the feature vector of the image in the article content. The image feature vector can be obtained through the corresponding convolutional neural network.

[0123] Since the text in an article can be considered to be composed of multiple sentences, and sentences can be considered to be composed of multiple words, the text feature vector here can refer to either the article word vector or the article sentence vector. The article word vector can be obtained using a pre-trained neural network model.

[0124] Step 404: Input the image feature vector and the text feature vector into the image-text multimodal sub-network to obtain the image-text evaluation result of the article content.

[0125] This step includes the following sub-steps:

[0126] 1. Generate text feature representations that integrate image information and image feature representations that integrate text information through the image-text multimodal sub-network;

[0127] Text feature representation refers to a feature vector that integrates image information based on the text feature vector.

[0128] Image feature representation refers to the feature vector that integrates text information based on the image feature vector.

[0129] 2. Through the text-image multimodal sub-network, the text feature representation and image feature representation are integrated to generate the text-image evaluation results of the article content.

[0130] The fusion method here can be to fuse through the corresponding neural network model.

[0131] For example, Figure 6 An exemplary structural diagram of a graphic-text multimodal subnetwork provided by an exemplary embodiment of the present application is shown. The graphic-text multimodal subnetwork can be divided into two parts, left and right, from an overall perspective.

[0132] like Figure 6 As shown in the figure, we will first introduce the left part. From top to bottom, the left part is the transformer encoder 61 (transformer refers to a neural network model based on the encoder + decoder architecture. The attention mechanism is applied in the transformer model and is used in the field of natural language processing), transformer encoder 62, and attention fusion layer 63.

[0133] The transformer encoder 61 takes as input the article word vectors W1 to Wn (where n represents the maximum number of words in a sentence) and outputs the article sentence vectors S1 to SL (where L represents the maximum number of sentences in the article). The transformer encoder combines the n article word vectors it inputs and outputs L article sentence vectors.

[0134] The transformer encoder 62 takes as input the sentence vectors S1 to SL and outputs text feature vectors H1 to Hm (where m represents the maximum number of images in the article). There is a one-to-one correspondence between text feature vectors and images in the article. The transformer encoder combines the L input sentence vectors and outputs m text feature vectors.

[0135] The input of the attention fusion layer 63 is the text feature vectors H1 to Hm, and the output is the text feature representation 66 of the fused image features. The attention fusion layer 63 can fuse the corresponding image information on the basis of the text feature vectors through the attention mechanism to obtain the text feature representation 64 of the fused image features. The specific fusion process can be referred to Figure 1 The related calculation methods in the sentence encoder + attention layer 12 and text encoder + attention layer 13 in the VistaNet model shown.

[0136] Next, the right part of the image-text multimodal sub-network is introduced: the right part includes the extraction layer 64 and the attention fusion layer 65.

[0137] The input of the extraction layer 64 is the m images in the article content, and the output is the image feature vectors M1 to Mm. The extraction layer 64 is a pre-trained feature extraction network that can extract the corresponding image feature vectors from the images.

[0138] The input of the attention fusion layer 65 is the image feature vectors M1 to Mm, and the output is the image feature representation 66 fused with text information. Similar to the attention fusion layer 63, the attention fusion layer 65 can fuse the corresponding text information on the basis of the image feature vector through the attention mechanism to obtain the image feature representation 64 fused with text features. Similarly, the specific fusion process can be referred to Figure 1The related calculation methods in the sentence encoder + attention layer 12 and text encoder + attention layer 13 in the VistaNet model shown.

[0139] The image-text multimodal sub-network also includes a fusion layer 67 and a multi-layer perceptron (MLP) 68. The input of the fusion layer 67 is the text feature representation 64 of the fused image information and the image feature 66 of the fused text information, and the output is the fused feature.

[0140] The input of the multilayer perceptron 68 is the fused features, and the output is the image and text evaluation results. The multilayer perceptron 68 is used to identify and classify the fused features, extract useful information therein, and obtain the image and text evaluation results.

[0141] In summary, the image-text feature evaluation results can assess the correlation between images and text within an article. This can address issues such as insufficient image quality, redundant image information assigned to the same text description, and repeated images taken from multiple angles. This enables the sub-network to obtain accurate evaluation results, improving the recognition of high-quality content. Furthermore, the image-text multimodal sub-network can also learn features where images and text don't match, demonstrating its exceptional performance in articles with emotional or inspirational content.

[0142] Step 405: extracting objective prior features from the multimodal feature information, where the objective prior features include at least one of statistical features, linguistic features, image quality features, and account features.

[0143] The statistical features include at least one of page height, image area, word count, and paragraph count.

[0144] Linguistic features include at least one of lexical diversity, syntactic diversity, rhetorical devices and poetic quotations.

[0145] The image quality feature includes at least one of image clarity, number of image channels, image size, and number of images.

[0146] Account characteristics include at least one of account level, account verticality, and consumption data such as account collections and likes.

[0147] Step 406: Input the objective prior features into the objective prior feature sub-network to obtain the objective prior evaluation results of the article content.

[0148] For example, Figure 7The following figure shows an exemplary structural diagram of an objective prior feature subnetwork provided by an exemplary embodiment of the present application. From top to bottom, the objective prior feature subnetwork is divided into an embedding layer 71, a feature cross layer 72, and a multilayer perceptron 74. The input of the embedding layer 71 is at least one of statistical features, linguistic features, image quality features, and account features, and the output is a continuous feature 1 to feature x (x is a positive integer). The function of the embedding layer 71 is to convert discrete statistical features, linguistic features, image quality features, and account features into features that can be represented by continuous vectors, while also reducing the dimensionality of the input features.

[0149] The input to feature crossover layer 72 is features 1 through x, and the output is the crossover total feature 73. Feature crossover layer 72 multiplies each of the input features 1 through x, assigns weights, and then sums the results to obtain the corresponding total feature 73. Total feature 73 can represent the input as a whole, including statistical features, linguistic features, image quality features, and account features.

[0150] The input of the multilayer perceptron 74 is the total features 73, and the output is the objective prior evaluation result. The multilayer perceptron 74 is used to identify and classify the total features 73, extract useful information therein, and obtain the objective prior feature evaluation result.

[0151] In summary, objective a priori evaluation results can be used to derive implicit objective experience results. For example, the influence of account authority on article content, the influence of rhetorical techniques used in an article on article content, and so on. These influences are difficult to notice but exist in reality. Objective a priori evaluation results can visualize these implicit objective experience results, making it easier to obtain reasonable evaluation results.

[0152] Step 407: Extract article word vectors from the multimodal feature information.

[0153] Step 408: Input the article word vector into the text sub-network to obtain the text evaluation result of the article content.

[0154] Figure 8 FIG2 shows an exemplary structural diagram of a text sub-network according to an exemplary embodiment of the present application. The text sub-network includes a transformer encoder 81, a transformer encoder 82 and a multi-layer perceptron layer 84.

[0155] The transformer encoder 81 takes as input the article word vectors W1 to Wn (where n represents the maximum number of words in a sentence) and outputs the article sentence vectors S1 to SL (where L represents the maximum number of sentences in the article). The transformer encoder combines the n article word vectors it inputs into a single L-valued sentence vector.

[0156] The transformer encoder 82 takes as input the sentence vectors S1 to SL and outputs a text feature vector 83, where the text feature vector H represents the feature vector corresponding to all the words in the article. The transformer encoder combines the L input sentence vectors and outputs the text feature vector 83.

[0157] The input of the multilayer perceptron layer 84 is the text feature vector 83, and the output is the text evaluation result. The multilayer perceptron layer 84 is used to identify and classify the input text feature vector, extract useful information therein, and obtain the text evaluation result.

[0158] In summary, the text evaluation result is a specific evaluation of the text. Since the input features are highly correlated with the text, the most accurate evaluation result of the text in the article content can be obtained, making the final evaluation result more accurate in terms of text.

[0159] Step 409: extracting the image feature vector and the text feature vector from the multimodal feature information. The text feature vector is a feature vector corresponding to all or part of the text in the article content.

[0160] Step 410: Input the image feature vector and the text feature vector into the typesetting sub-network to obtain the typesetting evaluation result of the article content.

[0161] Figure 9 The following is an exemplary structural diagram of a typesetting sub-network according to an exemplary embodiment of the present application. The typesetting sub-network includes: a long short-term memory neural network 91, an attention fusion layer 92, a CNN 93 and a multi-layer perception layer 95.

[0162] The LSTM neural network 91 and the attention fusion layer 92 work together. The input of the LSTM neural network 91 is the interleaved image feature vectors and text feature vectors V1 to VL. The input of the attention fusion layer 92 is the typesetting features that are a fusion of the image information and text feature vectors, and the output is the typesetting features of the text feature vector that is a fusion of the image information.

[0163] The input of CNN93 is the staggered image feature vector and text feature vectors V1 to VL, and the output is the layout feature of the image vector fused with text information. The output of attention fusion layer 92 and the output of CNN93 constitute the layout feature 94.

[0164] The input of the multilayer perceptron layer 95 is the typesetting features 94, and the output is the typesetting evaluation result. The multilayer perceptron layer 95 is used to identify and classify the input typesetting features, extract useful information therein, and obtain the typesetting evaluation result.

[0165] In summary, the typography evaluation results can provide a corresponding assessment of the layout of the article content. This is because the layout of the article content is also an implicit objective experience. The impact of typography on the article content is subtle but real. The typography evaluation results can visualize these implicit objective experience results, facilitating the development of reasonable evaluation results.

[0166] Step 411: Assign corresponding weight values ​​to multiple evaluation results through the image-text prior quality evaluation model and attention mechanism.

[0167] The weight value can be adjusted according to actual needs.

[0168] Optionally, the weight values ​​are determined by a pre-trained neural network.

[0169] Step 412: Using the image-text prior quality evaluation model and the weight values, a weighted average is performed on the multiple evaluation results to obtain a multi-dimensional evaluation result of the article content.

[0170] For example, Figure 10 As shown, Figure 10 The complete image-text prior quality evaluation model of this embodiment is shown. For details, please refer to Figures 5 to 9 The corresponding content will not be repeated here.

[0171] To sum up, this embodiment covers all factors that affect the evaluation of article content by integrating multiple evaluation results of multiple dimensions of article content. Even in complex scenarios, it can reasonably evaluate the article content and obtain the high-quality and low-quality parts of the article content, so that the obtained evaluation results are more in line with the actual situation, which is conducive to merchants recommending high-quality article content to users.

[0172] Four different sub-networks are used to obtain corresponding evaluation results. Users can select corresponding evaluation results for integration according to actual needs to meet actual needs.

[0173] Figure 11 FIG1 shows a flow chart of a method for training a multimodal sub-network of images and texts provided by an exemplary embodiment of the present application. Figure 2 The method is performed by the terminal 220 or the server 240 or other computer device shown, and includes the following steps:

[0174] Step 1101: Obtain a picture and text training set, which includes sample articles and real picture and text evaluation results corresponding to the sample articles.

[0175] There are many ways to obtain sample articles, such as downloading sample articles from the Internet, receiving sample articles sent by other terminals, obtaining sample articles from local storage, or obtaining sample articles input in real time. This application does not impose any restrictions on this.

[0176] The real image-text evaluation results are the evaluation results made by relevant technical personnel or reading users on the image-text relationship in the sample article.

[0177] For the specific process of steps 1102 to 1104 in this embodiment, reference may be made to steps 401 to 404 .

[0178] Step 1102: Extract sample image information and sample text information from the sample article.

[0179] Step 1103: extracting a sample image feature vector and a sample text feature vector according to the sample image information and the sample text information.

[0180] Step 1104: Input the sample image feature vector and the sample text feature vector into the image-text multimodal sub-network to obtain a predicted image-text evaluation result.

[0181] Step 1105: Train the image-text multimodal sub-network based on the error loss between the predicted image-text evaluation result and the actual image-text evaluation result.

[0182] Optionally, the network parameters in the graphic-text multimodal sub-network are corrected by an error back propagation algorithm.

[0183] Optionally, when the error loss is no greater than a threshold, the training of the image-text multimodal sub-network is completed. The threshold can be determined by the technician.

[0184] Optionally, when the number of iterations of the image-text multimodal subnetwork reaches a threshold, the training of the image-text multimodal subnetwork is completed.

[0185] In summary, this embodiment provides a training method for a graphic-text multimodal subnetwork, which can quickly construct a graphic-text multimodal subnetwork, and the obtained graphic-text multimodal subnetwork can accurately obtain graphic-text multimodal evaluation results of the article content.

[0186] Figure 12 FIG1 shows a flow chart of an objective prior feature sub-network training method provided by an exemplary embodiment of the present application. Figure 2 The method is performed by the terminal 220 or the server 240 or other computer device shown, and includes the following steps:

[0187] Step 1201: Obtain an objective priori training set, where the objective priori training set includes sample articles and true objective priori evaluation results corresponding to the sample articles.

[0188] The true objective prior evaluation results are the evaluation results made by relevant technical personnel or reading users on the objective prior features in the sample articles.

[0189] For the specific process of step 1202 to step 1204 in this embodiment, reference may be made to step 401 to step 402 and step 405 to step 406.

[0190] Step 1202: Extract sample image information and sample text information from the sample article.

[0191] Step 1203: Obtaining sample objective prior features of the sample article based on the sample image information and the sample text information.

[0192] Step 1204: Input the objective prior features of the sample into the objective prior feature sub-network to obtain the predicted objective prior evaluation results.

[0193] Step 1205: Train the objective priori feature sub-network according to the error loss between the predicted objective priori evaluation result and the true objective priori evaluation result.

[0194] Optionally, the network parameters in the objective prior feature sub-network are corrected by an error back-propagation algorithm.

[0195] Optionally, when the error loss is not greater than a threshold, the training of the objective prior feature sub-network is completed. The threshold can be determined by the technician.

[0196] Optionally, when the number of iterations of the objective prior feature sub-network reaches a threshold, the training of the objective prior feature sub-network is completed.

[0197] In summary, this embodiment provides a method for training an objective priori feature subnetwork, which can quickly construct an objective priori feature subnetwork, and the obtained objective priori feature subnetwork can accurately obtain an objective priori evaluation result of the article content.

[0198] Figure 13 A flow chart of a text sub-network training method provided by an exemplary embodiment of the present application is shown. Figure 2 The method is performed by the terminal 220 or the server 240 or other computer device shown, and includes the following steps:

[0199] Step 1301: Obtain a text training set, which includes sample articles and real text evaluation results corresponding to the sample articles.

[0200] The real text evaluation results are the evaluation results made by relevant technicians or reading users on the text in the sample article.

[0201] For the specific process of step 1302 to step 1304 in this embodiment, reference may be made to step 401 to step 402 and step 407 to step 408.

[0202] Step 1302: Extract sample text information from the sample article.

[0203] Step 1303: Extract the sample article word vector based on the sample text information.

[0204] Step 1304: Input the sample article word vector into the text sub-network to obtain the predicted text evaluation result.

[0205] Step 1305: Train the text sub-network based on the error loss between the predicted text evaluation result and the real text evaluation result.

[0206] Optionally, the network parameters in the text sub-network are corrected by an error back-propagation algorithm.

[0207] Optionally, when the error loss is no greater than a threshold, the training of the text sub-network is completed. The threshold can be determined by the technician.

[0208] Optionally, when the number of iterations of the text sub-network reaches a threshold, the training of the text sub-network is completed.

[0209] In summary, this embodiment provides a text sub-network training method, which can quickly construct a text sub-network, and the obtained text sub-network can accurately obtain text evaluation results for the content of the article.

[0210] Figure 14 A flow chart of a typesetting sub-network training method provided by an exemplary embodiment of the present application is shown. Figure 2 The terminal 220 or the server 240 shown in the figure performs the method, which includes the following steps:

[0211] Step 1401: Obtain a typesetting training set, which includes sample articles and actual typesetting evaluation results corresponding to the sample articles.

[0212] The actual typesetting evaluation results are the evaluation results made by relevant technical personnel or reading users on the typesetting of the sample articles.

[0213] For the specific process of step 1402 to step 1404 in this embodiment, reference may be made to step 401 to step 402 and step 407 to step 408.

[0214] Step 1402: Extract sample image information and sample text information from the sample article.

[0215] Step 1403: extracting a sample image feature vector and a sample text feature vector according to the sample image information and the sample text information.

[0216] Step 1404: Input the sample image vector and the sample text feature vector into the typesetting subnetwork to obtain the predicted typesetting evaluation result.

[0217] Step 1405: Train the typesetting sub-network based on the error loss between the predicted typesetting evaluation result and the actual typesetting evaluation result.

[0218] Optionally, the network parameters in the typesetting sub-network are corrected by an error back propagation algorithm.

[0219] Optionally, when the error loss is no greater than a threshold, the training of the typesetting sub-network is completed. The threshold can be determined by the technician.

[0220] Optionally, when the number of iterations of the objective prior feature sub-network reaches a threshold, the training of the typesetting sub-network is completed.

[0221] In summary, this embodiment provides a typesetting sub-network training method, which can quickly construct a typesetting sub-network, and the obtained typesetting sub-network can accurately obtain text evaluation results for article content.

[0222] Figure 15 An exemplary business architecture diagram of an exemplary embodiment of the present application is shown. The architecture diagram is divided into two parts: a low-quality filtering module 1501 and a high-quality identification module 1502.

[0223] Low-quality filtering module 1501 is used to filter low-quality and sub-low-quality content from article content. Low-quality content includes, but is not limited to, at least one of vulgar content, rumors, clickbait, and advertising and marketing. Vulgar content refers to content that is detrimental to social progress and the user's physical and mental development. Rumors refer to content in an article that is inconsistent with reality. Clickbait refers to content that doesn't match the title. Advertising and marketing refers to content that promotes or advertises a product. Sub-low-quality content includes, but is not limited to, at least one of: meaningless content, formulaic content, gossip, promotional content, patchwork content, negative impact content, gossip, and soft advertising. Meaningless content refers to content that is dispensable and provides no useful information to users. Formulaic content refers to content written according to a fixed template. Gossip refers to articles that make groundless speculation about people or events. Promotional content refers to articles that promote individuals, groups, or places. Patchwork content refers to articles that are spliced ​​together in whole or in part from other articles. Negative impact content refers to articles whose content may negatively impact individuals, groups, or society. "Sloppy writing" refers to articles that are unrefined and resemble spoken language. Advertisement articles are essentially advertisements, but it is difficult for users to directly determine whether an article is an advertisement based on its content. Optionally, the low-quality filtering module 1501 is implemented using a corresponding low-quality filtering neural network.

[0224] The quality identification module 1502 is used to extract the quality parts of the article content. The quality identification module includes: a feature extraction layer 1505, a feature fusion layer 1503, a multi-objective feedback layer 1504, and a logic decision layer 1506.

[0225] The feature extraction layer 1505 extracts corresponding feature vectors based on three aspects: image-text multimodality, image-text atomic capabilities, and image-text nested layout. The input of the feature extraction layer is the article content, and the output is the extracted feature vectors. Specifically, image-text multimodality is used to extract feature vectors corresponding to images within the article content, feature vectors corresponding to text, and feature vectors corresponding to the correlation between images and text. Image-text atomic capabilities extract corresponding feature vectors based on four aspects: linguistics, statistics, account, and article style. Linguistics includes lexical diversity, syntactic diversity, rhetorical devices, and poetic quotations. Statistics includes page height, image size and word count, number of paragraphs, and average image clarity and aesthetics. Account includes account level, account verticality (account verticality refers to the expertise of the account's published content in a specific field), and consumption data such as account favorites and likes. Article style includes practicality, positive energy, and professionalism.

[0226] The input to the feature fusion layer 1503 is the feature vector extracted from the article content, and the output is the prior quality content of the article content (prior quality content refers to the content obtained by predicting the quality parts of the article content without user feedback). The feature fusion layer fuses the input feature vectors and obtains the prior quality content based on the fused feature vectors.

[0227] The input of the multi-objective feedback layer 1504 is the fused feature vector and feedback information from multiple users, and the output is a posteriori high-quality content (a posteriori high-quality content refers to the content predicted after correcting the input feature vector based on user feedback information). The multi-objective feedback layer can correct the input feature vector based on user feedback information and fuse the corrected feature vector to obtain the a posteriori high-quality content. In another implementation method, the multi-objective feedback layer can first fuse the input feature vector to obtain high-quality content, and then correct the high-quality content based on user feedback information to obtain a posteriori high-quality content.

[0228] The input of the logic decision layer 1506 is the prior quality content and the a posteriori quality content, and the output is the quality content of the article content. The logic decision layer 1506 will comprehensively evaluate the prior quality content and the a posteriori quality content to obtain the quality content of the article content.

[0229] This embodiment innovatively breaks down the complex scenario of determining high-quality text and image content from multiple dimensions, including text and image multimodality, accounts, article layout experience, and article linguistic atomic features (such as the lexical diversity used in the article, whether the article uses diverse syntax such as metaphors and parallelism, and whether the article quotes ancient poetry). It also builds an integrated model that integrates the text and image modality subnetwork, the objective prior feature subnetwork, the text subnetwork, and the layout subnetwork, thereby completing the identification and determination of high-quality text and image content.

[0230] This model is mainly used in the task of judging the quality of graphic content in the content algorithm research and development center. The model accuracy rate reaches 94%, and the coverage rate of high-quality graphic content reaches 16%. The recommended weighted experiments were conducted on the identified high-quality graphic content on the browser and express side, which realized the priority recommendation of high-quality content to users, and achieved good business results on the business side. After using the graphic and text priori high-quality recognition algorithm described in this application to conduct high-quality content weighted recommendation experiments, the overall market click pv (page view, page views) on the browser side increased by 0.38%, the market exposure efficiency increased by 0.43%, the market CTR (Click-Through-Rate, click-through rate) increased by 0.394%, and the user time increased by 0.17%; at the same time, the DAU (Daily Active User) retention rate on the next day increased by 0.165%, the average share per person in the interactive indicator data increased by 1.705%, the average like per person increased by 4.215%, and the average comment per person increased by 0.188%.

[0231] The following is an embodiment of the device of the present application. For details not described in detail in the embodiment of the device, reference can be made to the corresponding records in the above method embodiment, and no further details will be given herein.

[0232] Figure 16 A schematic diagram of the structure of an article content evaluation device provided by an exemplary embodiment of the present application is shown. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device 1600 includes:

[0233] Extraction module 1601, used to extract image information and text information from the article content;

[0234] An evaluation module 1602 is configured to obtain multiple evaluation results of the article content based on the image information and the text information, wherein the multiple evaluation results include at least two of an image-text evaluation result, a text evaluation result, an objective priori evaluation result, and a typesetting evaluation result;

[0235] The evaluation fusion module 1603 is used to fuse the multiple evaluation results to obtain a multi-dimensional evaluation result of the article content.

[0236] In an optional design of the present application, the extraction module 1601 is also used to extract multimodal feature information from the image information and the text information.

[0237] The evaluation module 1602 is further configured to input the multimodal feature information into a priori high-quality evaluation model for images and texts to obtain the multiple evaluation results for the article content. The priori high-quality evaluation model for images and texts is configured to obtain the multiple multi-dimensional evaluation results based on the multimodal feature information.

[0238] In an optional design of the present application, the extraction module 1601 is further used to extract image feature vectors and article word vectors in the multimodal feature information, where the article word vectors are feature vectors corresponding to words in the article content.

[0239] The evaluation module 1602 is further configured to input the image feature vector and the article word vector into the image-text multimodal sub-network to obtain the image-text evaluation result of the article content.

[0240] In an optional design of the present application, the evaluation module 1602 is also used to generate a text feature representation of fused image information and an image feature representation of fused text information through the image-text multimodal sub-network; and to fuse the text feature representation and the image feature representation through the image-text multimodal sub-network to generate the image-text evaluation result of the article content.

[0241] In an optional design of the present application, the extraction module 1601 is also used to extract objective prior features in the multimodal feature information, and the objective prior features include at least one of statistical features, linguistic features, image quality features, and account features.

[0242] The evaluation module 1602 is further configured to input the objective priori features into the objective priori feature sub-network to obtain an objective priori evaluation result of the article content.

[0243] In an optional design of the present application, the extraction module 1601 is also used to extract article word vectors from the multimodal feature information.

[0244] The evaluation module 1602 is further configured to input the article word vector into the text sub-network to obtain the text evaluation result of the article content.

[0245] In an optional design of the present application, the extraction module 1601 is also used to extract image feature vectors and text feature vectors in the multimodal feature information, and the text feature vector is a feature vector corresponding to all or part of the text in the article content.

[0246] The evaluation module 1602 is further configured to input the image feature vector and the text feature vector into the typesetting sub-network to obtain the typesetting evaluation result of the article content.

[0247] In an optional design of the present application, the evaluation fusion module 1603 is also used to assign corresponding weight values ​​to the multiple evaluation results through the image-text prior high-quality evaluation model and the attention mechanism; and to perform weighted averaging on the multiple evaluation results through the image-text prior high-quality evaluation model and the weight values ​​to obtain the multi-dimensional evaluation results of the article content.

[0248] In an optional design of the present application, the device further includes: a training module 1604.

[0249] The training module 1604 is used to obtain a picture and text training set, which includes sample articles and real picture and text evaluation results corresponding to the sample articles; extract sample image information and sample text information from the sample articles; extract sample image feature vectors and sample text feature vectors based on the sample image information and the sample text information; input the sample image feature vectors and the sample text feature vectors into the picture and text multimodal sub-network to obtain predicted picture and text evaluation results; and train the picture and text multimodal sub-network based on the error loss between the predicted picture and text evaluation results and the real picture and text evaluation results.

[0250] In an optional design of the present application, the training module 1604 is also used to obtain an objective prior training set, which includes sample articles and real objective prior evaluation results corresponding to the sample articles; extract sample image information and sample text information in the sample articles; obtain sample objective prior features of the sample articles based on the sample image information and the sample text information; input the sample objective prior features into the objective prior feature sub-network to obtain predicted objective prior evaluation results; and train the objective prior feature sub-network based on the error loss between the predicted objective prior evaluation results and the real objective prior evaluation results.

[0251] In an optional design of the present application, the training module 1604 is also used to obtain a text training set, which includes sample articles and real text evaluation results corresponding to the sample articles; extract sample text information from the sample articles; extract sample article word vectors based on the sample text information; input the sample article word vectors into the text sub-network to obtain predicted text evaluation results; and train the text sub-network based on the error loss between the predicted text evaluation results and the real text evaluation results.

[0252] In an optional design of the present application, the training module 1604 is also used to obtain a typesetting training set, which includes sample articles and real typesetting evaluation results corresponding to the sample articles; extract sample image information and sample text information in the sample articles; extract sample image feature vectors and sample text feature vectors based on the sample image information and the sample text information; input the sample image vector and the sample text feature vector into the typesetting sub-network to obtain a predicted typesetting evaluation result; and train the typesetting sub-network based on the error loss between the predicted typesetting evaluation result and the real typesetting evaluation result.

[0253] In summary, this embodiment integrates multiple evaluation results from multiple dimensions of article content to provide a corresponding evaluation of the article content from multiple perspectives, making the final evaluation result more closely aligned with the user's actual reading experience. Furthermore, since the final evaluation result integrates evaluations from multiple dimensions, it can effectively reduce erroneous evaluation results and improve the robustness of the entire solution.

[0254] Figure 17 17 is a schematic diagram illustrating the structure of a computer device according to an exemplary embodiment. The computer device 1700 includes a central processing unit (CPU) 1701, a system memory 1704 including a random access memory (RAM) 1702 and a read-only memory (ROM) 1703, and a system bus 1705 connecting the system memory 1704 and the CPU 1701. The computer device 1700 also includes a basic input / output system (I / O system) 1706 for facilitating information transmission between various components within the computer device, and a mass storage device 1707 for storing an operating system 1713, application programs 1714, and other program modules 1715.

[0255] The basic input / output system 1706 includes a display 1708 for displaying information and an input device 1709 such as a mouse and keyboard for user input. The display 1708 and the input device 1709 are connected to the central processing unit 1701 via an input / output controller 1710 connected to the system bus 1705. The basic input / output system 1706 may also include an input / output controller 1710 for receiving and processing input from a variety of other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1710 also provides output to a display screen, printer, or other types of output devices.

[0256] The mass storage device 1707 is connected to the central processing unit 1701 via a mass storage controller (not shown) connected to the system bus 1705. The mass storage device 1707 and its associated computer-readable medium provide non-volatile storage for the computer device 1700. In other words, the mass storage device 1707 may include computer-readable media (not shown) such as a hard disk or a CD-ROM drive.

[0257] Without loss of generality, the computer device readable medium may include computer device storage media and communication media. Computer device storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer device readable instructions, data structures, program modules or other data. Computer device storage media include RAM, ROM, Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), CD-ROM, Digital Video Disc (DVD) or other optical storage, tape cassettes, magnetic tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer device storage media is not limited to the above-mentioned ones. The above-mentioned system memory 1704 and mass storage device 1707 can be collectively referred to as memory.

[0258] According to various embodiments of the present disclosure, the computer device 1700 may also be connected to a remote computer device on a network such as the Internet for operation. That is, the computer device 1700 may be connected to a network 1711 via a network interface unit 1712 connected to the system bus 1705, or the network interface unit 1712 may be used to connect to other types of networks or remote computer device systems (not shown).

[0259] The memory further includes one or more programs, which are stored in the memory. The central processing unit 1701 implements all or part of the steps of the above-mentioned article content evaluation method by executing the one or more programs.

[0260] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one instruction, at least one program, code set or instruction set is stored. The at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the article content evaluation method provided by the above-mentioned various method embodiments.

[0261] The present application also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set. The at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the article content evaluation method provided in the above method embodiment.

[0262] According to another aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the article content evaluation method described above.

[0263] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0264] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0265] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for evaluating article content, characterized in that: The method comprises: Extracting image information and text information from the article content; extracting multimodal feature information from the image information and the text information; Inputting the multimodal feature information into a graphic-text prior quality evaluation model to obtain multiple evaluation results of the article content, wherein the graphic-text prior quality evaluation model is used to obtain the multiple evaluation results in multiple dimensions based on the multimodal feature information; the multiple evaluation results include at least two of the graphic-text evaluation results, the text evaluation results, the objective prior evaluation results, and the typesetting evaluation results; Fusion of the multiple evaluation results to obtain a multi-dimensional evaluation result of the article content; The image-text prior quality evaluation model includes: an image-text multimodal subnetwork; the multiple evaluation results include: image-text evaluation results; The step of inputting the multimodal feature information into a priori quality evaluation model for images and texts to obtain multiple evaluation results for the article content includes: Extracting an image feature vector and a text feature vector from the multimodal feature information, wherein the text feature vector is a feature vector corresponding to all or part of the text in the article content; The image feature vector and the text feature vector are input into the image-text multimodal sub-network to obtain the image-text evaluation result of the article content.

2. The method according to claim 1, characterized in that Inputting the image feature vector and the text feature vector into the image-text multimodal sub-network to obtain the image-text evaluation result of the article content includes: Generate a text feature representation of fused image information and an image feature representation of fused text information through the image-text multimodal sub-network; The text feature representation and the image feature representation are fused through the image-text multimodal sub-network to generate the image-text evaluation result of the article content.

3. The method according to claim 1, characterized in that The image-text prior quality evaluation model further includes: an objective prior feature sub-network; the multiple evaluation results include: an objective prior evaluation result; The step of inputting the multimodal feature information into a priori quality evaluation model for images and texts to obtain multiple evaluation results for the article content includes: Extracting objective prior features from the multimodal feature information, where the objective prior features include at least one of statistical features, linguistic features, image quality features, and account features; The objective priori features are input into the objective priori feature sub-network to obtain an objective priori evaluation result of the article content.

4. The method according to claim 1, wherein The image-text prior quality evaluation model further includes: a text sub-network; the multiple evaluation results further include: a text evaluation result; Inputting the multimodal feature information into a priori quality evaluation model for images and texts to obtain the multiple evaluation results of the article content includes: Extracting article word vectors from the multimodal feature information; The article word vector is input into the text sub-network to obtain the text evaluation result of the article content.

5. The method according to claim 1, characterized in that The image-text prior quality evaluation model further includes: a typesetting sub-network; the multiple evaluation results include: a typesetting evaluation result; The step of inputting the multimodal feature information into a priori quality evaluation model for images and texts to obtain multiple evaluation results for the article content includes: Extracting an image feature vector and a text feature vector from the multimodal feature information, wherein the text feature vector is a feature vector corresponding to all or part of the text in the article content; The image feature vector and the text feature vector are input into the typesetting subnetwork to obtain the typesetting evaluation result of the article content.

6. The method according to any one of claims 1 to 5, characterized in that: The fusing of the multiple evaluation results to obtain a multi-dimensional evaluation result of the article content includes: Assign corresponding weights to the multiple evaluation results through the image-text prior quality evaluation model and attention mechanism; The plurality of evaluation results are weightedly calculated using the image-text prior quality evaluation model and the weight values ​​to obtain the multi-dimensional evaluation results of the article content.

7. The method according to claim 1 or 2, characterized in that The image-text multimodal sub-network is trained by the following method: Obtaining a picture-text training set, wherein the picture-text training set includes sample articles and real picture-text evaluation results corresponding to the sample articles; extracting sample image information and sample text information from the sample article; Extracting a sample image feature vector and a sample text feature vector according to the sample image information and the sample text information; Inputting the sample image feature vector and the sample text feature vector into the image-text multimodal subnetwork to obtain a predicted image-text evaluation result; The image-text multimodal subnetwork is trained according to the error loss between the predicted image-text evaluation result and the true image-text evaluation result.

8. The method according to claim 3, characterized in that The objective prior feature sub-network is trained by the following method: Obtaining an objective priori training set, wherein the objective priori training set includes sample articles and true objective priori evaluation results corresponding to the sample articles; Extracting sample image information and sample text information from sample articles; Obtaining a sample objective priori feature of the sample article according to the sample image information and the sample text information; Inputting the sample objective priori features into the objective priori feature sub-network to obtain a predicted objective priori evaluation result; The objective priori feature sub-network is trained according to the error loss between the predicted objective priori evaluation result and the true objective priori evaluation result.

9. The method according to claim 4, characterized in that The text sub-network is trained by the following method: Obtaining a text training set, wherein the text training set includes sample articles and real text evaluation results corresponding to the sample articles; Extracting sample text information from the sample article; Extracting sample article word vectors based on the sample text information; Inputting the sample article word vector into the text sub-network to obtain a predicted text evaluation result; The text sub-network is trained according to the error loss between the predicted text evaluation result and the true text evaluation result.

10. The method according to claim 5, characterized in that The typesetting sub-network is trained by the following method: Obtaining a typesetting training set, wherein the typesetting training set includes sample articles and real typesetting evaluation results corresponding to the sample articles; extracting sample image information and sample text information from the sample article; Extracting a sample image feature vector and a sample text feature vector according to the sample image information and the sample text information; Inputting the sample image vector and the sample text feature vector into the typesetting subnetwork to obtain a predicted typesetting evaluation result; The typesetting sub-network is trained according to the error loss between the predicted typesetting evaluation result and the actual typesetting evaluation result.

11. A device for evaluating article content, characterized in that: The device comprises: An extraction module, used to extract image information and text information from the article content; an evaluation module, configured to extract multimodal feature information from the image information and the text information; Inputting the multimodal feature information into a graphic-text prior quality evaluation model to obtain multiple evaluation results of the article content, wherein the graphic-text prior quality evaluation model is used to obtain the multiple evaluation results in multiple dimensions based on the multimodal feature information; the multiple evaluation results include at least two of the graphic-text evaluation results, the text evaluation results, the objective prior evaluation results, and the typesetting evaluation results; An evaluation fusion module, configured to fuse the multiple evaluation results to obtain a multi-dimensional evaluation result of the article content; The image-text prior quality evaluation model includes: an image-text multimodal subnetwork; the multiple evaluation results include: image-text evaluation results; the evaluation module is further used to: Extracting an image feature vector and a text feature vector from the multimodal feature information, wherein the text feature vector is a feature vector corresponding to all or part of the text in the article content; The image feature vector and the text feature vector are input into the image-text multimodal sub-network to obtain the image-text evaluation result of the article content.

12. A computer device, characterized in that: The computer device includes: a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the article content evaluation method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the article content evaluation method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the article content evaluation method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Article quality evaluation method, article recommendation method and corresponding devices

    CN111488931A

  • KR20200064198A