Book review generation method, system and equipment based on multi-modal data and medium

By constructing graph structures and using graph neural networks and recurrent neural networks to process multimodal data, high real-time and high accuracy e-book reviews are generated, and the subjectivity and logical relationship ignorance of book review generation in the existing technology is solved.

CN120296162AActive Publication Date: 2025-07-11UNICOM WOYUEDU TECH CULTURE CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510783554.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

When generating e-book reviews in the prior art, there are problems of strong subjectivity and ignoring the logical relationship between evaluations in different periods, resulting in insufficient real-time and accuracy of book reviews.

Method used

Through a multimodal data-based method, a graph structure is constructed, and a graph neural network and a recurrent neural network are used to extract text vectors from historical evaluation data, predict the current evaluation data, and generate book reviews with high real-time and high accuracy.

Benefits of technology

实现了更准确地反映电子书在不同时期的评价数据变化,提高了书评的实时性和准确性,避免了基于关键字查找导致的逻辑关系缺失。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296162A_ABST
    Figure CN120296162A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal data-based book review generation method, system and device and a medium, and the method comprises the steps: forming a vector matrix through taking each group of historical text vectors as row and column matrix elements, and forming a relation matrix through the interaction relation between every two groups of historical text vectors; then, by constructing a graph structure composed of a vector matrix and a relation matrix and extracting a plurality of current text vectors corresponding to a plurality of groups of historical text vectors from the graph structure, the method can fully mine the relation in each group of historical text vectors on a time sequence and the interaction relation among the plurality of groups of historical text vectors; according to the method, the evaluation data change of the electronic book in different periods can be better reflected, the problem of logical relationship missing caused by the evaluation data change in different periods due to search only based on keywords is avoided, the current evaluation data can be effectively predicted, and the real-time performance and accuracy of book review generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of e - book review processing, and in particular, to a method, system, device, and medium for generating book reviews based on multi - modal data. Background Art

[0002] An e - book is a digital publication that digitizes multimedia content such as text, pictures, sounds, and videos. An e - book review is an article that analyzes and comments on aspects such as the content, style, and value of an e - book. It usually contains detailed descriptions and evaluations of aspects such as the theme, plot, characters, language, and structure of the e - book, aiming to help readers understand the content and features of the e - book, as well as the author's viewpoints and styles, so as to guide readers to choose reading materials suitable for themselves or provide references for academic research.

[0003] After reading an e - book, users generate evaluations of the e - book in different ways on different platforms. For example, they upload their impressions in the form of text and pictures on the e - book reading platform, or upload evaluations in the form of videos on the short - video platform. For authors and e - book promoters, in order to expand the influence of e - book articles, they often generate book reviews to introduce and promote the e - book articles. The existing methods for generating book reviews mainly include the following two methods: The first method is to write manually, but manual writing is relatively subjective. The second method is to search for evaluations of the e - book on different platforms and generate a book review based on the evaluations. This method is more objective and efficient than the first method. However, most of this method is based on keyword search, and the content searched is the evaluations in all time periods after the e - book article is generated. In the process of generating a book review based on these evaluations, it often ignores the fact that the e - book evaluation changes over time due to multiple factors such as the audience group, social and cultural background, its own artistic value, and the e - book environment. The method based solely on keyword search lacks dynamic evolution and trend analysis of evaluation data and cannot provide book reviews with high real - time performance and high accuracy. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.

[0005] The main purpose of the embodiments of this disclosure is to propose a method, system, device, and medium for generating book reviews based on multi - modal data, which can obtain book reviews with high real - time performance and high accuracy.

[0006] The first aspect of the embodiments of this application proposes a method for generating a book review based on multi - modal data, and the method includes: Search for a first set of the target e - book based on keywords; the first set includes multiple groups of historical evaluation data of the target e - book, where any two groups of historical evaluation data belong to different - modality historical evaluation data, and each group of historical evaluation data includes evaluation data collected at multiple historical moments; Convert the multiple groups of evaluation data in the first set into multiple groups of historical evaluation texts, extract multiple groups of historical text vectors from the multiple groups of historical evaluation texts, form a second set with the multiple groups of historical text vectors, and sample the multiple groups of historical text vectors in the second set according to a time window, and each group of the sampled historical text vectors contains multiple historical text vectors; Generate a graph structure composed of a vector matrix and a relationship matrix based on the multiple groups of historical text vectors in the second set; where each matrix element in the vector matrix is composed of a historical text vector in the second set, and each row or each column of matrix elements is composed of a group of historical text vectors, and each matrix element in the relationship matrix represents the interaction relationship between any two rows or any two columns of matrix elements in the vector matrix; Predict multiple current text vectors corresponding to the multiple groups of historical text vectors in the second set according to the graph structure, and convert the multiple current text vectors into multiple current evaluation texts; Generate a book review of the target e - book according to the multiple current evaluation texts.

[0007] A method for generating a book review based on multi - modality data provided by an embodiment of the present disclosure has at least the following beneficial effects: By comprehensively using the multi - modality historical evaluation data of the target e - book, this method can more accurately reflect the logical relationship between evaluations in different periods, thereby generating a more real - time and accurate book review. Specifically, each group of historical text vectors in the multiple groups of historical text vectors is used as a row - column matrix element to form a vector matrix, and the interaction relationship between every two groups of historical text vectors is used to form a relationship matrix. Then, a graph structure composed of a vector matrix and a relationship matrix is constructed, and multiple current text vectors corresponding to the multiple groups of historical text vectors are extracted from the graph structure. This embodiment can fully explore the relationship within each group of historical text vectors in the time series and the interaction relationship between multiple groups of historical text vectors, can better reflect the change of evaluation data of the e - book in different periods, avoids the problem of missing logical relationships caused by the change of evaluation data in different periods due to only keyword - based search, can effectively predict the current evaluation data, and improves the real - time and accuracy of book review generation.

[0008] In some embodiments, the keyword is a sentence with the theme information in the target e - book.

[0009] In some embodiments, the subject information includes at least one of a title, an abstract, and a brief introduction.

[0010] In some embodiments, extracting multiple groups of historical text vectors from the multiple groups of historical evaluation texts includes: Extracting corresponding historical text vectors from each group of historical evaluation texts based on Doc2Vec.

[0011] In some embodiments, the different modalities include at least two of a text modality, an image modality, and an audio modality.

[0012] In some embodiments, predicting multiple current text vectors corresponding to the multiple groups of historical text vectors according to the graph structure includes: Constructing a first prediction model and a second prediction model, where the first prediction model includes any one of graph neural networks, and the second prediction model includes any one of recurrent neural networks; Extracting graph embedding features obtained after embedding the information in the relationship matrix into the vector matrix from the graph structure according to the first prediction model; Predicting a current text vector corresponding to each group of the historical text vectors from the graph embedding features according to the second prediction model.

[0013] In some embodiments, the training processes of the first prediction model and the second prediction model include the following procedures: Constructing an evaluation network and constructing multiple groups of historical text vectors to be trained; Generating a to-be-trained graph structure composed of a to-be-trained vector matrix and a to-be-trained relationship matrix according to the to-be-trained historical text vectors; the to-be-trained vector matrix is generated from the multiple groups of to-be-trained historical text vectors, and the elements in the to-be-trained relationship matrix are generated from the interaction relationships between every two matrix elements in the to-be-trained vector matrix; Inputting the to-be-trained graph structure into the first prediction model to obtain to-be-trained graph embedding features output by the first prediction model; Inputting the to-be-trained graph embedding features into the second prediction model to obtain multiple to-be-trained current text vectors predicted by the second prediction model for the multiple groups of to-be-trained historical text vectors; Constructing positive samples and negative samples; where the negative samples are obtained by splicing the to-be-trained historical text vectors and the to-be-trained current text vectors, and the positive samples are obtained by splicing the to-be-trained historical text vectors and preset correct labels; Inputting the positive samples and the negative samples into the evaluation model to determine the probabilities between the positive samples and the negative samples according to the evaluation model; According to the probability and a preset evaluation function, perform backpropagation on the first prediction model and the second prediction model until the training processes of the first prediction model and the second prediction model are completed.

[0014] A second aspect of the embodiments of the present application provides a book review generation system based on multimodal data. The system includes: An evaluation extraction unit, configured to find a first set of target e-books based on keywords; the first set includes multiple groups of historical evaluation data of the target e-book, where any two groups of historical evaluation data belong to different modalities of historical evaluation data, and each group of historical evaluation data includes evaluation data collected at multiple historical moments; A text vector extraction unit, configured to convert the multiple groups of evaluation data in the first set into multiple groups of historical evaluation texts, extract multiple groups of historical text vectors from the multiple groups of historical evaluation texts, form a second set with the multiple groups of historical text vectors, and sample the multiple groups of historical text vectors in the second set according to a time window, and each group of the sampled historical text vectors includes multiple historical text vectors; A matrix generation unit, configured to generate a graph structure composed of a vector matrix and a relationship matrix according to the multiple groups of historical text vectors in the second set; where each matrix element in the vector matrix is composed of a historical text vector in the second set, and each row or each column of matrix elements is composed of a group of historical text vectors, and each matrix element in the relationship matrix represents an interaction relationship between any two rows or any two columns of matrix elements in the vector matrix; A text vector extraction unit, configured to predict multiple current text vectors corresponding to the multiple groups of historical text vectors in the second set according to the graph structure, and convert the multiple current text vectors into multiple current evaluation texts; A book review generation unit, configured to generate a book review of the target e-book according to the multiple current evaluation texts.

[0015] A third aspect of the embodiments of the present application provides an electronic device, at least one controller and a memory for communicatively connecting with the at least one controller; the memory stores instructions executable by the at least one controller, and the instructions are executed by the controller to enable the controller to execute the multimodal data-based book review generation method as described in the first aspect.

[0016] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute the multimodal data-based book review generation method as described above.

[0017] It is understandable that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For the relevant descriptions, reference can be made to the above first aspect and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the related art. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a schematic flowchart of a book review generation method based on multimodal data provided by the present application; Figure 2 is a schematic diagram of a graph structure proposed in an embodiment of the present application; Figure 3 is a schematic flowchart of the training process of the overall system proposed in an embodiment of the present application; Figure 4 is a schematic diagram of the structure of a book review generation system based on multimodal data proposed in an embodiment of the present application; Figure 5 is a schematic diagram of the structure of an electronic device proposed in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0021] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0023] An e - book is a digital publication that digitizes multimedia content such as text, pictures, sounds, and videos. And an e - book review is an article that analyzes and comments on aspects such as the content, style, and value of an e - book. It usually includes detailed descriptions and evaluations of the theme, plot, characters, language, structure, etc. of the e - book, aiming to help readers understand the content and characteristics of the e - book, as well as the author's viewpoints and styles, so as to guide readers to choose suitable reading materials or provide references for academic research.

[0024] After reading an e - book, users generate evaluations of the e - book in different ways on different platforms. For example, they upload their reading impressions in the form of text and pictures on the e - book reading platform, or upload evaluations in the form of videos on the short - video platform. For authors and e - book promoters, in order to expand the influence of e - book articles, they often generate book reviews to introduce and promote the e - book articles. The existing ways of generating book reviews mainly include the following two methods: The first method is to write by hand, but writing by hand is quite subjective. The second method is to search for evaluations of the e - book on different platforms and generate a book review based on the evaluations. This method is more objective and efficient than the first method. However, most of this method is based on keyword search, and the content searched is the evaluations in all time periods after the e - book article is generated. In the process of generating a book review based on these evaluations, it often ignores the fact that the e - book evaluation changes over time due to multiple factors such as the audience group, social and cultural background, its own artistic value, and the e - book environment. The method based solely on keyword search lacks the dynamic evolution and trend analysis of evaluation data and cannot provide book reviews with high real - time performance and high accuracy.

[0025] This application relates to a method for generating a book review based on multi - modal data, which aims to solve the problems of real - time performance and accuracy in book review generation in the prior art. The prior art mainly relies on writing by hand or generating book reviews based on keyword search, which has problems such as strong subjectivity and ignoring the logical relationship between evaluations in different periods. Therefore, a new method is proposed to generate more accurate and real - time book reviews.

[0026] The method of this application includes the following steps S100 to S500: Step S100, based on keywords, search for a first set of data on the target e - book; the first set includes multiple groups of historical evaluation data on the target e - book, where any two groups of historical evaluation data belong to historical evaluation data of different modalities, and each group of historical evaluation data includes evaluation data collected at multiple historical moments.

[0027] Step S200: Convert multiple groups of evaluation data in the first set into multiple groups of historical evaluation texts, extract multiple groups of historical text vectors from the multiple groups of historical evaluation texts, form the multiple groups of historical text vectors into a second set, and sample the multiple groups of historical text vectors in the second set according to a time window so that each group of historical text vectors contains multiple historical text vectors.

[0028] Step S300: Generate a graph structure composed of a vector matrix and a relationship matrix based on the multiple groups of historical text vectors in the second set; wherein, each matrix element in the vector matrix is composed of a historical text vector in the second set, and each row or each column of matrix elements is composed of a group of historical text vectors, and each matrix element in the relationship matrix represents the interaction relationship between any two rows or any two columns of matrix elements in the vector matrix; Step S400: Predict multiple current text vectors corresponding to the multiple groups of historical text vectors in the second set according to the graph structure, and convert the multiple current text vectors into multiple current evaluation texts; Step S500: Generate a book review of the target e-book according to the multiple current evaluation texts.

[0029] Regarding step S110, the keyword can be a word or sentence with the theme information in the target e-book. The theme information includes at least one of the title, abstract, and introduction.

[0030] Among them, the acquisition of theme information can be achieved in various ways. For example, based on natural language processing technology (NLP), information such as the title, abstract, and introduction can be extracted from the metadata of the e-book. In addition, machine learning algorithms can also be used to automatically identify and extract these theme information. Through these methods, it can be ensured that the obtained theme information is accurate and representative, providing a reliable basis for the subsequent generation of book reviews.

[0031] The role of the keyword is to find the evaluation set of the target e-book from the entire network.

[0032] The keyword can be the main title. The keyword can also be the combination of the titles of each chapter.

[0033] Historical evaluation data of different modalities can include at least two of text modality, image modality, and audio modality. This means that the evaluation data is not limited to text, but can also include various forms of information such as pictures and audio. For example, the text evaluations, picture evaluations uploaded by users on e-book reading platforms (such as QQ Reading), or the video evaluations uploaded on short video platforms can all be part of the evaluation data. By combining data of multiple modalities, the evaluation information of the target e-book at different times can be more comprehensively reflected, thereby improving the accuracy and objectivity of the book review.

[0034] These evaluation data can be from the same user or different users, and there is no limitation here. For example: (1) User A has uploaded evaluation data for the target e-book at multiple time nodes on Platform 1; (2) User A has uploaded evaluation data for the target e-book at multiple time nodes on Platform 1 and Platform 2 respectively; (3) User A and User B have uploaded evaluation data for the target e-book on Platform 1; (4) User A and User B have both uploaded evaluation data for the target e-book on Platform 1 and Platform 2.

[0035] The evaluation data here are all historical evaluation data. Since the evaluation of e-books changes over time, the main idea of this application is to use these historical evaluation data to predict the current evaluation data, so that the current evaluation can reflect the law among historical evaluations, generate a book review for the target e-book based on the current evaluation data, and further enable the book review of the target e-book to reflect the change of evaluation data of the e-book at different times, avoiding the problem of missing logical relationships caused by only searching based on keywords.

[0036] The following introduces the generation of the first set and the second set: Taking m platforms (such as qq Reading, Douyin, Xiaohongshu) as an example, combine the historical evaluation data of the m platforms into an initial data set (i.e., the first set).

[0037] For step S200, then process the first set: Convert non-text evaluation data into text evaluations. For example, use image recognition technology to identify text information in image evaluation data and audio conversion technology to convert audio evaluation data into text data.

[0038] Then use technologies such as Word2Vec and Doc2Vec to vectorize the text to obtain historical text vectors: .

[0039] Among them, are respectively the sets of historical text vectors converted from the data collected by the first platform, the second platform, and the mth platform, and so on.

[0040] Then define time t, and use a window of length a to sample : , and these data form the second set.

[0041] Among them: , , up to .

[0042] In this embodiment, it is necessary to predict the current text vector at the t+1 moment, that is, to obtain: .

[0043] Regarding step S300, the following introduces how to generate the graph structure: Set up a third prediction model, where the third prediction model is mainly used to generate the relationship matrix, and generate the graph structure based on the vector matrix and the relationship matrix. The nodes of the graph structure are represented by the vector matrix, and the edges are represented by the relationship matrix. As Figure 2 shown.

[0044] Among them, the vector matrix (each row consists of a set of historical text vectors) is: .

[0045] The content to be solved is a current text vector corresponding to each set of historical text vectors, that is: .

[0046] Regarding the relationship matrix, it will be introduced in subsequent embodiments; Regarding step S400, the steps include the following process: (1) Construct a first prediction model and a second prediction model, where the first prediction model includes any kind of graph neural network, and the second prediction model includes any kind of recurrent neural network; (2) Extract the graph embedding features obtained after embedding the information in the relationship matrix into the vector matrix from the graph structure according to the first prediction model; (3) Predict a current text vector corresponding to each set of historical text vectors from the graph embedding features according to the second prediction model.

[0047] After obtaining a graph structure in the previous step, this embodiment uses the first prediction model, and the first prediction model can be any kind of graph neural network. For example, it is a graph convolutional neural network (GCN). The graph convolutional neural network extracts the graph embedding features with the first interaction information in the vector matrix and the second interaction information in the relationship matrix by using graph convolution operations. This graph embedding feature contains the relationship between each set of historical text vectors in the time series, and the interaction relationship between multiple sets of historical text vectors, can fully mine the logical relationship between the vectors in the vector matrix, determine the law of vector change over time, and then obtain the latest evaluation data.

[0048] After obtaining the graph embedding features, in this embodiment, a second prediction model is used to predict the current text vector at the current moment. The second prediction model can be any kind of recurrent neural network, for example, a gated recurrent unit (GRU). The gated recurrent unit (GRU) predicts a current text vector corresponding to each group of historical text vectors from the graph embedding features.

[0049] Specifically, the combination of the first prediction model and the second prediction model enables the effective extraction of graph embedding features with the first interaction information in the vector matrix and the second interaction information in the relationship matrix from the graph structure. Through the recurrent neural network, a current text vector corresponding to each group of historical text vectors is predicted from the graph embedding features. When processing the graph structure, the graph neural network can capture the complex relationships between nodes, while the recurrent neural network is good at processing time series data and can extract useful information from the graph embedding features.

[0050] It should be noted that since the graph neural network and the recurrent neural network are common tools in this field and their tools can be obtained on the Internet, they will not be elaborated here.

[0051] Before introducing the training processes of the first prediction model and the second prediction model, the generation of the relationship matrix by the third prediction model is introduced: The relationship matrix can reflect the interaction relationship between every two rows or every two columns of vector elements in the vector matrix. Therefore, the generation of the relationship matrix needs to be obtained in combination with the training processes of the first prediction model and the second prediction model.

[0052] The third prediction model consists of a fully connected layer and a convolutional model. First, the process of the third prediction model generating the initial relationship matrix includes: (1) Using the fully connected layer to map from Gaussian noise to a feature representation; (2) Using the convolutional model to generate the relationship matrix with dimensions.

[0053] The dimension is the same as the above m groups of historical evaluation data.

[0054] Then, the training processes of the first prediction model and the second prediction model are as follows: (1) Construct an evaluation model and construct multiple groups of historical text vectors to be trained; (2) Generate a to-be-trained graph structure consisting of a to-be-trained vector matrix and a to-be-trained relationship matrix according to the historical text vectors to be trained; the to-be-trained vector matrix is generated from multiple groups of historical text vectors to be trained, and the elements in the to-be-trained relationship matrix are generated from the interaction relationships between every two matrix elements in the to-be-trained vector matrix; (3) Input the graph structure to be trained into the first prediction model to obtain the graph embedding features to be trained output by the first prediction model; (4) Input the graph embedding features to be trained into the second prediction model to obtain multiple current text vectors to be trained corresponding to multiple historical text vectors to be trained predicted by the second prediction model; (5) Construct positive samples and negative samples; among them, the negative samples are obtained by splicing the historical text vectors to be trained and the current text vectors to be trained, and the positive samples are obtained by splicing the historical text vectors to be trained and the preset correct labels; the correct labels are set in advance; (6) Input the positive samples and negative samples into the evaluation model to distinguish the probabilities between the positive samples and negative samples according to the evaluation model; (7) According to the probability and the preset evaluation function, perform backpropagation on the second prediction model and the first prediction model until the training processes of the second prediction model and the first prediction model are completed.

[0055] The process of forward propagation includes: First, the third prediction model generates an initial graph structure (consisting of a vector matrix and an initial relationship matrix); then the first prediction model generates graph embedding features, and then the second prediction model generates the current text vector; finally, the evaluation model generates gradient parameters based on the output probability and the evaluation function. Among them, the evaluation model can be based on a min-max evaluation function that can be set to depth-first, and according to this function, the result output by the third prediction model can approximate the positive sample.

[0056] The process of backpropagation includes: Update the first prediction model, the second prediction model, and the third prediction model according to the gradient parameters. It can make the third prediction model generate a more accurate relationship matrix.

[0057] In some embodiments, since the negative samples are obtained by splicing the historical text vectors to be trained and the current text vectors to be trained, and the positive samples are obtained by splicing the historical text vectors to be trained and the preset correct labels, the positive and negative samples are also a time series. Therefore, it is set that the evaluation model can be composed of a long short-term memory network (LSTM) and a fully connected layer. The input data of the long short-term memory network (LSTM) is the positive and negative samples, and the output features can be used as the input data of the fully connected layer. Finally, the fully connected layer judges the probability.

[0058] After the training is completed, the first prediction model, the second prediction model, and the third prediction model can be used to generate the current evaluation text. It should be noted that the evaluation model does not participate in the generation process of the current vector during actual use.

[0059] Regarding step S500, in the case of generating multiple current text vectors as described above, complex interaction relationships between multiple groups of historical text vectors and the internal temporal relationships within the historical text vectors are captured among the multiple current text vectors. The multiple current evaluation texts generated from the multiple current text vectors can reflect the historical patterns of the evaluation of the target e-book.

[0060] Finally, high-real-time and high-accuracy book reviews can be generated based on the multiple current evaluation texts.

[0061] The beneficial effects of this application include: By comprehensively utilizing the multi-modal historical evaluation data of the target e-book, this method can more accurately reflect the logical relationships between evaluations in different periods, thereby generating more real-time and accurate book reviews. Specifically, each group of historical text vectors among multiple groups of historical text vectors is used as an element of a row-column matrix to form a vector matrix, the relationship matrix is formed by the interaction relationships between every two groups of historical text vectors, and then a graph structure including the vector matrix and the relationship matrix is constructed, and multiple current text vectors corresponding to multiple groups of historical text vectors are extracted from the graph structure. This embodiment can fully exploit the relationships in the time series within each group of historical text vectors and the interaction relationships between multiple groups of historical text vectors, can better reflect the changes in the evaluation data of the e-book in different periods, avoids the problem of missing logical relationships caused by the changes in the evaluation data in different periods due to only keyword search, can effectively predict the current evaluation data, and improves the real-time performance and accuracy of book review generation.

[0062] In some embodiments of this application, the Doc2Vec technology is used to extract the corresponding historical text vectors from each group of historical evaluation texts. Doc2Vec is an unsupervised learning algorithm that captures the semantic information of a document by mapping the entire document into a vector space of a fixed length. In the specific implementation process, first, multiple groups of historical evaluation texts need to be preprocessed, including operations such as removing stop words and word segmentation. Then, the preprocessed text data is used to train the Doc2Vec model, and the vector representation of each group of historical evaluation texts is generated through the model. In this way, multiple groups of historical evaluation texts can be converted into multiple groups of historical text vectors.

[0063] The application of the Doc2Vec technology can effectively convert text data into vector representations, which is convenient for subsequent calculation and processing. Secondly, the text vectors generated by Doc2Vec can better retain the semantic information of the text, which helps to improve the accuracy and quality of book review generation. Finally, the unsupervised learning characteristic of Doc2Vec enables it to process a large amount of historical evaluation text data, and has high applicability and scalability.

[0064] According to multiple current evaluation texts in step S500, generate a book review for the target e-book, including the following steps: Method 1: First, judge the similarity between multiple current evaluation texts and keywords.

[0065] Then, based on a similarity threshold, select several current evaluation texts from multiple current evaluation texts that are greater than the similarity threshold.

[0066] Finally, generate a book review for the target e-book based on several current evaluation texts. For example, after combining the current evaluation texts, a book review for the target e-book can be obtained.

[0067] Method 2: Use an evolutionary algorithm to generate a book review.

[0068] Different from the abstract, a book review is usually used as a recommendation. Therefore, a book review often contains the interest features required by the author or promoter to meet the personalized requirements.

[0069] First, define the interest information. The interest information needs to include the above keywords and words or short sentences customized by the author or promoter based on their own interests.

[0070] Then, from multiple current evaluation texts, judge multiple current evaluation texts with more interest information. For example, it can be judged by calculating the proportion of the interest information or calculating the total amount.

[0071] Then, generate a book review through an evolutionary algorithm. The process of the evolutionary algorithm is as follows: (1) Set the maximum number of evolutionary generations as T, set the initial population Y(0) to have N individuals, and an individual is a queue for generating a book review. Set a set, and the set includes multiple current evaluation texts. A new population is generated in each iteration.

[0072] (2) Individual evaluation: Calculate the fitness of each individual in the population P(i). The fitness value here can be set as the number of interest information.

[0073] (3) Selection, crossover, and mutation. Apply the mutation operator to the population.

[0074] (4) After the population P(i) is iterated through selection, crossover, and mutation, the next generation population P(i + 1) is obtained, and then the iteration process is repeated.

[0075] (5) Termination condition judgment: If t = T, then output the individual with the maximum fitness obtained during the evolutionary process as the optimal solution.

[0076] Interest information is set through Method 2. While retaining keywords, authors or promoters can add words or short phrases based on their own interests, enabling the selection of several current evaluation texts with more interest features. Subsequently, an evolutionary algorithm can be used to generate more excellent book reviews, enhancing the accuracy and adaptability of book review generation.

[0077] For example Figure 4 , an embodiment of the present application provides a book review generation system based on multi-modal data. The system includes: An evaluation extraction unit 1100 is configured to search for a first set of target e-books based on keywords; the first set includes multiple groups of historical evaluation data of the target e-book, where any two groups of historical evaluation data belong to different modalities of historical evaluation data, and each group of historical evaluation data includes evaluation data collected at multiple historical moments; A text vector extraction unit 1200 is configured to convert multiple groups of evaluation data in the first set into multiple groups of historical evaluation texts, extract multiple groups of historical text vectors from the multiple groups of historical evaluation texts, form a second set with the multiple groups of historical text vectors, and sample the multiple groups of historical text vectors in the second set according to a time window so that each group of historical text vectors contains multiple historical text vectors; A matrix generation unit 1300 is configured to generate a graph structure composed of a vector matrix and a relationship matrix based on the multiple groups of historical text vectors in the second set; wherein, each matrix element in the vector matrix is composed of a historical text vector in the second set, and each row or each column of matrix elements is composed of a group of historical text vectors, and each matrix element in the relationship matrix represents the interaction relationship between any two rows or any two columns of matrix elements in the vector matrix; A text vector extraction unit 1400 is configured to predict multiple current text vectors corresponding to the multiple groups of historical text vectors in the second set according to the graph structure, and convert the multiple current text vectors into multiple current evaluation texts; A book review generation unit 1500 is configured to generate a book review of the target e-book based on the multiple current evaluation texts.

[0078] It should be noted that the book review generation system based on multi-modal data provided in this embodiment and the above-mentioned book review generation method based on multi-modal data are based on the same inventive concept. Therefore, the relevant content of the above-mentioned book review generation method based on multi-modal data also applies to the content of the book review generation system based on multi-modal data. Therefore, it will not be elaborated here.

[0079] For example Figure 5 , an embodiment of the present application further provides an electronic device. This electronic device includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the book review generation method based on multimodal data described above in the present disclosure.

[0080] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0081] The electronic device of the embodiment of the present application will be introduced in detail below.

[0082] The processor 1600 can be implemented by means of a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention; The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present specification through software or firmware, the relevant program codes are stored in the memory 1700, and the processor 1600 is called to execute the book review generation method based on multimodal data of the embodiments of the present invention.

[0083] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 realize communication connections with each other inside the device through the bus 2000.

[0084] An embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned book review generation method based on multimodal data.

[0085] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices.

[0086] In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0087] The embodiments described in the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0088] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0090] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0091] In the description of the present application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0092] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0093] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0094] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0095] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0097] The above is a specific description of the preferred implementation of the embodiments of the present application. However, the embodiments of the present application are not limited to the above-mentioned implementation manners. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the embodiments of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the embodiments of the present application.

Claims

1. A book review generation method based on multimodal data, characterized in that The method includes: Based on keywords, searching for a first set of data regarding the target e - book; the first set includes multiple groups of historical evaluation data of the target e - book, where any two groups of historical evaluation data belong to different modalities of historical evaluation data, and each group of historical evaluation data includes evaluation data collected at multiple historical moments; Converting the multiple groups of evaluation data in the first set into multiple groups of historical evaluation texts, extracting multiple groups of historical text vectors from the multiple groups of historical evaluation texts, forming a second set with the multiple groups of historical text vectors, sampling the multiple groups of historical text vectors in the second set according to a time window, and each sampled group of historical text vectors contains multiple historical text vectors; Generating a graph structure composed of a vector matrix and a relationship matrix according to the multiple groups of historical text vectors in the second set; wherein, each matrix element in the vector matrix is composed of a historical text vector in the second set, and each row or each column of matrix elements is composed of a group of historical text vectors, and each matrix element in the relationship matrix represents the interaction relationship between any two rows or any two columns of matrix elements in the vector matrix; Predicting multiple current text vectors corresponding to the multiple groups of historical text vectors in the second set according to the graph structure, and converting the multiple current text vectors into multiple current evaluation texts; Generating a book review of the target e - book according to the multiple current evaluation texts.

2. The book review generation method based on multimodal data according to claim 1, wherein The predicting the multiple current text vectors corresponding to the multiple groups of historical text vectors according to the graph structure includes: Constructing a first prediction model and a second prediction model, where the first prediction model includes any kind of graph neural network, and the second prediction model includes any kind of recurrent neural network; Extracting, from the graph structure according to the first prediction model, graph embedding features obtained after embedding the information in the relationship matrix into the information in the vector matrix; Predicting, according to the second prediction model, a current text vector corresponding to each group of historical text vectors from the graph embedding features.

3. The book review generation method based on multimodal data according to claim 2, wherein The training process of the first prediction model and the second prediction model includes the following steps: Constructing an evaluation model and constructing multiple groups of historical text vectors to be trained; Generating a to - be - trained graph structure composed of a to - be - trained vector matrix and a to - be - trained relationship matrix according to the historical text vectors to be trained; The to - be - trained vector matrix is generated from the multiple groups of historical text vectors to be trained, and the elements in the to - be - trained relationship matrix are generated from the interaction relationships between every two matrix elements in the to - be - trained vector matrix; Inputting the to - be - trained graph structure into the first prediction model to obtain to - be - trained graph embedding features output by the first prediction model; Inputting the to - be - trained graph embedding features into the second prediction model to obtain multiple to - be - trained current text vectors predicted by the second prediction model corresponding to multiple groups of historical text vectors to be trained; Construct positive samples and negative samples; wherein, the negative samples are obtained by concatenating the historical text vectors to be trained and the current text vectors to be trained, and the positive samples are obtained by concatenating the historical text vectors to be trained and the preset correct labels; Input the positive samples and the negative samples into an evaluation model to determine the probability between the positive samples and the negative samples according to the evaluation model; According to the probability and a preset evaluation function, perform backpropagation on the first prediction model and the second prediction model until the training processes of the first prediction model and the second prediction model are completed.

4. The book review generation method based on multimodal data according to claim 1, wherein The keyword is a sentence with the theme information in the target e-book.

5. The book review generation method based on multi-modal data according to claim 4, wherein The theme information includes at least one of a title, an abstract, and a brief introduction.

6. The book review generation method based on multimodal data according to claim 1, wherein The extracting a plurality of sets of historical text vectors from the plurality of sets of historical evaluation texts includes: Extracting corresponding historical text vectors from each set of historical evaluation texts based on Doc2Vec.

7. The book review generation method based on multimodal data according to claim 1, wherein The different modalities include at least two of a text modality, an image modality, and an audio modality.

8. A book review generation system based on multimodal data, characterized in that, The system includes: An evaluation extraction unit for finding a first set of the target e-book based on a keyword; the first set includes a plurality of sets of historical evaluation data of the target e-book, where any two sets of historical evaluation data belong to historical evaluation data of different modalities, and each set of historical evaluation data includes evaluation data collected at multiple historical moments; A text vector extraction unit for converting the plurality of sets of evaluation data in the first set into a plurality of sets of historical evaluation texts, extracting a plurality of sets of historical text vectors from the plurality of sets of historical evaluation texts, forming a second set with the plurality of sets of historical text vectors, and sampling the plurality of sets of historical text vectors in the second set according to a time window, and each set of the sampled historical text vectors includes a plurality of historical text vectors; A matrix generation unit for generating a graph structure composed of a vector matrix and a relationship matrix according to the plurality of sets of historical text vectors in the second set; wherein, each matrix element in the vector matrix is composed of a historical text vector in the second set, and each row or each column of matrix elements is composed of a set of historical text vectors, and each matrix element in the relationship matrix represents the interaction relationship between any two rows or any two columns of matrix elements in the vector matrix; A text vector extraction unit for predicting a plurality of current text vectors corresponding to the plurality of sets of historical text vectors in the second set according to the graph structure, and converting the plurality of current text vectors into a plurality of current evaluation texts; A book review generation unit for generating a book review of the target e-book according to the plurality of current evaluation texts.

9. An electronic device, characterized in that, Includes: At least one controller and a memory for communicatively connecting with the at least one controller; The memory stores instructions executable by the at least one controller, and the instructions are executed by the controller to enable the controller to execute the multi-modal data-based book review generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the method for generating a book review based on multimodal data according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image generation method for keeping text content position relation

    CN115880398A

  • Collaborative filtering recommendation method, system and equipment based on deep learning and medium

    CN116226543A

  • Method, system and equipment for generating text abstract and storage medium

    CN117520535A

  • Method and device for predicting user evaluation information and electronic equipment

    CN117853175A

  • Text labeling method, system and equipment and storage medium

    CN117952068A