A method, an electronic device, and a storage medium for generating an abstract in the field of digital culture
By using vertical domain model and knowledge graph expansion technology in the field of digital culture, the problems of low efficiency and poor consistency in data abstract generation in the existing technology are solved, and efficient and accurate data description and value mining are achieved.
Patent Information
- Application Number
- CN202411737108.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-11-29
AI Technical Summary
In the prior art, data abstract generation methods in the field of digital culture are inefficient and have poor description consistency, and cannot deeply explore data value, and traditional tools cannot conduct in-depth analysis based on the semantics and context of the data.
A vertical domain model based on the base large model is adopted, combining multimodal data processing and digital cultural domain knowledge graphs, and high-quality abstract text is generated through multi-level recursive query and knowledge graph expansion.
It realizes efficient and automated data description, improves data understanding and usage effect, ensures the uniformity and flexibility of description, and can identify potential data associations and provide related backgrounds and application scenarios.
Smart Images

Figure CN119669458B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a method for generating abstracts in the digital culture field, an electronic device, and a storage medium. Background Art
[0002] In the digital culture industry and other industries, the data uploaded by users is usually raw data that has not been deeply processed, often lacking detailed descriptions, background information, and context associations. These raw data are usually difficult to directly use during analysis, display, and storage, and the potential value of the data cannot be fully explored. Currently, manual abstracts or template-based methods are generally used. For such abstract generation methods, the main problems faced by users are as follows:
[0003] (1) Manual operation is cumbersome: Manually expanding descriptions takes a long time. Especially when dealing with large-scale data sets, the efficiency is low and a large amount of human resources need to be invested.
[0004] (2) Limitations of templated tools: Although templated tools improve partial efficiency, the descriptions generated by them are often too single in content, unable to provide targeted descriptions according to the specific content of the data, lacking flexibility and depth.
[0005] (3) Poor description consistency: Due to the limitations of manual operations or templates, different batches of data may obtain inconsistent descriptions, affecting the accuracy during subsequent analysis and display.
[0006] (4) Lack of intelligence: Traditional tools cannot perform in-depth analysis based on the semantics, background, and context of the data, resulting in limited depth and breadth of the description information. Therefore, expanding the description of the data has become the key to solving this problem. By using AI (artificial intelligence) technology to expand the description of the data uploaded by users, descriptive content with rich semantics and detailed background can be automatically generated to help users better understand and use this data.
[0007] It can be seen that currently, the digital culture field needs an AI system that can automatically expand the data uploaded by users. This system should be able to generate comprehensive descriptions based on the original data, including the background, characteristics, possible uses, and associated information of the data. Users expect to reduce the time and workload of manual editing through this method, and improve the comprehensibility, usability, and display effect of the data. Summary of the Invention
[0008] In order to overcome the above defects, this application is proposed, that is, an intelligent abstract generation method that can perform multi-level recursive queries based on the folder hierarchy, file content, and tags to achieve efficient and accurate data retrieval. Users can quickly find the required data in a complex file structure through simple query operations, and can perform detailed classification and intelligent retrieval in combination with the file content and tag information.
[0009] Specifically, the first aspect of the present application provides a method for generating abstracts in the field of digital culture, including:
[0010] Obtain vertical domain multimodal data for which an abstract is to be generated;
[0011] Input the multimodal data into a vertical domain model to generate a first cross-modal label set corresponding to each of the multimodal data, where the vertical domain model is obtained based on a base large model;
[0012] Associate the first cross-modal label set with a knowledge graph in the field of digital culture;
[0013] Based on the knowledge graph in the field of digital culture, search for concept entities associated with the first cross-modal label set to form a set of concept entities;
[0014] According to a specified threshold and granularity, based on the knowledge graph in the field of digital culture, expand the set of concept entities to obtain an expanded set of entities, and update the knowledge graph in the field of digital culture based on the expanded set of entities;
[0015] Obtain a second cross-modal label set based on the vertical domain multimodal data, compare the expanded set of entities with the second cross-modal label set, and sort the expanded set of entities to obtain a sorted set of entities;
[0016] Generate an abstract text based on the sorted set of entities and the updated knowledge graph in the field of digital culture.
[0017] In one embodiment, obtaining the vertical domain model based on the base large model includes:
[0018] Transform the base large model to form a target large model, where the target large model sets a text encoder, an image encoder, an audio encoder, and a video encoder at the original input layer of the base large model to extract features from corresponding types of samples; set an input projector for projecting the extracted features into a common feature space; borrow the interaction between the self-attention mechanism and the feed-forward network of the base large model itself to achieve feature fusion; set a classifier after the original output layer of the base large model for generating cross-modal labels; set a knowledge extractor after the classifier for extracting entities and relationships from the labels;
[0019] Pre-train the target large model to obtain a pre-trained target large model;
[0020] Fine-tune and train the pre-trained target large model to obtain the vertical domain model.
[0021] In one embodiment, based on the digital cultural domain knowledge graph, concept entities associated with the first cross-modal tag set are found to form a concept entity set:
[0022] The concept entities are obtained by a knowledge extractor performing concept entity recognition on the tags associated with the first cross-modal tag set;
[0023] Whether the concept entities are similar to the entities in the digital cultural domain knowledge graph is determined through similarity comparison;
[0024] In the case where the concept entities are similar to the entities in the digital cultural domain knowledge graph, the concept entities are incorporated into the concept entity set.
[0025] In one embodiment, according to a specified threshold and granularity, based on the digital cultural domain knowledge graph, the concept entity set is expanded to obtain an expanded entity set, and the digital cultural domain knowledge graph is updated based on the expanded entity set, including:
[0026] During the process of forming the concept entity set, if the similarity of the concept entities to the entities in the digital cultural domain knowledge graph meets the specified threshold and granularity, the concept entities are incorporated into the concept entity set to obtain the expanded entity set.
[0027] In one embodiment, a second cross-modal tag set is obtained based on the vertical domain multi-modal data, the expanded entity set is compared with the second cross-modal tag set, and the expanded entity set is sorted to obtain a sorted entity set, including:
[0028] The vertical multi-modal data is sampled to obtain a second cross-modal tag set;
[0029] For each entity in the expanded entity set, the weight of the entity in the second cross-modal tag set is calculated;
[0030] The expanded entity set is sorted according to the frequency of occurrence of the entity in the second cross-modal tag set to obtain a sorted entity set.
[0031] In one embodiment, based on the sorted entity set and the updated digital cultural domain knowledge graph, summary text is generated, including:
[0032] According to the sorting in the sorted entity set, each entity in the sorted entity set is expanded according to a specified summary expansion granularity to obtain an entity expansion sequence;
[0033] For each entity in the entity expansion sequence, extract the corresponding relationships in the digital cultural domain knowledge graph, and generate summary text in combination with the entity expansion sequence.
[0034] In one embodiment, the method further includes:
[0035] Perform machine translation on the summary text to obtain a summary translation text in a specified language.
[0036] The second aspect of the present application provides a query method in the digital cultural domain, which is characterized by including:
[0037] Present a query entry on the human-computer interaction interface;
[0038] In response to a user inputting a multimodal data query request instruction at the query entry, obtain a query result;
[0039] In response to the user clicking on the query result, obtain a summary corresponding to the query result according to the method described in any one of the first aspect;
[0040] Present the summary on the human-computer interaction interface in the form of a thumbnail.
[0041] The third aspect of the present application provides an electronic device, including at least one processor and a memory. Among them, the memory stores computer execution instructions, and the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the digital cultural domain summary generation method described in any one of the first aspect, or the query method described in the second aspect.
[0042] The fourth aspect of the present application provides a computer-readable storage medium, in which multiple program codes are stored. It is characterized in that the program codes are suitable for being loaded and run by a processor to execute the digital cultural domain summary generation method described in any one of the first aspect, or the query method described in the second aspect.
[0043] It can be seen that compared with the prior art, the present solution has the following beneficial effects:
[0044] (1) Efficient and automated processing: Through the artificial intelligence large model, the system can automatically perform extended descriptions on the data uploaded by users, reducing the time and cost of manual operations, and is particularly suitable for processing large-scale data sets.
[0045] (2) Multimodal data support: The present solution can process different types of data (text, image, audio, video, etc.), and generate customized extended descriptions according to the specific data type, breaking through the limitations of traditional templated tools.
[0046] (3) Semantic enrichment and context association: Through semantic analysis and background expansion, the system can generate in-depth extended descriptions to help users obtain more valuable information from the data, improving the understanding and utilization of the data.
[0047] (4) Consistency and flexibility: The descriptions automatically generated by the system ensure the unity of the descriptions, avoiding potential inconsistencies that may occur in manual operations. In addition, the system supports users to customize the level of detail of the descriptions, enhancing flexibility.
[0048] (5) Intelligent association and value mining: Through the knowledge graph, the system can identify potential associations between the uploaded data and other data, and provide relevant backgrounds and application scenarios, enabling greater exploration of the value of the data uploaded by users.
[0049] Through this technology, users can quickly generate high-quality extended descriptions, reduce the complexity of data processing, and improve the efficiency and effectiveness of data presentation and utilization. Brief Description of the Drawings
[0050] Referring to the drawings, the disclosure of the present application will become more understandable. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present application. Among them:
[0051] Figure 1 Shows a flowchart of a method for generating abstracts in the digital culture field according to an embodiment of the present invention.
[0052] Figure 2 Shows a schematic diagram of a knowledge graph in the digital culture field according to an embodiment of the present invention.
[0053] Figure 3 Shows a data flow diagram according to an embodiment of the present application.
[0054] Figure 4 Shows a schematic diagram of grain size selection according to an embodiment of the present patent. Detailed Description of the Preferred Embodiments
[0055] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention and are not intended to limit the invention. In addition, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0056] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.
[0057] Some terms related to the present application will be explained here first.
[0058] Large language model: It refers to a model trained using deep learning technology that can process large-scale data.
[0059] Multimodal large language model: It can simultaneously process and understand various different modalities of data. A multimodal large language model can jointly model and predict using multiple modalities of data, thereby more comprehensively understanding and processing the input information.
[0060] Vertical domain: It refers to the requirements and characteristics specific to a particular industry or application field. These requirements are usually industry-specific, such as in the medical, financial, telecommunications, and other fields. Software architecture solutions in the vertical domain are typically designed to meet the business processes and regulatory requirements of a specific industry.
[0061] Base large model: It refers to a general large language model that has not undergone additional instruction fine-tuning or downstream task adaptation fine-tuning.
[0062] Vertical domain large model: It refers to a "specialized model" in a specific domain obtained by fine-tuning or adapting and fine-tuning the base large model using data samples from a specific domain.
[0063] Knowledge graph: It refers to a technical method that uses a graph model to describe knowledge and model the association relationships between all things in the world, consisting of nodes and edges. Among them, nodes represent entities, concepts, or attribute values, edges represent relationships, and entities and relationships form triples.
[0064] Knowledge graph in the digital culture field: It refers to a knowledge graph extracted from data of different file types such as text records, digital images, audio, and video in the digital culture field. Through the integration of corresponding entities, relationships, and attributes, a professional domain knowledge graph for different types of files is obtained.
[0065] Figure 2 Shows a schematic diagram of a knowledge graph in the digital culture field of an embodiment.
[0066] Such as Figure 2 As shown, Hongloumeng 201, Cao Xueqin 202, Lin Daiyu 203, the Central Ballet Troupe 204, and the Tianqiao Theatre 205 in the figure are all entities. These entities can have multiple attributes. For example, the attributes of Cao Xueqin 202 can include being from the Qing Dynasty and being a novelist, etc. The relationships between entities are as Figure 2 Shown, Cao Xueqin 202 wrote Hongloumeng 201, the Central Ballet Troupe 204 performed Hongloumeng 201 at the Tianqiao Theatre 205, and the protagonist of Hongloumeng 201 is Lin Daiyu 203, etc.
[0067] It should be noted that the digital culture field is different from general fields and has numerous proprietary terms or connotative relationships. Therefore, in the process of constructing a knowledge graph for the digital culture field, the dictionary file of this field can be used to enrich the comprehensiveness of the knowledge graph for the digital culture field, thereby improving the accuracy and efficiency of feature recognition and file search.
[0068] Specifically, the knowledge graph for the digital culture field can be a well - constructed proprietary domain knowledge graph, an incomplete proprietary domain knowledge graph, or even a brand - new proprietary domain knowledge graph. Since the concepts of a field, or rather, entities and terms, also develop and change with the times, the knowledge graph for the digital culture field cannot be fixed either and needs to be continuously enriched and updated. Therefore, when identifying tags for different types of data files, it is very likely that the identified tags cannot find corresponding entities in the knowledge graph for the digital culture field. At this time, the entities, attributes, and relationships corresponding to the new tags can be expanded into the knowledge graph for the digital culture field.
[0069] Generally speaking, the main innovation point of the present invention compared with the prior art is that in the current prior art, generally, the concept entities in the file are directly extracted through an artificial intelligence model, and the abstract text is generated through machine learning. During the machine - learning process, some concept entities may be omitted. To make up for this defect, in the present invention, after extracting the concept entities in the file, these extracted concept entities will be compared with the original file (or a sample of the original file) once to expand the concept entities, thereby improving the accuracy and integrity of the abstract generation.
[0070] In a first aspect, the present invention designs a method for generating an abstract in the digital culture field.
[0071] Figure 1 The flowchart showing the method for generating an abstract in the digital culture field according to an embodiment of the present invention is as follows. In one embodiment, as Figure 1 shown, the method for generating an abstract in the digital culture field includes:
[0072] S101: Obtain the vertical - domain multimodal data for which the abstract is to be generated.
[0073] Specifically, in the vertical - domain multimodal data for which the abstract is to be generated, the vertical domain, as described above, can be the requirements and characteristics of a specific industry or application field. For example, the digital culture field can form a vertical domain. Multimodal data refers to data with multiple modalities. For example, it can have one modality among text, image, audio, and video, or a combination of multiple modalities of data. It can be understood that the multimodal data obtained by the user in S101 can be only one of the above - mentioned modalities or a combination of multiple modalities of data.
[0074] S102: Input the multi-modal data into the vertical domain model to generate a cross-modal label set corresponding to each of the multi-modal data, where the vertical domain model is obtained based on a base large model.
[0075] In one embodiment, obtaining the vertical domain model based on a base large model includes:
[0076] Transform the base large model to form a target large model, where the target large model sets a text encoder, an image encoder, an audio encoder, and a video encoder at the original input layer of the base large model to extract features of corresponding types of samples; set an input projector for projecting the extracted features into a common feature space; borrow the interaction between the self-attention mechanism and the feed-forward network of the base large model itself to achieve feature fusion; set a classifier after the original output layer of the base large model for generating cross-modal labels; set a knowledge extractor after the classifier for extracting entities and relationships from the labels;
[0077] Pre-train the target large model to obtain a pre-trained target large model;
[0078] Fine-tune and train the pre-trained target large model to obtain the vertical domain model.
[0079] The following introduces the generation process of the vertical domain model:
[0080] For the classification of multi-modal data, existing general large language models are not satisfactory. However, large language models have strong semantic understanding capabilities. Therefore, the inventor chose to perform secondary development on an open-source base large model. For example, ChatGLM was used as the base large model. ChatGLM is an open language model based on the general language model GLM framework.
[0081] To meet the requirement of facilitating the classification and query of one's own multi-modal data, the present application constructs a new network architecture on the basis of the selected base large model. Refer to the appendix Figure 3 , Figure 3 is the network architecture according to an embodiment of the present application. In the dashed box 2 is the new network topology constructed in the present technical solution on the basis of the base large model.
[0082] The ChatGLM model is designed based on the Transformer architecture and consists of an encoder and a decoder part. ChatGLM is mainly used for processing natural language tasks, and its core advantage lies in understanding and generating text, rather than directly processing multimodal data. Therefore, to handle multimodal feature extraction, as shown in the figure, an additional encoder module is added to the 30th layer of the original input layer of the ChatGLM model, including a text encoder 31, an image encoder 32, an audio encoder 33, and a video encoder 34.
[0083] Among them, the text encoder 31, such as the BERT model, is used to extract the semantic features of text data; the image encoder 32, such as the CNN model, is used to extract the visual features of picture data; the audio encoder 32, such as Wave2Vec, is used to extract the audio features of audio data, such as Mel Frequency Cepstral Coefficients (MFCC); the video encoder 34, such as Two-Stream Networks, is used to extract the visual and motion features of video data.
[0084] The features of each modality extracted need to be projected into a common feature space for effective fusion. To this end, an input projector 35 is added to map the features of different modalities into a representation aligned with the text feature space, which can be achieved through a linear layer and an MLP.
[0085] For feature fusion, it is achieved by borrowing the interaction between the self-attention mechanism and the feed-forward network of the ChatGLM model itself, which is denoted by the label 36 in the figure. The self-attention mechanism allows the model to dynamically assign different weights between different words, thereby achieving feature fusion. The feed-forward network further performs a non-linear transformation on these fused features to enhance the expressive ability of the model.
[0086] The output layer 37 of the ChatGLM model generates a log probability distribution, which is a vector of raw scores for all possible classes. A classifier 38 is added after the output layer 37, for example, applying the softmax function to the output layer. The softmax function converts these raw scores into a probability distribution such that the sum of the probabilities of all classes is 1. In this way, each class has a predicted probability between 0 and 1. According to the output of the classifier 38, the class with the highest probability is selected as the final classification result. This step is usually achieved through the argmax function, which returns the index of the class with the highest probability.
[0087] So far, semantic tags have been assigned to the data of each modality respectively, and such tags have the property of "cross-modality". This is because there are encoders for the data of each modality and then feature fusion is performed, which comprehensively considers the data features of different modalities, thus making the generated cross-modal tags more accurate. That is to say, in this process, the data features of different modalities such as text, audio, pictures, and videos are integrated into a unified feature space, influencing each other and jointly determining the final cross-modal tags.
[0088] For example, when generating cross-modal tags for text, the constructed model will comprehensively consider the features of text and the features of audio, pictures, and videos. This can make the cross-modal tags of text not only contain the information of the text itself but also reflect the relevant information of multi-modal data such as audio, pictures, and videos. Similarly, for the generation of cross-modal tags for audio and videos, they will also be affected by the data features of other modalities.
[0089] For the tags of the data of each modality, their granularity and form may be different. For example, for the text modality, the generated cross-modal tags are usually regarded as one or more words because the main function of text is to express meaning, and usually, a concise word or phrase can convey information. For the audio or video modality, the generated cross-modal tags may tend to be at the phrase or sentence level because the data of these modalities usually contain a large number of audio frames or video frames, so longer sequences can be extracted from them, and then the corresponding cross-modal tags can be generated. Of course, this is not absolute. In some scenarios, such as when classifying complex audio or video content, longer or more complex tag sequences may be generated. At the same time, for the data of various modalities, the granularity and form of the tags can also be adjusted according to actual needs and application scenarios.
[0090] As Figure 3 shown, for the convenience of query, this application hopes to associate the cross-modal tags of each multi-modal data with the knowledge graph 5. In the ChatGLM network architecture, the ChatGLM model is a unidirectional language model mainly used to predict the next word, rather than performing entity recognition and relationship extraction on the given tags. Therefore, other techniques and methods need to be combined to achieve this. For this purpose, a knowledge extractor 39 is introduced into the Figure 3 network architecture.
[0091] In one example, an LSTM (Long Short-Term Memory network) is used as the knowledge extractor 39. The LSTM is a common sequence model that can be used to process sequence data. Since what needs to be processed are labels (words or sentences, essentially a text sequence), the LSTM is suitable. The LSTM model can be used to model the entity recognition and relationship extraction problems in the text sequence. Specifically, we can regard the text sequence as a sequence data and use the LSTM to capture the long-term dependencies in the sequence.
[0092] In the ChatGLM network architecture, we do not need to use the LSTM at every position, but can use it only at specific positions of interest. That is to say, outside the output layer of the ChatGLM model, an additional LSTM module can be added to model specific entity recognition and relationship extraction tasks. This module can share a part of the ChatGLM model, thereby reducing the number of model parameters and computational complexity.
[0093] Next, the newly constructed network architecture needs to be trained to finally obtain the vertical domain large model.
[0094] It should be noted that during the process of training the model using the samples in the digital culture field, the meanings of some labels have different semantics in different contexts. For example, the semantics of the generated label "Xichun" includes "cherish spring" and the name of a person. This may cause inaccurate return results when conducting retrieval queries in the future. Therefore, the ChatGLM model of this application can also output the semantics corresponding to the labels with multiple semantics. Furthermore, this application forms a multi-semantic mapping relationship between the multi-semantic labels and the corresponding multiple semantics and stores them in the multi-semantic mapping database 4, as Figure 3 shown. At the same time, mark the knowledge graph entities corresponding to the labels with multiple semantics.
[0095] S103: Associate the first cross-modal label set with the digital culture field knowledge graph.
[0096] Specifically, the digital culture field knowledge graph in the initial state can be a blank knowledge graph, or a knowledge graph under construction, or a relatively complete knowledge graph. That is to say, as data is continuously input, the digital culture field knowledge graph will become more complete.
[0097] After the labels in the first cross-modal label set are extracted, these labels in the set will be enriched into the digital culture field knowledge graph, making the entities, attributes, and relationships in the digital culture field knowledge graph more complete.
[0098] S104: Based on the digital culture field knowledge graph, search for concept entities associated with the first cross-modal tag set to form a concept entity set.
[0099] In one embodiment, based on the digital culture domain knowledge graph, searching for concept entities associated with the first cross-modal tag set to form a concept entity set includes:
[0100] Performing conceptual entity recognition on the tags associated with the first cross-modal tag set by a knowledge extractor to obtain the conceptual entity;
[0101] By comparing similarities, determining whether the conceptual entity is similar to an entity in the knowledge graph in the digital culture field;
[0102] In the case where the conceptual entity is similar to an entity in the digital culture field knowledge graph, the conceptual entity is included in the conceptual entity set.
[0103] Combination Figure 3 It can be seen that for each label in the first cross-modal label set, the concept entity corresponding to each label is identified by the knowledge extractor 39. For example, for a label "Dream of Red Mansions" in the first cross-modal label set, the concept of "Dream of Red Mansions" is extracted by the knowledge extractor 39, and a concept entity set such as {Dream of Red Mansions, Cao Xueqin, Ming and Qing novels} can be obtained. Compare each concept entity in {Dream of Red Mansions, Cao Xueqin, Ming and Qing novels} with the entity in the digital cultural field knowledge graph. For example, if it is found that the entity "Gao E" in the digital cultural field knowledge graph is similar to {Dream of Red Mansions, Cao Xueqin, Ming and Qing novels} and meets the specified similarity judgment criteria, then "Gao E" can be expanded into the concept entity set of {Dream of Red Mansions, Cao Xueqin, Ming and Qing novels}.
[0104] For similarity judgment, similarity judgment methods commonly used in machine learning can be used, such as cosine similarity, classification tree, neural network and other methods.
[0105] It is worth noting that when judging similarity, we can also pay attention to whether the entity has multiple semantics. If it has multiple semantics, there may be misjudgment. Therefore, when comparing similarity, we can use methods such as Figure 3 The multi-semantic mapping database 4 in the text is used to distinguish different semantics of entities. For example, the entity is "珍惜春", and the entity "珍惜春" has two semantics, one is "cherish spring", and the other is a character in the Dream of Red Mansions. At this time, the mapping relationship between the entity and the semantics should be called by the multi-semantic mapping database 4 to make a judgment.
[0106] S105: Based on the knowledge graph of the digital culture field, expand the set of concept entities according to the specified threshold and granularity to obtain an expanded entity set, and update the knowledge graph of the digital culture field based on the expanded entity set.
[0107] Specifically, continuing with the example in S104, for the expanded set of concept entities {Dream of the Red Chamber, Cao Xueqin, Ming and Qing Dynasties novels}, it can be further expanded according to the pre-set threshold and granularity. Here, the threshold generally refers to the standard value for judging similarity during the similarity comparison process. If it is greater than this threshold, it is considered similar; if it is less than this threshold, it is considered dissimilar. For example, if the threshold is set to 60%, those with a similarity greater than 60% are determined to be similar; those with a similarity less than 60% are determined to be dissimilar.
[0108] The role of setting the granularity is that the concept cannot be expanded without limit, but requires a certain boundary, and this boundary is not fixed. For example, sometimes users need to expand a relatively large amount of abstract data, and in this case, more concept entities are required, which means the granularity needs to be set larger; on the contrary, if users only need a brief abstract, the granularity can be set smaller.
[0109] Regarding the usage process of the granularity, reference can be made to Figure 4 , Figure 4 which shows the granularity selection process according to an embodiment. When the granularity is set to 1 (i.e., entities at one step away from "Dream of the Red Chamber"), the selected expanded entity set may include {Dream of the Red Chamber, Cao Xueqin, Lin Daiyu, Tianqiao Theater, Central Ballet Troupe}, which corresponds to all the concept entities within the inner ring 41 at this time; when the granularity is set to 2 (i.e., entities at two steps away from "Dream of the Red Chamber"), the selected expanded entity set may include {Dream of the Red Chamber, Cao Xueqin, Lin Daiyu, Tianqiao Theater, Central Ballet Troupe, Swan Lake}, which corresponds to all the concept entities within the outer ring 42 at this time. It can be seen that the number of concept entities when the granularity is set to 2 is more than that when the granularity is set to 1, which will further affect the number of abstracts that can be generated later.
[0110] In one embodiment, based on the knowledge graph of the digital culture field, expanding the set of concept entities according to the specified threshold and granularity to obtain an expanded entity set, and updating the knowledge graph of the digital culture field based on the expanded entity set includes:
[0111] During the process of forming the set of concept entities, if the similarity between the concept entity and the entities in the knowledge graph of the digital culture field meets the specified threshold and granularity, then incorporate the concept entity into the set of concept entities to obtain the expanded entity set.
[0112] Specifically, expanding similar entities in the knowledge graph of the digital culture field into the set of concept entities can actually be carried out simultaneously during the process of forming the set of concept entities. That is to say, in the knowledge graph of the digital culture field, when looking up the concept entities associated with each label in the first cross-modal label set, once an entity that meets the similarity requirements and also satisfies the specified threshold and granularity requirements is found, the entity in the knowledge graph of the digital culture field is directly expanded into the expanded entity set, which can greatly improve the efficiency of concept expansion.
[0113] For the newly added entities in the expanded entity set, if there is no corresponding concept in the knowledge graph of the digital culture field, these entities can be expanded into the knowledge graph of the digital culture field. At the same time, the relationships between these newly expanded entities and the original entities in the knowledge graph of the digital culture field can be realized by means of knowledge reasoning. For the realization of knowledge reasoning, reference can be made to the introduction in the second half of step S107.
[0114] S106: Obtain a second cross-modal label set based on the vertical domain multi-modal data, compare the expanded entity set with the second cross-modal label set, and sort the expanded entity set to obtain a sorted entity set.
[0115] In one embodiment, obtaining a second cross-modal label set based on the vertical domain multi-modal data, comparing the expanded entity set with the second cross-modal label set, and sorting the expanded entity set to obtain a sorted entity set includes:
[0116] Sample the vertical multi-modal data to obtain a second cross-modal label set;
[0117] For each entity in the expanded entity set, calculate the weight of the entity in the second cross-modal label set;
[0118] Sort the expanded entity set according to the frequency of occurrence of the entity in the second cross-modal label set to obtain a sorted entity set.
[0119] Figure 3 Shows a data flow diagram according to an embodiment of the present application. As Figure 3 Shown, in the knowledge graph 5 of the digital culture field, look up the concept entities associated with each label in the second cross-modal label set 6 to form a concept entity set 71, expand the concept entity set 71 to obtain an expanded entity set 72. Compare the expanded entity set 72 with the second cross-modal label set 6, and sort the expanded entity set 72 to obtain a sorted entity set 73.
[0120] The significance of comparing the extended entity set 72 with the second cross-modal label set 6 lies in that if only conventional means are used to expand the concept entity set with the help of the knowledge graph in the digital culture field, some important information in the original second cross-modal label set 6 may be missed, which will affect the accuracy of abstract expansion and cause concept deviation.
[0121] Continuing to take "Dream of the Red Chamber" as an example, in the process from the vertical domain multimodal data of the to-be-generated abstract to the concept entity set, the concept of "Gao E" may be lost. At this time, if the extended entity set 72 is compared with the second cross-modal label set 6, the concept of "Gao E" can be re-expanded into the extended entity set 72 to avoid the omission of the concept of "Gao E".
[0122] Specifically, the essence of the comparison is actually a process of comparing the extended entity set with the second cross-modal label set, which can be achieved through the following process:
[0123] First, according to different modalities, randomly extract labels from the vertical domain multimodal data to obtain the second cross-modal label set. For example, for text-type data, paragraphs or sentences in the text file can be randomly extracted, or the entire text file can be segmented to obtain the second cross-modal label set; for image-type data, several regions in the image can be randomly extracted, or the entire image can be image-recognized to obtain the second cross-modal label set; for audio-type data, several paragraphs in the audio file can be randomly captured, transcribed into text, and segmented to obtain the second cross-modal label set; and for video-type data, the second cross-modal label set can also be obtained by randomly capturing video frames and then performing image recognition.
[0124] Second, determine whether each label in the second cross-modal label set is similar to the entities in the extended entity set. Traverse the extended entity set to determine whether each label in the second cross-modal label set is in the extended entity set. This judgment process does not require exact identity. Instead, it can be determined by semantic comparison whether the maximum value of the semantic similarity between each label and each entity in the extended entity set reaches a specified threshold. If it exceeds this specified threshold, it is determined to be similar; if it is less than the specified threshold, it is determined to be dissimilar.
[0125] Third, expand the expanded entity set. For tags that are judged to be similar, it is considered that the concept corresponding to this tag is actually already in the expanded entity set. Even if it is forced to be added, it will increase the computational complexity when the content is expanded later. Therefore, for similar tags, choose not to operate. For dissimilar tags, the minimum tolerance set by the user can be used to determine whether this tag is useless for the subsequent summary expansion. In other words, when the maximum similarity between a tag and each entity in the expanded entity set is less than the specified threshold and greater than the minimum tolerance, the tag is expanded into the expanded entity set.
[0126] Fourth, sort according to the comparison results. In the process of expanding the entity set in the third step, although no operation is performed on similar tags, this does not mean that the entity is not important. Instead, it means that the concept is often mentioned and should be highlighted in the process of generating summaries. For such entities, you can choose to increase their weights so that they are at the front of the final sorted entity set. In this way, when generating summaries, these entities with higher weights and higher rankings will be given higher priority, whether in terms of the breadth of expansion or the position in the summary. It is worth noting that in order to improve efficiency, the sorting process can be carried out simultaneously with the process of expanding the entity set in the third step. That is to say, after finding similar entities in the expanded entity set, their weights can be increased at the same time and placed in a position that matches their weight.
[0127] For example, the augmented entity set currently includes {Lin Daiyu, Dream of the Red Chamber, Cao Xueqin}. After random label extraction of vertical domain multimodal data, the second cross-modal label set {The Story of the Stone, Dream of the Red Chamber, Grand View Garden} is obtained. It can be seen that "The Story of the Stone" and "Grand View Garden" are obviously "forgotten" labels after entering the vertical domain model. However, after comparing the second cross-modal label set with the augmented entity set, it is found that the similarity between "The Story of the Stone" and "The Dream of the Red Chamber" is greater than the specified threshold, so "The Story of the Stone" does not need to be included in the augmented entity set again. However, for "Grand View Garden", although its similarity with each entity in {Lin Daiyu, Dream of the Red Chamber, Cao Xueqin} is not as high as that of "The Story of the Stone", it is still greater than the minimum tolerance. Therefore, "Grand View Garden" is also included in the augmented entity set, and finally the augmented entity set is expanded to {Lin Daiyu, Dream of the Red Chamber, Cao Xueqin, Grand View Garden}. At the same time, since "The Story of the Stone" and "A Dream of Red Mansions" are similar, the weight of "A Dream of Red Mansions" will be increased, and the expanded entity set {Lin Daiyu, A Dream of Red Mansions, Cao Xueqin, Grand View Garden} will be sorted to obtain the sorted entity set {A Dream of Red Mansions, Lin Daiyu, Cao Xueqin, Grand View Garden}.
[0128] S107: Generate summary text based on the sorted entity set and the updated digital culture field knowledge graph.
[0129] In one embodiment, based on the sorted entity set and the updated digital culture domain knowledge graph, a summary text is generated, including:
[0130] According to the sorting in the sorted entity set, for each entity in the sorted entity set, expand it according to the specified summary expansion granularity to obtain an entity expansion sequence;
[0131] For each entity in the entity expansion sequence, extract the corresponding relationship in the digital culture domain knowledge graph, and combine the entity expansion sequence to generate a summary text.
[0132] Specifically, through the text generation technologies commonly used in current artificial intelligence (such as Transformer and pre-trained language models), the process of generating a summary can be satisfied. It should be noted that during the summary generation process, external knowledge can also be cited. At this time, specific knowledge encoders need to be set for these external knowledge. The choice of knowledge encoder can be based on the data structure of external knowledge. When introducing external knowledge such as pictures and videos, graph neural networks, convolutional neural networks, pre-trained language models, etc. can be correspondingly selected.
[0133] It should be noted that when generating the summary text, entities, attributes, and relationships will be used. Entities and attributes are actually relatively sufficient through the expansion of this technical solution. For relationships, especially for the entities supplemented from the second cross-modal label, the relationships between these entities and the original entities can be realized by means of knowledge reasoning. For example, logical-based reasoning, graph-based reasoning, machine learning-based reasoning, etc. can all be used to achieve knowledge reasoning. Currently, logical-based knowledge methods, such as first-order predicate logic and description logic; graph-based methods, such as Path Ranking Algorithm (PRA) and Association Rule Mining under Incomplete Evidence (AMIE); machine learning-based reasoning, such as using models like TransE and TransH, have all achieved relatively ideal results in entity relationship reasoning.
[0134] In one embodiment, the method further includes:
[0135] Perform machine translation on the summary text to obtain a summary translation text in a specified language.
[0136] Specifically, through the currently very mature machine translation technology, the summary text can be translated into a version in a specified language, such as an English version, a Japanese version, a French version, etc., so as to meet the more internationalized needs of different users.
[0137] Such asFigure 3 As shown, the abstract text 74 is the final generated result. It can be seen that the content within the dotted box 7 is actually the core of this technical solution. The key point of this technical solution lies in leveraging the knowledge graph in the digital culture field, and by matching the expanded entity set with the second cross-modal label set, the accuracy of abstract generation is improved.
[0138] The second aspect of the present invention further relates to a query method in the digital culture field, which is characterized by including:
[0139] Presenting a query entry on the human-computer interaction interface;
[0140] Responding to a user's multi-modal data query request instruction input at the query entry to obtain a query result;
[0141] Responding to the user clicking on the query result, and obtaining an abstract corresponding to the query result according to the method described in any one of the first aspects of the claims;
[0142] Presenting the abstract on the human-computer interaction interface in the form of a thumbnail.
[0143] The third aspect of the present invention further relates to an electronic device, including at least one processor and a memory. Among them, the memory stores computer-executable instructions, and the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the digital culture field abstract generation method described in any one of the first aspects, or the query method described in the second aspect.
[0144] The fourth aspect of the present invention further relates to a computer-readable storage medium, in which multiple program codes are stored. It is characterized in that the program codes are suitable for being loaded and run by a processor to execute the digital culture field abstract generation method described in any one of the first aspects, or the query method described in the second aspect.
[0145] The functions of each module in each system of the embodiments of the present invention can be referred to the corresponding descriptions in the above methods, and will not be elaborated here.
[0146] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0147] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations in which functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the involved functions or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0148] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained, for example, electronically by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0149] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0150] Those of ordinary skill in the art can understand that all or part of the steps carried out in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0151] In addition, in each of the embodiments of the present invention, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0152] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for generating an abstract in the field of digital culture, characterized in that, Including: Obtain vertical domain multimodal data for which an abstract is to be generated; Input the multimodal data into a vertical domain model to generate a first cross-modal label set corresponding to each of the multimodal data, where the vertical domain model is obtained based on a base large model; Associate the first cross-modal label set with a knowledge graph in the digital culture field; Based on the knowledge graph in the digital culture field, search for concept entities associated with the first cross-modal label set to form a set of concept entities; According to a specified threshold and granularity, based on the knowledge graph in the digital culture field, expand the set of concept entities to obtain an expanded entity set, and update the knowledge graph in the digital culture field based on the expanded entity set; Obtain a second cross-modal label set based on the vertical domain multimodal data, compare the expanded entity set with the second cross-modal label set, and sort the expanded entity set to obtain a sorted entity set; Generate an abstract text based on the sorted entity set and the updated knowledge graph in the digital culture field; Among them, obtaining the vertical domain model based on the base large model includes: Transform the base large model to form a target large model. In the target large model, a text encoder, an image encoder, an audio encoder, and a video encoder are set at the original input layer of the base large model to extract features from corresponding types of samples; an input projector is set to project the extracted features into a common feature space; the interaction of the self-attention mechanism and the feed-forward network of the base large model itself is used to achieve feature fusion; a classifier is set after the original output layer of the base large model to generate cross-modal labels; a knowledge extractor is set after the classifier to extract entities and relationships from the labels; Pre-train the target large model to obtain a pre-trained target large model; Fine-tune and train the pre-trained target large model to obtain the vertical domain model; Among them, obtaining a second cross-modal label set based on the vertical domain multimodal data, comparing the expanded entity set with the second cross-modal label set, and sorting the expanded entity set to obtain a sorted entity set includes: Randomly extract labels from the vertical domain multimodal data to obtain a second cross-modal label set; For each entity in the expanded entity set, calculate the weight of the entity in the second cross-modal label set; Sort the expanded entity set according to the weights of the entities in the second cross-modal label set to obtain a sorted entity set.
2. The method according to claim 1, wherein Based on the knowledge graph in the digital culture field, search for concept entities associated with the first cross-modal label set to form a set of concept entities: Use a knowledge extractor to perform concept entity recognition on the labels associated with the first cross-modal label set to obtain the concept entities; Judge whether the concept entities are similar to the entities in the knowledge graph in the digital culture field through similarity comparison; In the case where the concept entities are similar to the entities in the knowledge graph in the digital culture field, include the entities in the knowledge graph in the digital culture field in the set of concept entities.
3. The method according to claim 1, characterized in that, Based on the knowledge graph in the digital culture field, expand the set of concept entities according to the specified threshold and granularity to obtain an expanded entity set, and update the knowledge graph in the digital culture field based on the expanded entity set, including: During the process of forming the set of concept entities, if the similarity between the concept entity and the entities in the knowledge graph in the digital culture field meets the specified threshold and granularity, incorporate the concept entity into the set of concept entities to obtain the expanded entity set.
4. The method according to claim 1, wherein Generate a summary text based on the sorted entity set and the updated knowledge graph in the digital culture field, including: According to the sorting in the sorted entity set, expand each entity in the sorted entity set according to the specified summary expansion granularity to obtain an entity expansion sequence; For each entity in the entity expansion sequence, extract the corresponding relationships in the knowledge graph in the digital culture field, and combine with the entity expansion sequence to generate a summary text.
5. The method according to claim 1, characterized in that The method further includes: Perform machine translation on the summary text to obtain a summary translation text in a specified language.
6. A query method in the field of digital culture, characterized in that, Including: Present a query entry on the human-computer interaction interface; Respond to a multi-modal data query request instruction input by the user at the query entry to obtain a query result; Respond to the user clicking on the query result, and obtain the summary corresponding to the query result according to the method described in any one of claims 1-5; Present the summary on the human-computer interaction interface in the form of a thumbnail.
7. An electronic device, characterized in that, Including at least one processor and a memory, wherein the memory stores computer execution instructions, and the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the digital culture field summary generation method described in any one of claims 1 to 5, or the query method described in claim 6.
8. A computer-readable storage medium storing multiple program codes, characterized in that, The program code is suitable for being loaded and run by a processor to execute the digital culture field summary generation method described in any one of claims 1 to 5, or the query method described in claim 6.
Citation Information
Patent Citations
Document processing method and system based on natural language and knowledge graph
CN116501875A
Multi-modal model and method for fusing characters, images and audios
CN118861988A