Multi-modal abstract generation and output method and device, equipment and medium
By constructing the target domain knowledge graph and generating multimodal abstracts, the problems of multi-source data integration and personalized output are solved, and efficient information processing and personalized content presentation are achieved.
Patent Information
- Application Number
- CN202510276405.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
It is difficult for the existing technology to efficiently integrate multi-source data to generate multi-modal abstracts and realize personalized output, especially in the integration of Buddhist historical data and presentation of popular science content.
By collecting target field data and cleaning to obtain structured data, the correlation relationship between entities is extracted and constructing a target field knowledge graph. Generate text summary based on the knowledge graph, and match image data and audio data to integrate it to generate a multimodal summary. Build user portraits based on user behavior data and optimize the output of multimodal summary.
It realizes efficient integration of multi-source data and generation of multi-modal content, improves the comprehensiveness and accuracy of information processing, optimizes the extraction and analysis of entities and their association relationships, enhances the personalization and interactivity of content, and improves the efficiency and adaptability of knowledge summary.
Smart Images

Figure CN120216723A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a multi-modal summary generation and output method, device, equipment, and storage medium. Background Art
[0002] The existing technologies have significant deficiencies in information integration, summarization methods, audience adaptability, and personalized knowledge presentation, mainly reflected in difficulties in data integration, single summarization methods, poor popularization effects, and lack of personalized content presentation. Specifically, the information sources are scattered and the formats are not unified, resulting in great challenges in data integration and systematic processing; the summarization methods rely on manual processing, with large workload, low efficiency, and being easily affected by subjective factors; the popularization content is mostly highly academic professional knowledge, which is difficult for ordinary audiences to understand; the existing technologies also fail to output personalized information according to different user needs, restricting the breadth and depth of information dissemination.
[0003] In the field of cultural research, especially in Buddhist studies, the acquisition and collation of Buddhist historical materials face severe challenges. The resources of Buddhist literature are scattered and diverse, including a large number of ancient books and documents, as well as various forms such as academic papers, temple documents, religious scriptures, and online resources. The materials from different sources vary greatly in content, language, and format, and most are in an undigitized or unstructured state, making the integration of Buddhist historical information extremely complex. Especially in issues such as the history of Buddhism's introduction into China, the development of Buddhist sects, and the interpretation of Buddhist scriptures, the records in different ancient books and documents often vary and are distributed in different historical periods and regions, making it difficult to conduct systematic comparison and analysis. For example, regarding the time of Buddhism's introduction, the introduction route, and the acceptance situation in various places, different scriptures and historical records often have different statements, and the existing technologies are difficult to uniformly and effectively integrate and analyze this information. This not only makes the work of Buddhist historical research time-consuming and laborious but also affects the in-depth understanding of the development process of Buddhism.
[0004] In addition, there are also significant deficiencies in the summarization methods of Buddhist history. Currently, most Buddhist research and summarization rely on manual reading, induction, and summarization. This method is not only inefficient but also easily affected by the personal biases and knowledge backgrounds of researchers. For example, when summarizing the rise and fall of different Buddhist sects, different scholars may make different interpretations of the origin, development, and decline of the sects due to different focuses. The limitations of manual summarization make it impossible to comprehensively cover all relevant content in Buddhist literature and difficult to achieve information unity and consistency. With the continuous in-depth development of Buddhist studies, this manual summarization method not only cannot meet the requirements of rapid development in terms of time but also is difficult to adapt to the requirements of interdisciplinary and global research.
[0005] In terms of the effect of popular science, most of the existing popular science content on the history of Buddhism is targeted at professional researchers and the academic circle, making it difficult for the general public to understand. Many Buddhist books, lectures, and popular science articles use a large number of professional terms and original texts of Buddhist scriptures, introducing complex religious thoughts and philosophical concepts. This lacks sufficient popularity for ordinary readers, especially beginners, and affects their understanding and interest in Buddhism. In addition, the existing forms of Buddhist popular science are relatively single, mainly text-based, lacking interactivity and interestingness, and unable to attract more audiences, especially the interest of the younger generation. For example, regarding the evolution process of Buddhist doctrines, existing popular science books often focus on theoretical elaboration, lacking specific descriptions of the historical changes, figures' stories, and practical activities of Buddhism, resulting in poor dissemination effects of popular science content in practice.
[0006] At the same time, the existing ways of Buddhist popular science also lack personalized presentation. Different reader groups have different interest points and knowledge bases regarding the history of Buddhism, but the current popular science methods are difficult to provide customized content according to the specific needs of the audience. For example, beginners may be more in need of understanding the basic concepts and historical background of Buddhism, while in-depth researchers are concerned with the details of Buddhist philosophy and the development of different schools. The existing popular science materials on the history of Buddhism fail to make dynamic adjustments according to the interests and comprehension abilities of readers, so they cannot meet the needs of readers at different levels. In addition, most of the existing Buddhist education and popular science content is static, lacking interactivity and personalized learning path planning, and unable to provide personalized learning resources according to the learning progress and changing interests of readers.
[0007] In the field of medical and health, data integration and standardization are the core challenges currently faced. Medical records, clinical trial data, medical research materials, and treatment plans are usually distributed among different hospitals, research institutions, and medical platforms, and lack unified data standards and formats, resulting in huge difficulties in information integration. There are problems such as inconsistent formats, asymmetric data, and uneven quality among different data sources, greatly affecting the efficiency of data analysis and diagnosis and treatment optimization. For example, clinical data in the medical and health field usually adopts different recording standards, and different hospitals and clinics use different medical record templates, resulting in the inability to effectively connect and share the treatment process and historical data of the same patient on different platforms. The lack of unified standards for medical data poses quite a challenge for doctors and researchers in the selection of treatment plans and the research on disease prevention.
[0008] In the financial field, the sources of market data, policies and regulations, and economic reports are extensive and diverse. The data formats are not unified, and there is a lack of effective standardization processing, resulting in the problem of fragmented information in financial analysis. The data in the financial field includes real-time market data, economic indicators, stock market dynamics, industry reports, etc. However, most of this data comes from different exchanges, financial institutions, and government departments, and uses different data formats and standards. Since this data is not easily processed and integrated automatically, financial analysts and decision-makers usually need to spend a lot of time screening and summarizing information, thus affecting the efficiency and accuracy of decision-making. Summary of the Invention
[0009] The main objective of the present invention is to provide a multi-modal summary generation and output method, device, equipment, and storage medium, aiming to solve the technical problem that it is difficult to efficiently integrate multi-source data to generate a multi-modal summary and achieve personalized output in the prior art.
[0010] To achieve the above objective, the present invention provides a multi-modal summary generation and output method, including:
[0011] Collecting materials in the target field and performing cleaning processing on the materials in the target field to obtain structured data;
[0012] Extracting entities and the association relationships between entities from the structured data, and constructing a knowledge graph of the target field based on the entities and the association relationships between entities;
[0013] Generating a text summary containing the association information between the entities based on the knowledge graph of the target field;
[0014] Matching corresponding image data and audio data according to the text summary, and integrating the text summary, image data, and voice data to generate a multi-modal summary;
[0015] Constructing a user portrait according to the user behavior data, and outputting the multi-modal summary based on the user portrait.
[0016] Furthermore, to achieve the above objective, the present invention provides a multi-modal summary generation and output device, including:
[0017] A data processing module, configured to collect materials in the target field and perform cleaning processing on the materials in the target field to obtain structured data;
[0018] A knowledge graph construction module, configured to extract entities and the association relationships between entities from the structured data, and construct a knowledge graph of the target field based on the entities and the association relationships between entities;
[0019] A summary generation module, configured to generate a text summary containing the association information between the entities based on the knowledge graph of the target field;
[0020] A multi-modal fusion module, configured to match corresponding image data and audio data according to the text summary, and integrate the text summary, image data and voice data to generate a multi-modal summary;
[0021] A personalized output module, configured to construct a user profile based on user behavior data, and output the multi-modal summary based on the user profile.
[0022] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a multi-modal summary generation and output program stored in the memory and executable on the processor. When the multi-modal summary generation and output program is executed by the processor, the steps of the multi-modal summary generation and output method as described above are implemented.
[0023] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a multi-modal summary generation and output program is stored. When the multi-modal summary generation and output program is executed by a processor, the steps of the multi-modal summary generation and output method as described above are implemented.
[0024] Beneficial effects: The present invention relates to the technical field of data analysis and can be applied to business scenarios such as medical health, fintech, and cultural research. It discloses a multi-modal summary generation and output method, including: collecting data in the target field and performing cleaning processing to obtain structured data; extracting entities and the association relationships between entities from the structured data to construct a knowledge graph of the target field; generating a text summary containing entity association information based on the knowledge graph of the target field; matching corresponding image data and audio data for the text summary, and integrating the text summary, image data and voice data to generate a multi-modal summary; constructing a user profile based on user behavior data, and outputting the multi-modal summary based on the user profile. Through the integration and cleaning of multi-source data, the present invention improves the comprehensiveness and accuracy of information processing; through the construction of a knowledge graph, it optimizes the extraction and analysis of entities and their association relationships; through multi-modal fusion, it realizes the associated display of text, images and voices; through the optimization of the user profile for output, it improves the pertinence and personalization of information acquisition, thereby enhancing the efficiency and adaptability of knowledge summary. Description of the Drawings
[0025] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0026] Figure 1 It is a schematic diagram of an application environment of the multi-modal summary generation and output method in an embodiment of the present invention;
[0027] Figure 2Schematic flowchart of an embodiment of the multi-modal abstract generation and output method of the present invention;
[0028] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the multi-modal abstract generation and output device of the present invention;
[0029] Figure 4 Schematic diagram of a structure of a computer device in an embodiment of the present invention;
[0030] Figure 5 Schematic diagram of another structure of a computer device in an embodiment of the present invention. Detailed implementation manners
[0031] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] The multi-modal abstract generation and output method provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server through a network. The server can collect materials in the target field through the client and perform cleaning processing to obtain structured data; extract entities and the association relationships between entities from the structured data to construct a knowledge graph of the target field; generate a text abstract containing entity association information based on the knowledge graph of the target field; match the corresponding image data and audio data of the text abstract, and integrate the text abstract, image data, and voice data to generate a multi-modal abstract; construct a user portrait according to the user behavior data, and output the multi-modal abstract based on the user portrait. The present invention improves the comprehensiveness and accuracy of information processing through multi-source data integration and cleaning; optimizes the extraction and analysis of entities and their association relationships through knowledge graph construction; realizes the associated display of text, images, and voices through multi-modal fusion; optimizes the output through the user portrait to improve the pertinence and personalization of information acquisition, thereby improving the efficiency and adaptability of knowledge summary. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0033] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the multi-modal abstract generation and output method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0034] As Figure 2 shown, the multi-modal abstract generation and output method proposed by the present invention includes the following steps:
[0035] S10, collect data of the target domain and clean the data of the target domain to obtain structured data;
[0036] In this embodiment, the core objective of information collection and processing is to obtain data of the target domain from multiple data sources, and clean and structure the data to ensure the accuracy and operability of subsequent analysis. The multi-source nature of data brings problems such as data format, data quality, and information redundancy. To address these problems, multiple steps are required for data screening, format standardization, and error correction.
[0037] When collecting data of the target domain from multiple data sources, a distributed data collection architecture is adopted to support the efficient capture of massive data. The data sources can include digital libraries, academic paper databases, government websites, industry reports, social media platforms, etc. The data structures of different data sources may vary significantly. To ensure the integrity and consistency of the data, asynchronous parallel processing technology is used during the data collection process, enabling data from different sources to be transmitted and stored in independent channels. An incremental update strategy can also be applied during the data collection process to only capture the parts that have changed compared to the previous update, improving the data update efficiency and avoiding waste of storage and computing resources caused by repeated collection.
[0038] The collected data of the target domain includes various formats such as text, images, tables, and structured data. In the data cleaning step, different formats of data need to be processed separately. The processing of text data usually includes character denoising, encoding unification, typo correction, format conversion, etc. For example, when processing text extracted from historical documents, typos, garbled characters, or punctuation errors caused by OCR recognition may be encountered. An automatic character correction model based on deep learning can be used to correct the errors, and a language model can be combined for syntactic and semantic verification to improve the text quality. For image data, an image text recognition method based on a convolutional neural network is used to extract text information from the image, and content verification is performed by comparing with existing datasets to ensure the accuracy of recognition. For table data, the table structure needs to be parsed, field relationships need to be identified, and it needs to be converted into a parsable data format, such as JSON or a database table structure.
[0039] The cleaned data needs to be structured for subsequent analysis and processing. For text data, entity extraction can be performed through natural language processing technology to convert unstructured text into structured records. For table and database data, a standardized data model is used for storage to make it conform to a unified format specification. For image data, key features are associated with the corresponding text information through feature extraction to form a multi-modal data structure. At the data storage layer, a distributed database can be used for management to support the efficient access of data, and index optimization can be applied to improve the retrieval speed.
[0040] In different scenarios, the methods of data collection and cleaning can vary. For application scenarios with high real-time requirements, a streaming data processing framework such as Apache Kafka or Flink can be adopted to enable data collection and processing to be updated in a low-latency manner. In scenarios where data needs to be archived for a long time, a batch processing method can be used to collect data regularly and use the ETL (Extract-Transform-Load) process to organize and store the data. In addition, a data deduplication strategy can be combined, and hash comparison and feature matching algorithms can be used to eliminate redundant information during the data collection and cleaning stages to ensure the uniqueness of the data.
[0041] Data quality control can be optimized using a variety of technical means. For example, during data denoising, methods such as rule-based filtering and anomaly detection models can be used to eliminate low-quality or invalid information. During the standardization of data formats, predefined regular expressions and rule libraries can be used to convert fields such as text, dates, and units in different formats into a unified format. During data correction, an automatic error correction model based on machine learning can be applied. For example, an automatic semantic correction tool trained using a Transformer model can be used to improve the accuracy of text after OCR recognition.
[0042] In terms of structured data storage, a suitable data storage method can be selected according to different application requirements. For example, in scenarios where large-scale data needs to be analyzed quickly, a columnar database (such as Apache Parquet or ClickHouse) can be used to improve query efficiency. In business scenarios that require high-concurrency access, a NoSQL database (such as MongoDB, Elasticsearch) can be adopted to support flexible data retrieval and expansion.
[0043] Example illustration: In the field of healthcare, medical data comes from a wide range of sources, including electronic medical records, imaging data, clinical trial reports, gene sequencing data, etc. The data formats are diverse and the storage standards are not unified, which affects data analysis and applications. Automatically collecting medical data from multiple channels such as hospital databases, medical research institutions, and public health platforms, and performing format standardization, denoising, and error correction on the data can improve the integrity and accuracy of the data. For example, natural language processing techniques can be used to clean the unstructured text in doctors' medical records, remove redundant information, and extract key entities such as disease names, drugs, and treatment plans through medical term recognition methods. For medical images, deep learning models can be applied for image and text recognition, and the key information in CT and MRI reports can be extracted and converted into structured data, enabling doctors to analyze patients' health conditions more quickly and accurately, and improving the accuracy and efficiency of medical diagnosis.
[0044] In the financial field, there are numerous sources of financial market information, including stock market data, policies and regulations, economic analysis reports, corporate financial statements, etc. The data structure is complex and updated frequently. Traditional manual processing methods are difficult to ensure the integrity and timeliness of the data. By automatically collecting financial data from government agencies, financial media, securities companies, and market data providers, and performing intelligent cleaning and format conversion on the data, the timeliness and accuracy of the data can be improved. For example, using natural language processing methods to parse financial news, extract information such as market events, company names, and financial indicators, and uniformly convert them into structured data for financial analysts to conduct trend prediction and risk assessment. At the same time, for different formats of financial statements, table parsing methods can be used to automatically extract key data from the balance sheet, income statement, and cash flow statement, and convert them into computable numerical formats, thereby improving the efficiency of financial data processing and providing more efficient market decision-making support for investors.
[0045] In the field of Buddhism, Buddhist historical materials are scattered in different sources such as ancient books, temple inscriptions, academic papers, and forum discussions. The data formats are diverse and it is difficult to integrate them effectively. By obtaining Buddhist historical materials from channels such as digital resources in major libraries, official websites of Buddhist temples, and well-known Buddhist forums, and combining natural language processing and image recognition methods to clean the collected data, the data quality can be improved. For example, extracting text information from scanned ancient book images, automatically correcting OCR recognition errors through character correction and semantic analysis to ensure the accuracy of Buddhist scripture texts. At the same time, for Buddhist historical events, the time, place, and relationship between characters in the text can be automatically parsed, and structured data storage can be constructed, enabling researchers to quickly query the development of Buddhism in a certain historical period. For non-standard format documents such as inscriptions, computer vision methods are used to identify the content of the inscriptions and compare them with other Buddhist documents to construct a complete Buddhist knowledge system, enabling researchers to more systematically understand the context of the historical development of Buddhism.
[0046] By establishing a multi-source data collection and cleaning mechanism, the integrity and consistency of the data are improved, and problems such as inconsistent data formats and information redundancy are avoided from affecting subsequent analysis; through technical means such as natural language processing and character correction, the quality of text data is optimized, and the readability and accuracy of text information are improved; through image recognition and table parsing technologies, the efficient integration of multi-modal data is realized, ensuring that different types of data can be correlated with each other and used for knowledge extraction. Overall, the automation degree and processing efficiency of data processing are improved, laying a solid foundation for subsequent knowledge graph construction and intelligent analysis.
[0047] S20, extracting entities and the association relationships between entities from the structured data, and constructing a knowledge graph of the target domain based on the entities and the association relationships between entities;
[0048] In this embodiment, to construct a knowledge graph for the target domain, it is necessary to extract entities and the associated relationships between entities from structured data, and build a graph structure based on these relationships. This process involves multiple steps such as entity recognition, entity disambiguation, entity relationship extraction, relationship classification, and graph construction to ensure the integrity and accuracy of the knowledge graph.
[0049] In the entity recognition process, potential entities are first extracted from the text content of the structured data. A named entity recognition method based on a pre-trained language model is used to identify important concepts related to the target domain, such as person names, place names, events, times, organizations, etc. For tabular data and database data, keyword fields are extracted through field parsing and pattern matching techniques. For example, in historical research data, key entities such as "Xuan Zang", "Nalanda Temple", and "Tang Dynasty" can be extracted from the text, and attributes related to time and place can be parsed from the tabular data through pattern matching.
[0050] After entity recognition, entity disambiguation is required to ensure that the same entity is not misinterpreted as multiple independent entities due to different expressions. A context-based vector representation method can be used to map entities into a high-dimensional semantic space and compare them with the existing knowledge base. For example, "Buddhism" can refer to a religious system or a specific Buddhist scripture name, and the specific referent can be determined through context information to eliminate ambiguity.
[0051] In the entity relationship extraction stage, dependency syntactic analysis and semantic role annotation methods are used to parse the relationships between entities. For example, in the text "Xuan Zang studied at Nalanda Temple", it is parsed that there is a "studied at" relationship between "Xuan Zang" and "Nalanda Temple". In addition, for different types of structured data, different methods are used to extract relationship information. For tabular data, possible entity relationships such as "year - event" and "person - contribution" can be automatically recognized based on field names and data patterns. For knowledge base data, rule matching and knowledge reasoning techniques are used to further supplement missing relationships.
[0052] In the relationship classification stage, the extracted entity relationships are classified. Different relationship types are applicable to different knowledge structures, such as time association relationships (the correspondence between events and years), causal relationships (the impact of historical events), subordination relationships (the relationship between people and organizations), etc. A graph neural network (GNN) is used for relationship classification to improve the ability to identify complex relationships.
[0053] In the knowledge graph construction stage, a graph structure is constructed based on the extracted entities and relationships. Each entity is mapped to a node in the knowledge graph, and the relationship between each pair of entities is mapped to an edge. For example, when constructing a knowledge graph in the field of historical research, "Xuan Zang" can be used as one node, "Nalanda Temple" as another node, and the two are connected by the "studied at" relationship. At the same time, define the attributes of each relationship, such as timestamp, weight, etc., to support subsequent queries and analyses. For multimodal data, such as data containing images or videos, the visual features of entities can be further associated in the knowledge graph, so that when querying a certain historical figure, their portrait or sculpture data can be directly associated.
[0054] In different application scenarios, the extraction of entities and relationships and the construction methods of knowledge graphs can vary. For example, in the field of healthcare, entities mainly include disease names, drugs, treatment plans, clinical trial information, etc., and relationships can be "drug - indication", "disease - symptom", etc. In the financial field, entities include company names, industry categories, market data, etc., and relationships can be "enterprise - market performance", "company - policy impact", etc. Knowledge graphs in different fields may have differences in structure and semantics, so domain - adaptive methods need to be adopted, combining industry - specific language models in the knowledge extraction process to improve the accuracy of entity recognition and relationship extraction.
[0055] For real - time updated data, streaming processing technology can be combined, and an incremental update strategy can be adopted to only supplement newly emerging entities and relationships without affecting the existing knowledge structure. For example, in the process of constructing a knowledge graph in the news field, event detection technology can be used to automatically identify new hot events and dynamically add them to the existing knowledge graph.
[0056] For high - concurrency query requirements, graph databases (such as Neo4j, ArangoDB) can be used for knowledge graph storage to improve query efficiency. At the same time, combined with index optimization technology, the most relevant entity recommendations can be provided according to different relationship weights during query. For example, when a user queries a certain historical event, the relevant events closest in time and with the greatest impact to that event can be preferentially returned to improve the availability of information.
[0057] Through automated entity recognition, relationship extraction, and knowledge graph construction, the efficiency and accuracy of information processing have been improved, avoiding the subjectivity and time cost of manual knowledge system collation. The structured storage method of the knowledge graph makes the relationships between data clearer, supporting efficient information query and knowledge reasoning. Combining semantic analysis and graph neural network technology improves the accuracy of relationship classification, making the connections between different entities more rigorous and complete. Overall, it makes the knowledge extraction and organization of large - scale data more systematic, providing high - quality data support for subsequent intelligent analysis.
[0058] S30. Generate a text summary containing the association information between the entities based on the target domain knowledge graph.
[0059] In this embodiment, in order to generate a text summary from the target domain knowledge graph, it is necessary to extract the core content based on the entities and their association information in the knowledge graph, and construct a coherent text expression through natural language generation methods. This process involves four main links: path query of the knowledge graph, calculation of entity importance, generation of summary candidate sentences, and text optimization, to ensure that the generated summary not only has semantic integrity but also can accurately reflect the association information between entities.
[0060] First, perform a path query in the knowledge graph to identify relevant entities and their association information. The structure of the knowledge graph is usually composed of nodes and edges, where nodes represent entities and edges represent the relationships between entities. For example, in the study of Buddhist history, nodes can represent people (such as "Xuan Zang"), events (such as "journey to obtain scriptures"), locations (such as "Nalanda Temple"), etc., and edges can represent the relationships between these entities (such as "studied under", "translated", "established"). The path query can adopt methods such as depth-first search (DFS) or breadth-first search (BFS), combined with semantic relevance calculation, to filter out a set of highly relevant entities. For example, when querying the relevant information of "Buddhism in the Tang Dynasty", key entities such as "Xuan Zang", "translation of the Prajna Sutra", and "establishment of the Vinaya School" and their relationships can be extracted.
[0061] Based on the set of entities obtained from the path query, calculate the importance of each entity to determine the core information of the summary content. The calculation of entity importance can be based on multiple factors, including the connectivity of the entity in the knowledge graph (i.e., the number of associated entities), the text citation frequency of the entity (such as the number of times it appears in historical documents or research papers), the relevance of the entity in the query task, etc. For example, for the topic of "Xuan Zang's journey to obtain scriptures", the degree of association of "Xuan Zang" in the knowledge graph can be calculated, and combined with time information to confirm its influence in different historical stages, ensuring that entities with high influence are preferentially included in the summary content.
[0062] According to the selected core entities and their relationship information, generate candidate summary sentences. Adopt a method based on grammar templates, combined with the existing knowledge graph relationship types, to automatically generate semantically coherent sentences. For example, for the relationship "The Tripitaka Master (Xuan Zang) went to India to obtain Buddhist scriptures during the Tang Dynasty", it can be parsed based on the knowledge graph structure to get: "As a Buddhist master in the Tang Dynasty, Xuan Zang went to India to obtain Buddhist scriptures and translated important scriptures." This method can ensure that the generated summary content accurately reflects the structural information of the knowledge graph while maintaining readability.
[0063] To improve the readability and expression quality of text summaries, it is necessary to optimize the generated summary sentences. The optimization process can include syntactic correction, language style adjustment, removal of information redundancy, etc. For example, using a Transformer-based natural language generation model, such as BERT or GPT, to optimize the language fluency of the initially generated summary sentences to make the summary more natural. For example, the originally automatically generated summary "Xuan Zang, he translated the Prajna Sutra, which influenced the development of Buddhism." can be optimized to: "Xuan Zang translated the Prajna Sutra, having a profound impact on the development of Buddhism."
[0064] In different application scenarios, the generation methods of text summaries can vary. For example, in the field of healthcare, the summary needs to summarize around the patient's medical history, diagnosis, and treatment plan. In the financial field, the summary needs to automatically extract information such as market analysis and corporate dynamics and maintain logical coherence. In the field of Buddhist studies, the summary needs to take into account the chronological order of historical events, the influence relationships between figures, and the role of classical translations in the development of sects.
[0065] For summary generation with high real-time requirements, an incremental summary method based on streaming processing can be adopted, that is, as the knowledge graph is updated, the summary content is dynamically adjusted. For example, in financial market analysis, whenever new policy information is added to the knowledge graph, the summary content can be automatically updated to ensure that investors obtain the latest market dynamics. In Buddhist studies, as new literature is added to the knowledge system, the summary can automatically expand with new research findings and maintain the coherence of historical information.
[0066] For multi-modal summary generation, the image and speech information of the knowledge graph can be combined so that the summary not only contains text descriptions but also can be associated with images and speech. For example, in Buddhist knowledge popularization, based on the summary text, the corresponding Buddhist murals or sculptures can be automatically retrieved and relevant explanations can be generated, making the summary more interactive and visually expressive.
[0067] Example illustration: In the field of healthcare, when doctors review a patient's medical record, they often need to quickly understand the patient's past medical history, current diagnosis, and treatment plan. The summary generation method based on the knowledge graph can extract the patient's disease development path from the electronic medical record and medical research databases and generate a structured summary. For example, for a diabetic patient, a summary can be automatically generated: "The patient was diagnosed with type 2 diabetes five years ago, received insulin treatment in the past two years, and the most recent blood sugar control is good." This method can help doctors quickly understand the patient's condition and improve the diagnosis and treatment efficiency.
[0068] In the financial field, investors need to obtain core information about a certain market or company. However, manually reading a large number of market reports and financial data is often time-consuming and laborious. Through the summary generation of the knowledge graph, relevant data can be automatically extracted, and a structured market analysis summary can be generated. For example, when analyzing the market performance of a certain technology company, a summary can be automatically generated: "The company's revenue increased by 20% last year, and the R & D investment increased by 15% year-on-year. The main product lines were expanded to the fields of AI and autonomous driving." Such an automated summary can improve the efficiency of market analysis and help investors make decisions quickly.
[0069] In the field of Buddhism, the study of Buddhist history and scriptures involves a large number of ancient books, inscriptions, and academic papers. Researchers need to quickly master the core information of specific historical events or the evolution of sects. By using the summary generation method of the knowledge graph, the associated information between people, sects, and scriptures can be automatically extracted, and a clear historical context can be generated. For example, for the development of the Yogacara school, a summary can be automatically generated: "The Yogacara school originated from the Yogacara school in India. Xuanzang translated the Treatise on the Establishment of Consciousness Only in the Tang Dynasty, laying the foundation for the Chinese Yogacara system." In addition, by combining the visualization of historical maps and the propagation paths of sects, researchers can more intuitively understand the process of the development of Buddhism.
[0070] Through the summary generation based on the knowledge graph, it can be ensured that the generated text content accurately reflects the associated information between entities, avoiding the problems of easily missing key concepts or causing information fragmentation in traditional text summary methods. By using path query and entity importance calculation, it can be ensured that the summary content covers the core information in the knowledge graph, improving the accuracy and readability of the summary. Combining natural language generation technology to optimize the summary text makes the generated content more in line with the reading habits and improves the efficiency of information acquisition. Overall, it makes the organization and transmission of large-scale knowledge data more systematic, helping to improve the knowledge summary and information application capabilities in different fields.
[0071] S40, match the corresponding image data and audio data according to the text summary, and integrate the text summary, image data, and voice data to generate a multimodal summary;
[0072] In this embodiment, in order to generate a multimodal summary, key information needs to be extracted from the text summary, and the corresponding image data and audio data need to be matched so that it can be presented in multiple forms, improving the richness and interactivity of information expression. This process includes semantic analysis of the text summary, cross-modal data matching, temporal and semantic alignment, interactive content annotation, and final multimodal summary integration to ensure the content coordination among different data types.
[0073] First, perform semantic analysis on the text abstract to extract the core information for matching image and audio data. Use a semantic parsing model based on deep learning to identify key information such as entities, time, location, and events in the abstract, and utilize context information to judge the relative importance of these elements. For example, when describing the event of "Xuanzang traveling to India to obtain Buddhist scriptures in the Tang Dynasty", core entities such as "Xuanzang", "Tang Dynasty", "India", and "Buddhist scriptures" can be extracted and assigned different weights to preferentially select the most relevant image and audio data during subsequent matching.
[0074] Then, retrieve image data that matches the text abstract in the image database. Adopt an image indexing method based on a knowledge graph to keep the semantic information of the image data consistent with the text abstract. For example, in the field of Buddhist studies, if the abstract content involves "Nalanda Temple", historical images containing this location or photos of temple buildings are preferentially retrieved. If the abstract content involves "Xuanzang's journey to obtain scriptures", murals, sculptures, or historical maps related to it are retrieved. In addition, image content recognition technology can be used to analyze the image features in the database and perform similarity matching with the description of the text abstract to further improve the accuracy of the retrieval results.
[0075] When matching audio data, adopt methods of speech synthesis or speech retrieval to ensure that the audio content is consistent with the content of the text abstract. For an existing speech database, speech indexing technology can be used to extract the parts related to the abstract content from the recorded audio materials. For example, in the field of medical health, if the abstract content describes a treatment plan for a certain disease, medical expert explanation audio containing the introduction of this disease is preferentially matched. In the financial field, if the abstract involves the market performance of a certain enterprise, market analysis speech segments related to this enterprise are preferentially matched. If there is no suitable existing audio data, speech synthesis technology based on the Transformer architecture (such as Tacotron) is used to generate natural speech, enabling the text abstract to be presented in speech form, and combining prosody control technology to adjust parameters such as pitch, speech rate, and timbre to enhance the auditory experience.
[0076] After the matching is completed, it is necessary to perform temporal and semantic alignment on the text abstract, image data, and audio data to ensure the coordination of content in different modalities. Adopt a cross-modal attention mechanism to compare the key information in the text abstract with the content features of the image and audio, and adjust the display order of content in different modalities. For example, during the playback of a multimodal abstract, when the audio explains "Xuanzang's journey to the west to obtain scriptures", the corresponding historical map or mural is synchronously displayed, and Xuanzang's itinerary route is highlighted. For financial data, when the audio introduces the revenue growth of an enterprise, the corresponding financial statement image can be presented simultaneously, and visual annotations are added at key data points to enhance the understanding of the information.
[0077] To enhance the user's interactive experience, it is necessary to add interactive content annotations to image and audio data. For image data, define interactive regions in the areas corresponding to the text summaries, enabling users to click on specific regions to view detailed information. For example, in the field of Buddhism, when presenting a map of a Tang Dynasty temple, hotspots can be added to specific architectural parts, and clicking on them will expand the introduction of the building's history. In the field of healthcare, when presenting medical images, the lesion areas can be highlighted and detailed pathological analyses can be provided. In the financial field, when presenting a stock market trend chart, detailed data explanations can be provided for key time nodes.
[0078] Finally, integrate the text summaries, image data, and audio data to generate an interactive multimodal summary. Adopt a unified display format, with the text summary as the core content, and embed the matching image and audio data, allowing users to read the text, view pictures, and listen to voice explanations simultaneously. For example, in a financial news summary, the text part outlines market dynamics, the image part shows the market trend chart, and the audio part plays expert analysis and interpretation, enabling users to obtain information through multiple senses and improving the comprehensibility and immersion of the information.
[0079] In different application scenarios, the construction methods of multimodal summaries can vary. For example, in applications with high real-time requirements, streaming processing techniques can be adopted to enable the dynamic update of summary content. For example, in financial market analysis, when new economic data is released, the latest market data charts and voice interpretations can be automatically matched to keep the summary content always up-to-date.
[0080] In terms of content generation, different levels of multimodal summaries can be provided in combination with the user's reading preferences. For example, for patient education content in the field of healthcare, summaries of different complexities can be provided. For professional doctors, detailed medical images and interpretations are provided, while for ordinary patients, concise text summaries and popularized voice explanations are provided. In the field of Buddhism, the complexity of the summary content can be adjusted according to the user's background knowledge level. For example, for Buddhist researchers, complete historical maps, pictures of original documents, and detailed voice interpretations are provided, while for the general public, a simplified overview of historical events and interactive picture displays are provided.
[0081] For the storage and distribution of multimodal summaries, a hierarchical caching mechanism can be adopted to improve the content loading speed. For example, in mobile applications, the text summary and thumbnail images can be pre-loaded, and high-resolution pictures or full audio can be loaded after the user clicks on the interactive region to optimize the user experience. At the same time, intelligent recommendation algorithms can be combined to adjust the display order of multimodal summaries according to the user's historical browsing behavior, improving the personalized matching degree of the content.
[0082] By matching text summaries, image data, and audio data, a multi-modal summary is constructed, which improves the diversity and readability of information. The information is not limited to text expression but is presented through multiple sensory channels, enhancing the user's depth of understanding and information absorption efficiency. Cross-modal data matching technology ensures the content consistency between different data types and improves the accuracy and logic of the multi-modal summary. Through interactive content annotation, users can actively explore the information they are interested in, improving the operability of the summary and the user experience. Overall, a more efficient information transmission method is achieved, making information summarization and dissemination in multiple fields more intuitive, vivid, and accurate.
[0083] S50. Construct a user profile based on user behavior data and output the multi-modal summary based on the user profile.
[0084] In this embodiment, to optimize the output of the multi-modal summary, it is necessary to construct a user profile based on user behavior data so that the content and presentation form of the summary can adapt to the interests and needs of different users. This process involves the collection of user behavior data, the extraction of user interest tags, the modeling of user profiles, the personalized matching and hierarchical organization of summary content, and the interactive output based on user profiles to ensure that the summary can accurately meet the user's information needs and improve the efficiency and experience of information acquisition.
[0085] First, the collection of user behavior data is the basis for constructing a user profile. When users use the system, their behavior data includes browsing history, search records, reading duration, likes, comments, collections, forwards, and other interaction behaviors. For different types of behavior data, a hierarchical data storage method is adopted for management. For example, the behavior data accumulated over a long time is stored in a distributed database for long-term analysis of user interest trends, while the short-term behavior data is stored in a cache to support real-time personalized recommendations. For interaction behaviors related to multi-modal content such as videos and audios, data such as the user's playback progress, skipped content, and playback times can also be recorded to further optimize the modeling of user preferences.
[0086] In the user interest tag extraction stage, natural language processing and deep learning technologies are utilized to extract core interest points from user behavior data. For example, by analyzing users' search and reading records, the topics that users most frequently query are identified; by analyzing users' like and favorite behaviors, the preference intensity for a certain type of content is inferred. For example, in the field of Buddhism, if a user frequently searches for keywords such as "Yogacara School", "Chan School", "Xuanzang", etc., the interest tag of "Buddhist sects" can be extracted. In the field of medical and health, if a user browses a large number of articles about "diabetes" or "cardiovascular diseases", it can be inferred that they are interested in chronic disease management. In the financial field, if a user has long been concerned about content such as "tech stocks", "policy analysis", etc., it can be deduced that they are concerned about the trends in the stock market. Combining the time dimension of user behavior data, short-term interests and long-term interests can be further distinguished to ensure that the recommended multimodal summaries not only meet the current needs but also match the user's long-term concerns.
[0087] Based on the above-extracted interest tags and behavior characteristics, a personalized interest graph is constructed for user portrait modeling. A user interest modeling method based on graph neural networks is adopted, enabling the automatic learning and optimization of the association relationships between different interest points. For example, if a user is interested in "artificial intelligence" and also shows a high browsing frequency for "fintech", it can be speculated that they may be concerned about "the application of AI in the financial field", and relevant content can be preferentially provided during multimodal summary recommendation. To ensure the dynamic adaptability of the user portrait, an incremental update strategy is adopted, enabling new behavior data to continuously adjust the interest graph and optimize the accuracy of user preference prediction.
[0088] In the personalized matching and hierarchical organization stage, a match is made according to the content characteristics of the user portrait and the multimodal summary. The summary content is divided into a summary layer and a detailed layer according to the information granularity. The summary layer provides the core content and is suitable for quick browsing, while the detailed layer provides in-depth interpretation and is suitable for users with in-depth needs. For example, in the field of medical and health, for ordinary users, the summary layer can only display the basic symptoms and prevention measures of the disease, while the detailed layer provides in-depth information such as relevant research progress and drug mechanisms. In the field of Buddhism, the summary layer can briefly outline the development history of a certain sect, while the detailed layer can supplement its core doctrines, representative figures, and classic interpretations. In the financial field, the summary layer can present market dynamics and trends, and the detailed layer can supplement content such as data analysis and expert interpretations. Based on the interest tags in the user portrait, the system can intelligently select the information granularity suitable for the user, ensuring that the display of the summary is neither overloaded nor overly simplified.
[0089] Finally, in the interactive output stage based on the user profile, the presentation method of the summary is adjusted according to the user's interaction behavior habits. For example, for users who are accustomed to browsing image content, the visual elements in the multimodal summary can be preferentially displayed, and a function of clicking on the image to expand detailed information can be provided; for users who prefer voice explanations, the corresponding audio content can be automatically played when the summary is displayed; for users who like text reading, a way to expand the full text can be provided. In addition, to enhance the interactive experience, a personalized recommendation mechanism can be introduced in the summary. For example, multimodal summaries related to relevant topics can be recommended based on the user's browsing history, enabling users to explore the fields they are interested in more deeply.
[0090] In different application scenarios, the construction of the user profile and the personalized output method of the multimodal summary can vary. For example, in the field of healthcare, electronic health records (EHRs) and the user's health assessment data can be combined to generate more accurate health knowledge summaries, and the recommended content can be adjusted according to the user's medical history. For example, content related to a low-sugar diet can be recommended to diabetic patients. In the financial field, the user's investment preferences can be combined. For example, long-term stockholders and short-term traders have different demands for market information, so the hierarchical organization of the summary and the recommended content also need to be different. In the field of Buddhism, the user's historical browsing records and sect preferences can be combined. For example, for users interested in the Pure Land Sect, the summary content can preferentially display the historical development, representative scriptures, and cultivation methods of the Pure Land Sect.
[0091] For application scenarios with high real-time requirements, a user profile update mechanism based on stream computing can be adopted to enable the summary content to be dynamically adjusted as the user's behavior changes. For example, in real-time news push, if a user has recently frequently browsed the market dynamics in a certain field, the recommended summary content can tend to the latest news analysis in that field. When the user's interest changes, the recommended content is also adjusted accordingly to ensure that the summary always meets the user's needs.
[0092] In the output process of the multimodal summary, reinforcement learning can be used to optimize the recommendation strategy, enabling the system to continuously improve the matching effect of the summary according to the user's feedback. For example, if a user skips a certain type of summary content multiple times, the system can reduce the recommendation frequency of that type of summary. If a user stays on certain summaries for a long time, the recommendation weight of the relevant content can be increased. By continuously optimizing the recommendation logic, the personalized matching degree of the summary can be improved to make it more in line with the user's reading preferences.
[0093] Example illustration: In the field of healthcare, doctors and patients often need to integrate a large amount of information such as literature, images, and voice explanations when acquiring medical knowledge. Based on the multi-modal summary generation and output method, various types of data in the healthcare field can be integrated and a multi-modal summary can be generated to improve the efficiency of obtaining medical information. First, the system collects medical-related materials from channels such as medical databases, electronic health records (EHRs), scientific research paper databases, and picture archiving and communication systems (PACS), and cleans these data, removing irrelevant content, standardizing the text format, and converting medical image data into recognizable structured data. The cleaned data includes structured content such as disease names, symptom descriptions, treatment methods, drug information, and image data. After the data cleaning is completed, the system extracts entities from the structured data based on natural language processing (NLP) technology, such as diseases, drugs, treatment means, medical images, etc., and identifies the association relationships between them. For example, when studying the treatment plan for diabetes, "diabetes" can be extracted as an entity, and its association relationships with entities such as "insulin treatment", "blood glucose monitoring", and "diet management" can be established to construct a medical knowledge graph. Based on the constructed medical knowledge graph, the system can generate a text summary containing the association information between entities. For example, when a doctor queries the treatment plan for a certain type of cancer, the system can generate a summary: "The treatment methods for this cancer include surgery, radiotherapy, chemotherapy, and targeted therapy. The applicability of different plans depends on the patient's pathological stage and gene mutation status." Such summary content can help doctors quickly understand the core information of the treatment method. To enhance the visualization and audibility of medical information, the system matches relevant medical image data and voice data based on the text summary. For example, in the summary of coronary heart disease, the system can match coronary angiography images and provide voice explanations by cardiovascular experts, enabling doctors or patients to intuitively understand the manifestation form of the lesion. Finally, the system constructs a user profile based on user behavior data and optimizes the summary output according to the user profile. For example, for doctor users, the system preferentially recommends medical research summaries with high precision and provides detailed treatment guidelines; for ordinary patients, the system provides popularized medical knowledge summaries and matches interactive medical images, enabling patients to view interpretations of different disease sites by clicking on the highlighted areas. The system can also adjust the summary granularity according to the user's reading habits, such as providing in-depth pathological mechanism analysis for professional users and brief health management suggestions for ordinary users.
[0094] In the financial field, investors, market analysts and financial practitioners need to obtain market dynamics, corporate financial data and industry research reports. However, financial data comes from various sources, including financial news, financial reports, stock market trends, industry analysis, etc. How to efficiently integrate this information and generate valuable summaries is an important issue. First, the system collects financial data from financial news websites, stock exchange databases, corporate financial statements, policy announcements and other channels, and cleans the data, such as removing irrelevant information, standardizing financial indicators, and eliminating noise data, to generate structured data. The cleaned data includes company names, stock codes, financial data (such as revenue, net profit, market share), industry trends, etc. After the data is structured, the system extracts key entities from the data, such as companies, industries, policies, economic indicators, etc., and identifies the associations between them. For example, the system can identify the relationship between "a certain technology company-5G industry-market expansion-policy support" and construct a financial knowledge graph, so that investors can clearly see the relationship between market trends and corporate development. Based on the financial knowledge graph, the system generates a text summary. For example, in response to market dynamics, the system can generate a summary: "Recently, the 5G industry has been driven by policy support, the market share of related companies has increased, and a certain technology company's R&D investment in the field of 5G chips has increased by 20%." This kind of summary can help investors quickly understand the market situation. In addition, the system matches relevant chart data and voice data based on text summaries. For example, when describing "a company's stock price rose by 10%", the system can match the company's stock price chart and provide audio interpretations of financial analysts, allowing investors to analyze market trends from multiple angles. In order to optimize the personalized output of summaries, the system builds user portraits based on user behavior data. For example, for investors who have been paying attention to the technology industry for a long time, the system will give priority to summaries about industries such as artificial intelligence, 5G, and semiconductors, and match relevant policy interpretation videos; for short-term traders, the system will provide a high-frequency updated market analysis summary, and combine trend charts and voice broadcasts to achieve real-time information push.
[0095] In the fields of Buddhist studies and popular science, content such as Buddhist history, sect development, and classical studies involves a large number of ancient books, inscriptions, temple documents, and academic papers. Due to the scattered literature, complex language styles, and vast knowledge systems, traditional research methods are difficult to efficiently integrate this information. The system first collects Buddhist-related materials from channels such as digital ancient book databases, Buddhist temple official websites, academic paper databases, and Buddhist forums, and cleans the data, such as performing character recognition, modernizing classical Chinese, removing noisy text, etc., to obtain structured data. The cleaned data includes historical figures, sects, doctrines, classics, events, geographical locations, etc. After data collation, the system uses natural language processing technology to extract key entities from the data, such as "Xuan Zang", "Yogacara School", "Dharma Characteristics School", "Mahayana Sutras", etc., and identifies the relationships between them. For example, "Xuan Zang - studied under - Master Silabhadra", "Xuan Zang - translated - Treatise on the Establishment of Consciousness Only", "Yogacara School - originated from - Yogacara", etc. These entities and relationships are integrated into the Buddhist knowledge graph, enabling researchers to systematically understand the historical context of sect inheritance and classical translation. Based on the knowledge graph, the system can generate text summaries. For example, when a user queries the development of the Yogacara School, the system generates a summary: "The Yogacara School originated from the Yogacara in India. Xuan Zang translated Treatise on the Establishment of Consciousness Only in the Tang Dynasty, laying the foundation for the Chinese Yogacara system." Such summaries can help researchers quickly obtain the development process of Buddhism. To enhance the intuitiveness of information, the system matches relevant Buddhist images and audio data according to the text summary. For example, in the summary of "Xuan Zang's Journey to the West", the system can match the historical map of Xuan Zang's westward journey route and provide the audio of Buddhist scripture recitation, enabling users to understand Buddhist history more immersively. Finally, the system constructs a user portrait based on user behavior data and optimizes the personalized output of the summary. For example, for users interested in Buddhist art, the system preferentially recommends summary content related to Buddhist sculptures and murals; for scholars researching Buddhist philosophy, the system provides a detailed interpretation of classical texts, combined with voice reading and multimedia display, enabling them to deeply study the doctrines.
[0096] By constructing a user portrait based on user behavior data and using the user portrait to optimize the output of multi-modal summaries, the summary content can accurately match the user's interests, improving the efficiency of information acquisition and the personalized experience. Based on the hierarchical organization method of the user portrait, the summary can adapt to different users' needs for information depth, while avoiding problems of information overload or information insufficiency. Combined with interactive content display, the summary not only has high readability but also can optimize the display method according to the user's usage habits, enhancing the user's immersive experience. Overall, it further improves the intelligence level of information acquisition, increasing the accuracy and applicability of information presentation.
[0097] The present invention relates to the field of data analysis technology and can be applied to business scenarios such as medical health, fintech, and cultural research. It discloses a multi-modal summary generation and output method, including: collecting materials in the target field and performing cleaning processing to obtain structured data; extracting entities and the association relationships between entities from the structured data to construct a knowledge graph of the target field; generating a text summary containing entity association information based on the knowledge graph of the target field; matching the corresponding image data and audio data of the text summary, and integrating the text summary, image data, and voice data to generate a multi-modal summary; constructing a user portrait based on user behavior data and outputting the multi-modal summary based on the user portrait. Through the integration and cleaning of multi-source data, the present invention improves the comprehensiveness and accuracy of information processing; through the construction of a knowledge graph, it optimizes the extraction and analysis of entities and their association relationships; through multi-modal fusion, it realizes the associated display of text, images, and voices; through the optimization of output by the user portrait, it improves the pertinence and personalization of information acquisition, thereby enhancing the efficiency and adaptability of knowledge summary.
[0098] In one embodiment, the above S20 includes:
[0099] S201, performing part-of-speech tagging processing and named entity recognition processing on the text content of the structured data through a pre-trained language model to obtain a preliminary entity set;
[0100] S202, extracting supplementary entities from the table of the structured data and the results of image text recognition, and merging the supplementary entities with the preliminary entity set through an entity alignment module to generate a complete entity set;
[0101] S203, performing syntax and semantic analysis on the text content of the structured data to determine the causal relationship, subordination relationship, and time sequence between entities;
[0102] S204, mapping the entities in the complete entity set to the nodes of the knowledge graph of the target field, and mapping the causal relationship, subordination relationship, and time sequence between the entities to the edges of the knowledge graph of the target field;
[0103] S205, adding attribute descriptions to the nodes of the knowledge graph of the target field, and defining the types and weights of the edges of the knowledge graph of the target field.
[0104] In this embodiment, in order to construct the knowledge graph of the target field, it is necessary to extract entities and the association relationships between them from the structured data, and on this basis, establish the nodes and edges of the knowledge graph so that it can clearly express the logical relationship and hierarchical structure between entities. The whole process includes links such as entity recognition, entity merging, entity relationship parsing, knowledge graph construction, and attribute supplementation to ensure that the knowledge graph can accurately reflect the domain knowledge.
[0105] First, a pre-trained language model is used to process the text content of structured data and identify the entities therein. Through part-of-speech tagging and named entity recognition (NER) techniques, key entities related to the target domain are extracted from the text data. For example, in the field of healthcare, entities such as "diabetes", "insulin", and "blood glucose monitoring" are extracted; in the financial field, entities such as "market volatility", "central bank policies", and "stock indices" are extracted; in the field of Buddhism, entities such as "Xuan Zang", "Yogacara school", and "Sutra of the Lotus Flower of the Wonderful Dharma" are extracted. The pre-trained language model can utilize context information to improve the accuracy of entity recognition. For example, it can distinguish the different semantic meanings of "Huayan school" as the name of a Buddhist sect and "Huayan" as the name of a scripture.
[0106] Secondly, supplementary entities are extracted from the table and image OCR results of structured data, and these entities are aligned and integrated. Table data usually contains structured entity information, such as disease names, drug dosages, and diagnostic results in medical reports, while the OCR results of image data may contain non-standard text data such as inscriptions on steles and ancient books. For example, in Buddhist studies, the name of the eminent monk "Fa Zang" can be extracted from an ancient book image and compared with "Fa Zang (Huayan school)" in the existing database to avoid duplicate records. An entity alignment module is used to merge these supplementary entities with the text entities through methods such as string matching and semantic similarity calculation to generate a complete entity set.
[0107] Next, syntactic and semantic analysis is performed on the text content to determine the relationships between entities. The relationships can include causal relationships, subordination relationships, chronological order, etc. For example, in the medical field, the causal relationship between "hypertension" and "heart disease" can be analyzed; in the financial field, the relationship that "policy adjustment" leads to "market volatility" can be identified; in the field of Buddhism, relationships such as "Xuan Zang" studied under "Śīlabhadra" and translated "Vijñaptimātratāsiddhi Śāstra" can be identified. Techniques such as Dependency Parsing and Semantic Role Labeling are used to parse the logical relationships between entities in the text to ensure the accuracy of relationship extraction.
[0108] In the process of constructing a knowledge graph, the complete entity set is mapped to the nodes of the knowledge graph, and the relationships between entities are mapped to the edges of the knowledge graph. For example, in the field of Buddhism, the knowledge graph can have the following structure:
[0109] Nodes: "Xuan Zang", "Yogacara school", "Vijñaptimātratā school"
[0110] Edges: "Xuan Zang" → "Translate" → "Vijñaptimātratāsiddhi Śāstra", "Yogacara school" → "Origin" → "Vijñaptimātratā school"
[0111] Similarly, in the medical field, a "disease - symptom - treatment" knowledge graph can be constructed, and in the financial field, a "company - industry - market" knowledge graph can be constructed, enabling users to intuitively understand the associations between various entities.
[0112] To enhance the usability of the knowledge graph, it is also necessary to add attribute descriptions to the nodes and define the types and weights of the edges. For example, in the Buddhist knowledge graph, the node attributes of Xuanzang can include "birthplace", "major contributions", etc., and the edge weight of the "translation" relationship can be calculated based on the citation frequency of historical documents to represent the importance of this relationship. For the financial knowledge graph, attributes such as "market value" and "main business" can be added to the company nodes, and weights can be assigned to the "investment relationship" edges to measure the association strength between different enterprises. In this way, the knowledge graph can provide more accurate context information during querying and support reasoning analysis.
[0113] Through the above steps in this embodiment, entities and their relationships in the target domain data can be effectively extracted, a high - quality knowledge graph can be constructed, providing structured knowledge support for research and applications in different fields. The accuracy of entity recognition is improved through a pre - trained language model, the integrity of the knowledge graph is ensured using entity alignment and relationship parsing methods, and the usability of the knowledge graph is optimized through attribute descriptions and edge weights. Ultimately, the organization and retrieval of data are made more systematic, the efficiency of information query is improved, and support is provided for reasoning analysis based on the knowledge graph.
[0114] In one embodiment, the above S30 includes:
[0115] S301, extracting at least one association path composed of nodes and edges from the target domain knowledge graph, where each association path includes a starting entity, an ending entity, and a relationship chain connecting the starting entity and the ending entity;
[0116] S302, generating a position encoding matrix for each entity in the association path, where the position encoding matrix includes the hierarchical depth, time - axis coordinates, and spatial coordinates of the entity in the target domain knowledge graph;
[0117] S303, inputting the position encoding matrix into an encoder to determine the attention weights between entities in the association path through a position - aware self - attention mechanism;
[0118] S304, screening out key entities and key relationship chains from the association path according to the attention weights, and generating a set of candidate summary statement collections based on the key entities and key relationship chains;
[0119] S305, generating the text summary based on the set of candidate summary statement collections.
[0120] In this embodiment, in the process of generating a text summary based on the target domain knowledge graph, the key steps include the extraction of association paths, the calculation of entity position encoding, the calculation of weights based on the position-aware self-attention mechanism, the screening of key entities and key relationship chains, and the generation of the final text summary. This series of technical processes ensures that the summary can accurately reflect the entity associations in the knowledge graph and adapt to the knowledge expression requirements of different domains.
[0121] First, extract at least one association path composed of nodes and edges from the target domain knowledge graph. The selection of this path can be based on various strategies, such as semantic search based on the user's query intention, shortest path search based on path length, or dynamic path expansion based on semantic similarity. In the knowledge graph structure, each node represents an entity, and each edge represents the relationship between entities. Different path combinations can present various associations between entities, such as causal relationships, hierarchical relationships, chronological orders, etc. The selection of the path should ensure the integrity and coherence of the summary while avoiding information redundancy.
[0122] Next, generate a position encoding matrix for each entity, which contains information such as hierarchical depth, time-axis coordinates, and spatial coordinates. The hierarchical depth is used to measure the structural position of the entity in the knowledge graph. For example, an entity is located at the concept layer, instance layer, or specific event layer. The time-axis coordinates are used to express the time attributes of the entity, such as the occurrence time of historical events, the promulgation time of policies, the release time of corporate financial reports, etc. The spatial coordinates are used to describe geographical information and are applicable to knowledge graphs involving geographical locations, such as urban development, trade routes, religious propagation paths, etc. Through these encoded information, the spatio-temporal relationships of entities can be enhanced during the summary generation process, making the summary content more in line with factual logic.
[0123] Then, input the position encoding matrix into the encoder and calculate the attention weights between entities through the position-aware self-attention mechanism. The self-attention mechanism can calculate the importance of different entities in the input sequence, and the position-aware self-attention mechanism further combines time and space information, enabling the calculation of entity weights not only based on the semantics of the text itself but also considering its spatio-temporal distribution. For example, in the summary generation of a certain historical event, this mechanism can automatically adjust the importance of certain entities according to the time coordinates, making the events closer to the current time more prominent in the summary. In addition, based on the information of spatial coordinates, the system can adjust the focus of attention in the summary. For example, when studying global trade flows, it can give priority to highlighting economic entities related to the user's location.
[0124] After the entity weights are calculated, it is necessary to screen key entities and key relationship chains from the associated paths. The selection of key entities can be based on the attention weight ranking to ensure that the summary content covers the most informative entities. In addition, the screening process of key relationship chains needs to combine semantic analysis to ensure that the summary retains the most representative connections between entities. For example, when analyzing the causes of diseases, causal relationships should be preferred in the summary, while when analyzing sectarian inheritances, master-disciple relationships should be preferred. To improve the stability of screening, knowledge distillation techniques can be used so that the summary content can take into account knowledge from different sources while avoiding information bias.
[0125] Finally, a final text summary is generated based on the set of candidate summary sentences. The generation of text summaries can use rule-driven syntactic templates or deep learning-based generation models. The rule-driven method is suitable for highly structured information summaries, such as legal regulations interpretation, financial reports, etc., while the deep learning method is suitable for processing complex text information, such as news summaries, historical event reviews, etc. Combining these two methods can optimize the expression quality of the summary, making it both professional and improving readability.
[0126] In summary, by extracting associated paths from the knowledge graph, calculating positional encodings, adjusting entity weights based on the position-aware self-attention mechanism, screening key entities and relationship chains, and finally generating a text summary, the summary content is made more accurate, logically clear, and applicable to the knowledge expression needs of multiple fields.
[0127] In this embodiment, by extracting associated paths from the knowledge graph of the target domain, combining the positional encoding matrix to represent the hierarchical, temporal, and spatial information of entities, and using the position-aware self-attention mechanism to calculate the weights between entities, the screening of key entities and relationship chains is realized, and then a text summary with clear logic and highly condensed information is generated, improving the accuracy of information extraction, enhancing the spatio-temporal relevance of the summary, and optimizing the readability and domain adaptability of the summary content.
[0128] In one embodiment, the above S305 includes:
[0129] S3051, input the set of candidate summary sentences into a reinforcement learning model based on the proximal policy optimization algorithm, and optimize it through a reward function including entity coverage, temporal order consistency, and causal relationship matching degree, to generate an optimized set of candidate summary sentences;
[0130] S3052, screen out target summary sentences from the optimized set of candidate summary sentences according to the entity coverage threshold, temporal order consistency threshold, and causal relationship matching degree threshold;
[0131] S3053, determine the summary mode, and the summary mode includes a concise mode and a detailed mode;
[0132] S3054, when the abstract mode is the concise mode, extract from the target abstract statements the abstract statements that contain a single-layer relationship chain including a starting entity, an ending entity, and no intermediate entities, where the single-layer relationship chain is composed of a causal relationship, a subordination relationship, or a chronological relationship directly connecting the starting entity and the ending entity in the target domain knowledge graph;
[0133] S3055, when the abstract mode is the detailed mode, extract from the target abstract statements the abstract statements that contain a starting entity, an ending entity, an intermediate entity, a multi-layer relationship chain, and a geographical location description, where the multi-layer relationship chain is composed of two or more consecutive relationships and contains at least one intermediate entity;
[0134] S3056, combine the extracted abstract statements into the text abstract according to the chronological order or causal relationship of events in the target domain knowledge graph.
[0135] In this embodiment, during the generation process of the text abstract, it is necessary to screen based on the candidate abstract statement set using a reinforcement learning optimization strategy, and adjust the abstract structure and information granularity according to different abstract modes to ensure that the final abstract can accurately express the entity relationships in the knowledge graph and meet the needs of different users.
[0136] First, input the candidate abstract statement set into a reinforcement learning model based on the Proximal Policy Optimization (PPO) algorithm, and optimize it through a reward function that includes entity coverage, chronological order consistency, and causal relationship matching degree. Proximal Policy Optimization is a reinforcement learning method that can improve the optimization efficiency of the model while ensuring policy stability. In this method, entity coverage is used to measure whether the abstract covers the core entities in the knowledge graph, chronological order consistency ensures that the abstract statements conform to the chronological logic of event development, and causal relationship matching degree is used to measure whether the abstract content accurately reflects the causal associations in the knowledge graph. Through this optimization process, more logical and informationally complete abstract statements can be screened out.
[0137] Second, according to the entity coverage threshold, chronological order consistency threshold, and causal relationship matching degree threshold, screen the target abstract statements from the optimized candidate abstract statement set. The entity coverage threshold is used to determine whether the abstract contains enough key entities, the chronological order consistency threshold ensures that there are no logical reversals in the abstract, and the causal relationship matching degree threshold is used to exclude statements that do not conform to the knowledge logic. For example, in the medical field, if a certain abstract statement fails to correctly express the relationship between the cause and the symptom, the causal relationship matching degree is low, and this statement will be excluded.
[0138] Then, determine the summary mode, which includes a concise mode and a detailed mode. In different application scenarios, user requirements vary, so different levels of information expression methods are needed. For example, the concise mode is suitable for quick information acquisition, while the detailed mode is suitable for in-depth analysis and research.
[0139] In the concise mode, extract summary statements that contain a single-layer relationship chain with a starting entity, an ending entity, and no intermediate entities from the target summary statements. The single-layer relationship chain is composed of a causal relationship, a subordination relationship, or a chronological relationship that directly connects the starting entity and the ending entity in the target domain knowledge graph. For example, in the financial field, if querying "the reasons for market fluctuations", the system can extract a direct causal relationship such as "The central bank cuts interest rates → The stock market rises", without including complex intermediate variables.
[0140] In the detailed mode, extract summary statements that contain a starting entity, an ending entity, intermediate entities, a multi-layer relationship chain, and a geographical location description from the target summary statements. The multi-layer relationship chain is composed of two or more consecutive relationships and includes at least one intermediate entity. For example, in the field of Buddhism, if studying the inheritance relationship of the Yogacara school, the system can extract a multi-layer relationship chain such as "Yogacara school → Inheritance → Yogacara school → Development → Huayan school", and supplement the corresponding geographical location information, such as "The establishment of the Yogacara school was mainly influenced by the Buddhist system of Nalanda Temple in India and later spread to Chang'an in China." This detailed mode is suitable for scenarios that require in-depth analysis, such as academic research and industry reports.
[0141] Finally, according to the chronological order or causal relationship of events in the target domain knowledge graph, combine the extracted summary statements into the final text summary. Chronological arrangement is suitable for describing the development process of events, such as historical events and the development process of enterprises, while causal relationship sorting is suitable for analyzing influencing factors, such as medical diagnosis and market prediction. For example, in the medical field, if analyzing the complications of diabetes, the system can generate a causal chain of "Elevated blood sugar levels over a long period → Damage to blood vessels → Increased risk of cardiovascular disease", making the summary more in line with the pathological logic.
[0142] In this embodiment, the candidate summary statements are optimized through reinforcement learning, screened by combining entity coverage, chronological order consistency, and causal relationship matching degree, and the information granularity is dynamically adjusted according to different summary modes, so that the generated text summary can adapt to different user requirements while ensuring clear logic, improving the accuracy, hierarchical expression ability, and domain adaptability of the summary.
[0143] In one embodiment, the above S40 includes:
[0144] S401, Retrieve image data that matches the entity attributes in the text summary from a pre-built image database, where each image in the image database is associated with the entity attributes in the target domain knowledge graph;
[0145] S402, Input the text summary into a speech synthesis model, and generate speech data by combining the sentiment attribute labels of events in the target domain knowledge graph;
[0146] S403, Perform temporal and semantic alignment on the text summary, image data, and speech data through a cross-modal attention mechanism;
[0147] S404, Define interactive regions for the annotation regions in the aligned image data that correspond to the entity attributes of the text summary;
[0148] S405, Define synchronization markers for the time axis nodes of the aligned speech data, where the synchronization markers are associated with the coordinates of the corresponding annotation regions in the image data;
[0149] S406, Integrate the text summary, the image data containing interactive regions, and the speech data containing synchronization markers to generate an interactive multi-modal summary.
[0150] In this embodiment, during the process of matching image data and audio data based on the text summary and generating a multi-modal summary, it is necessary to perform image retrieval, speech synthesis, cross-modal alignment, interactive region definition, synchronization marker setting, and finally integration and output in sequence. Through these steps, the integration of text, image, and speech is achieved, making the summary content more intuitive and vivid, and providing an interactive function to enhance the user experience.
[0151] First, retrieve image data that matches the entity attributes in the text summary from a pre-built image database. Each image in the image database is associated with the entity attributes in the target domain knowledge graph to ensure that the retrieved images can accurately correspond to the key entities in the text summary. For example, in the medical field, if the text summary involves "the characteristics of lung X-ray images", the system can retrieve X-ray images with relevant lesion annotations from the database. In the financial field, if the text summary involves "the stock price trend of a certain enterprise", then a visualization chart showing the stock price changes of the enterprise can be retrieved. In the Buddhist field, if the text summary describes "Xuanzang's journey to the West to obtain Buddhist scriptures", the system can retrieve historical maps or images of temple sites related to his journey to enhance the visualization expression ability.
[0152] Secondly, input the text summary into the speech synthesis model and generate speech data in combination with the sentiment attribute tags of events in the target domain knowledge graph. The speech synthesis model can adopt the Tacotron series based on the Transformer architecture and combine with emotional speech synthesis technology to enable the synthesized speech to adjust the intonation according to the summary content. For example, in the medical field, when describing disease prevention advice, the speech can adopt a steady and instructive tone, while when describing emergency symptoms, the warning nature of the intonation can be enhanced. In the financial field, a somewhat exciting tone can be used when describing market prosperity, and a steady tone can be adopted when describing market risks. In the Buddhist field, a solemn and rhythmic tone can be used when explaining Buddhist doctrines, and when describing historical stories, the speech rate can be adjusted to make the narration more story-like.
[0153] Then, perform temporal and semantic alignment on the text summary, image data, and speech data through a cross-modal attention mechanism. The cross-modal attention mechanism can learn the correlation relationships between different modal data and ensure the content consistency of the multi-modal summary. For example, in the medical field, the system can ensure that the time points of the speech explanation correspond to specific regions of the medical images. For example, when describing the CT scan results, the diseased area can be highlighted. In the financial field, the speech interpretation can be synchronized with the stock market trend chart, enabling users to intuitively view key data points while listening to the market analysis. In the Buddhist field, when the speech tells about Xuanzang's study at Nalanda Temple, the system can synchronously display historical images of Nalanda Temple to enhance the immersive experience.
[0154] Next, define interactive regions for the annotation regions in the aligned image data that correspond to the entity attributes of the text summary. The interactive region refers to the region where the user can click or hover to obtain more information. For example, in medical images, the diseased area can be annotated so that when the user clicks, a detailed medical interpretation pops up. In financial data visualization, interactive regions can be defined for key data points (such as stock market sharp rise points, policy change points), enabling users to view detailed analysis. In Buddhist research, the locations of historical figure activities can be annotated in temple building images so that when the user clicks, they can obtain the historical background of that location.
[0155] Subsequently, define synchronization markers for the time axis nodes of the aligned speech data. The synchronization markers are used to associate the coordinate positions of the annotation regions in the image data, enabling the system to automatically highlight the corresponding image regions when the user listens to the speech explanation. For example, during the interpretation of a medical report, when the speech mentions "lung shadow", the system can automatically highlight the relevant region in the X-ray image. In financial analysis, when the speech mentions "the sharp decline of the stock market in the first quarter of 2023", the corresponding time point in the stock market trend chart can be highlighted synchronously. In the Buddhist field, when the speech interpretation mentions "Buddhist rituals recorded in Dunhuang murals", the key details of the mural can be synchronously displayed and specific parts can be highlighted to facilitate user understanding.
[0156] Finally, integrate the text summary, image data containing interactive regions, and voice data containing synchronization markers to generate an interactive multimodal summary. Such a multimodal summary can not only provide multi-angle displays of information but also enhance the user experience through interactive design. For example, in medical popular science, users can click on a disease name to view the detailed images and interpretations of that disease. In financial market analysis, users can slide the time axis to view the changes in market data at different time points. In Buddhist teaching, users can click on keywords in classical texts to view the corresponding image interpretations or historical background information.
[0157] In this embodiment, through cross-modal data matching, speech synthesis and temporal alignment, definition of interactive regions, and setting of synchronization markers, the integrated expression of text, images, and voices is achieved. The generated multimodal summary not only has the ability to organize information in a structured manner but also enhances the visualization and interactive experience, improves the user's understanding efficiency of complex information, and expands the application scenarios of multimodal content.
[0158] In one embodiment, the above S50 includes:
[0159] S501, obtain user behavior data and construct a user profile according to the user behavior data;
[0160] S502, extract interest tags from the user profile, match the interest tags with the entity attributes in the target domain knowledge graph, and generate a matching result;
[0161] S503, perform hierarchical organization processing on the multimodal summary according to information granularity to generate a hierarchical summary including a summary layer and a detailed layer;
[0162] S504, determine target content from the summary layer and / or the detailed layer of the hierarchical summary according to the matching result;
[0163] S505, sort the target content according to the time correlation information of events in the target domain knowledge graph to generate a temporally ordered hierarchical summary;
[0164] S506, add an interactive collapsible control to the temporally ordered hierarchical summary and output the temporally ordered hierarchical summary after adding the interactive collapsible control.
[0165] In this embodiment, in the process of constructing a user profile based on user behavior data and performing personalized output of multi-modal summaries, it mainly involves several key links such as user profile construction, interest tag extraction, summary hierarchical organization, matching target content, time-related sorting, and interaction control setting. Through these steps, the summary content can more accurately meet the user's needs and improve the pertinence and readability of information acquisition.
[0166] First, obtain user behavior data and construct a user profile based on the user behavior data. The user behavior data can be sourced from the user's historical browsing records, search keywords, like and favorite data, reading duration, interaction records, etc. After preprocessing these behavior data, they can be clustered and classified through a machine learning model to form a personalized user profile. For example, in the medical field, the user's search and reading records can be used to determine whether they are concerned about disease prevention, treatment plans, or health management. In the financial field, the user's behavior data can be used to determine whether they are short-term traders or long-term investors to provide suitable market analysis summaries. In the Buddhist studies field, the user's browsing data can be used to determine whether they are interested in Buddhist philosophy, Buddhist history, or Buddhist art, and adjust the summary content accordingly.
[0167] Secondly, extract interest tags from the user profile and match the interest tags with the entity attributes in the target domain knowledge graph to generate a matching result. Interest tags are abstract representations of the content that users focus on in the long or short term, such as "diabetes management", "new energy vehicle market", "Zen practice", etc. By matching the interest tags with the entities in the knowledge graph, the content most relevant to the user's needs can be filtered out. For example, in the medical field, if the user's interest tag is "cardiovascular health", the system can preferentially recommend summaries about hypertension, arteriosclerosis, etc. In the financial field, if the user's interest tag is "technology stock investment", the system can preferentially provide summaries about the development trends of the technology industry. In the Buddhist studies field, if the user's interest tag is "Pure Land School", the system can preferentially display the origin, development, and core scriptures of the Pure Land School.
[0168] Then, hierarchically organize and process the multi-modal summary according to the information granularity to generate a hierarchical summary containing a summary layer and a detailed layer. The summary layer provides the core points of the information and is suitable for quick reading, while the detailed layer provides more in-depth information and is suitable for users with in-depth research needs. For example, in the medical field, the summary layer summary may only include the main symptoms and treatment methods of a certain disease, while the detailed layer summary may include its pathogenesis, the latest clinical research, and the recommended medication plan. In the financial field, the summary layer may only contain an overall analysis of market trends, while the detailed layer may contain specific corporate financial statements and industry competition analysis. In the field of Buddhism, the summary layer can provide an overview of the development history of a certain sect, while the detailed layer can provide the evolution of doctrines in different periods and a detailed introduction of representative figures.
[0169] Next, according to the matching results, determine the target content from the summary layer and / or detailed layer of the hierarchical summary. The selection of the target content is based on the user's personalized needs. For example, for users who hope to quickly obtain information, the system can only provide the content of the summary layer, while for users who hope to conduct in-depth research, the system can provide the content of the detailed layer or a summary combining both. In the process of matching the target content, natural language understanding (NLU) technology and semantic similarity calculation of user interest tags can be used to ensure that the recommended summary content truly meets the user's needs.
[0170] Subsequently, sort the target content according to the time correlation information of events in the target domain knowledge graph to generate a chronological hierarchical summary. The time sorting can be based on the occurrence time of events, the timeline of research progress, or the market change trend. For example, in the medical field, if a user hopes to understand the treatment progress of a certain disease, the system can arrange relevant research results in chronological order, from early treatment methods to the latest clinical trial results. In the financial field, if a user hopes to understand the market performance of a certain enterprise, the system can organize the stock price changes, profit situations, and policy impacts of the enterprise in chronological order. In the field of Buddhism, if a user hopes to understand the process of Buddhism's introduction into China, the system can display the timeline information of Buddhism's spread via the Silk Road, sea routes, etc. in chronological order.
[0171] Finally, add an interactive collapsible control to the chronological hierarchical summary and output the chronological hierarchical summary with the interactive collapsible control added. The interactive collapsible control enables users to expand or collapse the summary content at different levels according to their needs, enhancing the flexibility of reading. For example, in a medical summary, users can click "Expand More" to view detailed medical research data. In the financial field, users can click on a time node to view the market performance and corporate news at that time. In the field of Buddhism, users can click on the collapsible control to expand the development history of Buddhism in a certain period and view the important figures and classic works in that period.
[0172] In this embodiment, a user portrait is constructed by obtaining user behavior data, and personalized summary output is performed in combination with the knowledge graph of the target domain, achieving accurate matching, hierarchical organization, and time series optimization of the summary content. At the same time, the interactive control enhances the user's ability to independently select information, enabling the summary to meet both the needs of quick browsing and in-depth research, and improving the applicability of the summary content and the user experience.
[0173] In one embodiment, the above S10 includes:
[0174] S101, collecting target domain materials from multiple data sources through a distributed data collection module, where the target domain materials include text material data and image material data;
[0175] S102, performing cleaning processing on the text material data using a natural language processing module, including character denoising, word segmentation, format conversion, and standardization processing;
[0176] S103, performing optical character recognition processing on the image material data to generate an optical character recognition result;
[0177] S104, integrating the processed text material data and the optical character recognition result in a preset format to generate the structured data.
[0178] In this embodiment, in the process of collecting target domain materials and performing cleaning processing to obtain structured data, four main links are involved: data collection, text cleaning, image processing, and data integration. Through these steps, the integrity, standardization, and applicability for subsequent processing of the data can be ensured.
[0179] First, target domain materials are collected from multiple data sources through a distributed data collection module. The target domain materials include text material data and image material data. The data collection module can use methods such as web crawlers, API interface calls, and database queries to obtain information from different sources. For example, in the medical field, data on disease diagnosis, treatment plans, and drug research can be collected from scientific research paper databases, electronic medical record systems, medical literature databases, etc. In the financial field, market analysis, financial data, and economic indicators can be collected from stock trading platforms, corporate announcements, industry reports, etc. In the Buddhist field, Buddhist-related materials can be collected from Buddhist canonical literature, historical inscriptions, temple literature, academic research papers, etc. Since the data formats from different sources are diverse, the data collection module needs to support structured data (such as database records), semi-structured data (such as JSON and XML formats), and unstructured data (such as PDF, pictures, and scanned documents).
[0180] Secondly, the text data is processed by the natural language processing module for character denoising, word segmentation, format conversion, and standardization. Character denoising includes removing unnecessary special characters and correcting OCR recognition errors. Word segmentation is used to split continuous text into meaningful word units to improve the accuracy of subsequent semantic analysis. For example, in Chinese text processing, word segmentation techniques based on statistical models or neural networks can be used to ensure that proper nouns (such as disease names, company names, Buddhist terms) are not wrongly split. Format conversion mainly unifies data from different sources, such as converting document formats like PDF and Word into plain text or structured formats. Standardization processing includes unifying time formats, unit conversions (such as currency units), and term consistency adjustments (such as standard naming of medical terms).
[0181] Then, the image data is processed by performing optical character recognition to generate the optical character recognition result. This process can use deep learning-based OCR (Optical Character Recognition) technology to extract text information from data such as scanned documents, handwritten notes, and stele images. For example, in the medical field, it can recognize handwritten text in pathology reports and convert it into searchable electronic data. In the financial field, it can extract key data from financial report screenshots or bank bills. In the Buddhist field, it can recognize text from image data such as ancient stele inscriptions and temple scriptures, enabling these historical materials to be digitally stored and retrieved. To improve recognition accuracy, the OCR module can combine context semantic analysis to automatically correct possible misrecognized characters.
[0182] Finally, the processed text data and the optical character recognition results are integrated according to a preset format to generate structured data. The data integration process involves steps such as information deduplication, data alignment, and relationship mapping. For example, in medical data processing, if data on a certain disease comes from multiple research institutions, duplicate information needs to be removed and associations established to form a complete case data. In financial data processing, it is necessary to align financial data from different sources in terms of time to ensure the timeliness and consistency of the data. In Buddhist data processing, it is necessary to compare and analyze similar records in different ancient books and classify and organize them to form a coherent historical record. Ultimately, all data is converted into standard structured formats such as database storage, JSON, or XML formats for subsequent knowledge extraction and modeling use.
[0183] Through distributed data collection, natural language processing, optical character recognition, and data integration, this embodiment realizes the efficient collection and cleaning of multi-source heterogeneous data, ensures the integrity, standardization, and usability of the data, improves the readability and consistency of the data, and provides high-quality input data for subsequent knowledge extraction, relationship analysis, and multi-modal summary generation.
[0184] In one embodiment, a multimodal summary generation and output device is provided. The multimodal summary generation and output device corresponds one-to-one to the multimodal summary generation and output method in the above embodiment. Referring to Figure 3 , Figure 3 FIG. is a schematic diagram of functional modules of a preferred embodiment of the multimodal summary generation and output device of the present invention. Data processing module 10, knowledge graph construction module 20, summary generation module 30, multimodal fusion module 40, and personalized output module 50. The detailed description of each functional module is as follows:
[0185] The data processing module 10 is configured to collect materials in the target field and perform cleaning processing on the materials in the target field to obtain structured data;
[0186] The knowledge graph construction module 20 is configured to extract entities and the association relationships between entities from the structured data, and construct a target field knowledge graph based on the entities and the association relationships between entities;
[0187] The summary generation module 30 is configured to generate a text summary containing the association information between the entities based on the target field knowledge graph;
[0188] The multimodal fusion module 40 is configured to match corresponding image data and audio data according to the text summary, and integrate the text summary, image data, and voice data to generate a multimodal summary;
[0189] The personalized output module 50 is configured to construct a user profile according to user behavior data, and output the multimodal summary based on the user profile.
[0190] In one embodiment, the knowledge graph construction module 20 is specifically configured to:
[0191] Perform part-of-speech tagging processing and named entity recognition processing on the text content of the structured data through a pre-trained language model to obtain a preliminary entity set;
[0192] Extract supplementary entities from the table of the structured data and the image text recognition result, and merge the supplementary entities with the preliminary entity set through an entity alignment module to generate a complete entity set;
[0193] Perform syntactic and semantic analysis on the text content of the structured data to determine the causal relationship, subordination relationship, and time sequence between entities;
[0194] Map the entities in the complete entity set to the nodes of the target field knowledge graph, and map the causal relationship, subordination relationship, and time sequence between the entities to the edges of the target field knowledge graph;
[0195] Add attribute descriptions to the nodes of the target domain knowledge graph, and define the types and weights of the edges of the target domain knowledge graph.
[0196] In one embodiment, the summary generation module 30 is specifically configured to:
[0197] Extract at least one association path composed of nodes and edges from the target domain knowledge graph, and each association path includes a starting entity, an ending entity, and a relationship chain connecting the starting entity and the ending entity;
[0198] Generate a position encoding matrix for each entity in the association path, where the position encoding matrix includes the hierarchical depth, time axis coordinates, and spatial coordinates of the entity in the target domain knowledge graph;
[0199] Input the position encoding matrix into the encoder to determine the attention weights between entities in the association path through the position-aware self-attention mechanism;
[0200] Screen out key entities and key relationship chains from the association path according to the attention weights, and generate a set of candidate summary sentences based on the key entities and key relationship chains;
[0201] Generate the text summary based on the set of candidate summary sentences.
[0202] In one embodiment, the summary generation module 30 is specifically configured to:
[0203] Input the set of candidate summary sentences into a reinforcement learning model based on the proximal policy optimization algorithm, and optimize it through a reward function including entity coverage, temporal order consistency, and causal relationship matching degree to generate an optimized set of candidate summary sentences;
[0204] Screen out target summary sentences from the optimized set of candidate summary sentences according to the entity coverage threshold, temporal order consistency threshold, and causal relationship matching degree threshold;
[0205] Determine the summary mode, where the summary mode includes a concise mode and a detailed mode;
[0206] When the summary mode is the concise mode, extract summary sentences from the target summary sentences that contain a single-layer relationship chain including the starting entity, the ending entity, and no intermediate entities, and the single-layer relationship chain is composed of a causal relationship, a subordination relationship, or a temporal order relationship directly connecting the starting entity and the ending entity in the target domain knowledge graph;
[0207] When the abstract mode is the detailed mode, extract the abstract statement containing the starting entity, the ending entity, the intermediate entity, the multi-layer relationship chain and the geographical location description from the target abstract statement. The multi-layer relationship chain is composed of two or more consecutive relationships and contains at least one intermediate entity;
[0208] Combine the extracted abstract statements into the text abstract according to the chronological order or causal relationship of events in the target domain knowledge graph.
[0209] In one embodiment, the multi-modal fusion module 40 is specifically configured to:
[0210] Retrieve image data matching the entity attributes in the text abstract from a pre-built image database, and each image in the image database is associated with the entity attributes in the target domain knowledge graph;
[0211] Input the text abstract into a speech synthesis model, and generate speech data in combination with the emotion attribute labels of events in the target domain knowledge graph;
[0212] Perform temporal and semantic alignment on the text abstract, image data and speech data through a cross-modal attention mechanism;
[0213] Define an interactive area for the annotation area corresponding to the entity attributes of the text abstract in the aligned image data;
[0214] Define a synchronization mark for the time axis nodes of the aligned speech data, and the synchronization mark is associated with the corresponding annotation area coordinates in the image data;
[0215] Integrate the text abstract, the image data containing the interactive area and the speech data containing the synchronization mark to generate an interactive multi-modal abstract.
[0216] In one embodiment, the personalized output module 50 is specifically configured to:
[0217] Obtain user behavior data and construct a user profile according to the user behavior data;
[0218] Extract interest tags from the user profile, match the interest tags with the entity attributes in the target domain knowledge graph, and generate a matching result;
[0219] Perform hierarchical organization processing on the multi-modal abstract according to the information granularity to generate a hierarchical abstract including a summary layer and a detailed layer;
[0220] Determine the target content from the summary layer and / or the detailed layer of the hierarchical abstract according to the matching result;
[0221] Sort the target content according to the time correlation information of events in the target domain knowledge graph to generate a chronological hierarchical summary;
[0222] Add an interactive collapsible control to the chronological hierarchical summary and output the chronological hierarchical summary with the interactive collapsible control added.
[0223] In one embodiment, the data processing module 10 is specifically configured to:
[0224] Collect target domain materials from multiple data sources through a distributed data collection module, where the target domain materials include text material data and image material data;
[0225] Perform cleaning processing on the text material data using a natural language processing module, including character denoising, word segmentation, format conversion, and standardization processing;
[0226] Perform image text recognition processing on the image material data to generate an image text recognition result;
[0227] Integrate the processed text material data and the image text recognition result in a preset format to generate the structured data.
[0228] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multi-modal summary generation and output method.
[0229] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a multi-modal summary generation and output method
[0230] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0231] Collect data in the target field and perform cleaning processing on the data in the target field to obtain structured data;
[0232] Extract entities and the association relationships between entities from the structured data, and construct a knowledge graph of the target field based on the entities and the association relationships between entities;
[0233] Generate a text summary containing the association information between the entities based on the knowledge graph of the target field;
[0234] Match corresponding image data and audio data according to the text summary, and integrate the text summary, image data, and voice data to generate a multi-modal summary;
[0235] Construct a user profile based on user behavior data, and output the multi-modal summary based on the user profile.
[0236] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are realized:
[0237] Collect data in the target field and perform cleaning processing on the data in the target field to obtain structured data;
[0238] Extract entities and the association relationships between entities from the structured data, and construct a knowledge graph of the target field based on the entities and the association relationships between entities;
[0239] Generate a text summary containing the association information between the entities based on the knowledge graph of the target field;
[0240] Match corresponding image data and audio data according to the text summary, and integrate the text summary, image data, and voice data to generate a multi-modal summary;
[0241] Construct a user profile based on user behavior data, and output the multi-modal summary based on the user profile.
[0242] It should be noted that for the functions or steps that can be realized by the above-mentioned computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0243] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0244] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0245] It should be noted that in the embodiments of this application, if there are software tools or components of other companies, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A multimodal summary generation and output method, characterized in that: The following steps are involved: Collecting target domain data, and cleaning the target domain data to obtain structured data; Extracting entities and relationships between entities from the structured data, and constructing a target domain knowledge graph based on the entities and relationships between entities; Generate a text summary containing association information between the entities based on the target domain knowledge graph; Matching corresponding image data and audio data according to the text summary, and integrating the text summary, image data and voice data to generate a multimodal summary; A user profile is constructed according to the user behavior data, and the multimodal summary is output based on the user profile.
2. The multimodal summary generation and output method according to claim 1, characterized in that: Extracting entities and relationships between entities from the structured data, and constructing a target domain knowledge graph based on the entities and relationships between entities, including: Performing part-of-speech tagging and named entity recognition on the text content of the structured data through a pre-trained language model to obtain a preliminary entity set; Extracting supplementary entities from the table of the structured data and the image text recognition results, and merging the supplementary entities with the preliminary entity set through an entity alignment module to generate a complete entity set; Performing grammatical and semantic analysis on the text content of the structured data to determine the causal relationship, affiliation and time sequence between entities; Mapping the entities in the complete entity set into nodes of the target domain knowledge graph, and mapping the causal relationships, affiliations and time sequences between the entities into edges of the target domain knowledge graph; Add attribute descriptions to the nodes of the target domain knowledge graph and define the types and weights of the edges of the target domain knowledge graph.
3. The multimodal summary generation and output method according to claim 1, characterized in that: Generating a text summary containing association information between the entities based on the target domain knowledge graph, including: Extracting at least one association path consisting of nodes and edges from the target domain knowledge graph, each association path comprising a starting entity, an ending entity, and a relationship chain connecting the starting entity and the ending entity; Generate a position encoding matrix for each entity in the association path, wherein the position encoding matrix includes the hierarchical depth, time axis coordinates and spatial coordinates of the entity in the target domain knowledge graph; Inputting the position encoding matrix into an encoder to determine the attention weights between entities in the association path through a position-aware self-attention mechanism; Filtering key entities and key relationship chains from the association path according to the attention weights, and generating a set of candidate summary sentences based on the key entities and key relationship chains; The text summary is generated based on the candidate summary sentence set.
4. The multimodal summary generation and output method according to claim 3, characterized in that: Generating the text summary based on the candidate summary sentence set includes: Inputting the candidate summary sentence set into a reinforcement learning model based on a proximal policy optimization algorithm, optimizing the set through a reward function including entity coverage, temporal consistency, and causal relationship matching, and generating an optimized candidate summary sentence set; Filtering a target summary sentence from the optimized candidate summary sentence set according to an entity coverage threshold, a time sequence consistency threshold, and a causal relationship matching threshold; Determine a summary mode, wherein the summary mode includes a concise mode and a detailed mode; When the summary mode is the concise mode, extracting a summary sentence including a start entity, an end entity, and a single-layer relationship chain without intermediate entities from the target summary sentence, wherein the single-layer relationship chain is composed of a causal relationship, a subordinate relationship, or a time sequence relationship directly connecting the start entity and the end entity in the target domain knowledge graph; When the summary mode is the detailed mode, extracting a summary sentence including a start entity, an end entity, an intermediate entity, a multi-layer relationship chain and a geographical location description from the target summary sentence, wherein the multi-layer relationship chain is composed of two or more continuous relationships and includes at least one intermediate entity; According to the time sequence or causal relationship of events in the target domain knowledge graph, the extracted summary sentences are combined into the text summary.
5. The multimodal summary generation and output method according to claim 1, characterized in that: Matching corresponding image data and audio data according to the text summary, and integrating the text summary, image data and voice data to generate a multimodal summary, including: Retrieving image data matching the entity attributes in the text summary from a pre-built image database, wherein each image in the image database is associated with the entity attributes in the target domain knowledge graph; Input the text summary into a speech synthesis model, and generate speech data in combination with the emotional attribute labels of events in the target domain knowledge graph; Performing temporal and semantic alignment on the text summary, image data, and speech data through a cross-modal attention mechanism; defining an interactive region for an annotated region in the aligned image data corresponding to an entity attribute of the text summary; Defining synchronization marks for the time axis nodes of the aligned speech data, wherein the synchronization marks are associated with corresponding marked area coordinates in the image data; The text summary, the image data including the interactive area, and the voice data including the synchronization mark are integrated to generate an interactive multimodal summary.
6. The multimodal summary generation and output method according to claim 1, characterized in that: Constructing a user profile according to the user behavior data, and outputting the multimodal summary based on the user profile, including: Acquire user behavior data, and build a user profile based on the user behavior data; Extracting interest tags from the user portrait, matching the interest tags with entity attributes in the target domain knowledge graph, and generating matching results; The multimodal summary is hierarchically organized according to information granularity to generate a hierarchical summary including a summary layer and a detailed layer; Determining target content from a summary layer and / or a detailed layer of the hierarchical summary according to the matching result; Sorting the target content according to the time association information of events in the target domain knowledge graph to generate a time-series hierarchical summary; An interactive foldable control is added to the temporal hierarchical summary, and the temporal hierarchical summary after the interactive foldable control is added is output.
7. The multimodal summary generation and output method according to claim 1, characterized in that: Collect target domain data and clean the target domain data to obtain structured data, including: Collecting target domain data from multiple data sources through a distributed data collection module, wherein the target domain data includes text data and image data; The text data is cleaned using a natural language processing module, including character denoising, word segmentation, format conversion and standardization; Performing image and text recognition processing on the image data to generate an image and text recognition result; The processed text data and the image character recognition results are integrated according to a preset format to generate the structured data.
8. A multimodal summary generation and output device, characterized in that: The multimodal summary generation and output device comprises: A data processing module is used to collect target domain data and clean the target domain data to obtain structured data; A knowledge graph construction module, used to extract entities and relationships between entities from the structured data, and to construct a target domain knowledge graph based on the entities and relationships between entities; A summary generation module, used to generate a text summary containing association information between the entities based on the target domain knowledge graph; A multimodal fusion module, used to match the corresponding image data and audio data according to the text summary, and integrate the text summary, image data and voice data to generate a multimodal summary; The personalized output module is used to construct a user portrait according to the user behavior data, and output the multimodal summary based on the user portrait.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a multimodal summary generation and output program stored in the memory and executable on the processor. When the multimodal summary generation and output program is executed by the processor, the steps of the multimodal summary generation and output method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a multimodal summary generation and output program, which, when executed by the processor, implements the steps of the multimodal summary generation and output method according to any one of claims 1 to 7.
Citation Information
Cited By
System and method for converting electric power notification text into 5G message
CN120429350A
A system and method for converting power notification text to 5G messages
CN120429350B
LLM output stability control method and system based on reinforcement learning
CN120804310A
Reinforcement learning-based llm output stability control method and system
CN120804310B
Meeting content intelligent generation processing method and system based on multi-modal large model
CN120873212A