Multi-modal large-model-driven figure knowledge graph construction method and multi-modal large-model-driven figure knowledge graph construction system
Through the multimodal large model, a multimodal character knowledge graph is constructed, which solves the problems of insufficient information and noise interference in the traditional knowledge graph, and realizes multi-dimensional fine portrayal and high-accurate representation of character characteristics.
Patent Information
- Application Number
- CN202510744974.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Traditional knowledge graph construction methods mainly rely on text information, with insufficient information dimensions, semantic ambiguity and noise interference, resulting in low accuracy of entity and relationship extraction, making it difficult to fully portray the multidimensional characteristics of characters.
The multimodal big model-driven method is adopted to integrate text, image and video data, and fusion processing is carried out through multimodal big model to build a multimodal character knowledge graph, use visual information to assist in verifying the accuracy of text content, and fill in the knowledge graph framework with multimodal generated data.
It realizes multi-dimensional and meticulous portrayal of character characteristics, improves the comprehensiveness, accuracy and robustness of the knowledge graph, reduces entity recognition errors caused by text noise, and builds a multimodal knowledge graph with clear structure and intuitive display.
Smart Images

Figure CN120409657A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graph modeling, and particularly to a method and system for constructing a person knowledge graph driven by a multi-modal large model. Background Art
[0002] With the popularization of the Internet and mobile terminals, a large amount of multi-modal data such as text, images, and videos is continuously generated, providing rich resources for knowledge representation and management. As an important branch of knowledge graphs, person knowledge graphs not only have extensive applications in fields such as social networks, public opinion analysis, and digital humanities, but also face the challenge of how to comprehensively and accurately depict the multi-dimensional characteristics of people.
[0003] Traditional knowledge graph construction methods mainly rely on text information for entity extraction and relationship inference, but this single-modal data has problems such as insufficient information dimensions, semantic ambiguity, and noise interference. Although text can express a large amount of semantic information, it often cannot completely describe the multi-dimensional characteristics of a person's appearance, emotions, behaviors, and environment. For example, the description of a person may only involve their occupation, nationality, or part of their experiences, while ignoring their visual image and behavioral characteristics. In addition, noise such as spelling mistakes, implicit information, and colloquial expressions in text data will directly affect the extraction accuracy of entities and relationships, resulting in semantic ambiguity or ambiguity in knowledge graphs constructed from pure text.
[0004] In response to the above problems, the industry has not yet proposed a better technical solution. Summary of the Invention
[0005] This application provides a method, system, storage medium, computer program product, and electronic device for constructing a person knowledge graph driven by a multi-modal large model, so as to at least solve the problems of information limitation and semantic ambiguity existing in constructing knowledge graphs with single text in the current related technologies.
[0006] In a first aspect, an embodiment of the present application provides a method for constructing a multi-modal large model-driven personal knowledge graph, including: obtaining a target person's name and multiple personal attribute description information, and constructing a knowledge graph framework; the knowledge graph framework includes a core node and multiple attribute nodes respectively connected to the core node, the core node is used to indicate the target person's name, and each attribute node is respectively used to indicate the corresponding personal attribute description information; obtaining corresponding target multi-source material data according to the target person's name, the target multi-source material data includes text materials, image materials and video materials that match the target person's name; for each of the personal attribute description information, extracting at least one attribute matching material segment corresponding to the personal attribute description information from the target multi-source material data, and performing fusion processing on each attribute matching material segment corresponding to the personal attribute description information based on a multi-modal large model to output corresponding attribute multi-modal generated data; the multi-modal generated data includes text generated data, image generated data and video generated data; filling each of the attribute multi-modal generated data into the corresponding attribute node in the knowledge graph framework according to the personal attribute description information to construct a multi-modal personal knowledge graph.
[0007] In a second aspect, an embodiment of the present application provides a multi-modal large model-driven personal knowledge graph construction system, including: a graph framework construction unit, configured to obtain a target person's name and multiple personal attribute description information, and construct a knowledge graph framework; the knowledge graph framework includes a core node and multiple attribute nodes respectively connected to the core node, the core node is used to indicate the target person's name, and each attribute node is respectively used to indicate the corresponding personal attribute description information; a multi-source material acquisition unit, configured to obtain corresponding target multi-source material data according to the target person's name, the target multi-source material data includes text materials, image materials and video materials that match the target person's name; an attribute data generation unit, configured to, for each of the personal attribute description information, extract at least one attribute matching material segment corresponding to the personal attribute description information from the target multi-source material data, and perform fusion processing on each attribute matching material segment corresponding to the personal attribute description information based on a multi-modal large model to output corresponding attribute multi-modal generated data; the multi-modal generated data includes text generated data, image generated data and video generated data; a graph attribute filling unit, configured to fill each of the attribute multi-modal generated data into the corresponding attribute node in the knowledge graph framework according to the personal attribute description information to construct a multi-modal personal knowledge graph.
[0008] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the method for constructing a multi-modal large model-driven personal knowledge graph according to any embodiment of the present application.
[0009] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and is characterized in that when the program is executed by a processor, the steps of the method for constructing a multi-modal large model-driven personal knowledge graph according to any embodiment of the present application are implemented.
[0010] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method for constructing a multi-modal large model-driven personal knowledge graph according to any embodiment of the present application are implemented.
[0011] Through a method and system for constructing a multi-modal large model-driven personal knowledge graph provided by the present application, at least the following technical effects can be produced:
[0012] (1) By introducing multi-source material data such as text, images, and videos, and performing fusion processing on the multi-modal data, the generated personal attribute descriptions can not only include structured information such as occupation and nationality, but also cover unstructured multi-modal information such as appearance features, expression actions, and speech intonations, making the personal knowledge graph more complete, three-dimensional, and diverse at the knowledge representation level. In addition, by using a multi-modal large model to perform semantic alignment and in-depth understanding of personal attribute materials in different modalities, the accuracy of text content can be verified by visual information, effectively reducing entity recognition errors caused by spelling mistakes, spoken language expressions, semantic ambiguities, etc. in the text, and significantly improving the accuracy and stability of entity and relationship extraction in the knowledge graph.
[0013] (2) By adopting attribute nodes based on personal attribute description information and combining a multi-modal generation data filling mechanism, each attribute information can be independently and carefully displayed, thereby realizing multi-dimensional fine characterization of personal features, constructing a knowledge graph model framework with good structural hierarchy and expandability, and also supporting subsequent expansion of the personal knowledge graph and dynamic attribute management and update.
[0014] Through the technical solution, by the drive of a multi-modal large model, a high-quality personal knowledge graph generated by multi-modal data fusion is not only clear and rigorous in logical structure, but also more intuitive in display, significantly improving the comprehensiveness, accuracy, and robustness of knowledge graph construction, and providing strong support for in-depth mining and intelligent application of multi-modal personal information. Brief Description of the Drawings
[0015] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 Shows a flowchart of an example of a method for constructing a multi-modal large model-driven personal knowledge graph according to an embodiment of the present application;
[0017] Figure 2 Shows according to Figure 1 An operational flowchart diagram of an example of step S110 in;
[0018] Figure 3 Shows according to Figure 1 An operational flowchart diagram of an example of step S120 in;
[0019] Figure 4 Shows an operational flowchart diagram of an example of extracting attribute-matching material fragments from target multi-source material data according to an embodiment of the present application;
[0020] Figure 5 Shows a structural block diagram of an example of a multi-modal large model-driven personal knowledge graph construction system according to an embodiment of the present application;
[0021] Figure 6 Is a structural schematic diagram of an embodiment of an electronic device of the present application. Detailed Embodiments
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0023] In the technical solutions of the present application, for the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved, etc., they all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0024] Figure 1 Shows a flowchart of an example of a method for constructing a multi-modal large model-driven personal knowledge graph according to an embodiment of the present application.
[0025] Regarding the execution entity of the method in the embodiments of the present application, it can be any controller or processor with computing or processing capabilities. Specifically, it can be implemented by a knowledge graph management platform or a knowledge graph management system. Through the high-quality knowledge graph generated by multi-modal data fusion, especially by constructing a knowledge graph attribute framework and performing extraction and fusion processing on multi-modal attribute matching material segments for character attributes, the semantic enhancement of character attributes is realized, enabling the knowledge graph to intuitively display multi-dimensional information of corresponding text, images, and videos for different character attribute nodes, and meeting the needs of diverse data presentation in fields such as digital humanities and public opinion analysis.
[0026] In some examples, it can be integrated and configured in an electronic device or a terminal in a software, hardware, or software-hardware combination manner, and the types of terminals or electronic devices can be diverse, such as mobile phones, tablets, or desktop computers, etc.
[0027] As Figure 1 shown, in step S110, the target person's name and multiple person attribute description information are obtained, and a knowledge graph framework is constructed.
[0028] Here, the knowledge graph framework includes a core node and multiple attribute nodes respectively connected to the core node. The core node is used to indicate the target person's name, and each attribute node is respectively used to indicate the corresponding person attribute description information.
[0029] In some embodiments, by user input or system automatic recognition, the name of the target person (such as "Zhang San") is determined, and natural language processing (NLP) technology can be used to perform named entity recognition (NER) on the input text to ensure the accuracy of the extracted target person's name.
[0030] Regarding the person attribute description information, on the one hand, it can be pre-set person feature types (such as occupation, native place, hobbies, etc.), so the standardized attribute structure can be agreed upon through system pre-set information. On the other hand, the person attribute description information can also be closely related to the target person. For example, some people have personalized representative deeds (such as winning major scientific research achievement awards, etc.). Through these deeds, a deeper characterization of the person can be achieved. Therefore, external data sources (such as encyclopedias, news reports, etc.) can also be used to expand the attribute description information related to the target person, and relationship extraction technology and semantic analysis technology can be used to convert unstructured text into structured person attribute descriptions.
[0031] In terms of the construction details of the knowledge graph framework, it can create a core node through the use of a knowledge graph modeling tool (such as Neo4j) to store the name of the target person, and create an attribute node for each attribute description information, and then store the relationships between the nodes and edges. Thus, through the knowledge graph framework, the scattered personal attribute information is organized in a structured form to construct a personal knowledge graph skeleton with a hierarchical structure and clear semantics, realizing the unified expression and management of multi-dimensional attributes of the person, and providing a clear target direction for subsequent data filling and semantic fusion. In addition, the knowledge graph framework supports dynamically adding new attribute nodes, can adapt to different task requirements, and improves the scalability and maintainability of the knowledge graph.
[0032] In step S120, the corresponding target multi-source material data is obtained according to the name of the target person, and the target multi-source material data includes text materials, image materials and video materials that match the name of the target person.
[0033] In some embodiments, based on the name of the target person, search engines, social media APIs, content databases and other channels are automatically called to retrieve and collect multi-source data related to the person. The data obtained includes but is not limited to: news texts, social dynamics, images (such as photos, portraits), video clips (such as interviews, speeches, documentaries). Furthermore, a multi-modal data preprocessing module is introduced to clean, format-convert and denoise the text data; perform image standardization, size normalization and content recognition on the image data; perform key frame extraction and speech-to-text processing on the video materials. A unified material index is constructed and labeled according to dimensions such as material type, time, and source, facilitating subsequent retrieval and matching. Thus, the comprehensive collection of multi-source heterogeneous information of the target person is realized, avoiding the one-sidedness or incompleteness of the person information caused by a single data source.
[0034] In step S130, for each personal attribute description information, at least one attribute matching material segment corresponding to the personal attribute description information is extracted from the target multi-source material data, and each attribute matching material segment corresponding to the personal attribute description information is fused based on a multi-modal large model to output corresponding attribute multi-modal generated data, and the multi-modal generated data includes text generated data, image generated data and video generated data.
[0035] In some embodiments, for each personal attribute description information (such as "occupation: scientist"), rules are designed or a classification model is trained to extract segments related to the attribute from the target multi-source material data. For example, in terms of text segments, sentences containing keywords such as "science" and "research" are extracted, and in terms of image segments, pictures related to laboratory scenes, scientific research equipment, etc. are extracted, and in terms of video segments, speech segments or experimental operation pictures of the target person at academic conferences are extracted.
[0036] Furthermore, the extracted text fragments, image fragments, and video fragments are input into a multimodal large model (such as CLIP or BLIP-2) for cross-modal fusion. Specifically, in the contrastive learning model of CLIP, it can map images and text into a common embedding space, and through contrastive learning, make the distance between semantically similar image-text pairs closer in the embedding space, while the distance between dissimilar pairs is farther. In this embodiment, an improved model structure based on CLIP is adopted. For the extracted text fragments and image fragments (or the frame image sequence corresponding to the video fragment), the improved CLIP can perform cross-modal fusion well, learn the semantic associations between cross-modal data, and make up for the deficiencies of traditional text extraction methods in semantic understanding and context expression. More details about the improvement of the model will be elaborated in combination with other examples below.
[0037] Through the embodiments of this application, a pre-trained multimodal large model is used to fuse and understand text / image / video content. Multimodal fusion can complement the key information in each data source, making the expression of each attribute more comprehensive and the semantics more accurate, reducing the extraction error caused by modal bias, and enhancing the multimodal expression ability of character attribute information, enabling abstract attributes to have intuitive visual and behavioral expressions.
[0038] In step S140, according to the character attribute description information, each attribute multimodal generated data is filled into the corresponding attribute node in the knowledge graph framework to construct a multimodal character knowledge graph.
[0039] In some embodiments, each attribute multimodal generated data is mapped to the corresponding attribute node in the knowledge graph framework. Exemplarily, the multimodal generated data (text, image, video) of "Occupation: Scientist" is filled into the corresponding attribute node, and a knowledge graph visualization tool (such as Gephi or D3.js) is used to present the relationship between the core node and the attribute node in a graphical manner. Each attribute node is attached with a link to the corresponding multimodal generated data, so that users can view the specific content after clicking, improving the interaction experience. Thus, through the visually presented multimodal character knowledge graph, such as the Huxiang culture knowledge graph with modern and contemporary Huxiang figures as the core, the structure is logically clear, and each attribute node enriches the connotation of the character portrait, meeting the needs of refined analysis.
[0040] Through the embodiments of the present application, a structured knowledge graph framework is constructed, realizing the efficient organization and retrieval of target person and their attribute information; by using multi-source data collection and semantic analysis technology, the comprehensiveness and accurate matching of materials are ensured; with the help of multi-modal large models to fuse multi-source segments such as text, images, and videos, high-quality and semantically consistent multi-modal data is generated, and through accurate matching and intelligent fusion, fine-grained extraction and semantic enhancement of each attribute of the person are realized; at the same time, through targeted retrieval and filling of the attribute nodes of the knowledge graph, a multi-modal person knowledge graph that is both structured and dynamically updated is accurately constructed, supporting dynamic updates and visual displays, significantly improving the timeliness of information and the user interaction experience.
[0041] Figure 2 shows an operation flow diagram of an example according to Figure 1 step S110 in.
[0042] As Figure 2 shown, in step S210, user input information is obtained, and the target person name is determined according to the user input information.
[0043] In some embodiments, the target person name is determined through user input information. In addition, to improve accuracy, the input name can also be standardized (such as disambiguation, format unification, etc.) to eliminate the potential confusion risks brought by aliases, abbreviations, or people with the same name.
[0044] In step S220, a person event recognition prompt word is constructed according to the target person name and a preset prompt word template, and the person event recognition prompt word is input into the large language model to trigger the large language model to output at least one person event keyword for the target person name.
[0045] Specifically, one or more prompt word templates for identifying person events are preset in the system to guide the large language model to generate keywords. Examples of the prompt word templates are as follows:
[0046] "Please list the important events related to [target person name]."
[0047] "What are the main achievements of [target person name]?"
[0048] "Describe the contributions of [target person name] in [field / industry]."
[0049] Furthermore, the target person's name is dynamically inserted into the prompt template to generate personalized prompts. For example, when "Zhang San" is input, the prompt "Please list the important events related to Zhang San." is generated. The generated prompt is input into a large language model (such as the GPT or Qwen series) to trigger its output of the key event keywords related to the target person. Exemplarily, taking the representative figure of Huxiang, "Yuan Longping", as an input example of the target person's name, keywords such as "hybrid rice", "three-line method", and "dream of having cool breezes under the rice plants" may be generated. Thus, through the preset prompt template and dynamic filling mechanism, high-quality event recognition prompts are quickly generated, automatically triggering the large language model to output event keywords related to the target person, enabling more accurate capture of the key events related to the person, and thereby determining more fine-grained and representative person attribute description information.
[0050] It should be noted that during the call process, configuration parameters such as the temperature parameter and response length of the large language model can be set to ensure that the large language model can both discover details and ensure the accuracy of the output. With the semantic reasoning ability of the large language model, diverse person event keywords can be extracted from a wide range of knowledge bases, effectively extracting the important event keywords related to the target person, which not only expands the person's background information but also provides high-quality event clues for subsequent attribute mapping.
[0051] In step S230, according to each person event keyword, the person attribute description information is determined.
[0052] In some embodiments, the person attribute description information can be directly defined according to the person event keywords. For example, "hybrid rice", "three-line method", and "dream of having cool breezes under the rice plants" are respectively used as the information related to the corresponding person attributes to trigger the generation of exclusive attribute nodes corresponding to major person events in the subsequent person knowledge graph, realizing the generation of personalized person knowledge graphs corresponding to different person names.
[0053] Regarding the implementation details of step S230, in some examples of the embodiments of the present application, corresponding customized person attribute description information is determined according to each person event keyword, and then the person attribute description information is determined according to at least one preset general person attribute description information and each customized person attribute description information.
[0054] Through the embodiments of the present application, customized person attribute description information is generated for each event keyword. The system can capture the specific attribute characteristics of the target person in different event backgrounds, reflecting personalized details. Moreover, the customized description is based on specific event keywords, which is closer to the actual context of the target person, improving the accuracy of semantic matching. In addition, by combining these customized information with the preset general attribute description information, standardized description is achieved, effectively balancing personalization and universality, and ensuring that the knowledge graph has a high generalization ability while expressing at a fine-grained level.
[0055] Figure 3 shows an operation flow schematic diagram of an example according to Figure 1 step S120 in
[0056] In step S310, the target person name is semantically matched with the material description information of each multi-source material in the multi-source material library to screen at least one initially matched multi-source material from the multi-source material library.
[0057] In some embodiments, the description information (including text title, tags, abstract, etc.) of each material in the multi-source material library is segmented, stop words are removed, and format normalization is performed, and the text information is converted into a high-dimensional semantic vector by using a pre-trained language model (such as BERT, etc.). Similarly, the obtained and standardized target person name is subjected to the same semantic encoding to generate a semantic vector of the target name. Furthermore, by using a cosine similarity or other distance metric method, the similarity score between the target person name vector and each material description vector is calculated. According to the set similarity threshold, materials with higher similarity are screened to ensure that at least one initially matched material is obtained, and at the same time, multiple candidate materials are sorted in descending order of similarity, which can greatly improve the accuracy and recall rate of the preliminary screening.
[0058] In step S320, based on the data noise scoring network, the data noise scores of each initially matched multi-source material are determined. The data noise scoring network includes a text data noise evaluation branch, an image data noise evaluation branch, and a video data noise evaluation branch connected in parallel.
[0059] In the multi-modal data noise scoring network, the text data noise evaluation branch uses a noise detection model to judge whether there are grammar errors, spelling mistakes, or information redundancy in the text content. The image data noise evaluation branch uses an image quality evaluation algorithm (such as quality evaluation based on a convolutional neural network) to detect problems such as image clarity, noise, and distortion. The video data noise evaluation branch evaluates whether there are noise factors such as blurring, frame loss, and sound interference in the video through key frame extraction and video quality detection algorithms. Here, each modal scoring branch independently scores the data of the corresponding modality, outputs the noise scores of text, image, and video, and aggregates them into a comprehensive noise score. Thus, through the parallel noise scoring network, the system can perform a detailed and quantitative noise evaluation on multi-modal materials.
[0060] Each branch in the data noise scoring network can be trained using a self-supervised contrast learning strategy. By constructing contrast samples of clean data and noisy data, the network automatically learns the feature differences between high-quality and low-quality data, thereby improving the scoring accuracy.
[0061] In step S330, the initial matching multi-source materials with corresponding data noise scores not exceeding the preset noise score threshold are determined as the first initial matching materials, and the initial matching multi-source materials with corresponding data noise scores exceeding the noise score threshold are determined as the second initial matching multi-source materials.
[0062] In some embodiments, for the noise scores independently output for each branch, the text, image, and video materials are screened respectively according to the preset noise level threshold. Through the clear noise threshold and classification mechanism, the system can automatically distinguish high-quality and low-quality materials, ensure that only high signal-to-noise ratio data directly participates in the subsequent construction, and at the same time process the low-quality data separately, improving the overall accuracy and reliability of the final multi-source material data.
[0063] It should be noted that in some application scenarios, the noise scores output by each branch can be weighted and fused to obtain a comprehensive multi-modal noise score. However, in this embodiment, it is more recommended to keep the modalities independent in order to adopt customized reconstruction strategies for different materials.
[0064] In step S340, the second initial matching multi-source materials are input into the material reconstruction network to generate reconstructed multi-source materials. The material reconstruction network includes parallel text data reconstruction branches, image data reconstruction branches, and video data reconstruction branches.
[0065] Here, the text reconstruction branch uses a text correction and generation model to correct spelling, optimize grammar, complete content, or generate summaries for the original text, thereby generating text data with clearer semantics and more standardized expressions. The image reconstruction branch uses image enhancement technologies, such as super-resolution reconstruction, denoising processing, and image repair algorithms, to process low-quality images and restore or enhance image details and clarity. The video reconstruction branch combines technologies such as key frame extraction, video denoising, picture quality enhancement, and speech recognition optimization to reconstruct low-quality videos and output video data with smooth pictures and clear content.
[0066] Specifically, each branch in the material reconstruction network independently reconstructs the low-quality material data and outputs the reconstructed text, image, and video data. Preferably, the noise scores of the reconstructed data can be evaluated again to ensure that the processed data meets the preset quality standards, and only the reconstructed data with noise scores not exceeding the threshold is selected as the effective output.
[0067] The material reconstruction network can effectively repair and improve the data quality of low-quality materials. By specifically reconstructing text, image, and video data, the noise interference is reduced, ensuring that even if the original data has noise problems, clear and accurate information can still be provided after reconstruction, providing high-quality support for knowledge graph construction.
[0068] In step S350, based on the first initial matching material and the target reconstructed multi-source material, the target multi-source material data is determined, and the target reconstructed multi-source material is the reconstructed multi-source material whose corresponding noise score does not exceed the noise score threshold.
[0069] Specifically, on the one hand, the first initial matching materials that meet the noise score requirements are directly incorporated into the target multi-source material data; on the other hand, after the second initial matching materials are processed by the reconstruction network, the noise scores are evaluated again, and the reconstructed data with scores not exceeding the noise threshold is selected as the target reconstructed multi-source material. Thus, by integrating high-quality original data and data improved through reconstruction, the finally generated target multi-source material data has high signal-to-noise ratio, integrity, and diversity, ensuring that subsequent multi-modal information fusion and knowledge graph construction can rely on accurate, rich, and clear data resources to effectively support the refined analysis and expression of personal information.
[0070] Regarding the specific details of the data noise scoring network, in some examples of the embodiments of the present application, the text data noise evaluation branch is used to extract the global semantic features of the text in the initial matching multi-source materials through the BERT model, map the extracted text features through the fully connected layer, and normalize them using the Sigmoid activation function, thereby obtaining the corresponding text noise score:
[0071] h T = BERT(T), Equation (1)
[0072] s T = σ(W T ·h T + b T ), Equation (2)
[0073] In the formula, h T represents the feature vector obtained by the text T passing through the BERT model, s T ∈[0,1] represents the text noise score, W T and b T respectively represent the weight matrix of the fully connected layer and the bias term of the fully connected layer of the text scoring branch; σ is the Sigmoid activation function, which is used to ensure that the output is normalized to [0,1].
[0074] It should be understood that the higher the score value, the more serious the noise, and the lower the score value, the higher the data quality.
[0075] In the text data noise evaluation branch, a pre-trained language model (such as BERT) is used to extract the global semantic features of the text material, and the [CLS] vector of BERT is used as the text representation. The extracted text features are mapped to a scalar noise score through a fully connected layer and normalized using the Sigmoid activation function, thereby obtaining the text noise score. Thus, it can effectively capture spelling, grammar errors, and ambiguity problems in the text, so as to accurately judge the text data during noise filtering.
[0076] The image data noise evaluation branch is used to extract the deep features of the images in the initial matching multi-source materials through a convolutional neural network. The feature maps of the convolutional layer are pooled into a fixed-length feature vector using global average pooling, and the pooled image features are mapped and normalized using the Sigmoid function through a fully connected layer, thereby obtaining the image noise score:
[0077] F I = CNN(I), Equation (3)
[0078] f I = GAP(F I ), Equation (4)
[0079] s I = σ(W I ·f I + b I ), Equation (5)
[0080] Where, F I represents the feature map of the image I extracted by the convolutional network, f I represents the image feature vector after global average pooling; CNN represents the convolutional neural network, GAP represents global average pooling; s I ∈ [0,1] represents the image noise score, W I and b I represent the weight matrix of the fully connected layer and the bias term of the fully connected layer of the image scoring branch, respectively.
[0081] In the image data noise evaluation branch, a pre-trained convolutional neural network is used to extract the deep features of the image, and global average pooling (Global Average Pooling, GAP) is used to pool the feature maps of the convolutional layer into a fixed-length feature vector to achieve global feature pooling. The pooled image features are mapped to a noise score through a fully connected layer and output as the image noise score through Sigmoid normalization. Thus, problems such as blurring and low resolution in the image can be accurately identified.
[0082] The video data noise evaluation branch is used to extract at least one first key frame of the video in the initially matched multi-source materials, and input them into the image data noise evaluation branch respectively to calculate the corresponding key frame noise scores, and aggregate the scores of all the first key frames by using the weighted average method, so as to obtain the video noise score:
[0083]
[0084] In the formula, f vj represents the global image feature vector of the j-th first key frame of the video, s vj ∈[0,1] represents the noise score of the j-th first key frame calculated by calling the image data noise evaluation branch, N v represents the total number of the first key frames extracted from the video, is a trainable parameter vector for calculating the importance of the first key frame, represents the transposed vector of, represents the sum of the importance score exponential values of N v first key frames in the video, a j represents the importance weight of the j-th first key frame, satisfying s v ∈[0,1] represents the video noise score.
[0085] In the video data noise evaluation branch, several representative frames are selected from the video by using the key frame extraction algorithm The image data noise scoring branch is applied to each key frame to independently calculate the noise score s vj of each frame. By using the weighted average method to temporally aggregate the scores of all key frames, the overall video noise score s v is obtained. Thus, through the temporal aggregation method, the frames with higher noise in the video get lower weights, so as to obtain a stable noise score reflecting the overall video quality, providing a basis for the reconstruction and quality control of video data. In addition, directly reusing the image data noise scoring branch to score each key frame effectively reduces the complexity and model data volume of the video data noise evaluation branch, and also reduces the demand for system operation resources.
[0086] Through the design of the above data noise scoring network, the cross-modal, independent, and accurate noise assessment of text, image, and video data is achieved. Specifically, specially designed branches are adopted for different data modalities, effectively capturing grammar errors and ambiguities in text, blurry or low-quality features in images, and dynamic changes in key frames in videos. Thus, through the cross-modal noise assessment method, the actual quality of each data type can be more accurately reflected. In addition, the attention weight mechanism introduced in the video branch ensures that the importance of each key frame can be adaptively adjusted, thereby obtaining a more stable and objective overall video noise score. This mechanism performs particularly well in complex video scenarios.
[0087] Regarding the implementation details of the material reconstruction network, in some examples of the embodiments of the present application, the text data reconstruction branch is used to adopt a pre-trained generation model as a conditional generator to generate a reconstructed text according to the input noisy text by maximizing the conditional generation probability:
[0088]
[0089] where T′ is the input noisy text, representing the text data after preprocessing; T cand is the candidate generated text, representing the text sequence that the generator may output; represents the finally generated reconstructed text, aiming to eliminate the noise errors in the original text and restore the semantics; P(T cand ∣T′; θ T ) is the conditional generation probability calculated by the pre-trained generation model, which represents the probability of generating text T T under the model parameters θ cand with the condition that the input is T′.
[0090] In the text data reconstruction branch, the pre-trained generation model can adopt the T5 (Text-to-Text Transfer Transformer) model as the generator to achieve text repair through conditional generation, so that the output text has both high grammar accuracy and can restore key information. In addition, semantic consistency detection is used to ensure the consistency of the reconstructed text and the original text in key information, thereby avoiding information deviation caused by over-generation. Thus, the finally output reconstructed text has higher accuracy and clarity, and can effectively remove spelling mistakes, grammar problems, and ambiguous expressions.
[0091] The image data reconstruction branch is used to adopt an image reconstruction generator based on a generative adversarial network to reconstruct the input noisy image to generate a reconstructed image:
[0092]
[0093] In the formula, I′ represents the input noisy image, which indicates the image data after preprocessing; G(·; θ I ) is the image reconstruction generator, and its model parameters are θ I ; represents the finally generated reconstructed image, aiming to repair the blurring and noise in the original image; L I represents the total loss of the image reconstruction generator, including the reconstruction loss and the adversarial loss; represents the pixel-level L2 loss, which is used to measure the reconstructed image and the target reference image I ref The gap between; I ref represents the reference image used for supervised training, and L adv represents the adversarial loss output by the discriminator of the generative adversarial network; λ adv is the weight coefficient of the adversarial loss, which is used to balance the influence between the reconstruction loss and the adversarial loss.
[0094] In the image data reconstruction branch, an image generation network (such as a generative adversarial network) is used to reconstruct the low-quality image. By taking advantage of the generative model, the low-quality input is repaired by learning the distribution of a large number of high-quality images, the problems such as blurring, low contrast or noise in the image are analyzed, the details of the image are automatically restored, and a clear and detail-rich reconstructed image is generated.
[0095] In addition, a pixel-level reconstruction term and an adversarial term (λ adv L adv ) are simultaneously introduced into the loss function of the image data reconstruction branch. The reconstruction term constrains the reconstructed image and the reference image I ref The consistency at the pixel level ensures the accurate restoration of the global structure and details; the adversarial term urges the generator to produce more visually realistic images through the discriminator of the adversarial network, thus effectively avoiding the "over-smoothing" of the image or the lack of texture details caused by simple pixel-level reconstruction. Thus, during the image generation or reconstruction process, both the main structure of the reference image can be accurately restored, and the visual details and realism can be improved through the adversarial learning mechanism.
[0096] The video data reconstruction branch is used to extract at least one second key frame from the input noisy video and input them into the image reconstruction branch respectively to generate corresponding reconstructed key frames, and use temporal modeling to smooth and process the temporal consistency of the reconstructed key frames to generate a reconstructed video:
[0097]
[0098] In the formula, represents I iThe reconstructed key frames generated after passing through the image reconstructor, I i ′ represents the i-th second key frame in the input noisy video, and M represents the total number of second key frames extracted from the input noisy video; h i and h i-1 respectively represent the hidden state representations of the i-th second key frame and the (i - 1)-th second key frame in the LSTM time series model, θ V represents the model parameters of the LSTM time series model; is the finally output reconstructed video, which represents the reconstructed video frame sequence.
[0099] In the video data reconstruction branch, each key frame is separately reconstructed using the method in the image reconstruction branch to ensure the improvement of the image quality of each frame. At the same time, through the reuse of the model structure, the complexity and model data volume of the video data reconstruction branch are reduced. In addition, due to the requirement of temporal continuity of the video, after reconstruction, the LSTM time series model is also used to further process all key frames to smooth the inter-frame transition and correct the frame skipping problem caused by motion or changes, ensuring the overall coherence of the video and conforming to the temporal characteristics of real videos.
[0100] Through the embodiments of the present application, by using the collaborative work of each branch in the material reconstruction network, specialized reconstruction is carried out for the noise characteristics of different modalities (text, image, video), so that the output data meets high standards in terms of grammar, vision, and timing. In addition, through modality-specific reconstruction, each branch adopts a targeted reconstruction strategy, realizing the dynamic adaptive repair of different types of data noise.
[0101] It should be noted that in the video branches of the above data noise scoring network and material reconstruction network, the extraction of key frames in the video is involved, and the types of key frame extraction algorithms suitable for them can be diverse. In some examples of the embodiments of the present application, the key frame extraction algorithm can be a combination of the preliminary screening based on the difference between optical flow and color histogram and the video depth semantic analysis to realize the recognition and extraction of key frames.
[0102] Specifically, in the preliminary screening based on the difference between optical flow and color histogram, in order to capture the frames with large scene or motion changes in the video, the difference detection of optical flow and color histogram is first performed on consecutive frames.
[0103] 1) Optical flow change detection
[0104] For consecutive frames I t and in the video, calculate the optical flow field OF t , and define the average optical flow amplitude:
[0105]
[0106] Where M t represents the average optical flow amplitude between frame t and frame t-1, which is used to measure the degree of motion change, and N pix represents the total number of pixels in the frame, and OF t (i) represents the optical flow vector of the i-th pixel between frame t and frame t-1.
[0107] When M t exceeds the preset threshold τ flow , it is considered that there is an obvious change in frame t, and it can be used as one of the candidate key frames.
[0108] 2) Color histogram difference
[0109] Calculate the color histograms H(I t ) and H(I t-1 ) of frame t and t-1, and use the chi-square distance d t to measure the difference:
[0110] d t =χ 2 (H(I t ), H(I t-1 )) , Equation (14)
[0111] Where d t represents the color histogram difference of frame t; χ 2 represents the chi-square distance function, and the larger the output value, the more obvious the change in color distribution.
[0112] If d t >τ hist , then frame t can also be marked as one of the candidate key frames.
[0113] Here, through the preliminary screening of optical flow and color histogram, the number of frames to be processed is reduced, thereby improving the efficiency of subsequent depth feature extraction and clustering.
[0114] In video depth semantic analysis, depth features are extracted from all candidate key frames in the video to obtain higher-level semantic depth features. According to the semantic depth features of the candidate frames, clustering is performed to distinguish different content patterns in the video, and then the most representative frame is selected from each cluster as the key frame.
[0115] Specifically, a pre-trained convolutional neural network is used to extract the depth features of each frame:
[0116] r t =CNN(I t ) , Equation (15)
[0117] Where r t represents the depth feature vector of frame t.
[0118] The features of all candidate frames are clustered using the K-means algorithm. Let the number of clusters be Z, and each cluster C z has a center as follows:
[0119]
[0120] where μ z represents the center vector of the z-th cluster, and |C z | represents the number of frames in cluster z.
[0121] In each cluster C z , the frame closest to the cluster center is selected as the representative key frame of the cluster:
[0122]
[0123] where t z represents the key frame index selected in cluster z.
[0124] After K-means clustering, the frames selected from Z clusters are used as the key frame set.
[0125] Here, by using deep feature extraction and K-means clustering, the representative information of different visual contents in the video can be accurately captured from the candidate frames, ensuring that the selected key frames can comprehensively reflect the main changes in the video.
[0126] Through the embodiments of the present application, the key frames are defined as the frames of the main visual changes and contents in the video. By combining optical flow, color histogram, depth features, and temporal attention mechanism, the stability and adaptability of key frame extraction in various video scenarios are improved. In addition, by dynamically setting the number of clusters Z and the temporal attention weights, the algorithm can adaptively adjust the number of key frames and the selection criteria according to the complexity of the video content, further improving the flexibility and data representativeness of the overall system.
[0127] Figure 4 FIG. shows an operation flow diagram of an example of extracting an attribute-matching material segment from target multi-source material data according to an embodiment of the present application.
[0128] As Figure 4 shown, in step S410, an attribute label group is determined according to each person's attribute description information, and a label classifier for the attribute label group is constructed. Each attribute label is used to indicate the corresponding person's attribute description information.
[0129] In some embodiments, based on the acquired person attribute description information (such as occupation, interests, events, etc.), historical data and domain knowledge are statistically analyzed to automatically extract key attribute words, forming a preliminary list of attribute tags, such as "Occupation: Scientist", "Nationality: China", "Achievement: Received the Nobel Prize", generating a set of semantic attribute tags.
[0130] Furthermore, machine learning or deep learning methods are used to construct a label classifier for the group of attribute tags. Specifically, material fragments containing relevant attribute descriptions are extracted from existing data, and corresponding attribute tags are annotated for them, and a suitable classification model (such as BERT, TextCNN or SVM) is selected as the basic architecture of the label classifier. Then, the labeled data set is used to train the label classifier so that it can accurately match the corresponding attribute tags according to the input semantic description.
[0131] In step S420, the target multi-source material data is semantically segmented to obtain multiple material candidate fragments, and each material candidate fragment has a corresponding semantic description of the material fragment.
[0132] In some embodiments, text segmentation can adopt sentence boundary detection and semantic topic segmentation techniques to divide long texts so that each paragraph or sentence has relatively independent semantic information. Image region segmentation can use object detection and image region segmentation algorithms to segment different regions (such as the human face, clothing, scene background) in an image to ensure that each region corresponds to an independent semantic description. Video scene segmentation segments the video through key frame extraction and scene change detection, divides the part with unified semantic content in continuous shots into independent segments, and at the same time extracts the corresponding audio text transcription information.
[0133] Regarding the details of generating the semantic description of the material candidate fragments, for each text fragment, a pre-trained language model can be used to generate a summary to extract the core semantics; image recognition technology can be used to generate corresponding image tags or descriptions; by combining the key frame images and the audio transcription results, a comprehensive semantic description of the video fragment can be generated through multi-modal fusion.
[0134] In step S430, based on the label classifier, the attribute tags matched by the semantic descriptions of each material fragment are determined, thereby obtaining the attribute matching material fragments of the corresponding person attribute description information.
[0135] In some embodiments, the semantic description vectors of each candidate material segment are input into a label classifier, such that the label classifier outputs a corresponding matching label probability distribution for each material segment. Here, it is allowed that one candidate segment matches multiple attribute labels, which fully reflects the possible multiple person attribute information in the material. Furthermore, threshold screening is performed on the probability of each attribute label output by the label classifier to determine which candidate segments meet the matching requirements of specific attributes. Thus, based on the automatic matching of the label classifier, the efficiency and accuracy of accurately screening content related to person attributes from a large number of candidate materials are significantly improved. In addition, through multi-label matching, multi-dimensional attribute information in one candidate segment can be captured simultaneously, ensuring the comprehensiveness and fine-grainedness of the person attribute description in the knowledge graph.
[0136] Through the embodiments of the present application, standard or customized person attribute description information is regarded as independent labels to form a label set. Each label represents a person attribute (such as "appearance", "occupation", "hobbies", "experiences", etc.). Based on this label set, a label classifier is constructed by using a pre-trained model or a deep neural network, and this classifier is used to determine whether an arbitrarily input material segment belongs to a certain attribute label. By converting the attribute matching task into label classification of people, a clear corresponding relationship is established between the person attribute description information and the material segment, thereby realizing automatic attribute matching.
[0137] Thus, by using the label classifier to take the person attribute description information as independent labels, it is possible to accurately identify segments highly relevant to each attribute in multi-source materials, realize fine-grained attribute matching, and significantly improve the accuracy and comprehensiveness of the attribute information in the knowledge graph.
[0138] In some examples of the embodiments of the present application, the label classifier includes a cascaded projection layer, a label hint fusion layer, and a label matching layer.
[0139] The projection layer is used to map the semantic description of the material segment to an intermediate representation space:
[0140] h1 = ReLU(W1x + b1), Equation (18)
[0141] In the formula, x represents the embedded representation of the semantic description of the material segment, W1 and b1 respectively represent the weight matrix and bias term of the projection layer, h1 represents the projection intermediate representation, and ReLU represents the ReLU activation function.
[0142] In the projection layer, the system uses linear transformation and ReLU activation to map the semantic description of the candidate material segment to a feature space more suitable for subsequent processing, ensuring that the input data is normalized to a certain dimension before entering the next layer, and strengthening the attention to "useful" features while suppressing irrelevant or noisy features.
[0143] The label hint fusion layer is used to encode the semantic information of each attribute label into a hint vector for feature alignment with the semantic description of the material segment in the semantic space:
[0144] h 2,c = ReLU(h1 + P c ), Equation (19)
[0145]
[0146] In the formula, l c represents the c-th attribute label, P represents the label hint matrix corresponding to the attribute label group, and P c represents the learnable hint vector corresponding to l c , h 2,c represents the feature representation after the fusion of the semantic description of the material segment and l c , n represents the total number of attribute labels in the attribute label group, and H2 represents the label fusion feature matrix.
[0147] In the label hint fusion layer, the system adds the intermediate representation of the candidate material segment to the hint vector of a certain character attribute label, and obtains the fusion representation after passing through the ReLU activation again. Reactivating again can be regarded as the fusion process of "label semantics" and "segment semantics". In addition, the fused vectors are stacked row by row or column by column to form a matrix containing the fusion representations of multiple character attribute labels, providing structured data support for multi-label joint prediction.
[0148] It should be noted that different from a simple single-layer linear classifier, in the embodiments of this application, by using two ReLU activations and the fusion of the label hint vector and the candidate segment features, the network's semantic capture ability for character attribute labels is significantly enhanced. While capturing the segment features, customized feature mining is performed on each label with the help of the label hint vector, greatly improving the fineness of attribute matching. In addition, through the learnable hint vector, the semantic information of each label can be directly aligned with the segment in the feature space, and it can quickly adapt to new attribute labels with a small amount of annotation or hint tuning, having stronger scalability and domain adaptability.
[0149] Based on the label matching layer, calculate the matching scores corresponding to each attribute label:
[0150]
[0151] In the formula, represents the transposed vector of the shared scoring weight vector, b 2,c represents the bias term of the attribute label l c , z c represents the semantic description of the material segment in l cThe unnormalized matching scores on
[0152] After obtaining the fused representation, through linear transformation and bias terms, the unnormalized matching scores between the candidate material segments and each label are calculated. The scoring vector (or matrix) maps the fused representation to a distinguishable "matching score" space, indicating the degree of correlation between the current segment and each attribute label.
[0153] Based on the label matching layer, each matching score is processed through the Sigmoid activation to obtain the matching probability of the semantic description of the material segment for each corresponding attribute label:
[0154] p c = σ(z c ), Equation (22)
[0155] In the formula, σ is the Sigmoid activation function, and p c represents the matching probability between the semantic description of the material segment and l c .
[0156] Here, through the Sigmoid function, the matching scores are normalized to the interval [0, 1]. The output value can be regarded as the confidence that "the candidate segment contains this attribute label", enabling the calculation of probabilities for each label separately during multi-label prediction and allowing multiple labels to be activated simultaneously.
[0157] According to the material candidate segments whose corresponding matching probabilities exceed the preset probability threshold, the attribute matching material segments of the corresponding attribute labels are determined.
[0158] Through the embodiments of the present application, compared with the traditional shallow methods that only rely on keyword matching or single modality, "segment features" and "label hints" are deeply fused in a unified semantic space, enabling more refined and accurate identification of person attribute information.
[0159] Regarding the implementation details of the multi-modal large model in step S130, in some embodiments, the multi-modal large model can adopt a pre-trained multi-modal model (e.g., the CLIP model), and this multi-modal model can be fine-tuned to improve the accuracy and quality of the construction of the multi-modal person knowledge graph.
[0160] More specifically, the CLIP model itself has been pre-trained on a large scale and has strong generation or reconstruction capabilities. Here, no additional fine-tuning is performed on its generation capabilities, but more focus is placed on optimizing the fusion of multi-modal information and cross-modal consistency.
[0161] Specifically, the fine-tuning training for the multi-modal large model includes:
[0162] Define the training dataset:
[0163]
[0164] In the formula, represents the training data set, and N s represents the total number of samples in the training data set, and u represents the u-th specific sample in the data set; is the text segment corresponding to the u-th sample, representing the text description segment of the specific attributes of the person; is the image segment corresponding to the u-th sample, representing the image material segment of the specific attributes of the person; is the video segment corresponding to the u-th sample, representing the video material segment of the specific attributes of the person; y (u) represents the true label of the person attributes corresponding to the u-th sample.
[0165] Obtain the embedding representations of each modality using a pre-trained multi-modal large model:
[0166]
[0167] In the formula, f T (·), f I (·) and f v (·) represent the text encoder function, the image encoder function, and the video encoder function respectively, and are the embedding representations corresponding to the u-th sample text segment, the sample image segment, and the sample video segment respectively.
[0168] Then, through the cross-modal fusion module, obtain the unified feature representation of the cross-modal:
[0169]
[0170] In the formula, represents the cross-modal fusion feature corresponding to the u-th sample, and CrossAttention(·) represents the cross-modal attention mechanism.
[0171] Here, through the cross-modal attention mechanism (CrossAttention), flexibly capture the correlation and semantic complementary relationship of the features of each modality, fuse the embedding features of the three different modalities of text, image, and video, and obtain a unified semantic feature representation. Thus, realize the efficient integration of text, image, and video information, and obtain a more accurate and comprehensive multi-modal feature representation.
[0172] Calculate the cross-modal consistency loss, including:
[0173] First, map the cross-modal fusion feature into the independent embedding spaces of the text, image, and video modalities respectively, denoted as:
[0174]
[0175] In the formula, g T (·), g I (·), g v (·) respectively represent the trainable network layers that map the fused cross-modal features to the text modality space, the image modality space, and the video modality space; and respectively represent the reconstructed text embedding vector, the reconstructed image embedding vector, and the reconstructed video embedding vector.
[0176] Then, using the real modality features as the supervision signal, the consistency between the reconstructed features and the actual features is constrained:
[0177]
[0178] In the formula, represents the cross-modal consistency loss, which is used to ensure the consistency between the generated data and the real data.
[0179] Here, the core idea of the cross-modal consistency loss is to constrain the generated multi-modal features (reconstructed features) to be similar to the real single-modal features, thereby significantly improving the semantic unity of the cross-modal generated content and ensuring that the fused features can still maintain consistency when mapped back to the text, image, and video modality spaces respectively.
[0180] Specifically, in the training process combined with the above formula (27), using the real modality features as the supervision signal, the fused features are pushed to be as close as possible to the real features when mapped to a single modality space. By minimizing the distance between the modality features, the model continuously optimizes the cross-modal consistency, making the generated results more coordinated and consistent in the semantic space. Thereby, the semantic relevance and consistency between the multi-modal generated data (text, image, video) are significantly improved, effectively avoiding semantic fragmentation or deviation in the generated content of each modality, and improving the overall quality of the generated data.
[0181] Calculate the multi-modal contrast loss, including:
[0182] In each training batch, select the text-image pairs and text-video pairs with the same attributes as the positive samples, and select the segments with different attributes as the negative samples, and then optimize the cross-modal matching degree by calculating the cosine similarity between the embedding representations:
[0183]
[0184] In the formula, represents the multi-modal contrast loss, sim(·) represents the vector similarity function, represents the vector similarity of the u-th text-image positive sample pair, represents the vector similarity of the u-th text-video positive sample pair, represents the sum of similarities between the u-th text segment and other image segments that do not correspond to the person’s attributes, It represents the sum of similarities between the u-th text segment and other video segments without corresponding person attributes.
[0185] The core of the multimodal contrastive loss design is to narrow the gap between data with the same attributes across different modalities, while simultaneously widening the gap between data with different attributes across modalities. Leveraging contrastive learning principles, training with text-image and text-video pairs as positive and negative examples enables the model to learn a high-quality cross-modal aligned feature space.
[0186] In the training process combined with the above formula (28), each training batch contains a large number of positive and negative sample pairs. By calculating the cosine similarity, the optimization of the distance between positive and negative samples is supervised. The contrastive learning mechanism is used to ensure high-quality alignment of feature embeddings in the cross-modal space, providing a solid and reliable basic feature space for the generation and use of cross-modal fusion features.
[0187] According to the cross-modal consistency loss and multimodal contrast loss, the total loss function is calculated:
[0188]
[0189] Where, represents the total loss function, α and β are hyperparameter weights used to control the relative importance of cross-modal consistency loss and multimodal contrast loss.
[0190] Through the comprehensive loss function provided in the embodiment of the present application, the two key requirements of multimodal consistency and multimodal feature alignment are taken into account, and the effects of cross-modal fusion and consistency generation are comprehensively optimized. In addition, the balance between losses is adjusted by weight parameters α, β, so that the training process is more targeted. For example, by setting α = 0.6, β = 0.4, the importance of cross-modal consistency is highlighted while maintaining the cross-modal feature alignment effect. Thus, it is effectively ensured that the modal data generated after cross-modal feature fusion are semantically unified in content and verify each other, avoiding information mismatch between modalities, improving the quality and consistency of multimodal information fusion, achieving higher quality, multimodally consistent generated data, and improving the accuracy and robustness of the final character knowledge graph.
[0191] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of combined actions. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0192] Figure 5 FIG. 4 shows a structural block diagram of an example of a multi-modal large model-driven character knowledge graph construction system according to an embodiment of the present application.
[0193] As Figure 5 shown, the multi-modal large model-driven character knowledge graph construction system 500 includes a graph framework construction unit 510, a multi-source material acquisition unit 520, an attribute data generation unit 530, and a graph attribute filling unit 540.
[0194] The graph framework construction unit 510 is configured to obtain a target character name and multiple character attribute description information, and construct a knowledge graph framework; the knowledge graph framework includes a core node and multiple attribute nodes respectively connected to the core node, the core node is used to indicate the target character name, and each attribute node is respectively used to indicate the corresponding character attribute description information.
[0195] The multi-source material acquisition unit 520 is configured to obtain corresponding target multi-source material data according to the target character name, and the target multi-source material data includes text materials, image materials, and video materials that match the target character name.
[0196] The attribute data generation unit 530 is configured to, for each of the character attribute description information, extract at least one attribute matching material segment corresponding to the character attribute description information from the target multi-source material data, and perform a fusion process on the respective attribute matching material segments corresponding to the character attribute description information based on a multi-modal large model to output corresponding attribute multi-modal generation data; the multi-modal generation data includes text generation data, image generation data, and video generation data.
[0197] The graph attribute filling unit 540 is configured to fill each of the attribute multi-modal generation data into the corresponding attribute nodes in the knowledge graph framework according to the character attribute description information to construct a multi-modal character knowledge graph.
[0198] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the steps of any one of the above-mentioned multimodal large model-driven person knowledge graph construction methods of the present application.
[0199] In some embodiments, an embodiment of the present application further provides a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the steps of any one of the above-mentioned multimodal large model-driven person knowledge graph construction methods.
[0200] In some embodiments, an embodiment of the present application further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the multimodal large model-driven person knowledge graph construction method.
[0201] Figure 6 FIG. is a schematic hardware structure diagram of an electronic device for executing the multimodal large model-driven person knowledge graph construction method provided by another embodiment of the present application. As Figure 6 shown, the device includes:
[0202] One or more processors 610 and a memory 620, Figure 6 Taking one processor 610 as an example.
[0203] The device for executing the multimodal large model-driven person knowledge graph construction method may further include: an input device 630 and an output device 640.
[0204] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected by a bus or other means, Figure 6 Taking connection by a bus as an example.
[0205] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for constructing a multi-modal large model-driven personal knowledge graph in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the method for constructing a multi-modal large model-driven personal knowledge graph in the above method embodiments.
[0206] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely provided relative to the processor 610, and these remote memories can be connected to the electronic device through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0207] The input device 630 can receive input digital or character information, and generate signals related to the user settings and function control of the electronic device. The output device 640 may include a display device such as a display screen.
[0208] The one or more modules are stored in the memory 620 and, when executed by the one or more processors 610, execute the method for constructing a multi-modal large model-driven personal knowledge graph in any of the above method embodiments.
[0209] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.
[0210] The electronic device in the embodiments of the present application exists in various forms, including but not limited to:
[0211] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aiming to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0212] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.
[0213] (3) Portable entertainment devices: Such devices can display and play multimedia content. This type of device includes: audio and video players, handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.
[0214] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.
[0215] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0216] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the related technologies, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0217] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for constructing a person knowledge graph driven by a multi-modal large model, comprising: Obtaining a target person name and multiple person attribute description information, and constructing a knowledge graph framework; The knowledge graph framework includes a core node and multiple attribute nodes respectively connected to the core node, the core node is used to indicate the target person name, and each attribute node is respectively used to indicate the corresponding person attribute description information; Obtaining corresponding target multi-source material data according to the target person name, the target multi-source material data includes text material, image material and video material that match the target person name; For each of the person attribute description information, extracting at least one attribute matching material segment corresponding to the person attribute description information from the target multi-source material data, and fusing the respective attribute matching material segments corresponding to the person attribute description information based on the multi-modal large model to output corresponding attribute multi-modal generated data; The multi-modal generated data includes text generated data, image generated data and video generated data; Filling each of the attribute multi-modal generated data into the corresponding attribute nodes in the knowledge graph framework according to the person attribute description information to construct a multi-modal person knowledge graph.
2. The method according to claim 1, wherein, The obtaining the target person name and multiple person attribute description information includes: Obtaining user input information, and determining a target person name according to the user input information; Constructing a person event recognition prompt word according to the target person name and a preset prompt word template, and inputting the person event recognition prompt word into a large language model to trigger the large language model to output at least one person event keyword for the target person name; Determining person attribute description information according to each of the person event keywords.
3. The method according to claim 2, wherein The determining person attribute description information according to each of the person event keywords includes: Determining corresponding customized person attribute description information according to each of the person event keywords; Determining person attribute description information according to at least one preset general person attribute description information and each of the customized person attribute description information.
4. The method according to claim 1, wherein The obtaining corresponding target multi-source material data according to the target person name includes: Performing semantic matching between the target person name and the material description information of each multi-source material in the multi-source material library to screen at least one initial matching multi-source material from the multi-source material library; Determining the data noise score of each of the initial matching multi-source materials based on a data noise scoring network; the data noise scoring network includes a parallel text data noise evaluation branch, an image data noise evaluation branch and a video data noise evaluation branch; Determining the initial matching multi-source materials corresponding to the data noise score not exceeding a preset noise score threshold as the first initial matching materials, and determining the initial matching multi-source materials corresponding to the data noise score exceeding the noise score threshold as the second initial matching multi-source materials; Input the second initial matching multi-source material into the material reconstruction network to generate reconstructed multi-source materials; the material reconstruction network includes a text data reconstruction branch, an image data reconstruction branch, and a video data reconstruction branch connected in parallel; Determine the target multi-source material data according to the first initial matching material and the target reconstructed multi-source material; the target reconstructed multi-source material is the reconstructed multi-source material whose corresponding noise score does not exceed the noise score threshold.
5. The method according to claim 1, wherein For each of the person attribute description information, extracting at least one attribute matching material segment corresponding to the person attribute description information from the target multi-source material data includes: Determine an attribute tag group according to each of the person attribute description information, and construct a tag classifier for the attribute tag group; each of the attribute tags is used to indicate the corresponding person attribute description information; Perform semantic segmentation on the target multi-source material data to obtain a plurality of material candidate segments, and each of the material candidate segments has a corresponding semantic description of the material segment; Based on the tag classifier, determine the attribute tags matched by each of the semantic descriptions of the material segments, so as to obtain the attribute matching material segments of the corresponding person attribute description information.
6. The method according to claim 5, wherein, The tag classifier includes a cascaded projection layer, a tag hint fusion layer, and a tag matching layer; The projection layer is used to map the semantic description of the material segment to an intermediate representation space: h1 = ReLU(W1x + b1), where x represents the embedded representation of the semantic description of the material segment, W1 and b1 respectively represent the weight matrix and bias term of the projection layer, h1 represents the projection intermediate representation, and ReLU represents the ReLU activation function; The tag hint fusion layer is used to encode the semantic information of each attribute tag into a hint vector to perform feature alignment with the semantic description of the material segment in the semantic space: h 2,c = ReLU(h1 + P c ) where P c represents the learnable prompt vector corresponding to l c , l c represents the c-th attribute label, and P represents the label prompt matrix corresponding to the attribute label group; h 2,c represents the feature representation after the fusion of the semantic description of the material segment and l c , n represents the total number of attribute labels in the attribute label group, and H2 represents the label fusion feature matrix; Calculate the matching scores corresponding to each attribute tag based on the tag matching layer: In the formula, represents the transposed vector of the shared scoring weight vector, and b 2,c represents the bias term of the attribute label l c , and z c represents the unnormalized matching score of the semantic description of the material segment on l c ; Based on the tag matching layer, perform Sigmoid activation processing on each of the matching scores to obtain the matching probability of the semantic description of the material segment for each corresponding attribute tag: p c = σ(z c ) where σ is the Sigmoid activation function, and p c represents the matching probability between the semantic description of the material segment and l c ; Determine the attribute matching material segments of the corresponding attribute tags according to the material candidate segments whose corresponding matching probabilities exceed the preset probability threshold.
7. The method according to claim 1, wherein, For the fine-tuning training of the multi-modal large model, it includes: Define a training data set: In the formula, represents the training dataset, and N s represents the total number of samples in the training dataset, and u represents the u-th specific sample in the dataset; is the text segment corresponding to the u-th sample, representing the text description segment of the specific attributes of the person; is the image segment corresponding to the u-th sample, representing the image material segment of the specific attributes of the person; is the video segment corresponding to the u-th sample, representing the video material segment of the specific attributes of the person; y (u) represents the true label of the person attributes corresponding to the u-th sample; Use the pre-trained multi-modal large model to obtain the embedded representations of each modality: where f T (·), f I (·), and f v (·) represent the text encoder function, the image encoder function, and the video encoder function respectively, and are the embedding representations corresponding to the u-th sample text segment, sample image segment, and sample video segment respectively; Then, through a cross-modal fusion module, obtain a unified feature representation across modalities: wherein represents the cross-modal fusion feature corresponding to the u-th sample, and CrossAttention(·) represents the cross-modal attention mechanism; Calculate the cross-modal consistency loss, including: First, map the cross-modal fusion features into independent embedding spaces of the text, image, and video modalities respectively, denoted as: where, g T (·), g I (·), g v (·) respectively represent trainable network layers that map the fused cross-modal features to the text modality space, the image modality space, and the video modality space; and respectively represent the reconstructed text embedding vector, the reconstructed image embedding vector, and the reconstructed video embedding vector; Then, using the real modality feature as a supervision signal, constrain the consistency between the reconstructed feature and the actual feature: In the formula, represents the cross-modal consistency loss, which is used to ensure the consistency between the generated data and the real data; Calculate the multi-modal contrast loss, including: In each training batch, select text-image pairs and text-video pairs with the same attribute as positive samples, and select segments with different attributes as negative samples, and then optimize the cross-modal matching degree by calculating the cosine similarity between the embedded representations: Wherein, represents the multi-modal contrast loss, and sim(·) represents the vector similarity function, represents the vector similarity of the u-th text-image positive sample pair, represents the vector similarity of the u-th text-video positive sample pair, represents the total similarity between the u-th text segment and image segments of other non-corresponding person attributes, represents the total similarity between the u-th text segment and video segments of other non-corresponding person attributes; Calculate the total loss function according to the cross-modal consistency loss and the multi-modal contrast loss: In the formula, represents the total loss function, and α, β are hyperparameter weights used to control the relative importance of the cross-modal consistency loss and the multi-modal contrast loss.
8. A person knowledge graph construction system driven by a multi-modal large model, including: The knowledge graph framework construction unit is used to obtain the target person's name and multiple pieces of person attribute description information, and construct a knowledge graph framework; The knowledge graph framework includes a core node and multiple attribute nodes respectively connected to the core node. The core node is used to indicate the target person's name, and each attribute node is respectively used to indicate the corresponding person attribute description information; The multi-source material acquisition unit is used to obtain the corresponding target multi-source material data according to the target person's name. The target multi-source material data includes text materials, image materials, and video materials that match the target person's name; The attribute data generation unit is used to, for each of the person attribute description information, extract at least one attribute matching material segment corresponding to the person attribute description information from the target multi-source material data, and fuse the respective attribute matching material segments corresponding to the person attribute description information based on a multi-modal large model to output corresponding attribute multi-modal generation data; The multi-modal generation data includes text generation data, image generation data, and video generation data; The knowledge graph attribute filling unit is used to fill each of the attribute multi-modal generation data into the corresponding attribute nodes in the knowledge graph framework according to the person attribute description information to construct a multi-modal person knowledge graph.
Citation Information
Patent Citations
Vision and text cross-modal matching method based on consensus embedding space and similarity
CN115935194A
Multi-modal data fusion method based on position sensitive optimization
CN119339193A
Brassica oleracea knowledge dynamic expression and interaction method and system based on multi-modal fusion
CN119862954A
Cited By
Industrial risk knowledge graph construction method and system based on large language model
CN120822594A
An industry risk knowledge graph construction method and system based on a large language model
CN120822594B