User portrait construction method and device based on multi-modal information, equipment and medium

CN122548338APending Publication Date: 2026-08-11CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本发明提供一种基于多模态信息的用户画像构建方法、装置、设备及介质,用以解决如何解决用户对多模态信息处理效率较低的技术问题

Benefits of technology

[0009]上述基于多模态信息的用户画像构建方法、装置、设备及介质,所实现的方案中,可以通过客户端获取并解析目标用户的多模态信息,提取各个模态下的实体、关系及事件要素,所述实体至少包括人物实体;对各个模态提取的实体、关系及事件要素相互之间进行交叉关联分析,生成目标用户的事件时间轴信息和关注点标签;对不同模态下提取的实体进行跨模态实体对齐处理,并基于预设的规则,对跨模态实体对齐处理后的实体在不同模态下的属性信息进行一致性校验;对一致性校验后存在冲突的属性信息进行消歧处理,并对消歧处理后的属性信息进行融合处理,生成融合后的目标用户的基础信息;基于所述目标用户的基础信息、所述事件时间轴信息以及所述关注点标签,构建所述目标用户的用户画像。本发明通过跨模态实体对齐处理,将不同模态下提取的实体进行跨模态映射,识别出多个信息源的同一实体,并对识别出的实体的属性信息进行一致性校验和消歧处理,提高了多模态信息融合的准确性;同时,将各个模态提取的实体、关系及事件要素进行关联整合,生成事件时间轴与关注点标签,将分散的过程信息转化为结构化用户画像,提升了用户档案的信息丰富度与多模态信息处理效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548338A_ABST
    Figure CN122548338A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of artificial intelligence technology and is applicable to the financial and medical fields. It discloses a method, apparatus, device, and medium for constructing user profiles based on multimodal information. The method includes: performing cross-correlation analysis on entities, relationships, and event elements extracted from various modalities to generate event timeline information and attention tags for the target user; performing cross-modal entity alignment processing on entities extracted from different modalities, and performing consistency verification on the attribute information of the aligned entities in different modalities based on preset rules; disambiguating conflicting attribute information after consistency verification, and fusing the disambiguated attribute information to generate fused basic information of the target user; and constructing a user profile of the target user based on the basic information, event timeline information, and attention tags. This improves the efficiency of users in processing multimodal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology and is applicable to the financial and medical fields. In particular, it relates to a method, apparatus, device and medium for constructing user profiles based on multimodal information. Background Technology

[0002] In professional service sectors such as finance, insurance, and healthcare, practitioners need to track and maintain client relationships over the long term. A large amount of key client information exists in multimodal formats, including voice debriefings, handwritten notes from interviews, instant messaging logs, medical reports, and policy documents. Existing client management tools, based on structured fields, only support recording and reminding users of preset information such as name, age, and intended products. They cannot automatically parse and structure the aforementioned multimodal information, requiring practitioners to spend a significant amount of time manually organizing it.

[0003] However, the aforementioned methods based on manual organization and structured field recording have several problems in practical applications. First, existing tools only offer attachment upload or text pasting functions, requiring manual reading and transcription of each entry into the corresponding fields, resulting in low efficiency and a high risk of omissions in process information archiving. Second, the time-consuming organization leads to information dispersion and loss. After interviews, practitioners need a considerable amount of time to summarize and archive, resulting in a large amount of process information being scattered across personal communication tools, photo albums, and paper notebooks, unable to be integrated into searchable and reusable archival assets. When clients communicate again or critical points arise, practitioners struggle to retrieve historical details in a timely manner, affecting response efficiency and professional judgment. Furthermore, attribute records for the same client may be inconsistent across different modalities, requiring practitioners to make comparisons and judgments manually, leading to low archive utilization. Therefore, how to solve the problem of low efficiency in processing multimodal information is a pressing technical issue that needs to be addressed. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and medium for constructing user profiles based on multimodal information, in order to solve the technical problem of low efficiency in processing multimodal information by users.

[0005] In a first aspect, the present invention provides a method for constructing user profiles based on multimodal information, comprising: Acquire and parse the multimodal information of the target user, and extract the entity, relationship and event elements under each modality, wherein the entity includes at least the person entity; Cross-correlation analysis is performed on the entities, relationships, and event elements extracted from each modality to generate event timeline information and attention tags for the target user; Cross-modal entity alignment is performed on entities extracted from different modalities, and the consistency of attribute information of entities after cross-modal entity alignment is verified in different modalities based on preset rules. Disambiguation is performed on conflicting attribute information after consistency verification, and the disambiguated attribute information is then fused to generate the fused basic information of the target user. Based on the target user's basic information, the event timeline information, and the attention tags, a user profile of the target user is constructed.

[0006] Secondly, the present invention provides a user profile construction device based on multimodal information, comprising: The acquisition module is used to acquire and parse the multimodal information of the target user, and extract entities, relationships and event elements under each modality, wherein the entities include at least person entities; The analysis module is used to perform cross-correlation analysis on the entities, relationships and event elements extracted from each modality, and generate event timeline information and attention tags for the target user. The processing module is used to perform cross-modal entity alignment processing on entities extracted from different modalities, and to perform consistency verification on the attribute information of the entities after cross-modal entity alignment processing in different modalities based on preset rules. The fusion module is used to disambiguate conflicting attribute information after consistency verification, and then fuse the disambiguated attribute information to generate the fused basic information of the target user. The construction module is used to construct a user profile of the target user based on the target user's basic information, the event timeline information, and the attention tags.

[0007] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described user profile construction method based on multimodal information.

[0008] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described user profile construction method based on multimodal information.

[0009] The aforementioned method, apparatus, device, and medium for constructing user profiles based on multimodal information can acquire and parse the multimodal information of a target user through a client, extracting entities, relationships, and event elements under each modality, wherein the entities include at least person entities; perform cross-correlation analysis on the entities, relationships, and event elements extracted from each modality to generate event timeline information and attention tags for the target user; perform cross-modal entity alignment processing on the entities extracted under different modalities, and perform consistency verification on the attribute information of the entities after cross-modal entity alignment processing under different modalities based on preset rules; perform disambiguation processing on the attribute information that conflicts after consistency verification, and perform fusion processing on the attribute information after disambiguation processing to generate fused basic information of the target user; and construct the user profile of the target user based on the basic information of the target user, the event timeline information, and the attention tags. This invention uses cross-modal entity alignment processing to perform cross-modal mapping on entities extracted from different modalities, identifying the same entity from multiple information sources. It also performs consistency verification and disambiguation processing on the attribute information of the identified entities, improving the accuracy of multimodal information fusion. At the same time, it integrates and associates entities, relationships, and event elements extracted from various modalities to generate event timelines and attention tags, transforming scattered process information into structured user profiles, thereby improving the information richness of user profiles and the efficiency of multimodal information processing. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of an application environment for a user profile construction method based on multimodal information in one embodiment of the present invention.

[0012] Figure 2 This is a flowchart illustrating a user profile construction method based on multimodal information in one embodiment of the present invention.

[0013] Figure 3 yes Figure 2 A schematic diagram of a specific implementation method for step S20.

[0014] Figure 4 yes Figure 2 A schematic diagram of a specific implementation method for step S30.

[0015] Figure 5 yes Figure 2A schematic diagram of a specific implementation method for step S50.

[0016] Figure 6 This is a schematic diagram of a user profile construction device based on multimodal information in one embodiment of the present invention.

[0017] Figure 7 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention.

[0018] Figure 8 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The user profile construction method based on multimodal information provided in this invention can be applied to, for example... Figure 1 In the application environment, Figure 1This is a schematic diagram of an application environment for a user profile construction method based on multimodal information according to an embodiment of the present invention; wherein, the client communicates with the server via a network. The server can obtain and parse the multimodal information of the target user through the client, extract entities, relationships, and event elements under each modality, wherein the entities include at least person entities; perform cross-correlation analysis on the entities, relationships, and event elements extracted from each modality to generate event timeline information and attention tags for the target user; perform cross-modal entity alignment processing on the entities extracted under different modalities, and perform consistency verification on the attribute information of the entities after cross-modal entity alignment processing under different modalities based on preset rules; perform disambiguation processing on the attribute information that conflicts after consistency verification, and perform fusion processing on the attribute information after disambiguation processing to generate fused basic information of the target user; and construct the user profile of the target user based on the basic information of the target user, the event timeline information, and the attention tags. This invention employs cross-modal entity alignment processing to perform cross-modal mapping on entities extracted from different modalities, identifying the same entity from multiple information sources. It also performs consistency verification and disambiguation on the attribute information of the identified entities, improving the accuracy of multimodal information fusion. Simultaneously, it integrates and associates entities, relationships, and event elements extracted from various modalities to generate event timelines and attention tags, transforming scattered process information into structured user profiles, thus enhancing the information richness of user profiles and the efficiency of multimodal information processing. The invention will be described in detail below through specific embodiments.

[0021] Please see Figure 2 As shown, Figure 2 This is a flowchart illustrating a user profile construction method based on multimodal information provided in an embodiment of the present invention. The user profile construction method based on multimodal information specifically includes the following steps: S10: Acquire and parse the multimodal information of the target user, extracting entities, relationships, and event elements under each modality. The entities include at least person entities. Specifically, in this embodiment of the invention, the entity refers to the specific object referred to in the multimodal information, including specific objects such as person entities, institutional entities, or product entities. The person entity refers to a natural person mentioned in the multimodal information; the institutional entity refers to an organization or unit mentioned in the multimodal information; and the product entity refers to a financial product, insurance product, or medical product mentioned in the multimodal information. The relationship refers to the social or business connection between entities, including at least one of kinship, professional relationship, agency relationship, or consulting relationship. The event element refers to key information describing the behavior or state change of an entity at a specific time and in a specific scenario, including at least the event occurrence time, event type, and entity identifiers participating in the event. By parsing the content of each modal information separately and extracting entities, relationships, and event elements, automatic structuring processing of multimodal unstructured information is achieved. This overcomes the shortcomings of existing tools that only support attachment uploads or text pasting and cannot automatically parse multimodal information, significantly reducing the workload of manual secondary processing and avoiding the omission of process information. Specifically, it includes the following steps S11-S14: S11: Acquire voice data, image data, and dialogue text data associated with the target user. Specifically, in this embodiment of the invention, in professional service scenarios such as financial advisory and healthcare management, key information related to clients naturally exists in multiple modalities: interview content is recorded in voice form, paper reports are stored in image form, and instant messaging generates dialogue text data. Information in a single modality can only reflect a part of the client and cannot form a complete understanding. This invention provides a complete information source foundation for subsequent automatic parsing by uniformly accessing three modalities: voice data, image data, and dialogue text data. For example, in a financial scenario, the acquired voice data can be a recording of a financial advisor's interview with a client, the image data can be a scanned copy of the client's asset certificate or risk assessment questionnaire, and the dialogue text data can be WeChat communication records between the advisor and the client. In a medical scenario, the voice data can be a recording of communication between a doctor and the patient's family, the image data can be a medical examination report or medical record photo uploaded by the patient, and the dialogue text data can be a text and image consultation record from an online consultation platform. This invention overcomes the limitations of existing tools that can only receive single text inputs or rely on manual uploads of attachments, and realizes the systematic collection and access of multimodal raw information, avoiding the problem of one-sided customer profiles caused by fragmented information sources.

[0022] S12: Perform speech recognition and transcription processing on the speech data to generate timestamped dialogue text, identify key paragraph information in the dialogue text, and extract entity, relationship, and event elements based on the identified key paragraph information. Specifically, in this embodiment of the invention, the speech data records real dialogues during the service process, carrying the customer's explicit needs, implicit concerns, and commitments reached by both parties. After converting the speech into timestamped text using automatic speech recognition technology, natural language processing technology is used to perform semantic segmentation and paragraph type identification on the text, distinguishing key paragraphs such as statement of needs, expression of objections, confirmation of commitments, and follow-up arrangements, and then extracting person entities, relationships between entities, and event elements from the key paragraphs. Among them, the entities extracted from the speech data refer to the specific people, institutions, or products mentioned in the dialogue; the relationships refer to the kinship, professional, or business entrustment relationships between entities reflected in the dialogue; and the event elements refer to the time of occurrence, event type, and entities involved in the event described in the dialogue. For example, in a financial scenario, after the recording of a financial advisor's interview is transcribed, the system can identify the objection paragraphs where the customer has concerns about the risks of equity products, and extract the customer entity, the risk concern relationship, and the interview event. In medical settings, after processing, recordings of conversations between doctors and family members can extract key event elements such as the patient's identity, the kinship between the patient and family members, mentions of drug allergies, and follow-up appointments. This invention automates the process of manually listening to and recording audio debriefings. The extracted entities, relationships, and event elements are all timestamped, providing a precise time reference for subsequent event timeline construction. This significantly reduces the workload of manual processing and avoids information omissions and time misalignments caused by recording delays.

[0023] S13: Perform optical character recognition (OCR) processing on the image data to extract person entities, relationship indicators, and indicator values ​​from the image data, and extract entity, relationship, and event elements from the image data based on the extraction results. Specifically, in this embodiment of the invention, image data such as medical examination reports, insurance policy documents, and handwritten notes contain a large amount of structured or semi-structured text information. This information is stored in pixel form and cannot be directly retrieved or used in calculations. OCR technology is used to extract the text in the image into editable text, and then layout analysis technology is used to identify the layout relationship and table structure between the text, thereby extracting key information such as person names, kinship indicators, and medical examination indicator values, and further extracting the corresponding entity, relationship, and event elements. Here, the entities extracted from the image data refer to the specific people, institutions, or products recorded in the document; the relationships refer to the attribution relationships or kinship relationships marked between fields in the document; and the event elements refer to the occurrence time and event type of events such as medical visits, medical examinations, and contract signings recorded in the document. For example, in a financial scenario, after parsing the scanned asset certificate provided by the customer, entity and indicator values ​​such as the customer's name, asset category, and amount can be extracted to generate the customer's asset profile elements. In medical settings, patient-uploaded medical examination report photos, after parsing, can extract entity and indicator values ​​such as patient name, names and values ​​of various examination indicators, and abnormal markers, providing a quantitative basis for constructing health profiles. This invention transforms image materials, which originally existed only as static attachments, into searchable and computable structured data, enabling the automatic extraction of key information contained in the images and its participation in profile construction, thereby improving the data richness and objectivity of customer profiles.

[0024] S14: Perform intent recognition on the dialogue text data, and extract entity, relationship, and event elements from the dialogue text data based on the intent recognition results. Specifically, in this embodiment of the invention, instant messaging chat records contain the customer's proactively expressed concerns, questions, and emotional tendencies, which are freely expressed and closely related to the context. Through intent recognition technology, semantic understanding of the dialogue text is performed to distinguish different intent types such as consultation, complaint, confirmation, and feedback. Based on the intent recognition results, the entity, relationship, and event elements in the corresponding statements are located, enabling the automatic capture of the customer's true intent and potential needs from unstructured dialogues. Entities extracted from the dialogue text data refer to the two parties in the dialogue and the specific people, products, or services mentioned in the dialogue. Relationships refer to the service relationship between the two parties or the comparison or attribution relationship between the entities mentioned in the dialogue. Event elements refer to the consultation, complaint, appointment, and other events described or triggered by the customer through the dialogue, and their occurrence time. For example, in a financial scenario, a customer asks in WeChat: "What's the difference between this product and the one I bought before?" The system identifies this as a product comparison intent, extracts the customer entity, the current product and historical product entities and comparison relationships, and generates corresponding concern tags. In a medical setting, a patient describes on an online consultation platform that they are experiencing dizziness and their blood pressure seems high. The system recognizes this as a symptom statement and extracts patient, symptom, and time information to provide incremental information for the event timeline of the health profile. This invention can automatically capture the customer's true intent and implicit needs from unstructured, free-flowing conversations, avoiding the inefficiency and subjective bias of manually reviewing large amounts of chat history. This ensures that the focus tags in the customer profile are supported by real interactive data.

[0025] S20: Perform cross-correlation analysis on the entities, relationships, and event elements extracted from each modality to generate event timeline information and focus tags for the target user. Specifically, in this embodiment of the invention, by performing correlation analysis on entities, relationships, and event elements within the same modality to generate event timeline information and focus tags, previously isolated information fragments are integrated into information with temporal relationships and thematic characteristics, enabling practitioners to quickly grasp the key event context and focus of the client. For example... Figure 3 The above, Figure 3 yes Figure 2 A schematic flowchart of a specific implementation of step S20 includes the following steps S21-S23: S21: Within the same modality, the extracted entity, relationship, and event elements are associated to generate an intramodal information unit. This intramodal information unit includes a set of interconnected entities, a relationship network, and an event description within the corresponding modality. Specifically, in this embodiment of the invention, the entity, relationship, and event elements extracted after parsing single-modal content are isolated when not associated. For example, the customer's name, their spouse's name, and the objection raised in the same audio recording lack explicit semantic connections. This invention integrates the originally scattered elements into an intramodal information unit with inherent logic by establishing attribution connections between entities and relationships, participation connections between relationships and events, and subject connections between events and entities within the same modality. This information unit includes a set of interconnected entities, a relationship network, and an event description within the corresponding modality. For example, in a financial scenario, after parsing a financial advisor's audio review, the following entities are extracted: customer entity, spouse entity, marital relationship, and the spouse's objection to investment risks. This invention associates these elements into a complete audio modal information unit, described as the customer and their spouse, who are a married couple, jointly participating in a meeting regarding objections to investment risks. In a medical setting, a recorded phone conversation between a doctor and a family member, after parsing, extracts the patient entity, the family member entity, the parent-child relationship, and the medication reaction event reported by the family member. This invention links these elements into a single voice modal information unit, describing the family member and patient as having a parent-child relationship, and the family member reporting the patient's adverse medication reaction to the doctor. This invention integrates fragmented information within a modal into a structured information unit with complete semantics, overcoming the limitation of existing technologies that can only store isolated fields and cannot express the logical relationships between elements.

[0026] S22: Arrange the event elements in the information units within each modality in chronological order to generate cross-modal event timeline information. Specifically, in this embodiment of the invention, each information unit within a different modality contains event elements, and these event elements have time attributes. An event within a single modality only reflects a local interaction that occurs within that modality and cannot present all the content of the customer interaction. This invention extracts the event elements from each information unit in a unified manner and arranges them across modalities based on the time of event occurrence, forming a unified event timeline information spanning multiple modalities such as voice, image, and dialogue text. This timeline information completely records all interaction nodes from initial contact, needs communication, solution presentation, objection handling to service completion. For example, in a medical scenario, the dialogue text modality information unit contains events where patients describe symptoms on an online consultation platform, the voice modality information unit contains events where doctors follow up by phone, and the image modality information unit contains events where patients upload follow-up reports. This invention arranges these events in chronological order to form the patient's treatment timeline, presenting a complete disease progression trajectory from initial diagnosis, follow-up, to re-examination. This invention breaks the siloed state of event information under a single modality, integrates event elements scattered in different modalities into a unified timeline view, enabling practitioners to obtain the entire interaction history of customers in one stop, avoiding the inefficient operation of repeatedly switching between multiple systems to search for information, and significantly improving information retrieval efficiency and customer response speed.

[0027] S23: Perform content analysis on the intramodal information units to extract the target user's focus topics, sentiment tendencies, and behavioral patterns, and generate focus tags for the target user. Specifically, in this embodiment of the invention, the intramodal information units not only include objective facts such as entities and relationships, but also contain the focus tendencies, emotional states, and behavioral patterns revealed by the customer during the interaction. This invention performs content analysis on the intramodal information units of each modality, extracts the topics that the customer repeatedly mentions or focuses on through topic recognition technology, judges the polarity and intensity of the customer's emotions in the dialogue through sentiment analysis technology, identifies the periodic patterns and decision preferences presented by the customer in the interaction through behavioral pattern analysis technology, and generates focus tags for the target user based on the analysis results, including focus topic tags, sentiment tendency tags, and behavioral pattern tags. For example, in financial scenarios, the system analyzes the content of multiple information units within a customer's modality, finding that customers repeatedly mention topics such as children's education and study abroad planning in voice interviews and WeChat chats, exhibiting anxiety about future expenditures emotionally and a cyclical pattern of proactively consulting about products at the beginning of each year. Based on this, tags are generated for: children's education fund planning, future expenditure anxiety, and proactive consultation behavior at the beginning of the year. In medical scenarios, the system analyzes the content of patients' online consultations and telephone follow-up recordings, finding that patients frequently focus on the relationship between blood pressure fluctuations and diet, show hesitation about medication plans, and have significantly higher consultation frequency on Monday mornings than at other times. Based on this, tags are generated for: blood pressure and diet-related concern, medication decision-making hesitation, and concentrated consultation behavior on Mondays. The automatically generated concern tags of this invention allow practitioners to quickly grasp the core concerns of customers without having to review original materials, providing a direct and usable decision-making basis for subsequent personalized service strategies.

[0028] S30: Perform cross-modal entity alignment processing on entities extracted from different modalities, and perform consistency verification on the attribute information of the entities after cross-modal entity alignment processing in different modalities based on preset rules. Specifically, in this embodiment of the invention, the attribute information refers to information fields describing entity characteristics. For example, if the entity is a person entity, the attribute information includes age, occupation, income, contact information, health status indicators, etc. The same entity may have the same or different attribute information extracted in different modalities. For example, the customer's self-reported age and occupation can be extracted from voice data, the customer's registered age and health indicators can be extracted from the physical examination report in image data, and the income and contact information mentioned by the customer can be extracted from the dialogue text data. Due to differences in information sources, recording time, or expression, the same attribute information of the same entity may differ in different modalities. By performing cross-modal entity alignment and consistency verification on entities extracted from different modalities, the same entity from different sources such as voice, images, and chat records can be identified, and conflicts in attribute information between different modalities can be automatically detected, solving the data inconsistency problem caused by the scattered storage of multi-source information. Figure 4 The above, Figure 4 yes Figure 2 A schematic flowchart of a specific implementation of step S30 includes the following steps S31-S33: S31: Based on the entity's name, relationship network, and associated event elements, calculate the similarity between entities in different modalities. Entities with similarity exceeding a preset threshold are identified as the same entity, and a cross-modal entity mapping relationship is established. Specifically, in this embodiment of the invention, entities extracted from different modalities may appear in different name forms, such as the full name in speech-to-text, the name on an ID card recognized by image OCR, or a title or nickname in dialogue text. It is necessary to determine whether they point to the same entity. This invention integrates three dimensions: entity name similarity, relationship network structural similarity, and spatiotemporal consistency of associated event elements to calculate the comprehensive similarity between entities in different modalities. When the comprehensive similarity exceeds a preset threshold, the entity in different modalities is identified as pointing to the same entity, and a cross-modal entity mapping relationship is established. For example, in a financial scenario, Zhang identified in the speech modality and Zhang identified on an asset certificate scan in the image modality both have a relationship network that includes a spouse, Li, and both associated events point to the same financial interview. The system calculates that the comprehensive similarity exceeds the threshold, identifies them as the same customer, and establishes a mapping relationship. In a medical setting, the patient's self-identification as "Old Wang" in the text-based dialogue modality and the patient's name "Wang Mou" on the medical examination report in the image modality both point to the same online consultation. The system determines they are the same patient and establishes a mapping relationship based on event consistency and partial name matching. This invention overcomes the limitation of single-modal entity recognition relying solely on name string matching by introducing relationship networks and event elements as auxiliary alignment features, significantly improving the accuracy of cross-modal entity alignment.

[0029] S32: For the same entity with established cross-modal entity mapping relationships, extract the attribute information of the corresponding entity in each modality. Specifically, in this embodiment of the invention, the same entity recorded in different modalities often carries attribute information of different dimensions. Voice data may contain the customer's self-reported occupation and income information, image data may contain age and health indicators recorded in authoritative documents, and dialogue text may contain contact information and address actively updated by the customer. This invention, for the same entity with established cross-modal entity mapping relationships, traverses the attribute information associated with the entity in each modality, extracts and summarizes it according to the modal source, forming a multi-source attribute information set for the entity. For example, in a financial scenario, for a customer named Zhang with an established mapping relationship, the system extracts his self-reported annual income of approximately 500,000 yuan and occupation as a middle-level manager in a company from the voice modality, extracts his registered age of 42 years old from the asset certificate in the image modality, and extracts his mentioned work unit as a certain technology company from the dialogue text modality, summarizing them into a complete multi-source attribute information set for the customer. This invention centralizes and aggregates the attribute information of the same entity that is scattered across various modalities, breaking through the inherent limitation of the limited dimensions of information coverage in a single modality, and enabling the basic information dimensions of customer profiles to be supplemented by cross-source data.

[0030] S33: Compare attribute information of the same entity in different modalities. When the difference in attribute information between different modalities exceeds a preset tolerance range, it is determined that the corresponding attribute information is in conflict, and a conflict marker is generated. Specifically, in this embodiment of the invention, the attribute information of the same entity in different modalities may differ due to different recording times, information sources, or different expressions. For example, the self-reported age in the voice recording is inconsistent with the age recorded by the official source in the image file. This invention compares the same attribute information of the same entity in different modalities one by one. When the difference in the value of the same attribute between different modalities exceeds the preset tolerance range corresponding to the attribute type, it is determined that the attribute information has a cross-modal conflict, and a conflict marker is generated for the attribute. The conflict marker includes the name of the conflicting attribute, the value of each modality, and the difference range. For example, in a financial scenario, for the annual income attribute of customer Zhang, the self-reported annual income in the voice modality is 500,000 yuan, while the annual income derived from the bank statement in the image modality is approximately 300,000 yuan. The difference exceeds the preset 20% tolerance range, and the system determines that the annual income attribute is in conflict and generates a conflict marker. This invention, through automated cross-modal attribute comparison and conflict detection, can proactively discover data inconsistencies and accurately locate conflicting attributes during the multi-source information fusion process.

[0031] S40: Disambiguation processing is performed on conflicting attribute information after consistency verification, and the disambiguated attribute information is then fused to generate fused basic information of the target user. Specifically, in this embodiment of the invention, by disambiguating and fusing conflicting attribute information after consistency verification, information contradictions are automatically resolved based on credibility weights, generating user basic information that is traceable to its source and reliable in value. This avoids the subjectivity and inefficiency of manually judging the authenticity of multi-source information, ensuring the accuracy and traceability of user profile basic data. Specifically, this includes the following steps S41-S42: S41: For attribute information carrying conflict markers, determine the credibility weight of the corresponding attribute information in each modality, and select the attribute information with the highest credibility weight as the disambiguation result. Specifically, in this embodiment of the invention, information sources in different modalities have different credibility levels. Information recorded in authoritative documents is generally superior to oral statements, recently updated information is generally superior to earlier records, and information confirmed by the customer is generally superior to third-party retellings. This invention determines the credibility weight of the corresponding attribute information in each modality for attribute information carrying conflict markers based on three dimensions: the type of information source, the timeliness of the information, and the identity of the information provider, and selects the attribute information with the highest credibility weight as the disambiguation result. For example, in a financial scenario, for the customer's annual income attribute, the self-reported amount in the voice modality is 500,000 yuan, while the bank statement in the image modality derives it as 300,000 yuan. The system determines that the bank statement is an authoritative financial document with a higher credibility weight than the oral statement, and selects 300,000 yuan as the disambiguation result. This invention automatically completes the disambiguation processing of conflict attributes based on multi-dimensional credibility assessment, avoiding the subjectivity and inefficiency of manually judging the authenticity of multi-source information.

[0032] S42: Merge the disambiguated attribute information with the non-conflicting attribute information to generate basic user information, and record the source modality, original value, and disambiguation processing log for each attribute. Specifically, in this embodiment of the invention, the disambiguation processing only resolves the conflicting attributes. The disambiguated attributes need to be merged with the originally non-conflicting attributes to form complete basic user information. Simultaneously, to ensure the traceability and auditability of the basic information, the source path, original value, and disambiguation processing process for each attribute need to be fully recorded. This invention merges the disambiguated attribute information with the non-conflicting attribute information to generate basic user information, and records the source modality identifier, original value, and disambiguation processing log for each attribute. The disambiguation processing log includes the conflict discovery time, the values ​​of the conflicting parties, the credibility weight allocation, and the final selection result. For example, in a financial scenario, the system merges the disambiguated annual income of 300,000 yuan with non-conflicting attributes such as age 42, occupation (mid-level manager), and employer (a certain technology company) to generate the merged basic information for customer Zhang. It also records the disambiguation processing log for the annual income attribute, noting the source of bank statements, the original voice statement value of 500,000 yuan, and the weighting comparison process. This invention overcomes the deficiency in existing technologies where the merged results cannot trace back to the original basis by ensuring that each attribute in the generated basic information has complete source tracing and disambiguation records.

[0033] S50: Based on the target user's basic information, the event timeline information, and the attention tags, construct a user profile for the target user. Specifically, in this embodiment of the invention, by constructing a user profile using basic user information, event timeline information, and attention tags, the organic integration of multi-dimensional information is achieved. This allows practitioners to simultaneously obtain the customer's basic attributes, behavioral patterns, and interests within a single profile, transforming a static profile into a dynamic customer view that can support service decisions. Figure 5 The above, Figure 5 yes Figure 2 A schematic flowchart of a specific implementation of step S50 includes the following steps S51-S53: S51: Based on the aforementioned basic user information, a basic information dimension for constructing a user profile is built, and based on the aforementioned event timeline information, a timeline dimension for constructing the user profile is built. The timeline dimension includes event nodes arranged chronologically, with each event node associated with the participating entities and event type of the corresponding event. This invention uses basic user information as the basic information dimension for the profile, including core attribute fields such as age, occupation, income, and health indicators confirmed through disambiguation fusion, along with their source identifiers. Simultaneously, the event timeline information is used as the timeline dimension for the profile. This dimension consists of event nodes arranged chronologically, with each event node including the event occurrence time, event type, and the identifiers of the entities participating in the event, thus forming a complete and traceable user interaction trajectory. In a financial scenario, the basic information dimension of customer Zhang's user profile includes attributes such as an annual income of 300,000 yuan, age 42, and occupation as a middle-level manager in a company. The timeline dimension includes event nodes such as the interview on March 5th, product comparison consultation on March 12th, and asset certificate submission on March 20th. Each node is associated with the customer and participating entities such as their spouse. This invention constructs basic information and event trajectories as two independent dimensions of a profile, enabling practitioners to quickly view all static attribute information of customers and trace the complete interaction process along the timeline, thus upgrading information retrieval from scattered queries to structured browsing.

[0034] S52: Constructing a tag dimension for the user profile based on the aforementioned attention tags, wherein the tag dimension includes user-focused topic tags, sentiment tags, and behavioral pattern tags. The attention tendencies, emotional states, and behavioral patterns revealed by customers during interactions are key bases for developing personalized service strategies, requiring extraction and solidification from intramodal information units into independent dimensions of the profile. This invention constructs a tag dimension for the user profile based on the attention tags generated in the preceding steps. This dimension includes three sub-dimensions: attention topic tags, sentiment tags, and behavioral pattern tags, respectively reflecting the customer's focus at the content level, emotional state at the psychological level, and behavioral patterns at the time level. For example, in a financial scenario, customer Zhang's tag dimension includes the attention topic tag of children's education fund planning, the sentiment tag of future expenditure anxiety, and the behavioral pattern tag of proactively seeking advice at the beginning of the year. These three tags together outline the customer's core needs profile. This invention makes the customer characteristics implicit in the original interaction data explicit into a structured tag system, enabling practitioners to quickly grasp the customer's core concerns and behavioral characteristics without relying on personal experience or reviewing original records.

[0035] S53: Establish cross-dimensional association indexes for the basic information dimension, the timeline dimension, and the tag dimension using entity identifiers to generate a multi-dimensional user profile. Specifically, in this embodiment of the invention, although the basic information dimension, the timeline dimension, and the tag dimension are constructed independently, they describe different aspects of the same user and have inherent logical connections with each other. For example, a topic tag in the tag dimension may have a causal relationship with a specific event node in the timeline dimension, and a change in an attribute in the basic information dimension may be synchronized with an event sequence in the timeline. This invention uses entity identifiers as association keys to establish cross-dimensional association indexes for attribute fields in the basic information dimension, event nodes in the timeline dimension, and tag entries in the tag dimension, making the three dimensions mutually anchored, linked, and searchable, ultimately generating a multi-dimensional user profile. For example, in a financial scenario, a multi-dimensional user profile of customer Zhang is established with cross-dimensional associations using the entity identifier Zhang. When practitioners query the children's education fund planning tag in the tag dimension, the system can automatically locate the event nodes in the timeline dimension where the customer repeatedly mentions children's education, and the customer's age and family structure attributes in the basic information dimension through the association index, thereby understanding the background basis for the formation of the tag. This invention integrates three originally independent dimensions into an organically linked overall profile through cross-dimensional correlation indexing, enabling practitioners to obtain relevant information from other dimensions from any dimension. This achieves a cognitive upgrade from single-point query to multi-dimensional linkage, significantly improving the practical value and decision support capabilities of user profiles.

[0036] In one embodiment of the present invention, after constructing a user profile of the target user based on the target user's basic information, the event timeline information, and the attention tags, the method further includes: S61: Sensitive information identification is performed on the attribute information in the user profile. Based on preset desensitization rules, the identified sensitive information is desensitized to generate a desensitized user profile. Specifically, in this embodiment of the invention, the user profile integrates the fusion results of multimodal information such as voice, image, and dialogue text, which may include legally protected sensitive information such as name, ID number, contact information, bank account, and health indicators. This step uses preset sensitive information identification rules to traverse and scan the attribute information in each dimension of the user profile, identify attribute fields belonging to preset sensitive types, and mask, replace, or generalize them based on the corresponding desensitization rules to generate a desensitized user profile that can be used in general business scenarios. For example, in a financial scenario, the system performs sensitive information identification on the user profile of customer Zhang, identifying the name, ID number, and mobile phone number as sensitive information. The surname is retained, the given name is replaced with an asterisk, the middle digits of the ID number are masked, and the middle four digits of the mobile phone number are desensitized to generate a desensitized user profile for daily business reference. This invention, through automated sensitive information identification and desensitization, preserves the complete usability of non-sensitive information in the profile while ensuring user privacy and data compliance, thus avoiding the problem of profiles being unusable in general business scenarios due to blanket permission restrictions.

[0037] S62: In response to an access request for the user profile, the system performs permission verification on the access request and records an access log, wherein the access log includes the access subject, access time, and access operation type. Specifically, in this embodiment of the invention, the user profile carries comprehensive user information, and its access behavior needs to be incorporated into a unified permission control system. Upon receiving an access request for the user profile, this step first extracts the requester's identity identifier, verifies whether it has access rights to the target profile, and allows access only after successful verification. Simultaneously, the identity of the accessing subject, access timestamp, and access operation types such as query, modification, and export are recorded in the access log, forming a complete access audit chain. In a financial scenario, a financial advisor requests access to the user profile of client Zhang through an internal system. The system verifies the advisor's identity identifier, confirms that it falls within the authorized service scope of the client, allows access after successful verification, and records the access log, which includes the advisor's employee number, access time, and query operation type. If a back-end administrator attempts to export customer profile data in batches, the system records the access log of its export operation, providing a basis for subsequent compliance review. This invention establishes a complete traceability chain for user profile access behavior through mandatory permission verification and full access log recording, effectively preventing unauthorized access and information leakage, while providing tamper-proof log evidence for compliance auditing and post-event accountability.

[0038] S63: When unauthorized access or abnormal changes to attribute information in the user profile are detected, an audit alarm is generated, and a data snapshot before and after the change is recorded. Specifically, in this embodiment of the invention, user profiles may face two types of security risks during continuous updates and use: one is external attackers attempting to access the profile through unauthorized means, and the other is abnormal changes to attribute information caused by internal authorized users or system anomalies. This step monitors the access behavior and attribute change operations of the profile in real time. When an unauthorized access attempt or an abnormal change in attribute information exceeding a preset threshold is detected, an audit alarm is automatically generated to notify security management personnel, and a data snapshot before and after the change is recorded to provide complete evidence for event tracing. In a financial scenario, the system detects that an IP address repeatedly attempts to access the user profile of customer Zhang using different employee accounts late at night, which is determined to be unauthorized access behavior. An audit alarm is generated, and the IP address, the attempting account, and the time series are recorded. At the same time, the system detects that the annual income attribute in Zhang's profile has been abnormally modified from 300,000 yuan to 1 million yuan, which is significantly different from the original value recorded in the disambiguation processing log, triggering an abnormal change alarm and saving a data snapshot before and after the modification. This invention upgrades the security protection of user profiles from passive recording to active monitoring and real-time alarms, enabling a response to the first occurrence of a security incident. At the same time, by saving snapshots before and after changes, it provides a complete basis for incident tracing and data recovery, significantly improving the overall security of the user profile system.

[0039] As can be seen, in the above scheme, cross-modal entity alignment processing is used to map entities extracted from different modalities across modalities, identify the same entity from multiple information sources, and perform consistency verification and disambiguation processing on the attribute information of the identified entities, thereby improving the accuracy of multimodal information fusion. At the same time, entities, relationships, and event elements extracted from each modality are associated and integrated to generate event timelines and attention tags, transforming scattered process information into structured user profiles, thereby improving the information richness of user profiles and the efficiency of multimodal information processing.

[0040] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0041] In one embodiment, a user profile construction apparatus based on multimodal information is provided, which corresponds one-to-one with the user profile construction method based on multimodal information described in the above embodiments. For example... Figure 6 As shown, Figure 6This is a schematic diagram of a user profile building device based on multimodal information according to an embodiment of the present invention. The device includes an acquisition module 61, an analysis module 62, a processing module 63, a fusion module 64, and a construction module 65. Detailed descriptions of each functional module are as follows: The acquisition module is used to acquire and parse the multimodal information of the target user, and extract entities, relationships and event elements under each modality, wherein the entities include at least person entities; The analysis module is used to perform cross-correlation analysis on the entities, relationships and event elements extracted from each modality, and generate event timeline information and attention tags for the target user. The processing module is used to perform cross-modal entity alignment processing on entities extracted from different modalities, and to perform consistency verification on the attribute information of the entities after cross-modal entity alignment processing in different modalities based on preset rules. The fusion module is used to disambiguate conflicting attribute information after consistency verification, and then fuse the disambiguated attribute information to generate the fused basic information of the target user. The construction module is used to construct a user profile of the target user based on the target user's basic information, the event timeline information, and the attention tags.

[0042] In one embodiment, the acquisition module 61 is specifically used for: Acquire voice data, image data, and conversation text data associated with the target user; The speech data is processed by speech recognition and transcription to generate a timestamped dialogue text. Key paragraph information in the dialogue text is identified, and entity, relationship and event elements are extracted based on the identified key paragraph information. The image data is processed by optical character recognition to extract the person entities, relationship indicators and index values ​​from the image data, and the entity, relationship and event elements in the image data are extracted based on the extraction results. The intent is identified in the dialogue text data, and entity, relationship and event elements are extracted from the dialogue text data based on the intent identification results.

[0043] In one embodiment, the analysis module 62 is specifically used for: Within the same modality, the extracted entities, relationships, and event elements are associated to generate intramodal information units. Each intramodal information unit contains a set of interconnected entities, a relationship network, and an event description within the corresponding modality. The event elements in the information units within each modality are arranged in chronological order to generate cross-modality event timeline information; Content analysis is performed on the information units within the modality to extract the target user's focus topics, sentiment tendencies, and behavioral patterns, and to generate the target user's focus tags.

[0044] In one embodiment, the processing module 63 is specifically used for: Based on the entity's name, relationship network, and associated event elements, the similarity between entities in different modalities is calculated. Entities with similarity exceeding a preset threshold are identified as the same entity, and cross-modal entity mapping relationships are established. For the same entity with cross-modal entity mapping relationships, extract the attribute information of the corresponding entity in each modality; The attribute information of the same entity under different modalities is compared. When the difference in attribute information under different modalities exceeds the preset tolerance range, it is determined that there is a conflict in the corresponding attribute information and a conflict mark is generated.

[0045] In one embodiment, the fusion module 64 is specifically used for: For attribute information carrying conflict markers, determine the confidence weight of the corresponding attribute information under each modality, and select the attribute information with the highest confidence weight as the disambiguation result; The disambiguated attribute information is merged with the attribute information that does not conflict to generate basic user information, and the source modality, original value and disambiguation processing log of each attribute information are recorded.

[0046] In one embodiment, the construction module 65 is specifically used for: The user profile is constructed based on the user basic information, and the user profile is constructed based on the event timeline information, wherein the timeline dimension includes event nodes arranged in chronological order, and each event node is associated with the participating entity and event type of the corresponding event. The tag dimensions for constructing user profiles based on the aforementioned attention tags include user-related topic tags, sentiment tags, and behavioral pattern tags. By establishing a cross-dimensional association index using entity identifiers for the basic information dimension, the time axis dimension, and the tag dimension, a multi-dimensional user profile is generated.

[0047] In one embodiment, the user profile construction device based on multimodal information is further configured to: Sensitive information is identified in the attribute information of the user profile, and the identified sensitive information is desensitized based on preset desensitization rules to generate a desensitized user profile. In response to an access request to the user profile, the access request to the user profile is subject to permission verification and an access log is recorded, wherein the access log includes the access subject, access time and access operation type. When unauthorized access or abnormal changes to attribute information in the user profile are detected, an audit alert is generated, and data snapshots before and after the change are recorded.

[0048] This invention provides a user profile construction device based on multimodal information. Through cross-modal entity alignment processing, entities extracted from different modalities are mapped across modalities to identify the same entity from multiple information sources. The device also performs consistency verification and disambiguation processing on the attribute information of the identified entities, thereby improving the accuracy of multimodal information fusion. Simultaneously, it integrates and associates entities, relationships, and event elements extracted from various modalities to generate event timelines and attention tags, transforming scattered process information into a structured user profile and improving the information richness of user profiles and the efficiency of multimodal information processing.

[0049] Specific limitations regarding the user profile construction device based on multimodal information can be found in the limitations of the user profile construction method based on multimodal information above, and will not be repeated here. Each module in the aforementioned user profile construction device based on multimodal information can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0050] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, Figure 7 This is a schematic diagram of a computer device according to an embodiment of the present invention. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a user profile construction method based on multimodal information on the server side.

[0051] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 8 As shown, Figure 8This is another structural schematic diagram of a computer device according to an embodiment of the present invention. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the client-side functions or steps of a user profile construction method based on multimodal information.

[0052] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain multimodal information of the target user, perform content parsing on each modal information, and extract entity, relationship and event elements under each modality. The entities include at least person entities. The system performs correlation analysis on the entities, relationships, and event elements of each modality to generate event timeline information and attention tags for the target user. Cross-modal entity alignment is performed on entities extracted from different modalities, and based on preset rules, the consistency of attribute information of the human entity after cross-modal entity alignment is verified in different modalities. Disambiguation is performed on conflicting attribute information after consistency verification, and the disambiguated attribute information is then fused to generate the fused basic information of the target user. Based on the fused basic information of the target user, event timeline information, and attention tags, a user profile of the target user is constructed.

[0053] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain multimodal information of the target user, perform content parsing on each modal information, and extract entity, relationship and event elements under each modality. The entities include at least person entities. The system performs correlation analysis on the entities, relationships, and event elements of each modality to generate event timeline information and attention tags for the target user. Cross-modal entity alignment is performed on entities extracted from different modalities, and based on preset rules, the consistency of attribute information of the human entity after cross-modal entity alignment is verified in different modalities. Disambiguation is performed on conflicting attribute information after consistency verification, and the disambiguated attribute information is then fused to generate the fused basic information of the target user. Based on the fused basic information of the target user, event timeline information, and attention tags, a user profile of the target user is constructed.

[0054] This invention achieves automatic conflict resolution and fusion generation of reliable basic information for multimodal information through cross-modal entity alignment and consistency verification; and transforms scattered process information into structured user profiles by automatically generating event timelines and attention tags, thereby improving the information richness of user profiles and the efficiency of multimodal information processing.

[0055] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0056] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0057] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0058] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0059] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for constructing user profiles based on multimodal information, characterized in that, include: Acquire and parse the multimodal information of the target user, and extract the entity, relationship and event elements under each modality, wherein the entity includes at least the person entity; Cross-correlation analysis is performed on the entities, relationships, and event elements extracted from each modality to generate event timeline information and attention tags for the target user; Cross-modal entity alignment is performed on entities extracted from different modalities, and the consistency of attribute information of entities after cross-modal entity alignment is verified in different modalities based on preset rules. Disambiguation is performed on conflicting attribute information after consistency verification, and the disambiguated attribute information is then fused to generate the fused basic information of the target user. Based on the target user's basic information, the event timeline information, and the attention tags, a user profile of the target user is constructed.

2. The user profile construction method based on multimodal information according to claim 1, characterized in that, The process of acquiring and parsing the target user's multimodal information, extracting entities, relationships, and event elements under each modality, wherein the entities include at least person entities, including: Acquire voice data, image data, and conversation text data associated with the target user; The speech data is processed by speech recognition and transcription to generate a timestamped dialogue text. Key paragraph information in the dialogue text is identified, and entity, relationship and event elements are extracted based on the identified key paragraph information. The image data is processed by optical character recognition to extract the person entities, relationship indicators and index values ​​from the image data, and the entity, relationship and event elements in the image data are extracted based on the extraction results. The intent is identified in the dialogue text data, and entity, relationship and event elements are extracted from the dialogue text data based on the intent identification results.

3. The user profile construction method based on multimodal information according to claim 1, characterized in that, The process involves cross-referencing the entities, relationships, and event elements extracted from each modality to generate event timeline information and attention tags for the target user, including: Within the same modality, the extracted entities, relationships, and event elements are associated to generate intramodal information units. Each intramodal information unit contains a set of interconnected entities, a relationship network, and an event description within the corresponding modality. The event elements in the information units within each modality are arranged in chronological order to generate cross-modality event timeline information; Content analysis is performed on the information units within the modality to extract the topics, sentiments, and behavioral patterns of the target user, and to generate tags representing the target user's points of interest.

4. The user profile construction method based on multimodal information according to claim 3, characterized in that, The process of performing cross-modal entity alignment on entities extracted from different modalities, and then performing consistency verification on the attribute information of the aligned entities in different modalities based on preset rules, includes: Based on the entity's name, relationship network, and associated event elements, the similarity between entities in different modalities is calculated. Entities with similarity exceeding a preset threshold are identified as the same entity, and cross-modal entity mapping relationships are established. For the same entity with cross-modal entity mapping relationships, extract the attribute information of the corresponding entity in each modality; The attribute information of the same entity under different modalities is compared. When the difference in attribute information under different modalities exceeds the preset tolerance range, it is determined that there is a conflict in the corresponding attribute information and a conflict mark is generated.

5. The user profile construction method based on multimodal information according to claim 1, characterized in that, The process involves disambiguating conflicting attribute information after consistency verification and then fusing the disambiguated attribute information to generate fused basic information for the target user, including: For attribute information carrying conflict markers, determine the confidence weight of the corresponding attribute information under each modality, and select the attribute information with the highest confidence weight as the disambiguation result; The disambiguated attribute information is merged with the attribute information that does not conflict to generate basic user information, and the source modality, original value and disambiguation processing log of each attribute information are recorded.

6. The user profile construction method based on multimodal information according to claim 1, characterized in that, The process of constructing a user profile for the target user based on the target user's basic information, the event timeline information, and the attention tags includes: The user profile is constructed based on the user basic information, and the user profile is constructed based on the event timeline information, wherein the timeline dimension includes event nodes arranged in chronological order, and each event node is associated with the participating entity and event type of the corresponding event. The tag dimensions for constructing user profiles based on the aforementioned attention tags include user-related topic tags, sentiment tags, and behavioral pattern tags. By establishing a cross-dimensional association index using entity identifiers for the basic information dimension, the time axis dimension, and the tag dimension, a multi-dimensional user profile is generated.

7. The user profile construction method based on multimodal information according to claim 1, characterized in that, After constructing the user profile of the target user based on the target user's basic information, the event timeline information, and the attention tags, the method further includes: Sensitive information is identified in the attribute information of the user profile, and the identified sensitive information is desensitized based on preset desensitization rules to generate a desensitized user profile. In response to an access request to the user profile, the access request to the user profile is subject to permission verification and an access log is recorded, wherein the access log includes the access subject, access time and access operation type. When unauthorized access or abnormal changes to attribute information in the user profile are detected, an audit alert is generated, and data snapshots before and after the change are recorded.

8. A user profile construction device based on multimodal information, characterized in that, include: The acquisition module is used to acquire and parse the multimodal information of the target user, and extract entities, relationships and event elements under each modality, wherein the entities include at least person entities; The analysis module is used to perform cross-correlation analysis on the entities, relationships and event elements extracted from each modality, and generate event timeline information and attention tags for the target user. The processing module is used to perform cross-modal entity alignment processing on entities extracted from different modalities, and to perform consistency verification on the attribute information of the entities after cross-modal entity alignment processing in different modalities based on preset rules. The fusion module is used to disambiguate conflicting attribute information after consistency verification, and then fuse the disambiguated attribute information to generate the fused basic information of the target user. The construction module is used to construct a user profile of the target user based on the target user's basic information, the event timeline information, and the attention tags.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of creating a user profile based on multimodal information as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of creating a user profile based on multimodal information as described in any one of claims 1 to 7.