Multimedia memoir generation method and device based on multi-dimensional emotion perception

By collecting and analyzing users' multi-dimensional emotional data in real time, dynamically adapting visual and auditory feedback, constructing immersive interactive scenarios, and capturing emotional peak materials, the problem of lack of emotional analysis in multimedia memoir generation tools is solved, and memoir generation with greater emotional value and immersion is achieved.

CN121858709APending Publication Date: 2026-04-14PING AN HEALTH CLOUD CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing multimedia memoir generation tools lack in-depth analysis of users' emotional information and cannot dynamically adapt to visual and auditory feedback, resulting in a lack of immersion and emotional expression in the interaction process.

Method used

By collecting users' facial images, voice, and physiological data in real time, multi-dimensional emotion analysis is performed, and visual and auditory feedback materials are dynamically adapted to construct immersive interactive scenes and extract audio and video clips of emotional peaks.

Benefits of technology

It enhances the immersiveness and user-friendliness of the interactive process, and the generated multimedia memoirs can fully present core information and incorporate real emotional moments, satisfying the deep-seated need for emotional and immersive memory preservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858709A_ABST
    Figure CN121858709A_ABST
Patent Text Reader

Abstract

The invention discloses a multimedia memoir generation method and device based on multi-dimensional emotion perception, relates to the technical field of artificial intelligence, can be applied to financial science and technology and medical health business scenarios, and comprises the steps that in the AI interaction process, multi-dimensional emotion related data expressed by user recall is collected in real time; comprehensively analyzing the multi-dimensional emotion related data, and determining a real-time emotion type and emotion intensity of the user; according to the real-time emotion type and the emotion intensity, corresponding visual feedback materials and auditory feedback materials are dynamically and adaptively generated, and an immersive interaction scene is constructed based on the visual feedback materials and the auditory feedback materials; in the immersive interaction scene, intercepting a plurality of audio and video material slices when the emotion of the user reaches a peak value; and integrating the plurality of audio and video material slices and the core content expressed by the user memory to obtain the multimedia memoir of the user. The emotion information of the user in the memory expression process can be effectively captured and fused.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and apparatus for generating multimedia memoirs based on multi-dimensional emotion perception. Background Technology

[0002] With the diversification of population structure and the deep integration of digital technology with the financial and medical fields, the need for preserving personal life trajectories and key experiences among various groups is gradually upgrading from simple emotional records to those with emotional value and practical relevance. In the financial sector, whether it's young people in the early stages of wealth accumulation, middle-aged people at the peak of their wealth, or middle-aged and elderly people facing wealth inheritance, their financial decision-making logic, asset planning process, and the transmission of family wealth concepts urgently need to be preserved through concrete carriers, becoming an important link connecting intergenerational financial cognition and accumulating personal financial experience. In the medical field, the health management trajectory, diagnosis and rehabilitation experience, and doctor-patient interaction details of users of different ages are not only supplements to personal health records, but also valuable resources for conveying life experiences and providing health references for others. Multimedia memoirs, with their advantages of integrating text, images, audio, and video materials, can simultaneously carry the emotional memories, financial experiences, and medical-related trajectories of various groups, becoming a comprehensive carrier with both emotional value and practical significance.

[0003] Currently, mainstream AI dialogue generation tools guide users to input details of their memories through text-based question-and-answer sessions. Then, based on preset templates, the text content is spliced ​​with fixed images, background music, and other materials to form a basic multimedia memoir. Some advanced products support voice or video capture functions, but their core is only the recording and storage of raw audio and video data. They do not conduct in-depth analysis of the user's state during the capture process. The visual presentation, background sound effects, etc., are all fixed settings, and they can only achieve simple content overlay. They lack dynamic adaptation to the user's expression process and cannot effectively capture and integrate the user's emotional information during the memory expression process. They cannot meet the deep needs of different groups for emotional and immersive memory retention. Summary of the Invention

[0004] In view of this, this application provides a multimedia memoir generation method and device based on multi-dimensional emotion perception, which can effectively capture and integrate the emotional information of users in the process of expressing memories, and meet the deep needs of different groups for emotional and immersive memory retention.

[0005] According to a first aspect of this application, a method for generating multimedia memoirs based on multi-dimensional emotion perception is provided, comprising: During AI interaction, multi-dimensional emotion-related data of user recall and expression are collected in real time. The multi-dimensional emotion-related data includes at least facial image data, voice data and physiological sign data. By comprehensively analyzing the multi-dimensional emotion-related data, the user's real-time emotion type and intensity can be determined; Based on the real-time emotion type and the emotion intensity, corresponding visual feedback materials and auditory feedback materials are dynamically adapted and generated, and an immersive interactive scene is constructed based on the visual feedback materials and the auditory feedback materials; In the immersive interactive scenario, multiple audio and video clips are captured when the user's emotions reach their peak. By integrating the multiple audio and video material slices and the core content of the user's recollections, the user's multimedia memoir is obtained.

[0006] According to a second aspect of this application, a multimedia memoir generation device based on multi-dimensional emotion perception is provided, comprising: The data acquisition module is used to collect multi-dimensional emotion-related data that the user recalls and expresses in real time during the AI ​​interaction process. The multi-dimensional emotion-related data includes at least facial image data, voice data, and physiological sign data. The analysis module is used to comprehensively analyze the multi-dimensional emotion-related data to determine the user's real-time emotion type and intensity. The module is used to dynamically adapt and generate corresponding visual and auditory feedback materials based on the real-time emotion type and the emotion intensity, and to construct an immersive interactive scene based on the visual and auditory feedback materials. The capture module is used to capture multiple audio and video clips when the user's emotions reach their peak in the immersive interactive scenario. The integration module is used to integrate the multiple audio and video material slices and the core content of the user's recollection to obtain the user's multimedia memoir.

[0007] According to a third aspect of this application, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described method for generating multimedia memoirs based on multi-dimensional emotion perception.

[0008] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described method for generating multimedia memoirs based on multi-dimensional emotion perception.

[0009] By employing the aforementioned technical solutions, the multimedia memoir generation method and apparatus based on multi-dimensional emotion perception provided in this application accurately determines the user's real-time emotion type and intensity by comprehensively analyzing multi-dimensional emotion-related data such as user facial images, voice, and physiological signs in real time. This allows for the dynamic adaptation and generation of corresponding visual and auditory feedback materials, constructing an immersive interactive scene. This effectively solves the problems of fixed audiovisual presentation and lack of dynamic adaptation in existing technologies, significantly enhancing the immersiveness and user-friendliness of the interaction process and fully mobilizing the recollection expression desires of different groups. Furthermore, by extracting audio and video clips from the user's emotional peak and integrating them with the core content of the memory, the generated multimedia memoir not only fully presents the core memory information but also incorporates genuine emotional moments and state imprints. This overcomes the shortcomings of existing technologies in effectively capturing and integrating user emotional information, meeting the deep-seated needs of different groups for emotional and immersive memory preservation, and significantly enhancing the emotional value and commemorative significance of the memoir.

[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a multimedia memoir generation method based on multi-dimensional emotion perception provided in an embodiment of this application is shown. Figure 2 A flowchart illustrating a multimedia memoir generation method based on multi-dimensional emotion perception, according to another embodiment of this application, is shown. Figure 3 The diagram shows a schematic representation of a multimedia memoir generation device based on multi-dimensional emotion perception, as provided in an embodiment of this application. Detailed Implementation

[0012] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0013] Currently, mainstream AI dialogue generation tools guide users to input details of their memories through text-based question-and-answer sessions. Then, based on preset templates, the text content is spliced ​​with fixed images, background music, and other materials to form a basic multimedia memoir. Some advanced products support voice or video capture functions, but their core is only the recording and storage of raw audio and video data. They do not conduct in-depth analysis of the user's state during the capture process. The visual presentation, background sound effects, etc., are all fixed settings, and they can only achieve simple content overlay. They lack dynamic adaptation to the user's expression process and cannot effectively capture and integrate the user's emotional information during the memory expression process. They cannot meet the deep needs of different groups for emotional and immersive memory retention.

[0014] To address the aforementioned technical problems, embodiments of the present invention provide a multimedia memoir generation method based on multi-dimensional emotion perception, such as... Figure 1 As shown, the method includes: Step 110: During the AI ​​interaction process, collect multi-dimensional emotion-related data in real time, including at least facial image data, voice data, and physiological signs data.

[0015] Among them, AI interaction refers to the interactive process between users and artificial intelligence systems around the content related to recall and expression, including two-way communication behaviors such as information transmission and command response; multi-dimensional emotion-related data refers to the collection of various information related to the user's emotional state from multiple different dimensions, covering relevant data at the visual, auditory, and physiological levels; facial image data is image information that reflects the user's facial state obtained through image acquisition devices; voice data is sound information that is recorded by audio acquisition devices during the user's recall and expression process; and physiological sign data is various indicators that reflect the user's physical physiological state and is an important auxiliary information reflecting changes in the user's emotions.

[0016] In this embodiment of the disclosure, during the interaction between the user and the artificial intelligence system regarding the content of the memory, the system relies on a preset acquisition module and associated devices to synchronously and in real time acquire multi-dimensional data related to the user's emotions. Specifically, it can capture facial image data of the user when expressing the memory through a camera, record the user's voice data through an audio acquisition device, and simultaneously connect to the wearable device worn by the user to encrypt and acquire the user's physiological data. The acquisition of these three types of data is carried out synchronously with the user's memory expression process, comprehensively covering the visual, auditory, and physiological levels, forming a complete emotion-related data acquisition chain, which can provide data support for subsequent emotion analysis and scene adaptation.

[0017] By collecting multi-dimensional emotion-related data such as facial images, voice, and physiological signs in real time during AI interaction, it is possible to comprehensively and timely capture the emotional association information when users recall and express themselves. This breaks the limitations of a single data dimension and provides a comprehensive and reliable data foundation for accurately determining the type and intensity of users' real-time emotions. This ensures that subsequent dynamic adaptation based on emotions, immersive scene construction, and extraction of emotional peak materials are more targeted, effectively solving the problems of one-sided and delayed emotional information capture in traditional solutions. This lays a solid data foundation for the final generation of emotional and immersive multimedia memoirs.

[0018] Step 120: Conduct a comprehensive analysis of multi-dimensional emotion-related data to determine the user's real-time emotion type and intensity.

[0019] Among them, real-time emotion type refers to the specific emotion category that the user is in at the current interaction moment (such as joy, longing, excitement, calmness, etc.); emotion intensity refers to the intensity of the user's current emotion type, which is a quantitative description of the emotion level.

[0020] In this embodiment of the disclosure, the collected multi-dimensional emotion-related data, including facial images, voice, and physiological signs, can be analyzed separately for each dimension. Specifically, key features can be extracted from facial images to identify visual emotion and confidence level. After analyzing auditory emotion features from voice data, auditory emotion and confidence level can be obtained through neural network model analysis. Physiological sign data can be compared parameter by parameter with a preset general physiological baseline to determine physiological emotion and confidence level. Then, dynamic weights are assigned to the emotion recognition results of the three dimensions according to the scene adaptation requirements. Finally, a multi-dimensional weighted fusion algorithm is used to calculate the weighted emotion recognition results and confidence level of each dimension in combination with the dynamic weights. The emotion type with a comprehensive confidence level greater than a preset threshold is selected as the user's real-time emotion type, and the emotion intensity is determined based on the numerical range of the comprehensive confidence level.

[0021] Step 130: Based on the real-time emotion type and intensity, dynamically adapt and generate corresponding visual and auditory feedback materials, and construct an immersive interactive scene based on the visual and auditory feedback materials.

[0022] Among them, visual feedback materials are content carriers used for visual feedback, which can present visual effects that match emotions through images; auditory feedback materials are content carriers used for auditory feedback, which can present auditory effects that match emotions through sound; immersive interactive scenarios are environments that create a sense of immersion and focus on the interaction process by integrating appropriate visual and auditory feedback.

[0023] In this embodiment of the disclosure, based on the determined real-time emotion type and intensity of the user, the system first retrieves initial visual materials matching the emotion type from a preset visual material library, and then adjusts the visual parameters of the initial visual materials, such as color tone, clarity, dynamic element type, and filter effects, according to the emotion intensity to generate visual feedback materials that fit the current emotion. At the same time, it retrieves initial sound effect materials matching the emotion type from a preset sound effect library, and adjusts the auditory parameters such as background music BPM, timbre selection, and ambient sound volume ratio according to the emotion intensity to obtain suitable auditory feedback materials. Finally, the adjusted visual feedback materials and auditory feedback materials are synchronously superimposed and output to jointly construct an immersive interactive scene that is highly consistent with the user's real-time emotional state.

[0024] By dynamically adapting and generating visual and auditory feedback materials based on the user's real-time emotional type and intensity, and constructing immersive interactive scenes, this replaces the traditional fixed interactive presentation method. It allows the interaction process to respond to the user's emotional changes in real time, enriching the sensory feedback dimensions of the interaction and creating an atmosphere that fits the user's expression state. This effectively enhances the immersion and personalization of the interaction, increases the user's willingness and enjoyment to participate in the interaction, and lays the scene foundation for subsequent extraction of emotional peak audio and video materials and generation of emotional memoirs, making the entire process of expressing memories warmer and more adaptable.

[0025] Step 140: In an immersive interactive scenario, capture multiple audio and video clips when the user's emotions reach their peak.

[0026] Among them, audio and video material slices are fragmented audio and video materials extracted from immersive interactive scenes, containing visual images and auditory content at specific times, used to preserve key emotional moments.

[0027] In the embodiments of this disclosure, in the constructed immersive interactive scenario, the system can first preset the peak threshold of emotional intensity corresponding to different emotional types and the associated multi-dimensional data triggering conditions, and synchronize the user's emotional intensity data and multi-dimensional emotional related data in real time. When it is detected that the user's emotional intensity exceeds the corresponding peak threshold and the multi-dimensional emotional related data meets the triggering conditions, the system automatically extracts audio and video segments of preset length as audio and video material slices. Then, it adds exclusive tags such as emotional type, timestamp, and intensity level to each slice and stores the tagged slices in the material database for subsequent integration with the core content of the memory.

[0028] By accurately capturing and tagging multiple audio and video clips at the peak of user emotions in immersive interactive scenarios, the most emotionally charged key moments in the user's recollection and expression process can be effectively captured. This breaks through the limitations of traditional solutions that only retain core text or original audio and video and lack emotional imprints. It can provide valuable emotional material support for the subsequent integration and generation of emotional multimedia memoirs, so that the final product can not only present the core content of the memory, but also restore the true emotional expression state, greatly enhancing the emotional value and commemorative significance of the memoir. At the same time, the tagging design can also lay an efficient foundation for the accurate matching and integration of subsequent materials.

[0029] Step 150: Integrate multiple audio and video material slices and the core content of the user's recollection to obtain the user's multimedia memoir.

[0030] Among them, the core content of the user's recollection is extracted from the interaction data between the user and AI, which can reflect the key information of the recollection (such as key events, relationships between people, time nodes, core viewpoints, etc.); the multimedia memoir is the final result carrier that integrates multiple media forms such as text, audio and video, dynamic visuals, and appropriate sound effects to comprehensively present the user's recollection content and emotional state.

[0031] In this embodiment of the disclosure, the core content of the user's recollection can be extracted from the interactive data by an AI semantic analysis algorithm and organized into a timeline. At the same time, multiple stored audio and video material slices are analyzed by built-in audiovisual parameters. Based on the timestamp tags of the slices, they are accurately matched with the text narrative paragraphs corresponding to the core content. Then, the text presentation style is adjusted to be consistent with the visual style of the slices, the clarity of the user's voice in the slices is optimized, and the playback duration and transition effects of the text and slices are adapted. Finally, according to the timeline logic of the core content, all the materials that have been collaboratively optimized are systematically integrated to form a complete multimedia memoir.

[0032] By integrating multiple audio and video clips capturing the core content of a user's recollections and moments of emotional peak, the limitations of traditional memoirs relying solely on text or single materials can be overcome. This allows the presentation of memories to include key information while preserving genuine emotional moments. Furthermore, through audiovisual parameter adaptation and temporal optimization, the coherence of the content and the coordination of the presentation can be ensured. This results in a multimedia memoir with a multi-dimensional and three-dimensional expressive effect, enriching the presentation of memories and strengthening the preservation of emotional imprints. This significantly enhances the vividness and commemorative value of the memoir, fully satisfying users' dual needs for preserving both the content of their memories and their emotional experiences.

[0033] In summary, the multimedia memoir generation method based on multi-dimensional emotion perception provided in this application accurately determines the user's real-time emotion type and intensity by collecting and comprehensively analyzing multi-dimensional emotion-related data such as user facial images, voice, and physiological signs in real time. It then dynamically adapts and generates corresponding visual and auditory feedback materials to construct an immersive interactive scene. This effectively solves the problems of fixed audiovisual presentation and lack of dynamic adaptation in existing technologies, significantly improving the immersiveness and user-friendliness of the interaction process and fully mobilizing the willingness of different groups to express their memories. Furthermore, by extracting audio and video clips at the peak of the user's emotions and integrating them with the core content of the memory, the generated multimedia memoir not only fully presents the core memory information but also incorporates real emotional moments and state imprints. This compensates for the shortcomings of existing technologies in effectively capturing and integrating user emotional information, meeting the deep needs of different groups for emotional and immersive memory preservation, and significantly enhancing the emotional value and commemorative significance of the memoir.

[0034] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, and to fully illustrate the implementation methods of this embodiment, this embodiment also provides another multimedia memoir generation method based on multi-dimensional emotion perception, such as... Figure 2 As shown, the method includes: Step 210: During the AI ​​interaction process, collect multi-dimensional emotion-related data in real time, including at least facial image data, voice data, and physiological signs data.

[0035] In this embodiment of the disclosure, during the interaction between the user and the AI ​​system regarding past memories, the system simultaneously collects data through a pre-set multi-module acquisition link. Specifically, it can use a camera to capture the user's facial dynamics in real time, accurately acquiring facial image data such as the distribution of facial feature points, subtle expression changes, lip movement frequency, and eye movement trajectory; it can record the user's voice content throughout the process using an audio acquisition device (such as a microphone), completely preserving voice data such as tone and speed during the expression of memories; simultaneously, it connects to wearable devices such as smartwatches and bracelets worn by the user, acquiring the user's physiological characteristics such as heart rate, body temperature, and blood pressure in real time under encrypted transmission. These three types of data are collected in parallel and summarized synchronously to form a complete set of emotion-related data covering visual, auditory, and physiological dimensions, providing comprehensive and accurate raw data support for subsequent emotion analysis.

[0036] Accordingly, the steps of the embodiment may include: capturing facial image data of the user recalling the expression in real time through a camera, the facial image data including at least facial feature points, micro-expressions, lip movement frequency and eye movement trajectory data; recording voice data of the user recalling the expression in real time through an audio acquisition device; and acquiring physiological sign data collected by the user's wearable device when the user recalls the expression, the physiological sign data including at least the user's heart rate, body temperature and blood pressure data.

[0037] Step 220: Conduct a comprehensive analysis of multi-dimensional emotion-related data to determine the user's real-time emotion type and intensity.

[0038] For embodiments of this disclosure, step 220 may include the following steps: Step 220-1: Extract key facial features from facial image data, identify emotion types based on key facial features, and obtain the visual dimension emotion recognition results and corresponding confidence scores.

[0039] In this embodiment of the disclosure, facial image data of the user recalling the expression can be obtained first. The image is then processed by a preset computer vision algorithm to accurately extract key facial features, including micro-expression details reflecting emotions (such as the curvature of the corners of the mouth and wrinkles around the eyes), lip movement frequency, eye movement trajectory, and facial feature point distribution. The extracted key facial features are then input into a pre-trained emotion classification model. The model identifies the corresponding emotion type by analyzing and comparing the feature data, and outputs the confidence value of the emotion type. Finally, the emotion recognition result based solely on the visual dimension and the corresponding reliability index are obtained, thus completing the entire visual dimension emotion recognition process.

[0040] Step 220-2: Analyze the auditory emotion features of the speech data. Use a neural network model to perform deep learning analysis on the auditory emotion features to obtain the emotion recognition results and corresponding confidence scores in the auditory dimension.

[0041] In this embodiment of the disclosure, voice data recorded in real time when the user recalls their expression can be acquired first. The voice data is then preprocessed using an audio processing algorithm to accurately extract auditory emotional features such as intonation fluctuations, speech rate, volume, pauses, and timbre changes. These feature data are then input into a pre-trained neural network model. The model uses a deep learning algorithm to perform multi-level mining, analysis, and comparison of the feature data. Based on the mapping relationship between features and emotion types, the corresponding emotion category is determined, and the confidence score of the emotion category is output. Finally, a complete emotion recognition result and reliability index based solely on the auditory dimension are formed, providing auditory data support for subsequent multi-dimensional emotion fusion analysis.

[0042] Step 220-3: Compare the physiological signs data with the preset general physiological baseline of the population parameter by parameter, calculate the deviation range and fluctuation stability of each indicator, and determine the emotion recognition results and corresponding confidence levels of the physiological dimensions based on the deviation range and fluctuation stability of each indicator and the preset deviation threshold.

[0043] The preset general physiological baseline for the population is a pre-set standard range of indicators that conforms to the normal physiological state of the target population (such as the elderly), and serves as a reference benchmark for measuring whether physiological data is abnormal. The parameter-by-parameter comparison is the process of comparing each real-time collected physiological sign data (such as heart rate, body temperature, and blood pressure) with the corresponding indicator range in the general physiological baseline. The deviation range is the difference or proportion between the real-time physiological sign data and the general physiological baseline, used to quantify the degree to which the physiological indicators deviate from the normal range. The fluctuation stability is the stability of the frequency, amplitude, and trend of changes in physiological sign data within a certain period of time, reflecting the fluctuation of physiological state. The preset deviation threshold is a pre-set critical value used to judge whether the deviation of physiological indicators reaches the standard for correlation with emotional changes. If the deviation exceeds the threshold, the physiological data change is considered to be related to the emotional state.

[0044] In this embodiment of the disclosure, physiological data such as heart rate, body temperature, and blood pressure collected by wearable devices during the user's recollection can be obtained first. These real-time data are compared one by one with the corresponding indicators in a pre-set general physiological baseline for the population. The deviation of each physiological indicator from the baseline is calculated by an algorithm (such as the percentage of heart rate exceeding the baseline or the degree of body temperature deviating from the baseline). At the same time, the frequency and amplitude of changes in physiological data over a certain period of time are analyzed to determine its fluctuation stability. Then, combined with a preset deviation threshold, the physiological data changes are comprehensively evaluated to determine whether they are related to a specific emotional state. Finally, the physiological dimension emotion recognition result (such as emotional excitement, stability, or fluctuation) and the corresponding confidence value are output, thus completing the physiological-level emotion determination process.

[0045] Step 220-4: Assign dynamic weights to the emotion recognition results of the three dimensions of vision, hearing and physiology according to the scene adaptation requirements.

[0046] Among them, the scenario adaptation requirement refers to the adaptation requirement determined based on the specific scenario type of the user's recollection (such as warm family memories, memories of important events, memories of emotional expression, etc.) and the core characteristics of emotional expression in that scenario. This directly affects the reference priority of the emotion recognition results in each dimension. The dynamic weight refers to the quantitative indicator that adjusts the proportion of the three dimensions of recognition results in the comprehensive emotion judgment in real time based on the scenario adaptation requirement and the reliability of the emotion recognition results in each dimension (such as the level of confidence). This makes the weight allocation more targeted.

[0047] For the embodiments of this disclosure, the specific scenario type of the user's recollection and the corresponding scenario adaptation requirements can be identified first (e.g., a warm family recollection scenario emphasizes the intuitiveness of emotional expression, while an important event recollection scenario emphasizes the stability of emotions). Then, the confidence scores of the emotional recognition results from the three dimensions of visual, auditory, and physiological senses are combined to analyze the reference value and reliability of each dimension's recognition results in the current scenario. Subsequently, the weight ratios of the three dimensions are adjusted in real time according to the scenario adaptation rules. Specifically, higher weights can be assigned to dimensions with high confidence scores and that fit the core requirements of the scenario, while the weight ratios of dimensions with low confidence scores or weak scenario adaptability can be reduced. This ultimately forms a dynamic weight allocation scheme that accurately matches the current scenario, providing a quantitative basis for the subsequent fusion analysis of multi-dimensional emotional data.

[0048] By assigning dynamic weights to the three dimensions of emotion recognition results according to scene adaptation requirements, the limitations of fixed weight allocation mode can be broken. This allows the reference value of each dimension of emotion recognition results to be accurately matched with scene characteristics and recognition reliability. It can effectively avoid the impact of single-dimensional data bias or insufficient scene adaptability on the comprehensive emotion judgment, and significantly improve the accuracy and scene adaptability of the overall emotion recognition. The dynamically adjusted weight allocation method makes the emotion judgment more in line with the actual scene of the user's recollection and expression, ensuring that the visual and auditory feedback materials and immersive interactive scenes generated based on comprehensive emotions are more targeted, and can further enhance the emotional resonance and immersion of the interaction process.

[0049] Step 220-5: Employ a multi-dimensional weighted fusion algorithm to perform weighted calculations on the emotion recognition results and confidence scores of each dimension based on dynamic weights. In the weighted emotion recognition results, select emotion types with a comprehensive confidence score greater than a preset threshold as the user's real-time emotion type, and determine the emotion intensity based on the numerical range of the comprehensive confidence score.

[0050] Among them, the multi-dimensional weighted fusion algorithm is a fusion calculation method for three dimensions of data: visual, auditory, and physiological. By assigning different weights to each dimension, it combines the recognition results and confidence scores for comprehensive calculation to output a more accurate unified conclusion. The comprehensive confidence score is a quantitative indicator that reflects the comprehensive reliability of the emotion recognition results of the three dimensions after weighted calculation, and it is the core basis for determining the final emotion type. The preset threshold is a pre-set critical value used to screen for valid comprehensive confidence scores. Only when the comprehensive confidence score exceeds this value is the corresponding emotion type recognized as valid.

[0051] In this embodiment of the disclosure, the emotion recognition results and corresponding confidence scores of the three dimensions of vision, hearing and physiology can be obtained first. Combined with the dynamic weights determined according to the needs of scene adaptation, the confidence score of each dimension is multiplied by the corresponding dynamic weight and then summed to obtain the comprehensive confidence score through a multi-dimensional weighted fusion algorithm. Then, the emotion types with a comprehensive confidence score greater than a preset threshold are selected and determined as the user's real-time emotion type. Then, the intensity level of the user's real-time emotion type is finally determined according to the preset numerical range standard and the range in which the current comprehensive confidence score is located.

[0052] By employing a multi-dimensional weighted fusion algorithm, combining dynamic weights to weight the emotion recognition results and confidence levels of each dimension, and filtering valid results with preset thresholds and determining emotion intensity according to numerical ranges, this approach effectively integrates emotional information from different dimensions. It highlights the reference value of highly adaptable dimensions, avoids the one-sidedness and bias of single-dimensional recognition, and significantly improves the accuracy, reliability, and scene adaptability of real-time emotion type and intensity determination. This provides accurate and relevant core decision-making basis for subsequent dynamic generation of visual and auditory feedback materials and construction of immersive interactive scenarios, ensuring that interactive responses are more targeted and emotionally resonant.

[0053] Step 230: Based on the real-time emotion type and intensity, dynamically adapt and generate corresponding visual and auditory feedback materials, and construct an immersive interactive scene based on the visual and auditory feedback materials.

[0054] For embodiments of this disclosure, step 230 may include the following steps: Step 230-1: Call the preset visual material library to determine the initial visual material that matches the real-time emotion type. Adjust the visual parameters of the initial visual material based on the emotion intensity to obtain visual feedback material. The visual parameters include at least the screen tone, clarity, dynamic element type and filter effect.

[0055] The preset visual material library is a collection of pre-built visual content with various themes (such as nostalgia, warmth, etc.), which stores basic visual templates that can be directly called, providing original visual resources for mood adaptation.

[0056] In this embodiment of the disclosure, the system first retrieves initial visual materials matching the user's determined real-time emotion type from a preset visual material library. Then, based on the user's current emotional intensity, it makes targeted adjustments to the core visual parameters of the initial visual materials, such as color tone, clarity, dynamic element types, and filter effects. Specifically, the higher the emotional intensity, the more closely the parameter adjustments match the expression needs of that emotion (e.g., when the emotion is pleasant and intense, the color tone is warmer and the dynamic elements are richer), ultimately generating visual feedback materials that are precisely adapted to the user's real-time emotional state.

[0057] By calling and matching initial visual materials from a preset visual material library, and dynamically adjusting core visual parameters such as color tone and clarity based on emotional intensity, real-time adaptation of visual materials to the user's emotional state can be achieved, breaking the limitations of traditional fixed visual presentation. Personalized adjustment of visual parameters makes visual feedback more in line with the emotional atmosphere of the user's expression, which can effectively enhance the immersion and emotional resonance of the interaction process. This not only increases the user's (especially the elderly) willingness to interact, but also lays a visual foundation that fits the emotions for the construction of subsequent immersive interaction scenarios, making the process of recalling and expressing memories more warm and appropriate.

[0058] Step 230-2: Call the preset sound effect library to determine the initial sound effect material that matches the real-time emotion type. Adjust the auditory parameters of the initial sound effect material based on the emotion intensity to obtain auditory feedback material. The auditory parameters include at least the background music BPM, timbre selection, and the volume ratio of ambient sound.

[0059] The preset sound effects library is a collection of pre-built sound effects resources in various styles (such as soothing, upbeat, nostalgic, etc.), including basic sound effects materials such as background music and ambient sounds, which can provide original resource support for auditory emotion adaptation.

[0060] In this embodiment of the disclosure, the system can first retrieve initial sound effect materials matching the user's real-time emotion type from a preset sound effect library, and then adjust the core auditory parameters of the initial sound effect materials in a targeted manner based on the user's current emotional intensity. Specifically, the higher the emotional intensity, the faster the background music BPM adaptability, the enhanced texture of the corresponding timbre, and the optimized proportion of ambient sound volume as needed (e.g., reducing the proportion of ambient sound when emotionally agitated). The lower the emotional intensity, the more suitable the soothing rhythm, the softer timbre, and the reasonable proportion of ambient sound are. Ultimately, auditory feedback materials that accurately match the user's real-time emotional state are generated, providing auditory support for building an immersive interactive scene.

[0061] By calling and matching initial sound effect materials from a preset sound effect library, and dynamically adjusting the background music BPM, timbre, and ambient sound volume ratio based on emotional intensity, real-time and accurate adaptation of auditory materials to the user's emotional state can be achieved, breaking the limitations of traditional fixed sound effect presentation. Personalized auditory parameter adjustments make auditory feedback more in line with the emotional atmosphere of the user's expression, which can effectively enhance the immersion and emotional resonance of the interaction process. This not only increases the user's willingness and comfort to participate in the interaction, but also works with visual feedback materials to create a highly emotionally resonant immersive scene, making the process of reminiscing more warm and infectious, laying a rich auditory foundation for the final generation of emotional multimedia memoirs.

[0062] Step 230-3: Overlay visual and auditory feedback materials to create an immersive interactive scene.

[0063] In this embodiment of the disclosure, visual feedback materials and auditory feedback materials generated based on the user's real-time emotional type and intensity can be obtained first. Through a preset synchronous output mechanism, the two types of materials can be precisely aligned on the timeline. For example, the dynamic changes of the visual images can be matched in real time with the rhythm of the background music and the intensity of the ambient sounds to avoid audio-visual asynchrony. Finally, they are presented in a coordinated manner to form an interactive scene that is highly consistent with the user's current emotional state and allows the user to immerse themselves in the expression of memories.

[0064] By overlaying visual and auditory feedback materials that are appropriate for the emotions, a synergistic fusion of visual and auditory sensory experiences can be achieved, breaking the limitations of single-sensory feedback and significantly enhancing the immersion and emotional relevance of the interactive scene. This multi-sensory collaborative presentation method makes it easier for users to immerse themselves in the memory and expression context, effectively reducing resistance and increasing the willingness to interact.

[0065] Step 240: In an immersive interactive scenario, capture multiple audio and video clips when the user's emotions reach their peak.

[0066] For embodiments of this disclosure, step 240 may include the following steps: Step 240-1: Preset the peak intensity threshold and triggering conditions for different emotion types.

[0067] Among them, the peak intensity threshold is a pre-set critical value standard for each emotion type, representing the strongest degree of the emotion. It is the core quantitative basis for judging whether the emotion is at its peak (such as a specific value corresponding to the comprehensive confidence level). The trigger condition is a pre-set prerequisite for starting the emotional peak audio and video slice capture function. It needs to be judged together with the emotional intensity threshold and multi-dimensional data characteristics to ensure that the operation is triggered only when the emotion truly reaches its peak.

[0068] In this embodiment of the disclosure, the system can first sort out the core emotion types commonly found in user recollections (such as pleasure, excitement, and longing), and preset corresponding peak intensity thresholds for each emotion type (such as setting the peak confidence threshold for "pleasure" to 0.85 or higher, and setting the peak confidence threshold for "excitement" to 0.88 or higher). At the same time, it combines the emotion recognition features of visual, auditory, and physiological dimensions to set collaborative triggering conditions. For example, when the comprehensive confidence of a certain emotion exceeds the corresponding peak intensity threshold, and the data of the corresponding visual dimension (such as facial micro-expression amplitude), auditory dimension (such as speech rate and tone fluctuation), and physiological dimension (such as heart rate deviation amplitude) all meet the preset collaborative judgment criteria, a complete triggering condition is formed, providing a clear start rule for subsequent extraction of emotional peak audio and video slices.

[0069] Step 240-2: When the emotional intensity is determined to be greater than the peak intensity threshold and the multi-dimensional emotion-related data meets the triggering conditions, a preset length of audio and video clip is extracted from the current immersive interactive scene as an audio and video material slice.

[0070] The preset length is a pre-defined standard for the duration of audio and video clips, ensuring that the extracted material fully preserves the emotional peak moment without being too long.

[0071] In this embodiment of the disclosure, the system can monitor the intensity of the user's emotions in real time during the process of the user recalling and expressing in an immersive interactive scene. At the same time, it can verify whether the visual, auditory, and physiological multi-dimensional emotion-related data meet the preset trigger conditions. When the intensity of the emotion is detected to exceed the peak intensity threshold of the corresponding emotion type, and the multi-dimensional emotion data all meet the trigger conditions (such as the micro-expression amplitude meets the standard, the fluctuation of voice tone meets the standard, and the deviation of physiological indicators is within the set range), the system can automatically extract audio and video segments in the current immersive interactive scene according to the preset duration standard. These segments are audio and video material slices that record the peak state of the user's emotions.

[0072] By using both the peak threshold of emotional intensity and the fulfillment of triggering conditions by multi-dimensional emotional data as dual judgment criteria, and then extracting audio and video clips of a preset length in the immersive interactive scenario, it is possible to effectively ensure that the extracted material accurately corresponds to the key moment with the most emotional tension in the user's recollection, avoiding invalid extraction or omissions caused by a single condition judgment. At the same time, the extracted clips integrate the adaptive audio-visual effects of the immersive scene with the user's real emotional expression, which can provide high-quality core emotional material for the subsequent integration and generation of emotional multimedia memoirs. This allows the final product to fully restore the scene atmosphere and expression state of the emotional peak moment, greatly improving the emotional authenticity and commemorative value of the memoir.

[0073] Step 240-3: Add emotion type, timestamp, and intensity level tags to each audio / video clip and store them in the media database.

[0074] Among them, the emotion type label is an identifier of the specific emotion category of the user corresponding to the audio and video material slice (such as joy, excitement, longing, calm, etc.), which is derived from the emotion recognition results after multi-dimensional fusion; the timestamp label is an identifier of the time information (accurate to the hour, minute and second) that accurately records the moment when the audio and video material slice was captured, which is used to clarify the time position of the moment of emotional peak in the overall recall expression process; the intensity level label is a graded identifier of the intensity of the corresponding emotion of the audio and video material slice (such as mild, moderate and severe), which is determined based on the numerical range of the comprehensive confidence level; the material database is a structured data storage carrier specifically used to store tagged audio and video material slices, which has the functions of material retrieval, retrieval and management, and provides data support for subsequent integration to generate memoirs.

[0075] In this embodiment of the disclosure, after capturing the audio-visual material slice at the peak of the user's emotion, the core attribute information corresponding to the slice can be automatically extracted. Specifically, based on the emotion type obtained from the previous multi-dimensional fusion analysis, the precise timestamp of the slice capture moment, and the emotion intensity level determined according to the comprehensive confidence level, these three types of information can be encapsulated into a unique tag and uniquely associated with the audio-visual material slice. Subsequently, the audio-visual material slices with complete tags are uniformly uploaded to a structured material database according to preset data storage rules, ensuring that the attribute information of each slice is traceable and inseparable from the material itself, facilitating subsequent on-demand retrieval and integration.

[0076] Step 250: Extract the core content of the user's recollection from the interaction data using AI semantic analysis algorithms.

[0077] Among them, the AI ​​semantic analysis algorithm is an artificial intelligence algorithm with natural language understanding capabilities. It can process unstructured natural language data by word segmentation, semantic parsing, keyword extraction, and core information extraction, accurately capturing the key meanings and core logic behind the language. The core content of the user's recollection is the key information extracted from the user's interaction data that constitutes the subject of the recollection. It covers the core events, important people, key time nodes, core viewpoints, key details of emotional connection in the recollection, etc., and is the core support for the content of the memoir.

[0078] In this embodiment of the disclosure, all language data generated during the user's interaction with AI (including text input, speech-to-text colloquial expressions, etc.) can be collected first and then input into a pre-trained AI semantic analysis algorithm. The algorithm accurately identifies and extracts core events (such as "the first time I took my child to Beijing in 1980"), key people (such as family members and colleagues), time nodes, core experience details, and core viewpoints through word segmentation, semantic logic sorting, and redundant information filtering. At the same time, redundant content such as repeated expressions and irrelevant chatter is ignored, and finally a structured and logically clear set of core content of user memories is formed, which provides core text support for subsequent integration with emotional audio and video material slices.

[0079] Step 260: Based on the timestamp tags of audio and video material slices and the timeline of core content, integrate multiple audio and video material slices with the corresponding text narrative paragraphs of core content to form the user's multimedia memoir.

[0080] For embodiments of this disclosure, step 260 may include the following steps: Step 260-1: Perform built-in audiovisual parameter analysis on each audio and video material slice to extract the slice visual parameters and slice auditory parameters contained in the audio and video material slice.

[0081] The built-in audiovisual parameter parsing is a technical process of decoding, recognizing, and extracting the visual and auditory presentation parameters contained in the audio and video material slices themselves. The visual parameters of the slices are the core elements that determine the visual presentation effect in the audio and video material slices, including at least the color tone, clarity, dynamic element type, and filter effect, which are the quantitative manifestation of the visual style of the slices. The auditory parameters of the slices are the core elements that determine the auditory presentation effect in the audio and video material slices, including at least the background music BPM, timbre selection, and the proportion of ambient sound volume, which are the quantitative manifestation of the auditory style of the slices.

[0082] In this embodiment of the present disclosure, the system can retrieve each stored audio and video material slice from the material database one by one, decode the slice through a preset audio and video parameter parsing tool, and simultaneously extract two types of core parameters built into the slice. These can include visual parameters such as screen color tone, dynamic element type, filter effect, and clarity, as well as auditory parameters such as background music BPM, timbre category, and ambient sound volume ratio, forming a unique audio-visual parameter set for each audio and video material slice.

[0083] By analyzing the built-in audiovisual parameters of each audio-visual material slice and extracting the visual and auditory parameters of the slice, the audiovisual presentation characteristics of each emotional peak slice can be accurately grasped. This provides reliable data for subsequent adjustments to the visual style of the text narrative paragraphs, optimization of the clarity of user speech in the slice's audio, adaptation of the playback duration and transition effects of the text and slice, and can effectively avoid problems such as audiovisual disjointness and style conflict that occur during the integration process. This ensures that the final multimedia memoir is highly coordinated and unified in terms of text, visuals, and hearing, which can significantly improve the presentation coordination and immersion of the finished product, making the expression of memories more coherent and impactful.

[0084] Step 260-2: Based on the slice visual parameters, adjust the presentation style of the corresponding text narrative paragraphs to ensure that the text visual parameters are consistent with the slice visual parameters.

[0085] Among them, text visual parameters are specific quantitative indicators that determine the presentation style of text narrative paragraphs, such as font style, text color, line spacing, font weight, background effects, etc., and are the core basis for text visual expression.

[0086] In this embodiment of the disclosure, the system can first retrieve the visual parameters of the parsed audio and video material slices, accurately associate them with the corresponding text narrative paragraphs containing the core content of the user's memories through timestamp tags, and then adjust the text visual parameters of the text narrative paragraphs based on the core features of the slice visual parameters (such as warm tones corresponding to warm-toned text, retro filters corresponding to retro fonts, and high definition corresponding to simple layouts). This may include text colors that match the hue, font types that fit the filter style, layouts that adapt to the resolution, and text effects corresponding to dynamic elements, ultimately ensuring that the presentation style of the text narrative paragraphs is highly consistent with the visual style of the audio and video material slices.

[0087] By adjusting the presentation style of corresponding text narrative paragraphs based on the visual parameters of the slices, the visual effects of text and audio-visual slices can be synergistically adapted, avoiding the visual discontinuity caused by the disconnect between text and visual styles. This can significantly improve the overall visual harmony and immersion of multimedia memoirs. A unified style of text and visuals can more accurately convey the emotional atmosphere corresponding to the memories, allowing users to have a coherent sensory experience during browsing, further strengthening the emotional imprint of the memories, making the finished memoirs more holistic and impactful, and enriching the three-dimensional dimensions of memory presentation.

[0088] Step 260-3: Analyze the user's speech clarity in the slice auditory parameters. If background noise interference is detected, dynamically reduce the proportion of ambient volume in the slice auditory parameters and simultaneously increase the strength of the user's speech signal.

[0089] Among them, user voice clarity refers to the degree to which the voice signal expressing the user's memories in the audio and video material slice can be identified, which is a key indicator for measuring whether the voice can accurately convey the memory information; background sound interference refers to the situation in which ambient sounds, background music and other non-user voice sounds in the slice affect the recognition of user voice due to problems such as excessive volume or frequency overlap.

[0090] In this embodiment of the present disclosure, the system can retrieve the segment auditory parameters of audio and video material slices, and use audio analysis algorithms to detect the clarity of the user's voice signal, accurately determine whether there is any identification interference caused by background sound effects (such as ambient sound, background music). If interference signals are detected that affect the normal understanding of the user's voice, the audio optimization mechanism is automatically activated to dynamically reduce the proportion of ambient volume in the segment auditory parameters. At the same time, noise reduction, volume gain and other technologies are used to simultaneously improve the signal strength of the user's voice, ensuring that the user's voice is clear and prominent in the adjusted auditory effect and is not affected by background sound effects.

[0091] Step 260-4: Adjust the playback duration and transition effects of the corresponding audio and video material slices according to the length and emotional rhythm of the text narrative paragraphs.

[0092] Among them, the cut-in and cut-out transition effects are the visual connection methods when audio and video material slices start playing (cut in) and end playing (cut out), such as fade in and fade out, fast cut in, gradient overlay, etc., which are used to improve the smoothness of the connection between the slices and text and other materials.

[0093] In this embodiment of the disclosure, the system can first accurately associate text narrative segments with corresponding audio and video material slices using timestamp tags, analyze the length of the text narrative segments (such as the number of words and the length of the segment) to determine the appropriate playback time range, and simultaneously analyze the emotional rhythm contained in the text (such as brisk, soothing, or intense). Then, the playback duration of the audio and video material slices is adjusted according to the text length. For example, if the text is long, the playback time of the slice is appropriately extended; if the text is short, it is shortened accordingly, ensuring that the text content and the slice visuals are presented synchronously. Furthermore, corresponding transition effects are matched according to the emotional rhythm, such as using fast transitions for brisk emotions and fade-in / fade-out effects for soothing emotions, ultimately achieving a high degree of consistency between the text narrative and the audio and video slices in terms of time and emotional rhythm.

[0094] By adjusting the playback duration and transition effects of audio and video clips according to the length and emotional rhythm of the text narrative paragraphs, the problems of asynchronous playback of text and clips and disconnection of emotional rhythm can be effectively solved, ensuring the coherence and coordination of the memoir content. The duration adaptation allows users to watch the corresponding emotional segments in their entirety while browsing the text, and the transition effects that match the emotional rhythm can enhance the smoothness of emotional transmission, further improving the immersiveness and appeal of the multimedia memoir. This makes the presentation of the memory not only supported by clear textual logic, but also retains the vividness of the emotional segments, making the overall content more cohesive and emotionally compelling.

[0095] Step 260-5: Based on the timestamp tags of the audio and video material slices and the timeline of the core content, integrate the multiple audio and video material slices and text narrative paragraphs after the above collaborative optimization and adjustment to form the user's multimedia memoir.

[0096] In this embodiment of the disclosure, the timeline of the core content and the timestamp tags of each audio and video material slice can be retrieved first. By accurately matching the timestamps with the timeline, the corresponding position of each optimized audio and video material slice in the overall memory context can be determined. Then, in chronological order, multiple audio and video material slices that have undergone visual style adaptation, auditory clarity optimization, and adjustment of duration and transition effects are sequentially integrated with the corresponding text narrative paragraphs to ensure that the playback of the slices and the presentation of the text proceed synchronously, and that the emotional atmosphere and content logic are highly consistent. Finally, a user multimedia memoir with a complete structure, coherent time sequence, and multi-dimensional presentation is formed.

[0097] In summary, the technical solution in this application, by collecting multi-dimensional emotion-related data such as user facial images, voice, and physiological signs in real time during AI interaction, accurately determines the user's real-time emotion type and intensity through multi-dimensional weighted fusion analysis, and then dynamically adapts and generates visual and auditory feedback materials to construct an immersive interactive scene. Simultaneously, it captures audio and video material slices at the peak of emotion, and combines them with the core content of the memory extracted by AI semantic analysis. Through audio-visual parameter collaborative optimization and temporal integration, a multimedia memoir is formed. This not only effectively solves the problems of mechanical and boring interaction process, one-sided emotion capture, and single sensory feedback in existing technologies, but also significantly improves the immersiveness and adaptability of user interaction. Furthermore, it allows the memoir to completely preserve the real emotional moments and state imprints, making up for the shortcomings of weak emotions and thin content dimensions in traditional products. It achieves a multi-dimensional and three-dimensional presentation of memory content and emotional expression, which can greatly enhance the commemorative value and emotional warmth of the memoir, and fully meet users' deep needs for emotional and immersive memory preservation.

[0098] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a multimedia memoir generation device based on multi-dimensional emotion perception, such as... Figure 3 As shown, the device includes: acquisition module 31, analysis module 32, construction module 33, interception module 34, and integration module 35.

[0099] The data acquisition module 31 can be used to collect multi-dimensional emotion-related data of the user's recollection and expression in real time during the AI ​​interaction process. The multi-dimensional emotion-related data includes at least facial image data, voice data and physiological sign data. Analysis module 32 can be used to comprehensively analyze multi-dimensional emotion-related data to determine the user's real-time emotion type and intensity; Module 33 can be used to dynamically adapt and generate corresponding visual and auditory feedback materials based on real-time emotion type and intensity, and to build immersive interactive scenes based on the visual and auditory feedback materials; The capture module 34 can be used to capture multiple audio and video clips when the user's emotions reach their peak in an immersive interactive scenario. The integration module 35 can be used to integrate multiple audio and video material slices and the core content of the user's recollection to obtain the user's multimedia memoir.

[0100] In some embodiments of this application, the acquisition module 31 can be specifically used to capture facial image data of the user recalling the expression in real time through a camera. The facial image data includes at least facial feature points, micro-expressions, lip movement frequency and eye movement trajectory data. It can also record voice data of the user recalling the expression in real time through an audio acquisition device. Furthermore, it can acquire physiological sign data collected by the user's wearable device when the user recalls the expression. The physiological sign data includes at least the user's heart rate, body temperature and blood pressure data.

[0101] In some embodiments of this application, the analysis module 32 can be specifically used to extract key facial features from facial image data, identify emotion types based on key facial features, and obtain the emotion recognition results and corresponding confidence scores in the visual dimension; to analyze auditory emotion features from speech data, perform deep learning analysis on auditory emotion features through a neural network model, and obtain the emotion recognition results and corresponding confidence scores in the auditory dimension; to compare physiological sign data with a preset general physiological baseline for the population parameter by parameter, calculate the deviation magnitude and fluctuation stability of each indicator, and determine the emotion recognition results and corresponding confidence scores in the physiological dimension based on the deviation magnitude and fluctuation stability of each indicator and a preset deviation threshold; to assign dynamic weights to the emotion recognition results in the visual, auditory, and physiological dimensions according to the scene adaptation requirements; to use a multi-dimensional weighted fusion algorithm to perform weighted calculations on the emotion recognition results and confidence scores of each dimension based on the dynamic weights, and to select the emotion types with a comprehensive confidence score greater than a preset threshold as the user's real-time emotion type from the weighted emotion recognition results, and to determine the emotion intensity based on the numerical range of the comprehensive confidence score.

[0102] In some embodiments of this application, the construction module 33 can be specifically used to call a preset visual material library to determine initial visual material that matches the real-time emotion type, adjust the visual parameters of the initial visual material based on the emotion intensity to obtain visual feedback material, the visual parameters of which include at least the screen tone, clarity, dynamic element type and filter effect; call a preset sound effect library to determine initial sound effect material that matches the real-time emotion type, adjust the auditory parameters of the initial sound effect material based on the emotion intensity to obtain auditory feedback material, the auditory parameters of which include at least the background music BPM, timbre selection and ambient sound volume ratio; and superimpose and output the visual feedback material and auditory feedback material to form an immersive interactive scene.

[0103] In some embodiments of this application, the interception module 34 can be used to preset peak intensity thresholds and triggering conditions corresponding to different emotion types; when it is determined that the emotion intensity is greater than the peak intensity threshold and the multi-dimensional emotion-related data meets the triggering conditions, an audio and video segment of preset length is intercepted in the current immersive interactive scene as an audio and video material slice; an emotion type, timestamp, and intensity level label are added to each audio and video material slice, and the slice is associated and stored in the material database.

[0104] In some embodiments of this application, the integration module 35 can be used to extract the core content of the user's recollection from the interactive data through AI semantic analysis algorithms; based on the timestamp tags of the audio and video material slices and the timeline of the core content, it integrates multiple audio and video material slices with the corresponding text narrative paragraphs of the core content to form the user's multimedia memoir.

[0105] In some embodiments of this application, the integration module 35 can be specifically used to perform built-in audiovisual parameter analysis on each audio-visual material slice, extracting the slice visual parameters and slice auditory parameters contained in the audio-visual material slice; based on the slice visual parameters, adjust the presentation style of the corresponding text narrative paragraph to make the text visual parameters consistent with the slice visual parameters; analyze the clarity of the user's voice in the slice auditory parameters, and if background sound interference is detected, dynamically reduce the proportion of ambient volume in the slice auditory parameters and simultaneously improve the strength of the user's voice signal; adjust the playback duration and cut-in / cut-out transition effects of the corresponding audio-visual material slice according to the length and emotional rhythm of the text narrative paragraph; and integrate multiple audio-visual material slices and text narrative paragraphs after the above-mentioned collaborative optimization and adjustment based on the timestamp tags of the audio-visual material slices and the timeline of the core content to form the user's multimedia memoir.

[0106] It should be noted that other corresponding descriptions of the functional units involved in the multimedia memoir generation device based on multi-dimensional emotion perception provided in this embodiment can be found in [reference]. Figure 1 and Figure 2The corresponding descriptions in [the document] will not be repeated here.

[0107] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method for generating multimedia memoirs based on multi-dimensional emotion perception is shown.

[0108] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0109] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 3 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The method for generating multimedia memoirs based on multi-dimensional emotion perception is shown.

[0110] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0111] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0112] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.

[0114] This invention, through real-time acquisition of multi-dimensional emotion-related data such as facial images, voice, and physiological signs during AI interaction, accurately determines the user's real-time emotion type and intensity through multi-dimensional weighted fusion analysis. It then dynamically adapts and generates visual and auditory feedback materials and constructs an immersive interactive scene, simultaneously capturing audio-visual material slices at emotional peaks. Combined with core memory content extracted through AI semantic analysis, and processed through audio-visual parameter optimization and temporal integration, a multimedia memoir is formed. This not only effectively solves the problems of mechanical and monotonous interaction processes, one-sided emotion capture, and singular sensory feedback in existing technologies, significantly improving the immersiveness and adaptability of user interaction, but also allows the memoir to fully preserve authentic emotional moments and state imprints, compensating for the shortcomings of traditional finished products in terms of weak emotion and thin content dimensions. It achieves a multi-dimensional and three-dimensional presentation of memory content and emotional expression, greatly enhancing the commemorative value and emotional warmth of the memoir, and fully satisfying users' deep needs for emotional and immersive memory preservation.

[0115] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.

[0116] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A method for generating multimedia memoirs based on multi-dimensional emotion perception, characterized in that, include: During AI interaction, multi-dimensional emotion-related data of user recall and expression are collected in real time. The multi-dimensional emotion-related data includes at least facial image data, voice data and physiological sign data. By comprehensively analyzing the multi-dimensional emotion-related data, the user's real-time emotion type and intensity can be determined; Based on the real-time emotion type and the emotion intensity, corresponding visual feedback materials and auditory feedback materials are dynamically adapted and generated, and an immersive interactive scene is constructed based on the visual feedback materials and the auditory feedback materials; In the immersive interactive scenario, multiple audio and video clips are captured when the user's emotions reach their peak. By integrating the multiple audio and video material slices and the core content of the user's recollections, the user's multimedia memoir is obtained.

2. The method according to claim 1, characterized in that, The real-time collection of multi-dimensional emotion-related data from user recollections includes: The camera captures facial image data of the user recalling and expressing themselves in real time. The facial image data includes at least facial feature points, micro-expressions, lip movement frequency and eye movement trajectory data. Record the user's voice data in real time while recalling their expression using audio acquisition equipment; Acquire physiological data collected by the user's wearable device when the user recalls and expresses themselves. The physiological data includes at least the user's heart rate, body temperature, and blood pressure.

3. The method according to claim 1, characterized in that, A comprehensive analysis of the multi-dimensional emotion-related data is performed to determine the user's real-time emotion type and intensity, including: Facial key features are extracted from the facial image data, and emotion type is identified based on the facial key features to obtain the visual dimension emotion recognition result and corresponding confidence level. The speech data is analyzed for auditory emotion features, and the auditory emotion features are analyzed by deep learning through a neural network model to obtain the emotion recognition results and corresponding confidence scores in the auditory dimension. The physiological signs data are compared with a preset general physiological baseline for the population, and the deviation range and fluctuation stability of each indicator are calculated. Based on the deviation range and fluctuation stability of each indicator and the preset deviation threshold, the emotion recognition result and corresponding confidence level of the physiological dimension are determined. Dynamic weights are assigned to the emotion recognition results across the visual, auditory, and physiological dimensions based on scene adaptation requirements. A multi-dimensional weighted fusion algorithm is adopted. Based on the dynamic weight, the emotion recognition results and confidence scores of each dimension are weighted and calculated. In the emotion recognition results after weighted calculation, the emotion types with a comprehensive confidence score greater than a preset threshold are selected as the user's real-time emotion types. The emotion intensity is determined according to the numerical range of the comprehensive confidence score.

4. The method according to claim 1, characterized in that, Based on the real-time emotion type and intensity, corresponding visual and auditory feedback materials are dynamically generated, and an immersive interactive scene is constructed based on the visual and auditory feedback materials, including: The system calls a preset visual material library to determine the initial visual material that matches the real-time emotion type. Based on the emotion intensity, the visual parameters of the initial visual material are adjusted to obtain visual feedback material. The visual parameters include at least the screen tone, clarity, dynamic element type, and filter effect. The system calls a preset sound effect library to determine the initial sound effect material that matches the real-time emotion type. Based on the emotion intensity, the auditory parameters of the initial sound effect material are adjusted to obtain auditory feedback material. The auditory parameters include at least the background music BPM, timbre selection, and the volume ratio of ambient sound. The visual feedback material and the auditory feedback material are superimposed to form an immersive interactive scene.

5. The method according to claim 1, characterized in that, In the immersive interactive scenario, multiple audio and video clips are extracted when the user's emotions reach their peak, including: Preset peak intensity thresholds and trigger conditions for different emotion types; When it is determined that the intensity of the emotion is greater than the peak intensity threshold, and the multi-dimensional emotion-related data meets the triggering condition, a preset length of audio and video clip is extracted from the current immersive interactive scene as an audio and video material slice. Add emotion type, timestamp, and intensity level tags to each audio / video clip and store them in the media database.

6. The method according to claim 1, characterized in that, The integration of the multiple audio and video material slices and the core content of the user's recollections yields the user's multimedia memoir, including: AI semantic analysis algorithms are used to extract the core content of users' recollections and expressions from interaction data; Based on the timestamp tags of the audio and video material slices and the timeline of the core content, the multiple audio and video material slices and the corresponding text narrative paragraphs of the core content are integrated to form the user's multimedia memoir.

7. The method according to claim 6, characterized in that, The method of integrating the timestamp tags of the audio and video material slices with the timeline of the core content, and combining the text narrative paragraphs corresponding to the multiple audio and video material slices with the core content, forms the user's multimedia memoir, including: For each audio-visual material slice, the built-in audio-visual parameters are analyzed to extract the slice's visual and auditory parameters. Based on the slice visual parameters, adjust the presentation style of the corresponding text narrative paragraphs to ensure that the text visual parameters are consistent with the slice visual parameters; The clarity of the user's voice in the sliced ​​auditory parameters is analyzed. If background noise interference is detected, the proportion of ambient volume in the sliced ​​auditory parameters is dynamically reduced, and the strength of the user's voice signal is increased simultaneously. Adjust the playback duration and cut-in / cut-out transition effects of the corresponding audio and video material slices according to the length and emotional rhythm of the text narrative paragraphs; The user's multimedia memoir is formed by integrating the timestamp tags of the audio and video material slices with the timeline of the core content, after the above-mentioned collaborative optimization and adjustment, and the text narrative paragraphs.

8. A multimedia memoir generation device based on multi-dimensional emotion perception, characterized in that, include: The data acquisition module is used to collect multi-dimensional emotion-related data that the user recalls and expresses in real time during the AI ​​interaction process. The multi-dimensional emotion-related data includes at least facial image data, voice data, and physiological sign data. The analysis module is used to comprehensively analyze the multi-dimensional emotion-related data to determine the user's real-time emotion type and intensity. The module is used to dynamically adapt and generate corresponding visual and auditory feedback materials based on the real-time emotion type and the emotion intensity, and to construct an immersive interactive scene based on the visual and auditory feedback materials. The capture module is used to capture multiple audio and video clips when the user's emotions reach their peak in the immersive interactive scenario. The integration module is used to integrate the multiple audio and video material slices and the core content of the user's recollection to obtain the user's multimedia memoir.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.

Citation Information

Cited By

  • Psychological state authenticity detection and intervention system based on eeg and psychological orthogonal basis group mapping

    CN122392952A