Video display method and device, electronic equipment and computer storage medium

By building a virtual image model based on user feature data and generating personalized video content, the problem of lack of personalization and emotional expression in accompanying assistant devices is solved, more natural situational interaction is achieved, and the user experience and the effectiveness of information transmission are improved.

CN120692425APending Publication Date: 2025-09-23BEIJING VISION WORLD TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510908819.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing accompanying assistant devices lack personalized and emotional expressions, making it difficult to achieve active and natural situational dialogues, which affects the user's interactive experience.

Method used

By acquiring user feature data to build a virtual image model, video content is generated in combination with preset events and automatically displayed when an event trigger is detected. The video content includes a virtual image and voice synthesis that is highly similar to the user, achieving personalized and emotional information delivery.

Benefits of technology

It improves the personalization and emotionality of information interaction, enhances user acceptance and interactive experience, and enhances the expression effect of reminders or caring content. It is particularly suitable for scenarios of companionship and health reminders for the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692425A_ABST
    Figure CN120692425A_ABST
Patent Text Reader

Abstract

The invention provides a video display method and device, electronic equipment and a computer storage medium, and the method comprises the steps: obtaining feature data of a user; constructing a virtual image model corresponding to the user according to the feature data; determining the video content of the target event according to the virtual image model and a preset target event, wherein the video content comprises prompt information related to the target event; and displaying the video content of the target event under the condition of detecting that the target event is triggered. According to the embodiment provided by the scheme, the individuation of information interaction and the effectiveness of information transmission can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video display technology, and in particular to a video display method, device, electronic device and computer storage medium. Background Art

[0002] Currently, accompanying assistant devices are widely used in scenarios such as home care, elderly care, and rehabilitation companionship. Common forms include voice assistants, smart speakers, and accompanying terminals with screens. These devices can generally assist users with daily management and emotional companionship to a certain extent. However, accompanying assistant devices in related technologies often lack personalized and emotional expression, and their passive response methods make it difficult to achieve active and natural contextual dialogue, which in turn affects the user's interactive experience. Summary of the Invention

[0003] The embodiments of the present application provide a video display method, device, electronic device, and computer storage medium, which can improve the personalization of information interaction and the effectiveness of information transmission. The above technical solutions are as follows:

[0004] In a first aspect, an embodiment of the present application provides a video display method, the method comprising:

[0005] Obtain user feature data;

[0006] Constructing a virtual image model corresponding to the user according to the feature data;

[0007] Determining the video content of the target event based on the virtual image model and the preset target event, wherein the video content includes prompt information related to the target event;

[0008] When the target event is detected to be triggered, the video content of the target event is displayed.

[0009] In a possible implementation, the feature data includes the user's appearance and voice features;

[0010] The aforementioned construction of the virtual image model corresponding to the aforementioned user based on the aforementioned feature data includes:

[0011] Performing three-dimensional modeling based on the shape features in the feature data to obtain a first virtual model that matches the shape features;

[0012] Performing speech modeling based on the sound features in the feature data to obtain a second virtual model that matches the sound features, wherein the second virtual model is a speech model for speech synthesis;

[0013] The first virtual model and the second virtual model are fused to obtain a virtual image model corresponding to the user.

[0014] In a possible implementation, the appearance features include the user's skeletal features and facial features;

[0015] The three-dimensional modeling process is performed based on the shape features in the feature data to obtain a first virtual model that matches the shape features, including:

[0016] Get the preset base model;

[0017] Adjusting the structure of the preset basic model based on the skeletal features in the feature data to obtain an initial virtual model corresponding to the appearance features;

[0018] Texture mapping is performed on the facial area of ​​the initial virtual model based on the facial features in the feature data to obtain a first virtual model that matches the appearance features.

[0019] In a possible implementation, the fusing of the first virtual model and the second virtual model to obtain the virtual image model corresponding to the user includes:

[0020] Determining target control nodes in the first virtual model, wherein the target control nodes include lip control nodes and expression control nodes;

[0021] Determining speech feature information based on the second virtual model;

[0022] Constructing a mapping relationship between the target control node and the voice feature information;

[0023] Based on the mapping relationship, the first virtual model and the second virtual model are fused to obtain a virtual image model corresponding to the user.

[0024] In a possible implementation, determining the video content of the target event based on the avatar model and the preset event includes:

[0025] Obtaining event information corresponding to the target event, including prompt content;

[0026] Inputting the prompt content into the second virtual model to generate event voice data corresponding to the target event;

[0027] Determine driving configuration parameters based on the event voice data, wherein the driving configuration parameters include lip-shaped driving parameters and expression driving parameters;

[0028] driving the target control node in the avatar model according to the driving configuration parameters to generate an action sequence synchronized with the event voice data;

[0029] The action sequence is synthesized with the event voice data to generate video content corresponding to the target event.

[0030] In a possible implementation, the target event is at least one of the following: a schedule reminder event, a geographic location trigger event, or a sensor linkage event;

[0031] When the target event is detected, the video content of the target event is displayed, including:

[0032] Through the preset event trigger detection mechanism, detect whether the trigger conditions of the above target events are met;

[0033] When it is detected that the triggering condition of the target event is met, it is determined that the triggering of the target event is detected, and the video content of the target event is displayed.

[0034] In a possible implementation, after displaying the video content of the target event, the method further includes:

[0035] Detect whether the target object responds to the operation;

[0036] If the target object is detected to have made the above-mentioned response operation within the preset time period, feedback is provided according to the response information corresponding to the above-mentioned response operation;

[0037] If the target object is not detected to perform the above-mentioned response operation within the above-mentioned preset time period, the video content of the above-mentioned target event is displayed again.

[0038] In a second aspect, an embodiment of the present application provides a video display device, comprising:

[0039] Acquisition module, used to obtain user feature data;

[0040] A construction module, configured to construct a virtual image model corresponding to the user according to the feature data;

[0041] a determination module, configured to determine the video content of the target event according to the virtual image model and the preset target event, wherein the video content includes prompt information related to the target event;

[0042] The display module is used to display the video content of the target event when the target event is detected to be triggered.

[0043] In a possible implementation, the feature data includes the user's appearance and voice features;

[0044] The above building blocks include:

[0045] A first processing unit is configured to perform three-dimensional modeling based on the shape features in the feature data to obtain a first virtual model that matches the shape features;

[0046] a second processing unit configured to perform speech modeling based on the sound features in the feature data to obtain a second virtual model that matches the sound features, wherein the second virtual model is a speech model for speech synthesis;

[0047] The fusion unit is used to fuse the first virtual model with the second virtual model to obtain a virtual image model corresponding to the user.

[0048] In a possible implementation, the appearance features include the user's skeletal features and facial features;

[0049] The first processing unit includes:

[0050] Get subunit, used to get preset basic model;

[0051] An adjustment subunit, configured to perform structural adjustment on the preset basic model based on the skeletal features in the feature data to obtain an initial virtual model corresponding to the appearance features;

[0052] The processing subunit is configured to perform texture mapping processing on the facial region of the initial virtual model based on the facial features in the feature data to obtain a first virtual model that matches the appearance features.

[0053] In a possible implementation, the fusion unit includes:

[0054] A first determining subunit is configured to determine a target control node in the first virtual model, wherein the target control node includes a lip control node and an expression control node;

[0055] A second determining subunit, configured to determine speech feature information based on the second virtual model;

[0056] A construction subunit, configured to construct a mapping relationship between the target control node and the voice feature information;

[0057] The fusion subunit is configured to fuse the first virtual model with the second virtual model based on the mapping relationship to obtain a virtual image model corresponding to the user.

[0058] In a possible implementation, the determining module includes:

[0059] An acquiring unit, configured to acquire event information corresponding to the target event, wherein the event information includes prompt content;

[0060] a generating unit, configured to input the prompt content into the second virtual model to generate event voice data corresponding to the target event;

[0061] A first determining unit is configured to determine driving configuration parameters based on the event voice data, wherein the driving configuration parameters include lip-shaping driving parameters and expression driving parameters;

[0062] a driving unit configured to drive the target control node in the avatar model according to the driving configuration parameters to generate an action sequence synchronized with the event voice data;

[0063] The synthesis unit is used to synthesize the above action sequence with the above event voice data to generate video content corresponding to the above target event.

[0064] In a possible implementation, the target event is at least one of the following: a schedule reminder event, a geographic location trigger event, or a sensor linkage event;

[0065] The above-mentioned display module includes:

[0066] A detection unit, configured to detect whether the triggering conditions of the target event are met through a preset event trigger detection mechanism;

[0067] The second determining unit is configured to determine that the triggering of the target event is detected when it is detected that the triggering condition of the target event is met, and to display the video content of the target event.

[0068] In a possible implementation, the apparatus further includes:

[0069] A detection module is used to detect whether the target object responds to the operation;

[0070] A feedback module is configured to provide feedback based on response information corresponding to the response operation when detecting that the target object has performed the response operation within a preset time period;

[0071] The display module is further configured to display the video content of the target event again when the target object is not detected to perform the response operation within the preset time period.

[0072] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory;

[0073] The above-mentioned memory stores a computer program, and the above-mentioned computer program is suitable for being loaded by the above-mentioned processor and executing the steps of the method provided by the first aspect of the embodiment of the present application or any possible implementation method of the first aspect.

[0074] In a fourth aspect, an embodiment of the present application provides a computer storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor and executing the steps of the method provided in the first aspect of the embodiment of the present application or any possible implementation of the first aspect.

[0075] The embodiment of the present application obtains the user's characteristic data, and constructs a virtual image model corresponding to the user based on the above characteristic data, determines the video content of the above target event based on the above virtual image model and the preset target event, wherein the above video content includes prompt information related to the above target event, and, when the trigger of the above target event is detected, displays the video content of the above target event. Thus, by obtaining the user's characteristic data, constructing a virtual image model that is highly similar to the user, and combining the preset event to generate video content containing prompt information, the video is automatically displayed when the event trigger is detected, thereby improving the personalized and emotional information transmission in the information interaction process, not only improving the user's acceptance and interactive experience, but also enhancing the expression effect of reminders or caring content. It is particularly suitable for scenarios such as companionship for the elderly and health reminders, helping to create a friendly and real interactive atmosphere and improving the effectiveness of information transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0077] Figure 1 A schematic structural diagram of a video display system provided by an exemplary embodiment of the present application;

[0078] Figure 2 A flowchart of a video display method provided by an exemplary embodiment of the present application;

[0079] Figure 3 A flowchart of another video display method provided as an exemplary embodiment of the present application;

[0080] Figure 4 A schematic structural diagram of a camera with a screen provided by an exemplary embodiment of the present application;

[0081] Figure 5 A flowchart of another video display method provided as an exemplary embodiment of the present application;

[0082] Figure 6 A schematic structural diagram of a video display device provided as an exemplary embodiment of the present application;

[0083] Figure 7 A schematic structural diagram of an electronic device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0084] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0085] The terms "first," "second," "third," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0086] Please refer to the following Figure 1 , which exemplarily shows a structural diagram of a video display system provided by an embodiment of the present application. Figure 1 As shown, the system includes a first device 110 and a second device 120. The first device 110 and the second device 120 can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0087] In some embodiments, the first device 110 and the second device 120 can both be smart cameras, smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart appliances, etc., but are not limited thereto.

[0088] Exemplarily, the first device 110 is a smart camera deployed in the residence of the target object (such as an elderly person), which is equipped with interactive components such as a camera, a microphone, and a speaker, and can serve as a care assistant device; the second device 120 can be a device with a display function in the residence, such as a smartphone, a smart watch, and other devices used by the target object, or a device with a display screen such as a television in the residence, which can perform remote data exchange and control interaction with the first device 110.

[0089] In some embodiments, the first device 110 is configured to perform the following steps: obtain user feature data; construct a virtual avatar model corresponding to the user based on the feature data; determine video content for the target event based on the virtual avatar model and a preset target event, wherein the video content includes prompt information related to the target event; and display the video content of the target event upon detecting that the target event is triggered. Specifically, the first device 110 can transmit the video content of the target event to the second device 120, so that the video content of the target event is displayed on the display screen of the second device 120.

[0090] Optionally, the first terminal device 110 may be installed with information collection devices such as a microphone and a camera, and may use the microphone and the camera to collect the user's characteristic data, or may obtain the user's characteristic data from other devices (such as the user's smartphone, tablet computer and other terminal devices).

[0091] In one embodiment, the video display system may also include only the first device 110 , in which components such as a camera, a display screen, a microphone, and a speaker are installed, that is, the first device 110 is a camera with a screen.

[0092] Optionally, the first device 110 is configured to perform the following steps: obtain user characteristic data; construct an avatar model corresponding to the user based on the characteristic data; determine video content for the target event based on the avatar model and a preset target event, wherein the video content includes prompt information related to the target event; and display the video content of the target event upon detecting that the target event is triggered. Specifically, the first device 110 displays the video content of the target event on its built-in display screen.

[0093] An exemplary embodiment of the present application provides a video display method. The video display method can be applied to the above-mentioned first device. For details, please refer to Figure 2 , which exemplarily shows a flow chart of a video display method provided in an embodiment of the present application. Figure 2 As shown, the video display method includes at least the following steps:

[0094] S201: Acquire user feature data.

[0095] The user can be the object for which the avatar model is to be generated and the provider of the original data. Specifically, the user can be a relative (such as a child) or friend with whom the target object wishes to have companionship, communication, or memory preserved digitally.

[0096] Optionally, the feature data includes appearance features and voice features of the user.

[0097] Among them, the above-mentioned appearance features are used to describe the user's external image, which may specifically include but are not limited to the user's skeletal features and facial features. The skeletal features can be the user's body structure information obtained through a depth camera, infrared scanning or somatosensory sensor, such as bone connection relationship, height ratio, joint position, etc., which can be used to drive the virtual character's actions; facial features may include facial structure, expression capture points, facial contour, skin color, hairstyle, etc., which can be used to build a three-dimensional facial model that matches the user's real appearance.

[0098] In addition, the sound features can be audio data collected by a microphone, which can be analyzed and processed to obtain the user's unique voiceprint features, timbre features, speaking speed, intonation and other parameters, which are used for subsequent speech synthesis or speech-driven models.

[0099] For example, during the data preparation phase, a user (such as the target subject's children or other family members or relatives) uploads multiple personal images of the user through the corresponding application or webpage. The multiple personal images can be video frame data contained in the video stream data or directly captured photo images; the multiple images contain multiple angles, multiple expressions, and multiple lighting scenes of the user, so that the first device can obtain a rich variety of appearance features. In addition, the user can also record an audio clip exceeding a preset length (such as 2 minutes), which is used to extract sound features. The audio content can include simple greetings, self-introductions, or examples of interactive language with the elderly at home.

[0100] In one embodiment, the user's feature data can be obtained by extracting features from a cloud device. The cloud device is connected to the first device through a network, and the cloud device can send the feature data to the first device through the network.

[0101] It is understandable that user feature data may also include age information, gender information, language preference information, personality label information, and other data. Age information is used to infer the age level of the user's appearance features, and can also be used to optimize modeling parameters such as voice intonation and expression frequency; gender information can be used as an auxiliary parameter to guide the facial proportions, hairstyle modeling, and voice feature selection of the three-dimensional model; language preference information can be used to optimize the intonation model in speech synthesis to make the generated speech more consistent with the user's regional language habits; personality label information can be used to customize the virtual image's tone, expression, and movement style, specifically "gentle" or "lively".

[0102] It should be noted that the user's characteristic data can be data extracted based on the audio, personal images, etc. provided by the user, or it can be data filled in or selected by the user on the relevant data collection page. This application does not make specific restrictions on this.

[0103] S202: Constructing a virtual image model corresponding to the user according to the feature data.

[0104] Among them, the virtual image model corresponding to the above-mentioned user can be a visual digital human image generated based on the user's characteristic data. The virtual image model can include a three-dimensional visual model and a speech synthesis model, and can interact with voice, action and expression in various scenarios.

[0105] S203: Determine the video content of the target event according to the virtual image model and the preset target event, where the video content includes prompt information related to the target event.

[0106] Among them, the target event can be a life interaction scene event preset by the user or automatically recognized by the system. Optionally, the above target event is at least one of the following: a schedule reminder event, a geographic location trigger event, and a sensor linkage event.

[0107] Optionally, the schedule reminder event can be an event triggered by a time plan or periodic task set in advance by the user, such as a reminder to take medicine, birthday wishes, wake-up time reminder, water drinking interval, etc. set by the user in the system.

[0108] For example, a schedule reminder event could trigger a "Good Morning Greetings" video at 8:00 every morning. A digital human (identical to the user's image) could smile and play the audio message "Mom / Dad, Good morning! A new day has begun. I hope you have a happy day." Another example could be playing a birthday greeting video at a preset time on the target person's birthday. In the video, the digital human voiced, "Dear Mom / Dad, happy birthday! I wish you happiness and longevity, and may you be healthy and happy every day."

[0109] In one embodiment, a geolocation trigger event may include an event triggered by a change in the location of a target subject (e.g., an elderly person). Specifically, this may be accomplished by capturing a video stream within the current field of view and determining the target subject's location information based on the video stream. The event is automatically triggered when the target subject is detected entering or leaving a specific location. For example, when the camera in the first device recognizes that the target subject is entering the kitchen area from the bedroom, it may be assumed that the target subject is about to drink water, eat, or perform other actions, and the corresponding video content may be automatically played, such as: "Mom, remember to drink a glass of warm water first."

[0110] Optionally, for the above-mentioned geographic location trigger events, multiple recognition mechanisms can be combined to improve the robustness of location information determination, such as combining the human posture and movement trajectory analysis of the target object in the video frame; image area division and background modeling; and auxiliary data fusion with external sensors such as door magnets, infrared, and thermal sensors (such as detecting door opening and closing, room temperature changes), etc.

[0111] In some embodiments, custom trigger areas can also be set, such as designating the kitchen, balcony, doorway, etc. as specific detection areas, and configuring corresponding target events and video content for each area to achieve intelligent companionship that is closer to real-life situations.

[0112] In some embodiments, geographic location triggering events may also include events triggered based on changes in the location of a user (such as the children of the target object). Specifically, the first device may be connected to a third device used by the user in communication, for example, by exchanging information through an intermediate server. The third device may obtain the user's geographic location information in real time based on an internal positioning module (such as a global positioning system or cellular base station positioning), and continuously detect the user's dynamic trajectory. When it is detected that the user has arrived at a preset home location (such as "home" or "company"), it can be determined that the user has completed the journey home from work, is about to arrive home, or has entered a certain important state, and then a notification message may be generated and sent to the intermediate server, so that the notification message can be forwarded to the first device through the intermediate server. After receiving the notification message, the first device automatically plays a video content containing the voice of "I'm home now, thank you for your hard work today" to the target object to inform the elderly (target object) that their children (users) have arrived home safely.

[0113] In the above embodiment, through remote location linkage, the system can achieve a warm and synchronous presentation of the children's life trajectory. Even if the person is far away, he or she can use the virtual image model to interact and communicate with the parents in a life-like manner, effectively alleviating the sense of lack of companionship.

[0114] Optionally, sensor linkage events can be events associated with changes in the status of terminal devices or home sensors, and can specifically include camera recognition, face detection, environmental sensing (such as temperature and humidity), door magnets, infrared, mattress sensors, etc. For example, if the camera recognizes that the target object has not been active in the room for more than a preset time threshold, it will automatically play a video content containing the voice message "Mom, how are you recently? Remember to rest." For another example, if the mattress sensor detects that the target object's wake-up time deviates from its historical wake-up time, the system will issue a voice greeting and remind it to pay attention to its health. For another example, if the door magnet sensor records that the target object enters or leaves the door, it will trigger a reminder to "be careful when going out, remember to bring your keys."

[0115] S204: When the target event is detected to be triggered, the video content of the target event is displayed.

[0116] For example, when a trigger condition related to a target event is detected, pre-generated video content based on that event is automatically called and displayed. This video content is presented by the user's corresponding avatar model, and combined with voice prompts, lip movements, and emotional expressions, it achieves natural voice and video interaction.

[0117] Optionally, the video content of the target event can be displayed in full screen on devices such as smart cameras, smart speakers with screens, or home tablets.

[0118] In an embodiment of the present application, by obtaining the user's characteristic data and constructing a virtual image model corresponding to the user based on the characteristic data, the video content of the target event is determined based on the virtual image model and the preset target event, wherein the video content includes prompt information related to the target event, and, when the trigger of the target event is detected, the video content of the target event is displayed. Thus, by obtaining the user's characteristic data, constructing a virtual image model that is highly similar to the user, and combining the preset event to generate video content containing prompt information, the video is automatically displayed when the event trigger is detected, thereby improving the personalized and emotional information transmission in the information interaction process, not only improving the user's acceptance and interactive experience, but also enhancing the expression effect of the reminder or care content. It is particularly suitable for scenarios such as companionship for the elderly and health reminders, helping to create a friendly and real interactive atmosphere and improving the effectiveness of information transmission.

[0119] An exemplary embodiment of the present application provides another video display method. The video display method can be applied to the above terminal device. For details, please refer to Figure 3 , which exemplarily shows a flow chart of a video display method provided in an embodiment of the present application. Figure 3 As shown, the video display method includes at least the following steps:

[0120] S301: Acquire user feature data.

[0121] Specifically, S301 is consistent with S201 and will not be repeated here.

[0122] S302: Performing three-dimensional modeling based on the shape features in the feature data to obtain a first virtual model that matches the shape features.

[0123] Optionally, the first virtual model may be the above-mentioned three-dimensional visual model, which may specifically include a three-dimensional facial network, a skeletal animation structure, texture mapping information, etc.

[0124] In some embodiments, in S302, the three-dimensional modeling process is performed based on the shape features in the feature data to obtain a first virtual model that matches the shape features, including:

[0125] S3021: Obtain a preset basic model.

[0126] The preset base model may be a standard human 3D model preloaded within the system, specifically a universal neutral model (e.g., a human mesh of standard proportions). The preset base model may include basic human structure, facial skeleton, control nodes, and an animation-driven bone layer.

[0127] It is understandable that the preset basic model, as the starting point of modeling, has a good topological structure, which facilitates subsequent personalized adjustments such as scaling, rotation, deformation, and mapping.

[0128] S3022: Structural adjustment is performed on the preset basic model based on the skeletal features in the feature data to obtain an initial virtual model corresponding to the appearance features.

[0129] Optionally, the skeletal features may be body proportion and skeleton information extracted from the user image or depth data, including height, shoulder width, limb length, joint position, etc.

[0130] It should be noted that the structural adjustment of the above-mentioned preset basic model based on the bone features in the above-mentioned feature data may be a deformation operation on the control points related to the bones in the basic model, such as lengthening the height of the model, adjusting the arm length or torso width, so that its overall shape is more in line with the body characteristics of the target object.

[0131] Optionally, the initial virtual model may refer to a personalized three-dimensional image model formed after the bone structure adjustment is completed, but does not yet include specific facial details and real texture mapping.

[0132] For example, the user's skeletal feature data, including key parameters such as height, limb proportions, and joint positions, can be first extracted through a posture estimation algorithm; then a mapping relationship is established between the extracted skeletal features and the standard skeleton of the basic model, and the differences in joint proportions and spatial positions are calculated; and based on the mapping relationship, the skeletal structure of the basic model is scaled and deformed, and the skeletal transformation is naturally transmitted to the three-dimensional mesh through a skinning algorithm; finally, an initial virtual model that matches the user's skeletal features is generated, providing a structural basis for subsequent facial modeling and action driving.

[0133] S3023: Performing texture mapping processing on the facial area in the initial virtual model based on the facial features in the feature data to obtain a first virtual model that matches the appearance features.

[0134] The facial area in the initial virtual model may be a three-dimensional grid area reserved in the initial virtual model for displaying the face.

[0135] Optionally, texture mapping is performed on the facial area in the above-mentioned initial virtual model based on the facial features in the above-mentioned feature data, which may refer to mapping the facial details in the two-dimensional image (such as eyes, lips, nose shape, skin color texture, etc.) to the corresponding positions of the three-dimensional model by projection or mapping, so that the model face visually restores the user's real appearance.

[0136] Through mapping processing, a complete three-dimensional virtual image with visual details such as the user's facial features, skin color, hairstyle, etc. can be generated, which is the first virtual model.

[0137] In other embodiments, a corresponding 3D facial network, skeletal network, and texture mapping information can be directly constructed based on the user's appearance features. Specifically, appearance features can be processed through multi-view image reconstruction and 3D face fitting to generate a 3D mesh model with facial contours, facial features, and geometric topology. Furthermore, skeletal features can be processed using pose estimation and human body modeling techniques to establish a skeletal network. Furthermore, corresponding texture mapping information can be generated based on the facial features. The 3D facial network, skeletal network, and texture mapping information are then fused to produce a first virtual model. For example, the 3D facial mesh can be bound to bone nodes in the skeletal network to achieve driven control of facial expressions. The texture mapping information can then be mapped to corresponding areas on the model surface to achieve realistic appearance restoration. Ultimately, a first virtual model is obtained that reproduces the user's appearance features, motion structure, and visual textures, providing foundational support for subsequent speech fusion, emotional expression, and video generation.

[0138] In the embodiment of the present application, the method of constructing a first virtual model based on appearance features can effectively extract the user's bone features and facial feature information, and combine three-dimensional modeling, bone binding and texture mapping technology to generate a personalized virtual image that highly restores the user's appearance and posture. This not only improves the modeling efficiency and restoration accuracy, but also provides a unified structural foundation for subsequent functions such as speech synthesis, emotional expression driving, and virtual video generation, thereby achieving a more natural and immersive human-computer interaction experience.

[0139] S303: Perform speech modeling based on the sound features in the feature data to obtain a second virtual model that matches the sound features. The second virtual model is a speech model used for speech synthesis.

[0140] Alternatively, the aforementioned sound features can be used to construct a text-to-speech (TTS) synthesis system using deep learning technology. Specifically, the sound features can be input into a speech modeling framework; the user's voiceprint is then embedded into the model as a control variable, ensuring that the synthesized speech has the same timbre as the user. Ultimately, a speech synthesis model is output that can accept arbitrary text and generate personalized speech.

[0141] It can be understood that the second virtual model is consistent with the audio sample uploaded by the user in terms of voice style. Compared with the first virtual model, the second virtual model is used to achieve personality restoration at the voice level.

[0142] S304: Fusing the first virtual model with the second virtual model to obtain a virtual image model corresponding to the user.

[0143] Among them, the above-mentioned first virtual model and the above-mentioned second virtual model are fused, that is, the visual modeling part (first virtual model) and the voice modeling part (second virtual model) are integrated, so that the virtual image not only has the user's three-dimensional appearance, but also can speak naturally and produce synchronous expressions, realizing a complete digital human model.

[0144] In some embodiments, in S304, the first virtual model and the second virtual model are merged to obtain a virtual image model corresponding to the user, including:

[0145] S3041: Determine voice feature information based on the second virtual model.

[0146] Among them, speech feature information includes data such as phoneme timing, pitch, speaking speed, intonation, strong and weak stress of the synthesized speech, which is used to drive facial expressions and lip synchronization.

[0147] S3042: Construct a mapping relationship between the target control node and the voice feature information.

[0148] Optionally, the target control node may be a key driving point in the first virtual model for controlling animation, such as a control bone node of lips, chin, eyes, eyebrows, and the like.

[0149] It can be understood that a mapping relationship is constructed between the above-mentioned target control node and the above-mentioned voice feature information, that is, the phonemes or rhythm features in the voice are formed into corresponding rules with the actions of the control nodes. For example: when the phoneme is "a", the mouth node is driven to open to a certain degree; when the tone is upward, the eyebrows are driven to rise, forming an expression linkage.

[0150] S3043: Based on the mapping relationship, the first virtual model and the second virtual model are merged to obtain a virtual image model corresponding to the user.

[0151] Optionally, based on the mapping relationship, the sound generated by the second virtual model and its driving parameters can be embedded into the first virtual model, so that the virtual image model can make voice in the video content and synchronously present natural lip changes and facial expressions during speaking.

[0152] In this embodiment of the present application, the fusion of the first and second virtual models enables deep integration of personalized speech synthesis results with the three-dimensional avatar model. This not only enables the avatar to output speech that closely matches the user's timbre, but also drives facial control nodes to synchronously generate mouth shape changes and facial expressions based on the phonemes, intonation, and emotional characteristics of the speech, thereby achieving precise coordination between speech and animation. This significantly enhances the avatar's anthropomorphic expressiveness and natural interaction, providing a reliable technical foundation for the subsequent generation of voice-driven videos and emotional expression.

[0153] S305: Acquire event information corresponding to the target event, where the event information includes prompt content.

[0154] The prompt content may be the voice text content in the video to be played by the avatar model. For example, the prompt content may be "Mom, remember to take your medicine."

[0155] In one embodiment, the event information corresponding to the target event may also include, but is not limited to, voice emotion parameters, facial expression tags, duration information or display time points, video background or scene parameters, etc. Among them, voice emotion parameters are used to control the tone style of the avatar model when synthesizing speech, such as gentleness, concern, joy, encouragement, etc.; facial expression tags are used to guide the types of facial expressions used by the avatar model, such as smiling, blinking, nodding, raising eyebrows, etc., to enhance the emotional expressiveness and interactive realism of the video; duration information or display time points are used to identify the target duration of the video content, or the playback time of the video content; video background or scene parameters can be used to set the video background style when the avatar model appears, such as home environment, cartoon scene, festival atmosphere, etc., to enhance visual diversity and immersion.

[0156] Optionally, the event information corresponding to the target event can be remotely set by the user through a related application in the third device.

[0157] S306: Input the prompt content into the second virtual model to generate event voice data corresponding to the target event.

[0158] Optionally, the second virtual model first performs a linguistic analysis on the prompt content, converting it into a phoneme sequence and annotating phonetic attributes such as stress, pauses, and intonation. This phoneme sequence can then be mapped into a sequence of intermediate acoustic feature parameters, including mel-spectrograms, pitch, and duration. These acoustic features are then fed into a neural network vocoder to synthesize the final event speech data. This process can incorporate parameters such as the user's voiceprint vector and intonation preferences, ensuring that the synthesized speech has highly personalized characteristics such as timbre, speaking rate, and tone that are highly consistent with the user.

[0159] S307: Determine driving configuration parameters based on the event voice data, where the driving configuration parameters include lip-shaping driving parameters and expression driving parameters.

[0160] The driving configuration parameters may be control data for driving the animation.

[0161] Optionally, lip-driven parameters are used to control the opening and closing of the mouth and changes in mouth shape. For example, the pronunciation characteristics of each phoneme (such as "a," "e," "m," and "b") are mapped to a corresponding mouth shape, forming a lip curve. This is then output as a set of time-varying mouth control values ​​(such as mouth opening and mouth corner positions) as lip-driven parameters.

[0162] Optionally, expression driving parameters are used to control facial expressions such as smiling and surprised in the eyes and eyebrows. Specifically, based on the intonation, pitch, and speech rate of the event voice data, combined with the semantic labels of the prompt (such as blessing, concern, and reminder), emotion labels (such as smiling, surprised, and peaceful) are generated and mapped to expression control parameters, such as eyebrow raising, eye curvature, and cheek lifting.

[0163] S308: driving the target control node in the avatar model according to the driving configuration parameters to generate an action sequence synchronized with the event voice data.

[0164] Optionally, the avatar model may be identified to determine the target control node in the avatar model.

[0165] In one embodiment, the system extracts speech feature information from generated event speech data, deriving driver configuration parameters that include lip shape changes and emotional expressions. These parameters are then mapped to target control nodes in the avatar model, such as the mouth, eyes, and eyebrows. Furthermore, through keyframe interpolation and timeline alignment, each control node is driven to generate an action sequence synchronized with the speech content. This allows the avatar to naturally display corresponding lip shape changes and emotional expressions during speech playback, achieving a highly coordinated dynamic performance of speech and movement, enhancing the realism and immersion of the interaction.

[0166] S309: Synthesize the action sequence and the event voice data to generate video content corresponding to the target event.

[0167] Optionally, the action sequence and event voice data are synthesized into a complete video file. This video is the final result of the digital human "appearing and speaking" and can be used for display.

[0168] The embodiment of the present application implements personalized speech synthesis and animation driving for the avatar based on the prompt content corresponding to the target event, combined with information such as voice emotion parameters, expression tags, display duration, and background scenes. The prompt is subjected to phonetic analysis by a second virtual model to generate event voice data, and further driven configuration parameters are extracted by combining voice rhythm and emotional characteristics. This allows precise control of the control nodes related to lip shape and expression in the avatar model, generating an action sequence that is highly synchronized with the voice, and ultimately synthesizing natural, smooth, and emotionally expressive video content. This effectively improves the coordination and anthropomorphism of the avatar in terms of language expression and visual presentation, and realizes an immersive interactive experience driven by multimodal integration.

[0169] S310: Detect whether the triggering condition of the target event is met through a preset event trigger detection mechanism.

[0170] In one embodiment, the event trigger detection mechanism includes at least one of the following: a time detection mechanism, an environment perception mechanism, and a user behavior recognition mechanism. The time detection mechanism may be based on an internal clock or calendar function, comparing the current time with the time conditions set for the target event to determine whether the triggering requirements are met; the environment perception mechanism may be based on connected sensors or devices (such as temperature and humidity sensors, light sensors, door magnets, infrared sensors, mattress monitors, etc.) to sense the current environmental conditions and determine whether to trigger an event; and the user behavior recognition mechanism may be based on devices such as cameras and microphones to recognize user actions or voice commands to determine whether the conditions for triggering an event are met.

[0171] S311: When it is detected that the triggering condition of the target event is met, it is determined that the triggering of the target event is detected, and the video content of the target event is displayed.

[0172] Optionally, when the event trigger detection mechanism determines that the trigger conditions of the target event have been met, the target event is deemed to have been successfully triggered, and an event trigger signal is generated, and the virtual image video content corresponding to the target event is called from the pre-generated video content library for display.

[0173] Optionally, the display method can be playback on a smart camera, playback on an external display screen, output on a smart speaker with a screen, etc. This application does not make any specific restrictions on this.

[0174] Figure 4 A schematic structural diagram of a camera with a screen provided by an exemplary embodiment of the present application is shown in FIG. Figure 4 The first device 400 is a camera with a screen, which includes a camera 410 and a display screen 420. In addition, it can also be provided with components such as a microphone and a sensor. When the first device 400 detects that the triggering conditions of the target event are met, it determines that the triggering of the above-mentioned target event is detected, and displays the video content of the above-mentioned target event through the display screen 420.

[0175] In the embodiment of this application, a preset event trigger detection mechanism is used to determine in real time whether the trigger conditions of the target event are met, and when the conditions are met, the video content corresponding to the event is automatically displayed, thereby realizing the intelligent and contextual presentation of the virtual image. This embodiment combines multi-dimensional perception methods such as time detection, environmental perception, and user behavior recognition, significantly improving the accuracy and adaptability of event triggering; combined with the video playback function of the virtual image, the system can push voice prompts and emotional interactive content in a natural and appropriate manner in the user's real life scenes, realizing a variety of intelligent interactive uses such as virtual companionship, health reminders, and caring greetings, thereby improving the practicality, responsiveness, and user experience of the virtual image model.

[0176] An exemplary embodiment of the present application provides another video display method. The video display method can be applied to the above terminal device. For details, please refer to Figure 5 , which exemplarily shows a flow chart of a video display method provided in an embodiment of the present application. Figure 5 As shown, the video display method includes at least the following steps:

[0177] S501: Acquire user feature data.

[0178] S502: Constructing a virtual image model corresponding to the user according to the feature data.

[0179] S503: Determine the video content of the target event according to the virtual image model and the preset target event, where the video content includes prompt information related to the target event.

[0180] S504: When the target event is detected to be triggered, the video content of the target event is displayed.

[0181] Specifically, S501-S504 are consistent with the above-mentioned S201-S204 and will not be repeated here.

[0182] S505: Detect whether the target object responds.

[0183] The response operation may be a reaction behavior of the target object to the video content, such as speaking, nodding, gesturing, approaching the camera, touching the screen, etc.

[0184] It is understandable that the purpose of detecting whether the target object responds is to detect whether the target object has made corresponding interactive behaviors when watching the video content displayed by the digital person (virtual child).

[0185] S506: If it is detected within a preset time period that the target object performs the response operation, feedback is provided according to the response information corresponding to the response operation.

[0186] The preset time period may be a pre-set time range for waiting for a response from the target object, such as 10 seconds, 15 seconds, etc., which is used to limit the time window for response judgment.

[0187] Optionally, a camera, microphone, or other sensors may be used to determine whether the target object has performed a predefined response operation.

[0188] In one embodiment, the response information may be the content of the corresponding response operation. For example, if the target object says "OK", the response information may be the voice message "OK".

[0189] In one embodiment, feedback is provided based on the response information corresponding to the above-mentioned response operation, and a reply can be made based on the response information through voice, video, animation, etc., to achieve a closed loop of interaction.

[0190] S507: If the target object is not detected to perform the response operation within the preset time period, the video content of the target event is displayed again.

[0191] If the target object is not detected to have made the above-mentioned response operation within the above-mentioned preset time period, it may indicate that the target object has not made any predefined response behavior within the specified time, which may be due to reasons such as not hearing clearly, not paying attention, or being unconscious. Then, the same video content can be replayed (or slightly adjusted) to reinforce the reminder or attract attention.

[0192] Furthermore, a playback threshold (for example, 3 times) can be set. When the number of repeated playbacks exceeds the playback threshold, visual attraction methods such as increasing the voice volume, adding vibration or light prompts, adding icon flashing, text highlighting, background animation, etc. can be used to improve the perception intensity.

[0193] The embodiment of the present application can detect in real time whether the target object has made a response operation after the video content of the target event is played, and realize intelligent feedback and interactive control based on the response result. On the one hand, by setting a preset time period, combining sensors such as cameras and microphones to detect behaviors such as speaking, nodding, and gestures, the recognition and judgment of natural user interactions can be realized; on the other hand, if no response is detected within the time period, the system can automatically repeat the prompt content and, if necessary, increase the reminder intensity by means of increasing the volume, enhancing the visual effect, etc., so as to effectively deal with situations where the user has not noticed or has not heard clearly. As a result, the intelligence level and fault tolerance of the interaction are improved, and the practicality and user experience of the virtual image system in scenarios such as daily companionship, health reminders, and voice interaction are enhanced.

[0194] Please refer to the following Figure 6 , which is a structural diagram of a video playback device provided by an exemplary embodiment of the present application. Figure 6 As shown, the video playback device 600 includes:

[0195] Acquisition module 601, used to acquire user feature data;

[0196] A construction module 602 is configured to construct a virtual image model corresponding to the user according to the feature data;

[0197] A determination module 603 is configured to determine the video content of the target event based on the virtual image model and the preset target event, wherein the video content includes prompt information related to the target event;

[0198] The display module 604 is configured to display the video content of the target event when the triggering of the target event is detected.

[0199] In a possible implementation, the feature data includes the user's appearance and voice features;

[0200] The building block 602 includes:

[0201] A first processing unit is configured to perform three-dimensional modeling based on the shape features in the feature data to obtain a first virtual model that matches the shape features;

[0202] a second processing unit configured to perform speech modeling based on the sound features in the feature data to obtain a second virtual model that matches the sound features, wherein the second virtual model is a speech model for speech synthesis;

[0203] The fusion unit is used to fuse the first virtual model with the second virtual model to obtain a virtual image model corresponding to the user.

[0204] In a possible implementation, the appearance features include the user's skeletal features and facial features;

[0205] The first processing unit includes:

[0206] Get subunit, used to get preset basic model;

[0207] An adjustment subunit, configured to perform structural adjustment on the preset basic model based on the skeletal features in the feature data to obtain an initial virtual model corresponding to the appearance features;

[0208] The processing subunit is configured to perform texture mapping processing on the facial region of the initial virtual model based on the facial features in the feature data to obtain a first virtual model that matches the appearance features.

[0209] In a possible implementation, the fusion unit includes:

[0210] A first determining subunit is configured to determine a target control node in the first virtual model, wherein the target control node includes a lip control node and an expression control node;

[0211] A second determining subunit, configured to determine speech feature information based on the second virtual model;

[0212] A construction subunit, configured to construct a mapping relationship between the target control node and the voice feature information;

[0213] The fusion subunit is configured to fuse the first virtual model with the second virtual model based on the mapping relationship to obtain a virtual image model corresponding to the user.

[0214] In a possible implementation, the determining module 603 includes:

[0215] An acquiring unit, configured to acquire event information corresponding to the target event, wherein the event information includes prompt content;

[0216] a generating unit, configured to input the prompt content into the second virtual model to generate event voice data corresponding to the target event;

[0217] A first determining unit is configured to determine driving configuration parameters based on the event voice data, wherein the driving configuration parameters include lip-shaping driving parameters and expression driving parameters;

[0218] a driving unit configured to drive the target control node in the avatar model according to the driving configuration parameters to generate an action sequence synchronized with the event voice data;

[0219] The synthesis unit is used to synthesize the above action sequence with the above event voice data to generate video content corresponding to the above target event.

[0220] In a possible implementation, the target event is at least one of the following: a schedule reminder event, a geographic location trigger event, or a sensor linkage event;

[0221] The display module 604 includes:

[0222] A detection unit, configured to detect whether the triggering conditions of the target event are met through a preset event trigger detection mechanism;

[0223] The second determining unit is configured to determine that the triggering of the target event is detected when it is detected that the triggering condition of the target event is met, and to display the video content of the target event.

[0224] In a possible implementation, the apparatus further includes:

[0225] A detection module is used to detect whether the target object responds to the operation;

[0226] A feedback module is configured to provide feedback based on response information corresponding to the response operation when detecting that the target object has performed the response operation within a preset time period;

[0227] The display module is further configured to display the video content of the target event again when the target object is not detected to perform the response operation within the preset time period.

[0228] The division of the modules in the above-described video playback device 600 is for illustration only. In other embodiments, the video playback device can be divided into different modules as needed to perform all or part of the functions of the above-described video playback device. The various modules in the video playback device provided in the embodiments of this specification can be implemented in the form of a computer program. The computer program can be run on a terminal or server. The program modules comprising the computer program can be stored in a memory of the terminal or server. When the computer program is executed by a processor, it implements all or part of the steps of the video playback method described in the embodiments of this specification.

[0229] See next Figure 7 , which is a structural diagram of an electronic device provided by an exemplary embodiment of the present application. Figure 7 As shown, the electronic device 700 may include: a processor 710 and a memory 720 , and may also include a user interface 730 , a network interface 740 and a communication bus 750 .

[0230] The processor 710 may include one or more processing cores. The processor 710 utilizes various interfaces and circuits to connect various components within the electronic device 700. It executes instructions, programs, code sets, or instruction sets stored in the memory 720, and accesses data stored in the memory 720 to perform various functions and process data within the electronic device 700. Optionally, the processor 710 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 710 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system and applications; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 710 and may be implemented as a separate chip.

[0231] Among them, the memory 720 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 720 includes a non-transitory computer-readable storage medium. The memory 720 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 720 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a receiving function, a control function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 720 may also be optionally at least one storage device located away from the aforementioned processor 710. As Figure 7 As shown, the memory 720 as a computer storage medium may include an operating system, a network communication module, a user interface module, and program instructions.

[0232] Optionally, the communication bus 750 is used to achieve connection and communication between these components. The user interface 730 may include a display screen (Display), a camera (Camera), and may also include a standard wired interface and a wireless interface; the network interface 740 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0233] exist Figure 7 In the electronic device 700 shown, the processor 710 may be configured to call program instructions stored in the memory 720 and specifically perform the following operations:

[0234] Obtain user feature data;

[0235] Constructing a virtual image model corresponding to the user according to the feature data;

[0236] Determining the video content of the target event based on the virtual image model and the preset target event, wherein the video content includes prompt information related to the target event;

[0237] When the target event is detected to be triggered, the video content of the target event is displayed.

[0238] In a possible implementation, the feature data includes the user's appearance and voice features;

[0239] The aforementioned construction of the virtual image model corresponding to the aforementioned user based on the aforementioned feature data includes:

[0240] Performing three-dimensional modeling based on the shape features in the feature data to obtain a first virtual model that matches the shape features;

[0241] Performing speech modeling based on the sound features in the feature data to obtain a second virtual model that matches the sound features, wherein the second virtual model is a speech model for speech synthesis;

[0242] The first virtual model and the second virtual model are fused to obtain a virtual image model corresponding to the user.

[0243] In a possible implementation, the appearance features include the user's skeletal features and facial features;

[0244] The three-dimensional modeling process is performed based on the shape features in the feature data to obtain a first virtual model that matches the shape features, including:

[0245] Get the preset base model;

[0246] Adjusting the structure of the preset basic model based on the skeletal features in the feature data to obtain an initial virtual model corresponding to the appearance features;

[0247] Texture mapping is performed on the facial area of ​​the initial virtual model based on the facial features in the feature data to obtain a first virtual model that matches the appearance features.

[0248] In a possible implementation, the fusing of the first virtual model and the second virtual model to obtain the virtual image model corresponding to the user includes:

[0249] Determining target control nodes in the first virtual model, wherein the target control nodes include lip control nodes and expression control nodes;

[0250] Determining speech feature information based on the second virtual model;

[0251] Constructing a mapping relationship between the target control node and the voice feature information;

[0252] Based on the mapping relationship, the first virtual model and the second virtual model are fused to obtain a virtual image model corresponding to the user.

[0253] In a possible implementation, determining the video content of the target event based on the avatar model and the preset event includes:

[0254] Obtaining event information corresponding to the target event, including prompt content;

[0255] Inputting the prompt content into the second virtual model to generate event voice data corresponding to the target event;

[0256] Determine driving configuration parameters based on the event voice data, wherein the driving configuration parameters include lip-shaped driving parameters and expression driving parameters;

[0257] driving the target control node in the avatar model according to the driving configuration parameters to generate an action sequence synchronized with the event voice data;

[0258] The action sequence is synthesized with the event voice data to generate video content corresponding to the target event.

[0259] In a possible implementation, the target event is at least one of the following: a schedule reminder event, a geographic location trigger event, or a sensor linkage event;

[0260] When the target event is detected, the video content of the target event is displayed, including:

[0261] Through the preset event trigger detection mechanism, detect whether the trigger conditions of the above target events are met;

[0262] When it is detected that the triggering condition of the target event is met, it is determined that the triggering of the target event is detected, and the video content of the target event is displayed.

[0263] In a possible implementation, after displaying the video content of the target event, the method further includes:

[0264] Detect whether the target object responds to the operation;

[0265] If the target object is detected to have made the above-mentioned response operation within the preset time period, feedback is provided according to the response information corresponding to the above-mentioned response operation;

[0266] If the target object is not detected to perform the above-mentioned response operation within the above-mentioned preset time period, the video content of the above-mentioned target event is displayed again.

[0267] The present application also provides a computer-readable storage medium having instructions stored therein that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of the above-described embodiments. If the various components of the video display device are implemented as software functional units and sold or used as independent products, they may be stored in the computer-readable storage medium.

[0268] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The above-mentioned computer program product includes one or more computer instructions. When the above-mentioned computer program instructions are loaded and executed on a computer, the above-mentioned process or function according to the embodiment of the present application is generated in whole or in part. The above-mentioned computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The above-mentioned computer instructions can be stored in a computer-readable storage medium or transmitted by the above-mentioned computer-readable storage medium. The above-mentioned computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The above-mentioned computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The above-mentioned available media can be magnetic media (for example, floppy disks, hard disks, tapes), optical media (for example, digital versatile discs (DVDs)), or semiconductor media (for example, solid state disks (SSDs)).

[0269] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. The technical features of this embodiment and the implementation scheme can be combined in any manner unless they conflict.

[0270] The above embodiments are merely preferred embodiments of the present application and are not intended to limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements made to the technical solutions of the present application by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present application.

Claims

1. A video display method, characterized in that: include: Obtain user feature data; constructing a virtual image model corresponding to the user according to the feature data; Determining the video content of the target event according to the virtual image model and the preset target event, wherein the video content includes prompt information related to the target event; When the target event is detected to be triggered, the video content of the target event is displayed.

2. The method according to claim 1, wherein The characteristic data includes the appearance characteristics and voice characteristics of the user; The step of constructing a virtual image model corresponding to the user according to the feature data includes: Performing three-dimensional modeling based on the shape features in the feature data to obtain a first virtual model that matches the shape features; Performing speech modeling processing based on the sound features in the feature data to obtain a second virtual model that matches the sound features, wherein the second virtual model is a speech model for speech synthesis; The first virtual model and the second virtual model are fused to obtain a virtual image model corresponding to the user.

3. The method according to claim 2, wherein The appearance features include the user's skeletal features and facial features; The performing of three-dimensional modeling based on the shape features in the feature data to obtain a first virtual model matching the shape features includes: Get the preset base model; Performing structural adjustment on the preset basic model based on the skeletal features in the feature data to obtain an initial virtual model corresponding to the appearance features; Texture mapping is performed on the facial area in the initial virtual model based on the facial features in the feature data to obtain a first virtual model that matches the appearance features.

4. The method according to claim 2, wherein The fusing the first virtual model and the second virtual model to obtain a virtual image model corresponding to the user includes: Determining target control nodes in the first virtual model, wherein the target control nodes include lip control nodes and expression control nodes; determining speech feature information based on the second virtual model; Constructing a mapping relationship between the target control node and the voice feature information; Based on the mapping relationship, the first virtual model and the second virtual model are fused to obtain a virtual image model corresponding to the user.

5. The method according to claim 4, wherein The step of determining the video content of the target event based on the virtual image model and the preset event includes: Acquire event information corresponding to the target event, wherein the event information includes prompt content; Inputting the prompt content into the second virtual model to generate event voice data corresponding to the target event; Determine driving configuration parameters based on the event voice data, wherein the driving configuration parameters include lip-shaped driving parameters and expression driving parameters; driving the target control node in the avatar model according to the driving configuration parameters to generate an action sequence synchronized with the event voice data; The action sequence is synthesized with the event voice data to generate video content corresponding to the target event.

6. The method according to claim 1, wherein The target event is at least one of the following: a schedule reminder event, a geographic location trigger event, or a sensor linkage event; The step of displaying the video content of the target event upon detecting that the target event is triggered includes: Detecting whether the triggering conditions of the target event are met through a preset event trigger detection mechanism; When it is detected that the triggering condition of the target event is met, it is determined that the triggering of the target event is detected, and the video content of the target event is displayed.

7. The method according to claim 1, wherein After presenting the video content of the target event, the method further includes: Detect whether the target object responds to the operation; If the target object is detected to have performed the response operation within a preset time period, feedback is provided according to the response information corresponding to the response operation; If the target object is not detected to perform the response operation within the preset time period, the video content of the target event is displayed again.

8. A video display device, characterized in that: include: Acquisition module, used to obtain user feature data; A construction module, configured to construct a virtual image model corresponding to the user according to the feature data; a determination module, configured to determine the video content of the target event according to the avatar model and a preset target event, wherein the video content includes prompt information related to the target event; The display module is used to display the video content of the target event when the trigger of the target event is detected.

9. An electronic device, characterized in that: include: processor and memory; The memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual image generation method and device, terminal and storage medium

    CN112337105A

  • Virtual face generation method

    CN113781610A

  • Metacosmic emotion accompanying virtual human implementation method and system based on neural network

    CN115494941A

  • Method and system for generating personalized virtual human

    CN116400806A

  • Multi-modal man-machine interaction method and device based on 3D virtual human and medical self-service machine

    CN118444819A