A method and apparatus for fusing video multi-modal content for character visualization

By integrating multimodal video content into a character visualization method, video data is extracted and aligned to generate video summaries and player-style visualizations. This solves the problem of difficulty in understanding the expressive skills and body language of characters in videos in existing technologies, and enables intuitive multidimensional analysis and understanding.

CN117033699BActive Publication Date: 2026-02-06INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202310976862.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-04
Publication Date
2026-02-06
Estimated Expiration
2043-08-04

AI Technical Summary

Technical Problem

Existing video character visualization technologies struggle to quickly understand a person's expressive skills and body language, and cannot provide intuitive, multi-dimensional analysis.

Method used

By integrating multimodal video content into a person visualization method, the original video data is extracted, modal features are extracted and aligned, and video summaries and player-style visualizations are generated. Facial expressions, eye gaze, gestures, and positional movements are combined to achieve person visualization.

Benefits of technology

It enables rapid understanding of the body language and expressive techniques of people in videos, provides intuitive multi-dimensional analysis, supports users in noticing subtle changes in body language, and has good comprehensibility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033699B_ABST
    Figure CN117033699B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of fusion video multimodal content character visualization method and device.The method includes: extracting the original data of target video under each modality;According to the minimum scale of the original data extracted under each modality, align the original data under each modality;Based on the original data under each modality after alignment, in the scale range set, according to the different modal characteristics, modal characteristic data is extracted;Based on the extracted modal characteristic data, for the visual form of video abstract, character visualization is carried out, and for the visual form of video player, character visualization is carried out;Based on the visual form of video player, video playback and the interactive realization of visual element in the process of playing are carried out.The present application innovatively uses the visual form of multimodal character abstract of fusion video content, and innovatively uses the visual form of video player for enhancing character non-verbal expression, can promote the quick understanding of video character body language and expression skill.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of visualization, and particularly relates to a method and device for visualizing characters by fusing multi-modal content of videos. BACKGROUND

[0002] Video character visualization technology aims to visualize and analyze characters in videos using computer graphics and computer vision technology. Meanwhile, different modalities of videos convey a lot of information, which affects the audience's experience and understanding of the video from multiple aspects. For example, to understand the effectiveness of a speech video, information from multiple modalities such as speech, image, and emotion needs to be combined. Visualizing characters in videos based on multi-modal information can help users quickly understand the video content, including the expression skills of the characters in the video, and find related information between videos. On the other hand, by combining interactive technology, users can interact with the visualized content to further analyze and understand the video content.

[0003] Existing methods have made many-sided researches on video character visualization technology and multi-modal understanding of videos. For example, Chinese patent application CN115880441A discloses a method and system for generating 3D visualized simulation characters, which can analyze dynamic content of faces and limbs of a basic character model based on VR model by obtaining character video data of a teaching process, but it is difficult for users to quickly understand the expression skills and body language of the characters in the video. Chinese patent application CN115690274A discloses a method for quickly generating 3D character animation based on videos, which only animates the posture of the characters in the video and does not involve further understanding of the video content. Chinese patent application CN113743271A discloses a method and system for visual analysis of effectiveness of video content based on multi-modal emotion, which analyzes the effectiveness of the video content by means of a visualization method, but it cannot analyze the body language of the characters in the video from the perspective of multi-modalities.

[0004] In summary, existing methods have limited ability to quickly understand the expression skills and body language of characters, and cannot provide users with intuitive and multi-dimensional understanding and analysis of video characters. SUMMARY

[0005] The present application proposes a novel method for visualizing characters by fusing multi-modal content of videos, which includes two forms of video summary and video player, and focuses on the body language of the characters in the video, including facial expressions, eye gaze, gestures, and position movements. The method aims to facilitate quick understanding of the body language and expression skills of the characters in the video.

[0006] The technical solution adopted by the present application includes the following steps:

[0007] A method for visualizing characters by fusing multi-modal content of videos, comprising the following steps:

[0008] extracting original data of the target video in each modality;

[0009] aligning the original data in each modality according to the minimum scale of the original data extracted in each modality;

[0010] extracting modality feature data according to different modality characteristics within a set scale range based on the aligned original data in each modality;

[0011] performing character visualization for the visual form of the video summary based on the extracted modality feature data;

[0012] performing character visualization for the visual form of the video player based on the extracted modality feature data;

[0013] performing video playing and interactive implementation of visual elements in the playing process based on the visual form of the video player.

[0014] Further, the modalities include at least one of video, image, audio, etc.

[0015] Further, the extracted original data of the image modality includes stage position data, eye gaze data, facial emotion data, and human body key point data.

[0016] Further, the extracted original data of the audio modality includes the pitch and tone of the video.

[0017] Further, the stage position data is extracted by the following steps:

[0018] 1) performing human target detection and target tracking from each video image frame of the target video;

[0019] 2) calculating the coordinate position of the center point of the human target detection frame relative to the horizontal and vertical coordinates of the target video interface or relative to the three-dimensional coordinate system of the stage.

[0020] Further, the eye gaze data is extracted by the following steps:

[0021] 1) performing face recognition and positioning from each video image frame of the target video;

[0022] 2) extracting eye gaze data in each face image.

[0023] Further, the facial emotion data is extracted by the following steps:

[0024] 1) performing face recognition and positioning from each video image frame of the target video;

[0025] 2) using clustering method or face recognition method to screen all face images appearing in the target video;

[0026] 3) Using facial emotion calculation method, extracting arousal and valence data in each face image, obtaining continuous emotional intensity data of facial expression;

[0027] 4) Identifying the emotion category of all face images, obtaining discrete emotional category data of facial expression.

[0028] Further, the human key point data is extracted by the following steps:

[0029] 1) Human target detection and target tracking are performed on each video image frame of the target video;

[0030] 2) Facial key point data and body key point data are extracted based on a single human target.

[0031] Further, the character face visualization in the form of video summary is performed by the following steps:

[0032] 1) Based on the extracted eye gaze data, the direction and position of eye gaze of each frame are calculated and arithmetically averaged within a set scale range.

[0033] 2) Based on the extracted facial emotion data, the distribution of discrete emotional types, the arithmetic average of valence and arousal are calculated within a set scale range.

[0034] 3) Based on the extracted pitch data, the arithmetic average of the pitch is calculated within a set scale range.

[0035] 4) According to the arithmetically averaged eye gaze direction and position coordinates, the polar coordinates based on the eye center point are calculated, and the eye elements in the face visualization are generated using the polar coordinates.

[0036] 5) Different regions of the face visualization are mapped to the distribution of different categories of discrete emotional types.

[0037] 6) The arousal of emotion is mapped by the convexity degree of the turning curve of the face periphery.

[0038] 7) The valence of emotion based on emotion reflects the positive and negative degree of emotion, and the arithmetic average of facial emotion valence is mapped to the up and down bending degree of the mouth element in the face visualization.

[0039] 8) According to the arithmetic average of the pitch, it is mapped to the width of the mouth element in the face visualization element.

[0040] Further, in the face element, the emotional distribution except for neutral emotion is displayed by the annular region of the ring chart, and the region inside the ring chart represents neutral emotion.

[0041] Further, the setting scale range represents a continuous speech video segment selected by the user.

[0042] Further, the visualization of the character pose in the form of a video summary is performed by the following steps:

[0043] 1) Based on the extracted pose data, key point alignment and coordinate standardization of the pose data are performed.

[0044] 2) Based on the aligned and standardized pose data, a fixed number of key pose frames are extracted.

[0045] 3) The fixed number of key pose frames are visualized, and the attribute values of the generated visualization elements correspond to their pose data.

[0046] Further, the attribute values of the generated visualization elements include color, transparency, stroke thickness, etc.

[0047] Further, the visualization of the character stage position in the form of a video summary is performed by the following steps:

[0048] 1) Based on the extracted stage position data, standardization is performed according to the screen size of the target video.

[0049] 2) The standardized stage position data is visualized in the form of a scatter plot, and the attribute values of the scatter plot elements correspond to their stage position data.

[0050] Further, the attribute values of the scatter plot elements include color, transparency, radius size, etc.

[0051] Further, the visualization of the character in the form of a video player is performed by the following steps:

[0052] 1) Based on the extracted character key point data, key point alignment with the video frame is performed frame by frame.

[0053] 2) According to the setting of different time scales, the character key point data corresponding to consecutive time segments are combined.

[0054] 3) Using the visualization technology of masks, the character key point group is mapped on the mask, and the mask is superimposed on the corresponding video frame.

[0055] Further, the character key point data includes facial key points and body key points.

[0056] Further, different time scales include single frame and multiple frames.

[0057] Further, the interactive function in the form of a video player is realized by the following steps:

[0058] 1) Based on the video player, the play and pause function of the video player is realized in the form of character visualization.

[0059] 2) Based on the character visualization element, the corresponding interactive event of each region is realized.

[0060] A character visualization device fusing video multi-modal content, comprising:

[0061] A multi-modal feature data extraction module is used to extract the original data of the target video under each modality, align the original data under each modality according to the minimum scale of the original data extracted under each modality, and extract modal feature data within a set scale range based on the aligned original data under each modality according to different modal characteristics.

[0062] A video character visualization module is used to perform character visualization for the visualization form of the video summary based on the extracted modal feature data, perform character visualization for the visualization form of the video player based on the extracted modal feature data, and perform video playing and interactive implementation of the visualization element in the playing process based on the visualization form of the video player.

[0063] Compared with the prior art, the present application has the following advantages and positive effects:

[0064] 1. The present application innovatively uses the visualization form of multi-modal character summary fusing video content, which can cover more multi-modal information compared with the traditional expression form, and intuitively displays the actual significance of the visualization data through the personification visualization form, facilitating the user to quickly understand the video content.

[0065] 2. The present application innovatively uses the visualization form of the video player enhancing character non-verbal expression, which has good understandability compared with the traditional video playing form by superimposing the extracted non-verbal information on the video content, can support the user to perceive the small body changes, and can support the user to observe the rules of body, eye contact and other changes over time by adjusting the time scale of the key visualization, thereby assisting the user to better understand the non-verbal expression of the video character.

[0066] 3. The present application proposes a complete data extraction and data visualization process, which collects multi-modal data through an algorithm and automatically realizes data visualization, can be conveniently integrated into other data analysis processes, and has good expansibility. BRIEF DESCRIPTION OF DRAWINGS

[0067] Figure 1 The flowchart of the method of the present application.

[0068] Figure 2 The design diagram of the video character visualization based on the video summary.

[0069] Figure 3 Video character visualization based on video player. DETAILED DESCRIPTION

[0070] In order to make the person in the technical field better understand the present application, the video character multi-modal visualization method provided by the present application is described in detail below with reference to the drawings, but does not constitute a limitation on the present application.

[0071] As shown in the figure, the implementation steps of the method of the present application are as follows: Figure 1

[0072] (1) Extract human key points (including facial key points, body key points), eye gaze, facial emotion and stage position from video images frame by frame;

[0073] (2) Extract video sound, extract tone and pitch data based on video sound;

[0074] (3) Align the various data obtained from the image and audio according to the minimum scale of the extracted data;

[0075] (4) According to the extracted facial emotion, body key points, eye gaze, sound (tone and pitch) and stage position five kinds of multi-modal data, establish the mapping relationship between multi-modal data and video summary visualization design, generate video summary form visualization elements according to the mapping relationship;

[0076] (5) According to the extracted facial key points and body key points, establish the mapping relationship between the key points and the video image frames, and generate the video player form visualization elements according to the mapping relationship.

[0077] Exemplarily, the eye gaze, human key points, stage position and facial emotion data extracted from the video all need the current video frame to contain the video subject, if the current video frame does not contain the video subject, step (1) is skipped.

[0078] Exemplarily, the tone and pitch data extracted from the video sound both need the current video sound to correspond to the video subject, if the current video sound does not belong to the video subject, step (2) is skipped.

[0079] Exemplarily, the facial emotion data is to extract two types of data, discrete emotion categories and continuous emotion intensity, from the video modal using the emotion computing model, and the image sequence, audio and other modalities in the video are extracted according to their respective modalities, and there are different extraction scales such as frames, segments and sentences, which all need to be aligned according to the minimum scale.

[0080] Exemplarily, the video summary form visualization design currently mainly covers three parts of facial visualization, posture visualization and stage position visualization.​

[0081] For example, in facial visualization design, a ring chart is used to display the distribution of discrete emotion types, and the degree of emotional arousal is mapped by the convexity of the curves at the outer edge of the ring chart. Based on the arithmetic mean of facial emotional arousal, the formula generated in polar coordinates is:

[0082]

[0083]

[0084] Where θ n It is the polar angle of the nth inflection point, r n It is the polar radius of the nth inflection point, Δ r This is the mapping value of the arithmetic mean of arousal, where r is a preset constant and m is the number of inflection points. Based on the distribution of emotions other than neutral emotions, the proportion of each emotion is mapped onto an arc, and the endpoints of this arc are mapped onto a broken curve composed of m inflection points. The closed shape formed by the endpoints of the arc and their mapped endpoints on the broken curve is filled with the color corresponding to the emotion. Simultaneously, based on the arithmetic mean of the eye gaze direction and position coordinates, its polar coordinates based on the eye center are calculated, and the eyeball portion of the eye gaze element is generated based on these coordinates. Then, based on the arithmetic mean of facial emotional valence, it is mapped to the vertical curvature of the mouth element.

[0085] For example, in pose visualization design, the pose data is first aligned with the chest keypoint as the center and standardized using shoulder width as the unit of coordinates. Next, using the cosine similarity of upper body keypoints between two frames as a distance metric, a Gaussian mixture model is used to classify the standardized pose data within a set scale into ten categories. Finally, the classified data is mapped to pose visualization elements, with the attributes of the pose visualization elements matching their corresponding pose data.

[0086] For example, in the stage position visualization design, the stage position of the video characters is normalized to a specified visualization space and displayed in the form of a scatter plot. Each scatter point corresponds to the stage position at a certain moment. The scatter points can be set with transparency to avoid occlusion and improve visibility.

[0087] This example extracts multimodal data from a presentation video. The following describes the algorithms and tools used in this example from the perspective of different modalities. The specific implementation of this invention is not limited to the algorithms and corresponding tools described.

[0088] (1) Facial emotion: face recognition and localization from video image frames, clustering the faces using DBSCAN algorithm (Reference: M. Ester, H.-P. Kriegel, J. Sander, and X. Xu. A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, KDD’96, p. 226-231. AAAI Press, 1996.) to find all the face images of the speaker appearing in the video, then using AffectNet (Reference: A. Mollahosseini, B. Hasani, and M. H. Mahoor. Affectnet: A database for facial expression, valence, and arousal computing in the wild. IEEE Trans. Affect. Comput., 10(1):18-31, Jan. 2019. doi:10.1109 / TAFFC.2017.2740923) to extract arousal and valence data from the face, using open source method (Reference: O. Arriaga, M. Valdenegro-Toro, and P. Ploger. Real-time convolutional neural networks for emotion and gender classification. arXiv preprint arXiv:1710.07557, 2017.) to recognize the emotion category of the face image;

[0089] (2) Facial key points, eye gaze vector: OpenFace, an open-source toolbox (Reference: T. Baltrusaitis, A. Zadeh, Y.C. Lim, and L.-P. Morency. Openface 2.0: Facial behavior analysis toolkit. In 2018 13th IEEE International Conference on Automatic Face Gesture Recognition (FG 2018), pp. 59-66, 2018.) is used to perform face recognition and localization from video image frames and extract facial key points and eye gaze vectors (Reference: E. Wood, T. Baltruaitis, X. Zhang, Y. Sugano, P. Robinson, and A. Bulling. Rendering of eyes for eye-shape registration and gaze estimation. In 2015 IEEE International Conference on Computer Vision (ICCV), pp. 3756-3764, 2015. doi: 10.1109 / ICCV.2015.428).

[0090] 3) Stage position, posture key points: First, the Faster R-CNN method in the open-source toolbox MMPose (Reference: M. Contributors. Openmmlab pose estimation toolbox and benchmark. https: / / github.com / open-mmlab / mmpose, 2022.) toolbox is used, and ResNet-50-FPN is used as the backbone network for human target detection. The stage position of the video character is represented by calculating the horizontal and vertical coordinates of the human target detection box center point relative to the target video interface. Then, the HRNet (Reference: K. Sun, B. Xiao, D. Liu, and J. Wang. Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp. 5693-5703, 2019.) pre-trained model trained on the COCO dataset is used to extract 2D pose skeleton data based on human target detection.

[0091] (3) Pitch and tone: The open-source Python library Parselmouth (Reference: Y. Jadoul, B. Thompson, and B. de Boer. Introducing Parselmouth: A Python interface to Praat. Journal of Phonetics, 71: 1-15, 2018.) is used to extract the tone and pitch of the video sound from the video.

[0092] As shown in Figure 2 , the speech video is used as the video data resource, and the facial emotion, eye gaze, body posture, and stage position of the video character are used as the basic data for visual element encoding.Figure 2 In (A), the above-mentioned basic data and the encoding relationship mapping table of visual elements, Figure 2 In (B), the color coding of the discrete emotional categories of the face, for example, a bluish-green color can represent a lower valence and a negative emotional bias, and a reddish-yellow color can represent a higher valence and a positive emotional bias. According to the set time region, the corresponding data is found in the basic data, and is processed respectively to obtain the discrete emotional category distribution of the facial emotion, the average of the valence and the arousal, the average of the eye gaze vector, the distribution of the stage position data, the key frame of the posture data, and the average of the pitch in the set time region. Then, the corresponding visual elements are generated according to the above-mentioned data and combined to obtain the video character summary. Figure 2 In (C), a schematic diagram of generating a video summary in different time scale intervals is shown. It can be seen that the video summary generation results in different scales.

[0093] As shown in Figure 3 The eye gaze, facial key points, and posture key points of the video character are used as the basic data, and the real-time and time period body language masks can be generated according to the set type. A1 is the bounding box of the video character, A3 is a face mask fitted with the facial key points, A2 is a ray mask of the eye gaze, the transparency of which gradually decreases along the ray direction, and A4 is a skeleton mask of the upper body of the video character, which is composed of the skeleton data of the previous 6 frames including the currently played video frame, and the transparency thereof decreases with the change of time. Meanwhile, hovering over the mask to call up the detail panel is supported, and B is the user hovering over the mask in different regions, which can trigger the corresponding data detail panel.

[0094] Another embodiment of the present application provides a character visualization device fusing multi-modal content of a video, which comprises:

[0095] A multi-modal feature data extraction module is configured to extract original data of a target video in each modality, align the original data in each modality according to the minimum scale of the original data extracted in each modality, and extract modal feature data in a set scale range according to different modal characteristics based on the aligned original data in each modality.

[0096] A video character visualization module is configured to perform character visualization in a visual form of a video summary based on the extracted modal feature data, perform character visualization in a visual form of a video player based on the extracted modal feature data, and perform video playing and interactive implementation of visual elements in a playing process based on the visual form of the video player.

[0097] The specific implementation process of each module is described above in the description of the method of the present application.

[0098] Another embodiment of the present application provides a computer device (computer, server, smart phone, etc.) comprising a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program comprising instructions for performing the steps of the method of the present application.

[0099] Another embodiment of the present application provides a computer readable storage medium (such as ROM / RAM, magnetic disk, optical disk) storing a computer program, the computer program being executed by a computer to implement the steps of the method of the present application.

[0100] The video character multi-modal visualization method and electronic device of the present application are described in detail above, but it is obvious that the specific implementation form of the present application is not limited thereto. Various obvious changes made by those skilled in the art without departing from the spirit of the method of the present application and the scope of the claims are within the protection scope of the present application.

Claims

1. A method for visualizing people by integrating multimodal video content, characterized in that, Includes the following steps: Extract raw data from the target video in each modality; Align the original data for each modality based on the minimum scale of the original data extracted from each modality; Based on the original data of each mode after alignment, modal feature data are extracted according to different modal characteristics within a set scale range; Based on the extracted modal feature data, we visualize people in a way that suits the visualization format of video summaries; Based on the extracted modal feature data, character visualization is performed for the visualization format of the video player; Based on the visual form of a video player, implement video playback and the interaction of visual elements during playback; The visualization of people in the form of video summaries includes the following steps: visualizing people's faces in the form of video summaries. Based on the extracted eye gaze data, within a set scale range, the direction and position of eye gaze in each frame are calculated and the arithmetic average is performed. Based on the extracted facial emotion data, within a set scale range, the arithmetic mean of the distribution, valence, and arousal of its discrete emotion types is calculated. Based on the extracted pitch data, the arithmetic mean of the pitches is calculated within a set scale range; Based on the arithmetic average of the gaze direction and position coordinates, calculate its polar coordinates based on the eye center point, and use these polar coordinates to generate eye elements in facial visualization. Different regions of a face can be used to map the distribution of different discrete emotion types. The degree of emotional arousal is mapped by the convexity of the curves around the face. Emotion-based efficacy reflects the positive or negative degree of emotion, and the arithmetic mean of facial emotional valence is mapped to the degree of vertical curvature of the mouth element in facial visualization. The arithmetic mean of the pitches is mapped to the width of the mouth element in the facial visualization.

2. The method according to claim 1, characterized in that, The modality includes at least one of video, image, and audio; the extraction of raw data of the target video in each modality includes the raw data of the extracted image modality including stage position data, eye gaze data, facial emotion data, and human body key point data, and the raw data of the extracted audio modality including the pitch and tone of the video.

3. The method according to claim 2, characterized in that, The following steps are used to extract the stage position data, eye gaze data, facial emotion data, and human body key point data: Extract stage position data: Perform human target detection and target tracking from each video image frame of the target video; calculate the horizontal and vertical coordinates of the center point of the human target detection box relative to the target video interface or the coordinate position relative to the stage's three-dimensional coordinate system; Eye gaze data extraction: Perform face recognition and localization from each video image frame of the target video; extract eye gaze data from each face image; Extracting facial emotion data: Performing face recognition and localization from each video image frame of the target video; using clustering or face recognition methods to filter out all facial images appearing in the target video; Arousal and valence data are extracted from each face image using facial emotion computing methods to obtain continuous emotion intensity data of facial expressions; emotion categories are identified for all face images to obtain discrete emotion category data of facial expressions. Extracting human key point data: Human target detection and target tracking are performed from each video image frame of the target video; Facial and body key point data are extracted based on individual human targets.

4. The method according to claim 1, characterized in that, Visualize human poses in the form of video summaries using the following steps: Based on the extracted pose data, key point alignment and coordinate standardization of the pose data are performed; Based on the aligned and normalized pose data, a fixed number of key pose frames are extracted. A fixed number of pose keyframes are visualized, and the attribute values ​​of the generated visualization elements correspond to their pose data. Furthermore, the stage positions of characters are visualized in the form of video summaries through the following steps: Based on the extracted stage position data, it is standardized according to the screen size of the target video; The standardized stage position data is visualized in the form of a scatter plot, with the attribute values ​​of the scatter plot elements corresponding to their stage position data.

5. The method according to claim 1, characterized in that, Visualize a person in the form of a video player using the following steps: Based on the extracted key point data of the person, the key points are aligned with the video frame by frame. Based on different time scale settings, the key point data of the characters in the corresponding continuous time segments are combined; By using masking visualization technology, key points of the person are mapped onto the mask, and the mask is overlaid on the corresponding video frame.

6. The method according to claim 1, characterized in that, To implement interactive functionality in the form of a video player, follow these steps: Based on the visual representation of characters in a video player, implement the play / pause function of the video player; Based on the visual elements of the character, interactive events corresponding to each area are implemented.

7. A device for visualizing people by fusing multimodal video content using the method described in any one of claims 1 to 6, characterized in that, include: The multimodal feature data extraction module is used to extract the raw data of the target video in each modality. Based on the minimum scale of the raw data extracted in each modality, the raw data in each modality is aligned. Based on the aligned raw data in each modality, modal feature data is extracted within a set scale range according to the characteristics of different modalities. The video character visualization module is used to visualize characters based on extracted modal feature data in a visual format for video summaries; to visualize characters based on extracted modal feature data in a visual format for video players; and to implement video playback and interactive visualization elements during playback based on the visual format of the video player.

8. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video content validity visual analysis method and system based on multi-modal emotion

    CN113743271A

  • System and method for quickly generating any 3D character animation based on video

    CN115690274A

  • 3D visual simulation character generation method and system

    CN115880441A

  • Method and device for detecting human-object interaction relationship in video

    CN112464875A

  • Multi-modal video emotion visualization method and device based on spiral and text

    CN113743267A