Front-end processing method of digital human video stream, equipment and medium

Through long link channels and cascaded layout video container technology, combined with edge computing optimization, the delay and adaptability problems in the front-end processing of digital human video streams are solved, achieving an efficient and seamless video streaming playback experience.

CN120358388AActive Publication Date: 2025-07-22HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510829890.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-22
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In the prior art, the front-end processing solution of digital human video streams has performance bottlenecks, especially in the problems of high network latency, large differences in equipment performance, poor adaptability of diversified terminals, and insufficient fault tolerance and dynamic adaptation, resulting in poor user experience.

Method used

By building a long link channel to receive segmented video stream data, two video containers with a cascade layout are loaded and played alternately, combined with edge computing nodes to optimize data compression and cache, dynamically adjust the segmentation points and playback parameters of segmented video streams to achieve seamless connection and resource optimization.

Benefits of technology

It significantly reduces transmission delay, eliminates playback gaps and lags, improves user interaction immersion, adapts to the smooth operation needs of terminals with different performance, and ensures the continuity of video streams and efficient playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358388A_ABST
    Figure CN120358388A_ABST
Patent Text Reader

Abstract

The invention discloses a front-end processing method and device for a digital human video stream, and a medium, and relates to the technical field of video stream processing. The method comprises the following steps: based on a pre-constructed long link channel, receiving segmented digital human video stream data generated by segmentation of a server, and caching the segmented digital human video stream data to a front end queue according to a transmission sequence; monitoring the playing progress of the current playing video container, and when the playing progress reaches a preset playing progress threshold value, extracting the digital human video stream data of the next segment from the front end queue and loading the digital human video stream data to another video container; and when it is monitored that the playing of the current playing video container is ended, switching to another video container, and triggering the other video container to play the pre-loaded next segment of digital human video stream data, so as to realize seamless connection of the segment of digital human video stream data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video stream processing, and in particular, to a front-end processing method, device, and medium for digital human video streams. Background Art

[0002] With the wide application of digital human technology in fields such as live streaming, intelligent customer service, and virtual conferences, the real-time generation and efficient playback of digital human video streams have become the core links to improve the user experience. In the existing technology, digital human video streams usually rely on the centralized processing and transmission of the backend server, resulting in significant performance bottlenecks in the receiving, assembling, and playback processes of the front-end devices. The backend centralized processing mode needs to transmit the complete video stream data to the front-end via the network, and the high-latency problem is particularly prominent in real-time interaction scenarios, making it difficult to meet the user's demand for instant feedback. At the same time, the transmission of complex video stream data requires extremely high network bandwidth stability, and problems such as stuttering and data loss are likely to occur when the network fluctuates or the device performance is insufficient, seriously affecting visual coherence.

[0003] Currently, most front-end processing solutions adopt a single video container sequential loading and playback mechanism, lacking support for preloading and dynamic switching of segmented data. During playback, the front-end needs to wait for the current video segment to be completely loaded and played before starting to load the next segment, resulting in obvious playback gaps and a jerky transition. In addition, the existing technology does not fully consider the impact of device performance differences on rendering efficiency. Low-end devices are prone to resource overload when parsing high-complexity video streams, while high-end devices cannot fully utilize their hardware advantages to achieve better performance. This processing method is difficult to adapt to diverse terminal environments and limits the large-scale application of digital human technology.

[0004] Furthermore, the existing methods have significant deficiencies in terms of fault tolerance and dynamic adaptation. When the video stream loading fails or the playback is interrupted, the traditional solutions need to reload the entire data queue, resulting in a break in the user experience. At the same time, the fixed-layout video container cannot adapt to diverse display scenarios such as mobile vertical screens and split screens, causing waste of the display area or picture occlusion. Although edge computing technology has been partially introduced to relieve the backend pressure, its cooperation mechanism with the front-end is not yet mature, and the caching strategy and dynamic resource allocation lack intelligent design, making it difficult to effectively reduce the transmission delay and improve the system robustness. Summary of the Invention

[0005] Embodiments of this application provide a front-end processing method, device, and medium for digital human video streams to solve the above technical problems.

[0006] On the one hand, embodiments of this application provide a front-end processing method for digital human video streams, including: Based on a pre - constructed long - link channel, receive the segmented digital human video stream data generated by the server through segmentation, and cache the segmented digital human video stream data into the front - end queue in the transmission order. Through two video containers with a stacked layout, alternately load and play the segmented digital human video stream data in the front - end queue. Monitor the playback progress of the currently playing video container. When the playback progress reaches the preset playback progress threshold, extract the next segmented digital human video stream data from the front - end queue and load it into the other video container. When it is monitored that the playback of the currently playing video container ends, switch to the other video container, and trigger the other video container to play the pre - loaded next segmented digital human video stream data to achieve seamless connection of the segmented digital human video stream data.

[0007] In an implementation manner of the present application, before receiving the segmented digital human video stream data generated by the server through segmentation based on a pre - constructed long - link channel, the method further includes: Receive, through the server, an intelligent question - answering request initiated by the user based on the front - end device, generate a corresponding answer text, perform semantic analysis on the answer text, and identify key semantic nodes with context association. Real - time monitor the network bandwidth of the front - end device, and dynamically adjust the maximum duration threshold of the segmented digital human video stream data according to the monitoring result. If the network bandwidth is lower than the preset threshold, shorten the length of the segmented digital human video stream data according to a preset ratio to reduce the amount of data transmitted each time. According to the key semantic nodes and in combination with the adjusted maximum duration threshold, dynamically adjust the segmentation points of the digital human video stream so that each segmented digital human video stream data contains a complete semantic unit. According to the segmentation points, split the answer text into multiple statements, generate video stream segments corresponding to each statement, and associate action parameters and expression parameters of the digital human with each video stream segment to obtain segmented digital human video stream data.

[0008] In an implementation manner of the present application, perform sentiment analysis on each statement at the server, identify the emotion type corresponding to the statement, and generate an associated emotion feature marker for the segmented digital human video stream data corresponding to the statement. Parse the emotion feature markers of the segmented digital human video stream data on the front - end device, and based on the marker types in the parsing results and according to the predefined playback parameter mapping rules, adjust the video playback speed and video playback volume corresponding to the segmented digital human video stream data. During the video playback process, apply the adjusted video playback speed and video playback volume in real - time to ensure that the emotional expression of the played video is synchronized with the video content.

[0009] In an implementation of the present application, two video containers in a stacked layout are used to alternately load and play the segmented digital human video stream data in the front-end queue, specifically including: Control the display hierarchy of the two video playback containers through the z-index property. When the current playback video container loads and plays the current segmented digital human video stream data in the front-end queue, set the display hierarchy of the current playback video container to the visible state and set the display hierarchy of the other video container to the hidden state; After the other video container completes the preloading of the next segmented digital human video stream data in the front-end queue, set the display hierarchy of the current playback video container to the hidden state; When activating the other video container, synchronously update the display hierarchy of the other video container to the visible state and trigger the playback instruction of the other video container.

[0010] In an implementation of the present application, after caching the segmented digital human video stream data into the front-end queue in the transmission order, the method further includes: Dynamically compress the action parameters and expression parameters of the digital human in the segmented digital human video stream data on the edge computing node according to the hardware performance parameters of the front-end device; Bind the optimized action parameters and expression parameters to the segmented digital human video stream data to generate an optimized video stream segment adapted to the hardware performance parameters of the front-end device; Through the edge computing node, locally cache and store the preset video stream segments that are frequently used, and generate a temporary access link, so that the front-end device preferentially pulls the video stream segments from the local cache of the edge computing node. If not hit, a data acquisition request is sent to the server; Through the collaborative caching mechanism between the edge computing node and the front-end device, and according to the network bandwidth fluctuation, dynamically adjust the distribution ratio of the segmented digital human video stream data between the edge computing node and the front-end device to reduce the network transmission delay.

[0011] In an implementation of the present application, it further includes: When it is detected that the current playback video container fails to load or the playback is interrupted, switch to the other video container and try to reload the same segmented digital human video stream data; If the other video container also fails to load, send a regeneration request to the server to obtain an alternative version of the segmented digital human video stream data; After the successful loading of the alternative version, insert the alternative version into the original position of the digital human video stream data of the faulty segment in the front-end queue and continue playing, and mark the digital human video stream data of the faulty segment as low priority to delay the retry loading of the digital human video stream data of the faulty segment.

[0012] In one implementation manner of the present application, it further includes: When the front-end device plays the segmented digital human video stream data, the gaze area information of the digital human image in the current playing video container is obtained in real time; the gaze area information is collected by an eye movement tracking device pre-deployed in the front-end device; According to the gaze area information, it is judged whether the user continuously gazes at a specific part of the digital human image for more than a preset gaze duration threshold. If so, pause the video playback of the current playing video container, and send the user's gaze behavior information and the playback progress information of the currently playing segmented digital human video stream data to the server; Through the server, analyze the currently playing segmented digital human video stream data, extract the detailed information content related to the specific part, and generate a corresponding extended explanation video stream segment according to the detailed information content, so as to send the extended explanation video stream segment and the corresponding playback insertion point information to the front-end device; Through the front-end device, in the current playing video container, insert and play the extended explanation video stream segment starting from the playback progress position corresponding to the playback insertion point information, and after the extended explanation video stream segment is played, continue to play the remaining unplayed part in the segmented digital human video stream data.

[0013] In one implementation manner of the present application, it further includes: Monitor the screen direction change of the front-end device or the split-screen operation triggered by the user. If the front-end device switches to the portrait mode, adjust the two video containers in the stacked layout to a side-by-side layout, and dynamically adjust the size ratio of the two video containers; In the split-screen mode, hide the other video container in the two video containers, pause the preloading operation of the other video container, and keep the currently playing video container for playing to adapt to the display area.

[0014] On the other hand, an embodiment of the present application further provides a front-end processing device for a digital human video stream, and the device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a front-end processing method for a digital human video stream as described above.

[0015] On the other hand, an embodiment of the present application also provides a non-volatile computer storage medium storing computer-executable instructions, which when executed, implement a front-end processing method for a digital human video stream as described above.

[0016] The embodiments of the present application provide a front-end processing method, device, and medium for a digital human video stream, including at least the following beneficial effects: Receiving the segmented video stream data generated by the server-side segmentation through a pre-constructed long link channel, avoiding the frequent handshake overhead caused by traditional polling or short connections, and significantly shortening the transmission time of data from the server-side to the front-end. The segmented data is cached in order in the front-end queue. Combining the dynamic monitoring of the playback progress and the preloading mechanism, the video stream can reach the front-end quickly with smaller data units and complete rendering, effectively reducing the playback delay perceived by users, especially suitable for real-time interaction scenarios. Using two video containers in a stacked layout to alternately load and play the segmented data, and by preloading the next video segment and instantaneously switching the container level when the playback is completed, completely eliminating the loading gap and picture stuttering problems in the traditional single-container playback mode. The dual-container cooperation mechanism not only saves the cache resource occupancy of the front-end device, but also optimizes the utilization rate of hardware resources through parallel loading and playback, adapting to the smooth operation requirements of different performance terminals. By listening to the playback end event and automatically triggering the container switch, it ensures that multiple video streams present a continuous and seamless playback effect from the user's perspective, enhancing the interactive immersion. Description of the Drawings

[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic diagram of an application scenario of a front-end processing method for a digital human video stream provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a front-end processing method for a digital human video stream provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of a collaborative caching method for an edge computing node and a front-end device provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of a front-end processing device for a digital human video stream provided by an embodiment of the present application; Figure 5Schematic diagram of the internal structure of a front-end processing device for a digital human video stream provided by an embodiment of the present application. Detailed implementation manners

[0018] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0019] The following will describe in detail the technical solutions provided in each embodiment of the present application with reference to the drawings.

[0020] A front-end processing method for a digital human video stream provided by an embodiment of the present application can be applied to an application environment as Figure 1 shown. As Figure 1 shown, the application environment may include: a front-end device 101 carried by a user, a communication network 102, a server 103, and a database 104. The communication network 102 can serve as a data transmission channel to provide a communication link for communication between the front-end device 101 and the server 103. After the server 103 segments the digital human video stream to generate segmented digital human video stream data, the segmented digital human video stream data is transmitted to the front-end device 101 carried by the user based on the communication network 102 and stored in the front-end queue of the front-end device 101.

[0021] Among them, the database 104 is connected to the server 103 and can be used to store the original digital human video stream and the segmented digital human video stream data generated after segmentation. The database 106 can be integrated on the server, or placed on the cloud or other network servers.

[0022] Figure 2 Schematic flowchart of a front-end processing method for a digital human video stream provided by an embodiment of the present application.

[0023] The implementation of the analysis method involved in the embodiments of the present application can be a terminal device or a server, and the present application does not make special restrictions on this. For the convenience of understanding and description, the following embodiments will be described in detail using a server as an example.

[0024] It should be noted that the server can be a single device or a system composed of multiple devices, that is, a distributed server, and the present application does not make specific limitations on this.

[0025] As Figure 2 shown, a front-end processing method for a digital human video stream provided by an embodiment of the present application includes: Step 201: Based on the pre - constructed long - link channel, receive the segmented digital human video stream data generated by the server - side segmentation, and cache the segmented digital human video stream data into the front - end queue in the transmission order.

[0026] It should be noted that the digital human video stream in the embodiments of this application (Digital Human VideoStream) refers to the video data stream generated by a virtual character (digital human) generated by a computer in real - time or offline.

[0027] In this embodiment, the long - link channel is implemented through the WebSocket protocol, establishing a persistent two - way communication connection with the server - side, avoiding the additional overhead caused by the frequent handshakes of traditional HTTP short - connections. It can be understood that the persistence of the long - link channel enables the server - side to push the segmented digital human video stream data to the front - end device in real - time, significantly reducing the transmission delay. It should be noted that when the server - side generates the segmented digital human video stream data, it performs semantic analysis on the answer text input by the user, identifies the key semantic nodes related to the context, such as the end of an interrogative sentence, emotional turning points, and segments the video stream based on a dynamically adjusted maximum duration threshold, where the maximum duration threshold is determined according to the real - time monitored front - end network bandwidth. For example, if the network bandwidth is low, the server - side will shorten the duration of a single - segment video stream to ensure that the data volume of each segment adapts to the current transmission conditions. If the network bandwidth is good, the server - side will appropriately increase the duration of a single - segment video stream, thereby increasing the content transmitted in a single - segment video stream.

[0028] Specifically, the segmented digital human video stream data includes video frames, audio streams, digital human action parameters, and expression parameters. The action parameters are such as gesture trajectories, and the expression parameters are such as the amplitude of a smile. After receiving this segmented digital human video stream data through the long - link channel, the front - end device stores the segmented digital human video stream data into the front - end queue in the transmission order using the first - in - first - out (FIFO) strategy. Exemplarily, the front - end queue is managed using a circular buffer structure. When the queue capacity reaches the preset upper limit, the earliest received segment is automatically removed to release memory resources.

[0029] In this embodiment, when the server - side receives a request initiated by the front - end triggered by the user, such as an intelligent chat question - and - answer, asks a question, processes the question to get an answer to the question, splits the answer to the question into multiple sentences according to the sentence - segmentation logic, and then the algorithm generates specific actions and facial forms for the digital human for each sentence and pays attention to the continuity of the actions with the previous sentence, and pushes the generated video stream to the front - end client in the form of a long - link (WebSocket).

[0030] The front end stores the segmented digital human video stream data pushed by the server through the long link channel in the form of a queue. When it is detected that there is only one piece of data in the front end queue, the data can be filled into the encapsulated video container component and loaded and played.

[0031] Figure 3 A flow chart of a collaborative caching method between an edge computing node and a front-end device provided in an embodiment of the present application. Figure 3 , 301. On the edge computing node, according to the hardware performance parameters of the front-end device, the action parameters and expression parameters of the digital human in the segmented digital human video stream data are dynamically compressed.

[0032] In this embodiment, the deep integration of data compression, cache management and dynamic distribution is achieved through the collaborative optimization mechanism of edge computing nodes and front-end devices. It can be understood that edge computing nodes refer to computing units deployed at the edge of the network, such as CDN nodes or local servers close to front-end devices, which have real-time processing and caching capabilities. After the segmented digital human video stream data is cached to the front-end queue in the order of transmission, the edge computing node dynamically adjusts the data compression strategy according to the hardware performance parameters of the front-end device, such as CPU computing power, GPU rendering capability, and memory capacity. Exemplarily, if the front-end device is a low-performance mobile terminal, the edge computing node will downsample the action parameters and expression parameters of the digital human to reduce the amount of data to adapt to the device processing capability, action parameters such as limb movement trajectory, gesture frequency, and expression parameters such as facial muscle movement amplitude. If it is a high-performance PC terminal, the original parameters are retained to ensure rendering accuracy. It should be noted that the hardware performance parameters are actively reported by the front-end device during initialization, or dynamically obtained by the edge computing node through probe requests, such as WebRTC performance detection.

[0033] 302. Bind the compressed action parameters and expression parameters with the segmented digital human video stream data to generate optimized video stream segments adapted to the hardware performance parameters of the front-end device.

[0034] In this embodiment, the edge computing node rebinds the compressed action parameters and expression parameters to the original video stream segment to generate an optimized video stream segment. For example, for a video stream containing a digital human waving, the edge computing node can only retain the hand trajectory data of the key frames and remove redundant intermediate frames, thereby reducing the computing resources required for rendering. It should be noted that the optimized video stream segmentation format must be compatible with the decoder of the front-end device, such as H.264 encoding, to avoid additional conversion overhead. This process is coordinated with the dynamic segmentation logic of the server. The server divides the video stream according to semantic units, and the edge computing node is further optimized according to device performance to form a two-level adaptation mechanism.

[0035] 303. Through the edge computing node, the locally cached preset video stream segments frequently used are stored, and a temporary access link is generated, so that the front-end device preferentially pulls the video stream segments from the local cache of the edge computing node. If not hit, a data acquisition request is sent to the server.

[0036] In this embodiment, the edge computing node analyzes historical access data to identify frequently used video stream segments, such as the digital human action segments corresponding to common questions and answers, and stores them in the local cache. Exemplarily, for the product function introduction answers frequently appearing in the intelligent customer service scenario, the edge computing node prestores the corresponding video stream segments. It should be noted that the cache policy adopts the Least Recently Used (LRU) algorithm for management, and low-frequency data is automatically eliminated when the cache space is insufficient. For the cached segments, the edge computing node generates a temporary access link, such as a CDN URL with a time-limited signature. The front-end device directly pulls the data through this link, avoiding repeated requests to the server. If not hit, the front-end device falls back to the server to obtain the original data, and at the same time the edge computing node records the access frequency of this segment to update the cache policy.

[0037] 304. Through the collaborative caching mechanism between the edge computing node and the front-end device, and according to the network bandwidth fluctuation, the data distribution ratio of the segmented digital human video stream between the edge computing node and the front-end device is dynamically adjusted to reduce the network transmission delay.

[0038] In this embodiment, the collaborative caching mechanism means that the edge computing node and the front-end device dynamically adjust the data source ratio according to the real-time network conditions. For example, when the network bandwidth is low, the edge computing node increases the push ratio of the locally cached segments and reduces the amount of data pulled from the server; when the network bandwidth is sufficient, the front-end device is allowed to mix and pull the edge cache and server data to ensure the freshness of the content. Specifically, the edge computing node dynamically calculates the distribution ratio by monitoring the network bandwidth fluctuation, such as based on the TCP congestion window size or the throughput reported by the front-end. Exemplarily, if it is detected that the network delay exceeds the threshold, the edge computing node will preferentially distribute the cached lightweight segments and pause the transmission of non-critical data, such as high-precision expression parameters.

[0039] Step 202. Monitor the playback progress of the current playback video container. When the playback progress reaches the preset playback progress threshold, extract the next segmented digital human video stream data from the front-end queue and load it into another video container.

[0040] The front-end device uses two video containers for layered layout. Taking the data in the front-end queue as the data source, the two video containers are respectively assigned values. For example, the data of the first segmented digital human video stream in the queue is for the upper-level video container in the layout, and the data of the second segmented digital human video stream is for the lower-level video container in the layout.

[0041] In this embodiment, after listening to the playback progress of the currently playing video container, the display hierarchy of the two video playback containers is controlled through the z-index property. Exemplarily, in the initial state, the z-index value of container A is 2, which is the visible state, and the z-index value of container B is 1, which is the hidden state. It can be understood that the design of the layered layout makes the two containers visually overlap in the same display area, and the user only perceives the playback screen of the top-level video container. It should be noted that when container A plays the current segmented digital human video stream data, container B preloads the next segmented digital human video stream data, and after container A finishes playing, the display hierarchy of container A is set to the hidden state, and the display hierarchy of container B is promoted to the visible state, thereby triggering container B to play the preloaded next segmented digital human video stream data.

[0042] In this embodiment, if the emotional feature is used as one of the indicators for generating the digital human video stream, it will affect the generation speed of the digital human video stream. However, if there is no emotion in the voice, it will reduce the user experience. To reduce this impact, after the digital human video stream is generated, it is segmented into multiple sentences according to the sentence-breaking logic based on the question answer, and each sentence corresponds to a video stream. At the same time, the emotional feature of each sentence is judged and marked on the corresponding video stream.

[0043] After receiving the video stream, the front-end device identifies the emotional feature mark corresponding to the video stream, and each emotional feature mark corresponds to a video playback speed and a video playback volume.

[0044] For example: when happy, the video playback speed and video playback volume of the video stream will be moderately increased; when sad, the video playback speed and video playback volume of the video stream will be moderately decreased; when there is no emotional change, the video playback speed and video playback volume of the video stream will be played normally.

[0045] In this embodiment, when container A loads and plays the first segmented digital human video stream data in the front-end queue, container B is in a standby state. The resource loading situation is monitored through the loadedmetadata event of the video, and the playback progress of the video is monitored through timeupdate to dynamically and alternately switch between the upper and lower video containers. When the playback progress of container A triggers the preloading condition, container B starts to load the next segmented digital human video stream data. Exemplarily, in a digital human live broadcast scenario, when container A plays the digital human explanation video corresponding to the current statement, container B has pre-loaded the action data of the next statement, such as gesture changes. It should be noted that the alternating loading mechanism of the dual containers collaborates with the fault tolerance strategy. If a certain segment fails to load, the system can immediately switch to the standby container to attempt reloading to avoid playback interruption.

[0046] In this embodiment, after a video stream data is loaded into the video container, this data will be mounted to the video container with the topmost z-index level. When the loaded callback is triggered, the play of the current playing video container is executed, and the playback progress is monitored through onTimeupdate. When the playback progress reaches 90%, it is detected whether there is subsequent data in the data queue. If there is no subsequent data, the playback ends or waits for the corresponding form. If there is subsequent data, the data is sequentially fetched and filled into the second video player for loading. It should be noted that if the resource of a single video stream is large, this threshold can be appropriately reduced.

[0047] In this embodiment, the detection of loading failure or playback interruption is achieved through the error event and stalled event of the front-end video container. Exemplarily, when the video container fails to load segmented data due to network packet loss or data corruption, the error event will be triggered; if the playback freezes due to insufficient buffering for more than the preset duration, it is marked as an interrupted state through a custom timer. It should be noted that once a failure of the current playing container (such as container A) is detected, the system immediately elevates the z-index level of the other video container (such as container B) to the visible state and attempts to reload the same segmented data. Specifically, when reloading, the optimized lightweight version is preferentially pulled from the edge computing node. If the edge cache misses, it falls back to the server to request the original data.

[0048] If the container B still fails to load the same segmented data, the front-end device sends a regeneration request to the server. This request carries the unique identifier of the faulty segment and the cause of the failure. The unique identifier can be, for example, the segment ID or the timestamp, and the cause of the failure can be, for example, network timeout or data verification failure. After receiving the request, the server generates an alternative version according to the type of the failure. Exemplarily, for a loading failure caused by insufficient bandwidth, the server can generate a low-resolution version or a simplified video stream that only retains the key action frames; if it is due to data corruption, the segment is re-rendered and a redundant check code is attached. It should be noted that the server quickly reconstructs the content based on the original semantic nodes to ensure that the alternative version is semantically coherent with the context.

[0049] After the alternative version is successfully loaded, the front-end device inserts it into the original position of the faulty segment in the front-end queue and immediately triggers playback. Exemplarily, in a virtual meeting scenario, if a segment of the digital human's answer fails to load due to network jitter, the alternative version will play seamlessly, avoiding interruption of the conversation. The alternative version can be, for example, a simplified video that only contains basic lip-sync animations. At the same time, the original faulty segment is deleted.

[0050] Step 203: When it is detected that the playback of the current video container ends, switch to another video container and trigger the other video container to play the pre-loaded next segmented digital human video stream data to achieve seamless connection of the segmented digital human video stream data.

[0051] This application listens for the video playback completion node through the ended event, and controls playback and stop through the play event and the stop event respectively.

[0052] In this embodiment, after the ended event is triggered when the playback of the current video container ends, the play event of another video container is immediately triggered, and the z-index values of the two video containers are swapped to achieve the switching of the display levels of the two video containers. When the playback of another video container reaches 90% of the threshold, the data in the front-end queue is detected to fill the first video container, and the corresponding two levels are switched, and so on, to achieve the orderly connection and playback of multiple video streams in the queue. Finally, the seamless connection of multiple video streams in the front-end vision is achieved, just like playing a single video, and the performance is better.

[0053] In this embodiment, the eye tracking device refers to an infrared camera or a dedicated sensor module integrated in the front-end device, such as AR glasses and high-end mobile terminals, which calculates the projection area of the line of sight on the screen in real time by capturing the user's eye movement trajectory and pupil focus position. For example, in the digital human teaching scene, when the user looks at the gesture area of the digital human in the video container, the eye tracking device samples the line of sight coordinates at a millisecond frequency and converts it into the gaze area information relative to the digital human image through a coordinate mapping algorithm, such as the left hand holding area and facial expression area. It should be noted that this information is encapsulated in JSON format, including fields such as area boundary coordinates, gaze start timestamp and duration.

[0054] The front-end device determines whether the user continues to gaze at a specific part for more than a preset time threshold based on the gaze area information. Specifically, the system maintains a sliding time window, such as the line of sight data in the last few seconds. If the cumulative gaze time in the same area exceeds the threshold, the interaction logic is triggered. For example, in a virtual shopping guide scenario, if the user continues to gaze at the product held by the digital person for more than the threshold, the system determines that the user needs to obtain detailed information about the product. At this time, the front-end device immediately pauses the video stream of the current playback container, and sends the user's gaze behavior information and the progress information of the current playback segment, such as timestamp and segment ID, to the server based on the long link channel. It should be noted that when playback is paused, the container level remains visible, but the video frame is frozen to maintain a static display of the picture.

[0055] It is understandable that after receiving the gaze behavior information, the server performs a multimodal analysis on the currently playing segmented digital human video stream data. For example, if the user's gaze area is the drug packaging held by the digital human, the server extracts the text information on the packaging, such as ingredients and usage, through image recognition, and combines voice recognition technology to parse the drug-related explanation content in the current segment. Subsequently, the server generates an extended explanation video stream segment, which may include a rotating display of a 3D model of the drug, a detailed instruction manual pop-up window, or supplementary explanation audio. It should be noted that the server ensures that the extended segment is semantically coherent with the original content based on the context of the original answer text, such as inserting a side effect description after the drug introduction.

[0056] Specifically, the server returns the generated extended interpretation video stream segment and playback insertion point information to the front-end device, such as the start timestamp of the "drug display" node in the original segment. The front-end device inserts the extended segment at the corresponding position of the current playback container to form a mixed playback stream. For example, in an online education scenario, when the user looks at the formula derivation area explained by the digital human, the extended segment displays a step-by-step calculation animation of the formula in a picture-in-picture format. After the extended segment is played, the system automatically continues to play the remaining content from the insertion point of the original segment. It should be noted that the extended segment can be directly superimposed and played in the current container without enabling a backup container to avoid layout conflicts.

[0057] In this embodiment, the screen orientation change is detected by a gyroscope or an orientation sensor of the front-end device, and the generated event (such as orientationchange) triggers the layout adjustment logic. Exemplarily, when the mobile device switches from landscape to portrait, the system obtains the screen aspect ratio change information in real time. The split-screen operation refers to a multi-window display mode actively triggered by the user (such as the Android split-screen function), and the front-end device listens for the split-screen state through the operating system API (such as WindowManager). It should be noted that this listening mechanism runs independently of the eye-tracking logic, but shares the same event handling thread to avoid resource conflicts.

[0058] Specifically, when it is detected that the device switches to the portrait mode, the system adjusts the two video containers in the stacked layout to a side-by-side layout. Exemplarily, the original stacked containers A (z-index: 2) and B (z-index: 1) are repositioned to be vertically arranged: container A occupies the upper part of the screen (such as 60% of the height), and container B occupies the lower part (such as 40% of the height), and both of their widths are adapted to the screen width. It can be understood that when dynamically adjusting the size ratio, the system calculates the container size based on the available screen area to ensure that the display ratio of the digital human image is not distorted. It should be noted that the hierarchical control logic fails temporarily in this scenario because both containers are in the visible state, but the preloading mechanism still operates according to the original rules, that is, container B continues to load the subsequent segmented data to prepare for the possible landscape restoration.

[0059] In this embodiment, if the user triggers a split-screen operation, such as shrinking the application window to half screen, the system immediately hides the standby video container (such as container B) and pauses its preloading task to release computing resources. Exemplarily, in the split-screen mode, the size of the current playing container (container A) is dynamically compressed to the split-screen area, such as 50% of the screen width, and redundant UI elements (such as control buttons) are removed to maximize the display area. It should be noted that after hiding container B, the unloaded segmented data in the front-end queue is still cached in order, but the preloading is only resumed when the split-screen exits. In addition, the edge computing node preferentially pushes lightweight segments in this mode to ensure low resource consumption in the split-screen scenario.

[0060] Figure 4 It is a schematic structural diagram of a front-end processing device for a digital human video stream provided by an embodiment of the present application. As Figure 4As shown in the figure, a front-end processing device for a digital human video stream disclosed in this application includes a server module, a WebSocket module, and a front-end module. The server module includes a video stream segmentation generation mechanism, so that on the WebSocket server, the digital human video stream is segmented through the video stream segmentation generation mechanism. The WebSocket module includes a long connection channel, specifically, the long connection channel is implemented through the WebSocket protocol, and then through the long connection channel, the segmented digital human video stream data is pushed to the front-end device in real time. The front-end module includes a front-end queue for caching the received segmented digital human video stream data and pushing it to the WebSocket client.

[0061] It should be noted that the front-end device in this application uses two video containers for layered layout, and uses a time event listener to monitor the playback progress of the currently playing video container. Then, according to the playback progress, the segmented digital human video stream data in the front-end queue is alternately played, and a natural video stream is rendered in real time to improve the user's viewing experience.

[0062] In some of the embodiments, before receiving the segmented digital human video stream data generated by the server segmentation based on the pre-constructed long connection channel, it further includes: Receiving, by the server, an intelligent question-and-answer request initiated by the user based on the front-end device, generating a corresponding answer text, and performing semantic analysis on the answer text to identify key semantic nodes associated with the context; Real-time monitoring of the network bandwidth of the front-end device, and dynamically adjusting the maximum duration threshold of the segmented digital human video stream data according to the monitoring results; Dynamically adjusting the segmentation point of the digital human video stream according to the key semantic nodes and the adjusted maximum duration threshold, so that each segmented digital human video stream data contains a complete semantic unit; Cutting the answer text into multiple sentences according to the adjusted segmentation point, generating a video stream segment corresponding to each sentence, and associating action parameters and expression parameters of the digital human with each video stream segment to obtain segmented digital human video stream data.

[0063] In some of the embodiments, it further includes: Performing sentiment analysis on each sentence on the server, identifying the emotion type corresponding to the sentence, and generating an associated emotion feature mark for the segmented digital human video stream data corresponding to the sentence; Parsing the emotion feature marks of the segmented digital human video stream data on the front-end device, and adjusting the video playback speed and video playback volume corresponding to the segmented digital human video stream data according to the mark type in the parsing result and based on the predefined playback parameter mapping rules; During the video playback process, the adjusted video playback speed and video playback volume are applied in real time to ensure that the emotional expression of the played video is synchronized with the video content.

[0064] In some of these embodiments, after monitoring the playback progress of the currently playing video container, it further includes: Controlling the display levels of two video playback containers through the z-index property, so that when the currently playing video container loads and plays the current segmented digital human video stream data in the front-end queue, the display level of the currently playing video container is set to the visible state, and the display level of the other video container is set to the hidden state; After the currently playing video container finishes playing the current segmented digital human video stream data, the display level of the currently playing video container is set to the hidden state; When activating the other video container, the display level of the other video container is synchronously updated to the visible state, and a playback instruction for the other video container is triggered.

[0065] In some of these embodiments, after caching the segmented digital human video stream data into the front-end queue in the transmission order, it further includes: Dynamically compressing the action parameters and expression parameters of the digital human in the segmented digital human video stream data on the edge computing node according to the hardware performance parameters of the front-end device; Binding the compressed action parameters and expression parameters to the segmented digital human video stream data to generate an optimized video stream segment adapted to the hardware performance parameters of the front-end device; Through the edge computing node, locally caching and storing the preset video stream segments with high frequency usage, and generating a temporary access link, so that the front-end device preferentially pulls the video stream segment from the local cache of the edge computing node. If not hit, a data acquisition request is sent to the server; Through the collaborative caching mechanism between the edge computing node and the front-end device, and according to the network bandwidth fluctuation, dynamically adjusting the distribution ratio of the segmented digital human video stream data between the edge computing node and the front-end device to reduce the network transmission delay.

[0066] In some of these embodiments, it further includes: In the case of detecting that the currently playing video container fails to load or the playback is interrupted, switch to the other video container and attempt to reload the same segmented digital human video stream data; If the other video container also fails to load, send a regeneration request to the server to obtain an alternative version of the segmented digital human video stream data; After the alternative version is successfully loaded, insert the alternative version into the original position of the faulty segmented digital human video stream data in the front-end queue and continue playing, and delete the faulty segmented digital human video stream data.

[0067] In some of these embodiments, it further includes: When the front-end device plays the segmented digital human video stream data, it obtains in real time the information on the area of the user's gaze on the digital human image in the current video container being played; According to the gaze area information, it determines whether the user continuously gazes at a specific part of the digital human image for more than a preset gaze duration threshold. If so, it pauses the video playback of the current video container being played and sends the user's gaze behavior information and the playback progress information of the currently played segmented digital human video stream data to the server; Through the server, it analyzes the currently played segmented digital human video stream data, extracts the detailed information content related to the specific part, and generates a corresponding extended explanation video stream segment according to the detailed information content, so as to send the extended explanation video stream segment and the corresponding playback insertion point information to the front-end device; Through the front-end device, in the current video container being played, it inserts and plays the extended explanation video stream segment starting from the playback progress position corresponding to the playback insertion point information, and after the extended explanation video stream segment is played, it continues to play the remaining unplayed part of the segmented digital human video stream data.

[0068] In some of these embodiments, it further includes: It monitors the change of the screen orientation of the front-end device or the split-screen operation triggered by the user. If the front-end device switches to the portrait mode, it adjusts the two video containers in a stacked layout to a side-by-side layout and dynamically adjusts the size ratio of the two video containers; In the split-screen mode, it hides the other video container in the two video containers, pauses the preloading operation of the other video container, and retains the current video container being played to adapt to the display area.

[0069] The above are the method embodiments proposed in this application. Based on the same inventive concept, the embodiments of this application also provide a front-end processing device for a digital human video stream, and its structure is as Figure 5 shown.

[0070] Figure 5 This is the internal structure schematic diagram of a front-end processing device for a digital human video stream provided by the embodiments of this application. As Figure 5 shown, the device includes: At least one processor; And a memory communicatively connected to at least one processor; Wherein, the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that at least one processor can: Receive the segmented digital human video stream data generated by the server segmentation based on a pre-constructed long link channel, and cache the segmented digital human video stream data into the front-end queue in the transmission order. Monitor the playback progress of the current playback video container. When the playback progress reaches the preset playback progress threshold, extract the next segmented digital human video stream data from the front-end queue and load it into another video container. When it is monitored that the playback of the current playback video container ends, switch to another video container, and trigger the other video container to play the pre-loaded next segmented digital human video stream data to achieve seamless connection of the segmented digital human video stream data.

[0071] The embodiment of the present application also provides a non-volatile computer storage medium, storing computer-executable instructions, which can be executed to: Receive the segmented digital human video stream data generated by the server segmentation based on a pre-constructed long link channel, and cache the segmented digital human video stream data into the front-end queue in the transmission order. Monitor the playback progress of the current playback video container. When the playback progress reaches the preset playback progress threshold, extract the next segmented digital human video stream data from the front-end queue and load it into another video container. When it is monitored that the playback of the current playback video container ends, switch to another video container, and trigger the other video container to play the pre-loaded next segmented digital human video stream data to achieve seamless connection of the segmented digital human video stream data.

[0072] Each embodiment in the present application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0073] The device and medium provided by the embodiment of the present application correspond one-to-one with the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.

[0074] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0075] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0076] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0078] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0079] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0080] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0081] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

Claims

1. A front-end processing method for a digital human video stream, characterized in that, The method includes: Based on a pre-constructed long link channel, receiving the segmented digital human video stream data generated by the server through segmentation, and caching the segmented digital human video stream data into the front-end queue in the transmission order. Listening to the playback progress of the current playback video container, and when the playback progress reaches the preset playback progress threshold, extracting the next segmented digital human video stream data from the front-end queue and loading it into another video container. When it is monitored that the playback of the current playback video container ends, switching to the other video container, and triggering the other video container to play the pre-loaded next segmented digital human video stream data to achieve seamless connection of the segmented digital human video stream data.

2. The front-end processing method of a digital human video stream according to claim 1, wherein Before receiving the segmented digital human video stream data generated by the server through segmentation based on the pre-constructed long link channel, the method further includes: Receiving, by the server, an intelligent question-and-answer request initiated by the user based on the front-end device, generating a corresponding answer text, and performing semantic analysis on the answer text to identify key semantic nodes associated with the context. Real-time monitoring the network bandwidth of the front-end device, and dynamically adjusting the maximum duration threshold of the segmented digital human video stream data according to the monitoring result. Dynamically adjusting the segmentation point of the digital human video stream according to the key semantic nodes and the adjusted maximum duration threshold, so that each segmented digital human video stream data contains a complete semantic unit. Segmenting the answer text into multiple sentences according to the adjusted segmentation point, generating a video stream segment corresponding to each sentence, and associating action parameters and expression parameters of the digital human with each video stream segment to obtain the segmented digital human video stream data.

3. The front-end processing method of a digital human video stream according to claim 1, wherein, The method further includes: Performing sentiment analysis on each sentence on the server, identifying the emotion type corresponding to the sentence, and generating an associated emotion feature marker for the segmented digital human video stream data corresponding to the sentence. Parsing the emotion feature marker of the segmented digital human video stream data on the front-end device, and adjusting the video playback speed and video playback volume corresponding to the segmented digital human video stream data according to the marker type in the parsing result and based on the predefined playback parameter mapping rule. Applying the adjusted video playback speed and video playback volume in real time during video playback to ensure that the emotional expression of the played video is synchronized with the video content.

4. The front-end processing method of a digital human video stream according to claim 1, characterized in that After listening to the playback progress of the current playback video container, the method further includes: Controlling the display hierarchy of the two video playback containers through the z-index attribute, so that when the current playback video container loads and plays the current segmented digital human video stream data in the front-end queue, setting the display hierarchy of the current playback video container to the visible state and setting the display hierarchy of the other video container to the hidden state. After the current playback video container finishes playing the current segmented digital human video stream data, setting the display hierarchy of the current playback video container to the hidden state. When activating the other video container, synchronously updating the display hierarchy of the other video container to the visible state and triggering the playback instruction of the other video container.

5. A front-end processing method for a digital human video stream according to claim 1, characterized in that After caching the segmented digital human video stream data into the front-end queue according to the transmission order, the method further includes: Dynamically compressing the action parameters and expression parameters of the digital human in the segmented digital human video stream data on the edge computing node according to the hardware performance parameters of the front-end device; Binding the compressed action parameters and expression parameters to the segmented digital human video stream data to generate an optimized video stream segment adapted to the hardware performance parameters of the front-end device; Through the edge computing node, locally cache and store the preset video stream segments that are frequently used, and generate a temporary access link, so that the front-end device preferentially pulls the video stream segment from the local cache of the edge computing node. If not hit, a data acquisition request is sent to the server; Through the cooperative caching mechanism between the edge computing node and the front-end device, and according to the network bandwidth fluctuation, dynamically adjust the distribution ratio of the segmented digital human video stream data between the edge computing node and the front-end device to reduce the network transmission delay.

6. The front-end processing method of a digital human video stream according to claim 1, wherein The method further includes: In the case of detecting that the current playing video container fails to load or the playback is interrupted, switch to another video container and try to reload the same segmented digital human video stream data; If the other video container also fails to load, send a regeneration request to the server to obtain an alternative version of the segmented digital human video stream data; After the alternative version is successfully loaded, insert the alternative version into the original position of the faulty segmented digital human video stream data in the front-end queue and continue playing, and delete the faulty segmented digital human video stream data.

7. A front-end processing method for a digital human video stream according to claim 1, characterized in that, The method further includes: When the front-end device plays the segmented digital human video stream data, real-time obtain the gaze area information of the user on the digital human image in the current playing video container; According to the gaze area information, determine whether the user continuously gazes at a specific part of the digital human image for more than a preset gaze duration threshold. If so, pause the video playback of the current playing video container, and send the user gaze behavior information and the playback progress information of the currently playing segmented digital human video stream data to the server; Through the server, analyze the currently playing segmented digital human video stream data, extract the detailed information content related to the specific part, and generate a corresponding extended explanation video stream segment according to the detailed information content, so as to send the extended explanation video stream segment and the corresponding playback insertion point information to the front-end device; Through the front-end device, in the current playing video container, insert and play the extended explanation video stream segment starting from the playback progress position corresponding to the playback insertion point information, and after the extended explanation video stream segment is played, continue to play the remaining unplayed part of the segmented digital human video stream data.

8. The front-end processing method of a digital human video stream according to claim 1, wherein The method further includes: Monitor the screen direction change of the front-end device or the split-screen operation triggered by the user. If the front-end device switches to the portrait mode, adjust the two video containers with a stacked layout to a side-by-side layout, and dynamically adjust the size ratio of the two video containers. In the split-screen mode, hide the other video container in reserve among the two video containers, pause the preloading operation of the other video container, and keep the currently playing video container for playing to adapt to the display area.

9. A front-end processing device for a digital human video stream, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a front-end processing method for a digital human video stream according to any one of claims 1-8.

10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, a front-end processing method for a digital human video stream according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Subsequent video clip playing control platform and method

    CN114979694A

  • Method and device for realizing short video acceleration based on H5 multiple players

    CN116320587A

  • Video data source information loading method based on 5G network

    CN116801002A

  • Video preloading method and device, equipment and storage medium

    CN117061833A

  • Segmented rendering digital human video interaction method and device, terminal and storage medium

    CN118301413A