High-concurrency AI teaching-assistant digital human interaction system
By migrating the image rendering process of digital humans to the client and using front-end rendering technology, the high requirements for server computing and bandwidth in traditional digital human synthesis technology are solved, and a high concurrency, low cost and smooth user experience is achieved.
Patent Information
- Application Number
- CN202510262276.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional digital human synthesis technology relies on high computing power and high bandwidth networks on the server side, resulting in increased system complexity, limited user interaction experience and high operating costs.
Migrate the image rendering process of digital people to the client, adopts front-end rendering technology, and uses the browser's graphics rendering pipeline and hardware acceleration mechanism to achieve efficient image drawing.
It reduces the demand for back-end computing resources, reduces network bandwidth usage, improves the system's concurrent processing capabilities and user experience, and reduces operating costs.
Smart Images

Figure CN120219583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital human AI, and particularly relates to a high-concurrency AI teaching assistant digital human interaction system. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, the application scope of digital human synthesis technology has been continuously expanding and has now penetrated into many fields such as entertainment, education, medical care, and customer service. This technology can not only generate extremely realistic virtual characters but also greatly enrich the user experience, making it more vivid and interactive.
[0003] Moreover, digital human synthesis technology can also customize virtual characters with specific appearances, personalities, and even skills according to the needs of users and different scenarios. In the entertainment field, this technology is used to create special effects characters in movies, NPCs (non-player characters) in games, and virtual idols, etc., bringing an unprecedented immersive experience to the audience. In education, by creating image spokespersons for historical figures or scientific concepts, the learning process can be made more vivid and interesting, helping students better understand and remember knowledge. In the medical and health industry, virtual assistants developed using digital human technology can not only provide 24 / 7 non-stop consultation services but also assist doctors in diagnosing diseases and give personalized health management suggestions to patients. In addition, in the customer service field, AI-driven digital humans have become one of the important tools for many enterprises to improve service quality. They can quickly respond to customer inquiries and effectively solve common problems, thereby improving work efficiency and reducing operating costs. It is worth noting that with the continuous maturity and development of technology, the application scenarios of digital human synthesis will further expand to more fields in the future. For example, using virtual anchors to broadcast weather forecasts or sports event results in news reports; or in personal assistant services, by customizing exclusive image and voice features, users can experience more considerate services. In short, with relevant research and technological breakthroughs, digital human synthesis technology will bring infinite possibilities to us and profoundly change people's lifestyles and social interaction patterns.
[0004] However, traditional digital human synthesis technology often relies on the powerful computing power of the server side, which not only increases the complexity of the system but also limits the user's instant interaction experience. Since such methods usually require the support of high-bandwidth network connections and high-performance computing resources, the implementation cost increases significantly. More importantly, in the case of unstable network conditions, this highly dependent approach may seriously affect the user experience and thus reduce the overall satisfaction. Therefore, it is crucial to develop a more efficient and easy-to-use digital human synthesis solution to improve the user experience and promote the further development of this field. Summary of the Invention
[0005] The present invention provides a high-concurrency AI teaching assistant digital human interaction system. By migrating the digital human image rendering process to the client side, that is, in the user's browser, it effectively solves the technical problems of high consumption of AI computing power resources and high network bandwidth requirements in traditional digital human interaction systems, reduces operating costs, and improves the concurrency processing ability of the system.
[0006] The present invention provides a high-concurrency AI teaching assistant digital human interaction system, including a server side and a client side. The server side is connected to the client side. The client side uses front-end rendering for digital human image rendering, and its process is as follows:
[0007] When the digital human is initialized, the client side first loads all the action pictures of the digital human into the memory, creates a canvas, starts an animation loop based on the browser's requestAnimationFrame API, and calls the function for drawing images on the canvas in the requestAnimationFrame callback, which is executed 25 times per second to achieve a continuous animation effect.
[0008] Further, each action of the digital human adopts a random drawing method. The underlying implementation of the function for drawing images depends on the canvas element in HTML5. The canvas is obtained in JavaScript and the 2D rendering context is obtained, and the browser's graphics rendering pipeline and hardware acceleration mechanism are utilized to achieve efficient image drawing.
[0009] Further, the image of the digital human is an animated cartoon image. In digital human interaction, it is divided into two states: the digital human listening and the digital human speaking. The digital human speaking state is the state where the digital human's mouth opens and closes when playing voice, and the digital human listening state is the state where the digital human's mouth is closed when not playing voice.
[0010] Further, the digital human makes different action pictures according to 4 actions, which are single-hand introduction, both-hands introduction, standing and speaking, and standing without speaking;
[0011] Each action is a group of consecutive png action pictures, and each action is played at a speed of 25 frames per second, and the playing duration of each action is set.
[0012] Further, three actions are set for the digital human speaking state, which are: single-hand introduction, with a playing duration of 3 seconds; both-hands introduction, with a playing duration of 3 seconds; standing and speaking, with a playing duration of 3 seconds;
[0013] One action is set for the digital human listening state: standing without speaking, with a playing duration of 2 seconds.
[0014] Further, when the digital human is in the listening state, it repeats the action of standing still without speaking, and the digital human has the actions of body swaying and blinking;
[0015] When the digital human is in the speaking state, set the playback order of single - hand introduction, double - hand introduction, and standing and speaking according to preset rules, and perform repeated loop playback using the playback order until the voice playback ends and it enters the listening state of the digital human.
[0016] Further, the preset rules are as follows:
[0017] Set the standing and speaking action as A, the single - hand introduction action as B, and the double - hand introduction action as C, and set the playback order to satisfy that an A must be interspersed before or after B or C, B or C must be interspersed before and after A, and it is randomly determined whether B or C appears in the position of B or C until the digital human finishes speaking.
[0018] The beneficial effects of the present invention are as follows:
[0019] The present invention uses client - side rendering technology to move the image rendering process of the digital human from the server - side to the client - side. By optimizing the architecture design, the system can support a large number of users to have high - quality interactive experiences simultaneously without causing performance degradation or response latency, reducing the demand for backend computing resources and the use of network bandwidth. This not only improves the user experience but also makes large - scale deployment possible. Due to reducing the demand for high - performance servers and high - bandwidth networks, the operating costs are significantly reduced, making the application of digital human interaction technology more economically feasible, especially important for educational institutions and individual users with limited budgets. The system design takes into account the performance under different network environments and can maintain good operating status even under low - bandwidth or unstable network conditions, ensuring the wide applicability and reliability of the service. By improving the response speed and reducing latency, it provides users with a more fluent, natural, and immersive interactive experience, enhancing user participation and satisfaction. Description of the Drawings
[0020] Figure 1 It is a schematic structural diagram of the high - concurrency AI teaching assistant digital human interaction system of the present invention.
[0021] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments
[0022] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0023] The present invention provides a high-concurrency AI teaching assistant digital human interaction platform. By migrating the digital human image rendering process to the client side, that is, in the user's browser, the problems existing in the traditional digital human interaction system are effectively solved. For example, in the traditional system, the digital human image usually needs to be rendered on the server side, which not only consumes a large amount of AI computing resources, but also places extremely high requirements on network bandwidth, resulting in a sharp increase in operating costs and limiting the concurrent processing ability of the system.
[0024] As Figure 1 shown, the present invention provides a high-concurrency AI teaching assistant digital human interaction system, including a server side and a client side. The server side is connected to the client side, and the client side uses front-end rendering to render the digital human image. The process is as follows:
[0025] When the digital human is initialized, the client side first loads all the action pictures of the digital human into the memory, creates a canvas, and starts an animation loop based on the browser's requestAnimationFrame API. In the requestAnimationFrame callback, a function for drawing images on the canvas is called 25 times per second to achieve a continuous animation effect. In order to achieve a more natural effect of the digital human's speaking action, each action of the digital human is randomly drawn. The underlying implementation of the function for drawing images depends on the canvas element in HTML5. The canvas is obtained in JavaScript, and the 2D rendering context is obtained. The browser's graphics rendering pipeline and hardware acceleration mechanism are used to achieve efficient image drawing.
[0026] By adopting the front-end rendering technology, our system significantly reduces the demand for backend computing resources, enabling more users to enjoy high-quality interactive experiences simultaneously without worrying about server overload or response latency. In addition, this method reduces the dependence on a stable and high-speed network connection, ensuring a good user experience even in poor network conditions. In this way, both educational institutions and individual users can enjoy a smoother and more efficient digital human teaching assistant service at a lower cost, promoting the wide application and development of artificial intelligence technology in the education field.
[0027] In one embodiment, the digital human image is an animated cartoon image, not a realistic style. Therefore, the problem of lip alignment when the digital human speaks can be ignored, and only a change in the opening and closing of the mouth is required when speaking.
[0028] In digital human interaction, there are two states: the digital human listening and the digital human speaking. The digital human speaking is the state where the digital human plays voice, and the digital human listening is the state where the digital human does not play voice. For example, after the digital human stops speaking, it is in the digital human listening state.
[0029] After the digital human image prototype is determined, the digital human creates different action pictures according to 4 actions. The 4 actions are single-handed introduction, two-handed introduction, standing and speaking, and standing without speaking. Each action is a set of consecutive png action pictures, and each action is played at a speed of 25 frames per second, and the playing duration of each action is set.
[0030] Specifically, there are three actions for the state where the digital human speaks, which are respectively:
[0031] Single-handed introduction: 75 frames, playing duration is 3 seconds;
[0032] Two-handed introduction: 75 frames, playing duration is 3 seconds;
[0033] Standing and speaking: 75 frames, playing duration is 3 seconds.
[0034] There is one action for the state where the digital human listens:
[0035] Standing without speaking, 50 frames, playing duration is 2 seconds;.
[0036] In one embodiment, when the digital human is in the listening state, the action of "standing without speaking" is repeatedly looped. Note that although the digital human does not speak, it should not be in a static state, but has small body shakes and blinking actions, which makes it more realistic and lively.
[0037] When the digital human is in the speaking state, the actions of "single-handed introduction", "two-handed introduction", and "standing and speaking" are repeatedly looped until the speaking stops and it enters the listening state of the digital human.
[0038] Set the playing order of the 3 actions as follows:
[0039] Let "standing and speaking" be A, "single-handed introduction" be B, and "two-handed introduction" be C. Then the playing order of the 3 animations needs to satisfy: A must be interspersed before or after B or C, B or C must be interspersed before and after A, and it is randomly determined whether B or C appears in the position of B or C until the digital human finishes speaking. For example, the playing orders are ABACABACA, ABACABABA, ACACABAB.
[0040] Interspersing "single-handed introduction" and "two-handed introduction" in the middle of "standing and speaking" is because if there is only "standing and speaking", it will seem very dull. "Single-handed introduction" is the action of the digital human stretching out and spreading the right hand, and "two-handed introduction" is the action of the digital human stretching out and spreading both hands.
[0041] The present invention uses client - side rendering technology to move the rendering process of the digital human's image from the server - side to the client - side (in the user's browser). This method significantly reduces the demand for backend computing resources, decreases the use of network bandwidth, thereby improving the system's concurrent processing ability and overall efficiency. It is applicable to the scenario of student teaching assistant digital humans. Since each student needs to interact with their own AI teaching assistant digital human simultaneously, high requirements are imposed on the concurrent ability and low - bandwidth occupation ability of the digital human. The reason why general digital human systems consume a large amount of AI computing power resources and network bandwidth is that rendering the expressions, lip - sync, etc. of digital humans requires a large amount of AI computing power, and the client usually does not have such hardware conditions. Therefore, it needs to be rendered on the server - side and then the video is transmitted to the client, which also consumes a large amount of network bandwidth. The present invention simplifies the expressions and lip - sync of digital humans to achieve front - end rendering of digital humans, greatly reducing resource consumption. At the same time, compared with pure voice interaction without digital humans, it greatly increases the interest of interaction.
[0042] The present invention has the following effects:
[0043] 1. High concurrent ability: By optimizing the architecture design, the system can support a large number of users to have high - quality interactive experiences simultaneously without causing performance degradation or response latency. This not only improves the user experience but also makes large - scale deployment possible.
[0044] 2. Cost - effectiveness: Due to reducing the demand for high - performance servers and high - bandwidth networks, the present invention significantly reduces the operation cost, making the application of digital human interaction technology more economically feasible, especially important for educational institutions and individual users with limited budgets.
[0045] 3. Adaptability and flexibility: The system design takes into account the performance under different network environments and can maintain good operating conditions even under low - bandwidth or unstable network conditions, ensuring the wide applicability and reliability of the service.
[0046] 4. Improvement of user experience: By increasing the response speed and reducing latency, the present invention provides users with a smoother, more natural and immersive interactive experience, enhancing user participation and satisfaction.
[0047] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article or method including such an element.
[0048] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A high-concurrency AI teaching assistant digital human interaction system, characterized in that: It includes a server and a client, the server is connected to the client, and the client uses front-end rendering to render the image of the digital human, and the process is as follows: When the digital human is initialized, the client first loads all the action pictures of the digital human into the memory, creates a canvas, starts the animation loop based on the browser's requestAnimationFrame API, calls the function of drawing images on the canvas in the requestAnimationFrame callback, and executes it 25 times per second to achieve a continuous animation effect.
2. The high-concurrency AI teaching assistant digital human interaction system according to claim 1 is characterized in that: Each action of the digital human is drawn randomly. The underlying implementation of the image drawing function relies on the canvas element in HTML5. The canvas and 2D rendering context are obtained in JavaScript, and the browser's graphics rendering pipeline and hardware acceleration mechanism are used to achieve efficient image drawing.
3. The high-concurrency AI teaching assistant digital human interaction system according to claim 1 is characterized in that: The image of the digital human is an animated cartoon image. In the digital human interaction, there are two states: digital human listening and digital human speaking. The digital human speaking is a state in which the digital human opens and closes its mouth when playing voice, and the digital human listening is a state in which the digital human keeps its mouth tightly closed when not playing voice.
4. The high-concurrency AI teaching assistant digital human interaction system according to claim 1 is characterized in that: The digital human produces different action pictures according to four actions, and the four actions are introduction with one hand, introduction with both hands, standing and speaking, and standing and not speaking; Each action is a set of continuous PNG action pictures. Each action is played at a speed of 25 frames per second, and the playing time of each action is set.
5. The high-concurrency AI teaching assistant digital human interaction system according to claim 4 is characterized in that: There are three actions for setting the digital human's speaking state, namely: one-handed introduction, with a playback duration of 3 seconds; two-handed introduction, with a playback duration of 3 seconds; standing and speaking, with a playback duration of 3 seconds; The digital human's listening state is set to have one action: standing without speaking, and the playing time is 2 seconds.
6. The high-concurrency AI teaching assistant digital human interaction system according to claim 5 is characterized in that: When the digital human is in a listening state, the action of standing still without speaking is repeated, and the digital human has body shaking and eye blinking actions; When the digital human is in the speaking state, the playback sequence of single-handed introduction, double-handed introduction, and standing speaking is set according to preset rules, and the playback sequence is used for repeated loop playback until the voice playback ends and the digital human enters the listening state.
7. The high-concurrency AI teaching assistant digital human interaction system according to claim 6 is characterized in that: The preset rules are: The standing speaking action is set as A, the one-hand introduction action is set as B, and the two-hand introduction action is set as C. The playback order is set to meet the requirement that A must be inserted before and after B or C, and B or C must be inserted before and after A. B or C is randomly determined to appear in the position of B or C until the digital person finishes speaking.
Citation Information
Patent Citations
Digital human rendering method and device, storage medium and electronic equipment
CN113886551A
Digital object rendering method and device, electronic equipment and nonvolatile storage medium
CN117786074A
Model rendering method and device, equipment and storage medium
CN117853635A
Rendering method and device based on WEB security isolation system and medium
CN118733918A
Animation compositor for digital avatars
US20240386644A1