Sublimation AI processor-based digital human video generation method and related equipment
By using the Ascend AI processor-based method in the digital video generation process, dynamically adjusting resource deployment and task splitting, the problems of task concurrent processing and resource optimization management in digital video generation are solved, and efficient resource utilization and task execution are achieved.
Patent Information
- Application Number
- CN202510211116.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-25
AI Technical Summary
There are problems in the process of digital human video generation, including concurrent processing of tasks and resource optimization management, resulting in low resource utilization and low task execution efficiency.
The digital human video generation method based on Ascend AI processor is adopted, sensitive word detection and speech synthesis are performed through task forwarding services, the deployment method of the main video synthesis service is dynamically adjusted, and the task is split into subtasks and distributed to the sub-video synthesis service, and the video is uploaded after verifying the integrity of the synthesis frame number.
Effectively split and allocate tasks, optimize resource utilization, improve the execution efficiency of digital human synthesis tasks, ensure the accuracy and completeness of synthesis results, and improve resource scheduling efficiency and task completion reliability.
Smart Images

Figure CN119967254A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of programming technology, and in particular to a method for generating digital human videos based on an Ascend AI processor and related equipment. Background Art
[0002] With the rapid development of digital technology, the application of digital humans (or virtual humans) in multiple fields has gradually become popular, covering industries such as entertainment, education, medical care, and customer service. Digital humans can accurately simulate human appearance, behavior, and emotions, making the interaction between people and computers more natural and smooth, and improving the efficiency of interaction. This trend not only opens up innovative business models for enterprises, but also brings a new experience to users. Especially in the context of the continuous maturity of virtual reality (VR) and augmented reality (AR) technologies, the application of digital human technology is becoming more and more extensive, especially in the fields of digital marketing, social media, and online education. The potential has attracted the attention and attention of more and more organizations.
[0003] The importance of concurrent execution of AI tasks is closely related to the characteristics of large video memory of the Ascend processor. As the demand for AI applications continues to increase, concurrent execution of tasks has become the key to improving computing efficiency and accelerating data processing. Due to its large-capacity video memory, the Ascend AI processor can deploy multiple services simultaneously on the same graphics card, optimize resource utilization, and reduce waiting time between tasks. This not only improves the performance of the overall system when processing multiple AI tasks, but also ensures efficient parallel execution of different tasks without sacrificing computing power. In this way, enterprises can improve business processing capabilities and shorten response times to better meet the rapid needs of the market and customers. Summary of the invention
[0004] The technical problem to be solved by the present invention is: a digital human video generation method and related equipment based on the Ascend AI processor are proposed, aiming to solve the problems of concurrent task processing and resource optimization management in the process of digital human video generation.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: a method for generating digital human video based on Ascend AI processor, comprising the following steps:
[0006] S10, the user terminal submits a digital human synthesis task, including selecting a digital human model, inputting synthesis text or uploading audio;
[0007] S20, the task forwarding service detects sensitive words in the input content and generates synthesized audio through the speech synthesis service;
[0008] S30, sending the digital human model parameters and synthesized audio to the main video synthesis service;
[0009] S40, dynamically adjust the deployment method of the main video synthesis service based on the available video memory and quantity of the Ascend AI processor;
[0010] S50, the main video synthesis service splits the task into several subtasks and distributes them to the sub-video synthesis services;
[0011] S60, the sub-video synthesis service executes the sub-task and returns the result, and the main video synthesis service verifies the integrity of the synthesized frame number;
[0012] S70, the synthesized video is compressed and uploaded to the storage server, and the task status is updated.
[0013] Furthermore, the sensitive word detection in step S20 includes:
[0014] S21. When a user uploads audio, sensitive words are detected after it is converted into text through automatic speech recognition;
[0015] S22, when the user inputs the synthetic text, directly perform sensitive word detection;
[0016] S23, generating synthetic audio after the detection is passed.
[0017] Furthermore, the dynamic adjustment in step S40 includes:
[0018] The main video synthesis service is dynamically deployed based on the available video memory size and number of processors of the Ascend AI processor in the server, and each main service is associated with several sub-video synthesis services.
[0019] Furthermore, the task splitting in step S50 includes:
[0020] The main video synthesis service obtains the number of currently available sub-services n, splits the total task into n+1 sub-tasks, assigns the first n sub-tasks to the sub-services, and the last sub-task is executed by the main service.
[0021] Furthermore, the frame number integrity verification in step S60 includes:
[0022] The main video synthesis service counts the number of synthesized frames returned by the subtask. If the sum is equal to the target total number of frames, the task is considered successful.
[0023] Furthermore, the subtask execution in step S60 includes:
[0024] The sub-video synthesis service calculates the required number of synthesis frames based on the audio duration and video frame rate, generates video frames or images, and returns them to the main service.
[0025] Furthermore, the synthetic video compression in step S70 includes:
[0026] The main video synthesis service integrates the synthesis results of all subtasks in frame sequence and uploads them to cloud storage or file server through network protocol.
[0027] The present invention also provides a digital human video generation device based on the Ascend AI processor, comprising:
[0028] The task submission module is used for the user terminal to submit the digital human synthesis task, including selecting the digital human model, inputting the synthesis text or uploading the audio;
[0029] The audio synthesis module is used by the task forwarding service to detect sensitive words in the input content and generate synthesized audio through the speech synthesis service;
[0030] A material sending module, used to send digital human model parameters and synthesized audio to the main video synthesis service;
[0031] A dynamic adjustment module, which is used to dynamically adjust the deployment mode of the main video synthesis service based on the available video memory and quantity of the Ascend AI processor;
[0032] The task splitting module is used to split the main video synthesis service task into several subtasks and distribute them to the sub-video synthesis services;
[0033] The service verification module is used for the sub-video synthesis service to execute sub-tasks and return results, and the main video synthesis service to verify the integrity of the synthesised frames;
[0034] The video compression module is used to compress the synthesized video and upload it to the storage server, and update the task status.
[0035] The present invention also provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, a digital human video generation method based on an Ascend AI processor as described in any one of the above items is implemented.
[0036] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, can implement the method for generating a digital human video based on an Ascend AI processor as described in any of the above items.
[0037] The beneficial effect of the present invention is that: through the technical solution of the present invention, under the coordinated work of multiple main video synthesis services and sub-video synthesis services, it is possible to effectively split and allocate tasks, optimize resource utilization, and improve the execution efficiency of digital human synthesis tasks. In addition, by comparing the number of frames of the synthesis result with the target number of frames, the accuracy and completeness of the digital human synthesis task are ensured. The implementation of this solution significantly improves the resource scheduling efficiency and the reliability of task completion during the synthesis process. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The specific structure of the present invention is described in detail below in conjunction with the accompanying drawings.
[0039] Figure 1 This is an application scenario diagram of a digital human video generation method based on an Ascend AI processor according to an embodiment of the present invention;
[0040] Figure 2 This is a flow chart of a method for generating a digital human video based on an Ascend AI processor according to an embodiment of the present invention;
[0041] Figure 3 This is a flow chart of sensitive word detection according to an embodiment of the present invention;
[0042] Figure 4 This is a block diagram of a digital human video generation device based on an Ascend AI processor according to an embodiment of the present invention;
[0043] Figure 5 A schematic block diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0046] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0047] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0048] like Figure 1 As shown, the digital human video generation method based on the Ascend AI processor provided in the embodiment of the present invention is applied in the following Figure 1In the application environment, the terminal communicates with the server through the network. The user terminal submits a digital human synthesis task, including selecting a digital human model, inputting a synthesized text or uploading an audio; the task forwarding service detects sensitive words on the input content, and generates a synthesized audio through the speech synthesis service; the digital human model parameters and the synthesized audio are sent to the main video synthesis service; the server dynamically adjusts the deployment mode of the main video synthesis service based on the available video memory and quantity of the Ascend AI processor; the main video synthesis service splits the task into several subtasks and distributes them to the sub-video synthesis service; the sub-video synthesis service executes the subtask and returns the result, and the main video synthesis service verifies the integrity of the synthesized frame number; the synthesized video is compressed and uploaded to the storage server, and the task status is updated. Through the technical solution of the present invention, under the collaborative work of multiple main video synthesis services and sub-video synthesis services, it is possible to effectively split and allocate tasks, optimize resource utilization, and improve the execution efficiency of digital human synthesis tasks. In addition, by comparing the number of frames of the synthesis result with the target number of frames, the accuracy and completeness of the digital human synthesis task are ensured. The implementation of this solution significantly improves the resource scheduling efficiency and the reliability of task completion during the synthesis process. The terminal may be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server may be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0049] like Figure 2 As shown, an embodiment of the present invention is: a method for generating a digital human video based on an Ascend AI processor, comprising the following steps:
[0050] S10. The user terminal submits a digital human synthesis task, including selecting a digital human model, inputting synthesis text, or uploading audio.
[0051] S20: The task forwarding service detects sensitive words in the input content and generates synthesized audio through the speech synthesis service.
[0052] like Figure 3 As shown, in a specific embodiment, the sensitive word detection in step S20 includes:
[0053] S21. When a user uploads audio, sensitive words are detected after it is converted into text through automatic speech recognition;
[0054] S22, when the user inputs the synthetic text, directly perform sensitive word detection;
[0055] S23, after the detection is passed, a synthetic audio is generated.
[0056] In this embodiment, when a user uploads audio, the task forwarding service uses automatic speech recognition (ASR) to convert the audio into text, and performs sensitive word detection on the converted text. If sensitive words are detected, the digital human synthesis task fails, otherwise the detection passes and proceeds to the next step;
[0057] When the user inputs the synthesized text, the task forwarding service detects sensitive words in the text. If sensitive words are detected, the digital human synthesis task fails. Otherwise, the detection passes and proceeds to the next step.
[0058] If the detection passes, the synthesized text or the converted text is sent to an idle speech synthesis service for synthesis to obtain synthesized audio.
[0059] S30: Send the digital human model parameters and synthesized audio to the main video synthesis service.
[0060] In this embodiment, the task forwarding service sends the digital human model parameters selected by the user and the user-uploaded audio or synthesized audio that has passed the detection to the currently idle main video synthesis service.
[0061] S40, dynamically adjust the deployment method of the main video synthesis service based on the available video memory and quantity of the Ascend AI processor.
[0062] In a specific embodiment, the dynamic adjustment in step S40 includes:
[0063] The main video synthesis service is dynamically deployed based on the available video memory size and number of processors of the Ascend AI processor in the server, and each main service is associated with several sub-video synthesis services.
[0064] In this embodiment, according to the actual business scenario, the deployment method of the main video synthesis service is dynamically adjusted according to the available video memory size of the server's Ascend AI processor and the number of Ascend AI processors to ensure the rational allocation of resources.
[0065] S50: The main video synthesis service splits the task into several subtasks and distributes them to the sub-video synthesis services.
[0066] In a specific embodiment, the task splitting in step S50 includes:
[0067] The main video synthesis service obtains the number of currently available sub-services n, splits the total task into n+1 sub-tasks, assigns the first n sub-tasks to the sub-services, and the last sub-task is executed by the main service.
[0068] In this embodiment, after the idle main video synthesis service receives the digital human synthesis task, the status of the main video synthesis service changes to "busy", and the status of the digital human synthesis task in the database is updated to "executing".
[0069] The main video synthesis service calculates the total number of frames N that need to be synthesized based on the duration of the user-uploaded audio or synthesized audio that has passed the detection and the frame rate (FPS) of the target synthesized video. frames :
[0070] N frames =T audio ×fps
[0071] The total number of video frames N of the target synthetic video is known frames , input the starting frame number S of the video index , end frame number E index , the final generated frame sequence T.
[0072] The generation of the final composite video frame sequence T is based on the following steps and mathematical relationships:
[0073] The forward sequence Foward is from S index To E index ,Right now:
[0074] Foward=[S index ,S index +1,S index +2,…,E index ];
[0075] This is a package containing E index -S index A sequence of +1 elements.
[0076] The reverse sequence is from E index -1 to S index +1,
[0077] Reverse=[E index -1,E index -2,E index -3,…,S index +1];
[0078] This is a package containing E index -S index A sequence of elements.
[0079] Combining the forward sequence and the reverse sequence, we get the initial frame sequence T:
[0080] T initial =Forward+Reverse
[0081] If the length of the generated frame sequence is less than the given total number of frames N frames , then continue splicing the forward and reverse sequences until the total length is at least equal to N frames .
[0082] After each splicing, the length L of the new frame sequence T becomes:
[0083] L=Len(T)=n·len(T initial ), where n is the number of splicing;
[0084] Until L ≥ N frames Finally, if the length of T exceeds N frames , then cut off the excess part to ensure that the total length is exactly N frames .
[0085] T=T[:N frames ];
[0086] The final generated frame sequence T satisfies the following conditions:
[0087] T initial is the initial frame sequence;
[0088] len(T initial )=2·(E index -S index +1)-1;
[0089] When the number of splicing times is n, len(T)≥N frames , and then cut off the excess part;
[0090]
[0091] The main video composition service obtains the number of currently available sub-video composition services by communicating with the health interface of the sub-video composition service. n , and dynamically generate the final frame sequence calculated in the digital human synthesis task according to the actual situation T It is split into several sub-digital human synthesis tasks, the first n digital human synthesis sub-tasks are sent to the sub-video synthesis interface, and the last digital human synthesis sub-task is sent to the main video synthesis interface, where n≥0.
[0092] The number of sub-video synthesis interfaces and main video synthesis interfaces is N:
[0093] N = n + 1 (n ≥ 0);
[0094] The basic length of each subframe sequence is:
[0095]
[0096] in It is the integer part of L divided by N, indicating the minimum length of each subframe sequence.
[0097] The remaining element r is the remainder when L is divided by n:
[0098] r = L mod N;
[0099] The remaining elements are evenly distributed among the first r sub - lists such that the lengths of these sub - lists are one element more than those of the other sub - lists.
[0100] When i < r, the length of the i - th sub - frame sequence is:
[0101] sublist_size i = size + 1, (i < r);
[0102] When i ≥ r, the length of the i - th sub - frame sequence is:
[0103] sublist_size i = size, (i ≥ r);
[0104] The elements of each sub - frame sequence are extracted from T[s index :e index .
[0105] For the starting position s of the i - th sub - frame sequence index is:
[0106]
[0107] The ending position e of the sub - frame sequence index = s index + sublist_size i ;
[0108] Each sub - video synthesis service that receives the sub - digital human synthesis task executes the task and notifies the main video synthesis service of the synthesis result after the task execution ends.
[0109] S60. The sub - video synthesis service executes the sub - task and returns the result, and the main video synthesis service verifies the integrity of the synthesized number of frames.
[0110] In a specific embodiment, the verification of the integrity of the number of frames in step S60 includes:
[0111] The main video synthesis service counts the number of synthesized frames returned by the sub - task. If the sum is equal to the target total number of frames, the task is determined to be successful.
[0112] In a specific embodiment, the execution of the sub - task in step S60 includes:
[0113] The sub - video synthesis service calculates the required number of synthesized frames according to the audio duration and the video frame rate, generates video frames or images, and then returns to the main service.
[0114] In this embodiment, the main video synthesis service checks the synthesis result and determines whether the digital human synthesis task is successful based on the number of synthesized images or the sum of the number of sub-digital human synthesized video frames and the total number of frames that need to be synthesized.
[0115] S70, the synthesized video is compressed and uploaded to the storage server, and the task status is updated.
[0116] In a specific embodiment, the synthetic video compression in step S70 includes:
[0117] The main video synthesis service integrates the synthesis results of all subtasks in frame sequence and uploads them to cloud storage or file server through network protocol.
[0118] In this embodiment, when the synthesis result of the sub-video synthesis service is consistent with the total number of frames of the target synthesized video, the main video synthesis service regards the digital human synthesis task as successfully executed, compresses the synthesized video and uploads it to the file storage server, and updates the status of the digital human synthesis task in the database to "success", and the status of the main video synthesis service becomes "idle".
[0119] The technical effects of the technical solution of this application include:
[0120] Through the reasonable allocation of tasks and dynamic resource scheduling, the efficiency of digital human video generation has been greatly improved. This efficiency improvement is not only reflected in the speed of video generation, but also in the full utilization of resources, avoiding the waste of hardware resources.
[0121] Through effective task management and resource scheduling, the stability and reliability of the system are enhanced. The dynamic resource allocation strategy ensures that the system can run stably even under high load conditions, reducing the risk of service interruption caused by resource competition.
[0122] By ensuring the high quality of the synthesis results, the user experience is improved. The result verification technology ensures the high accuracy and completeness of the video synthesis, and users can obtain video content that meets their expectations, thereby improving user satisfaction with the digital human video generation system.
[0123] The system is designed with scalability in mind and can be expanded as user needs grow. This scalability means that the system can adapt to different business scales and demand changes without the need for frequent hardware replacement or large-scale system upgrades.
[0124] like Figure 4 As shown, the present invention also provides a digital human video generation device based on the Ascend AI processor, comprising:
[0125] The task submission module 10 is used for the user terminal to submit a digital human synthesis task, including selecting a digital human model, inputting synthesis text or uploading audio;
[0126] The audio synthesis module 20 is used for the task forwarding service to detect sensitive words in the input content and generate synthesized audio through the speech synthesis service;
[0127] The material sending module 30 is used to send the digital human model parameters and the synthesized audio to the main video synthesis service;
[0128] A dynamic adjustment module 40, used to dynamically adjust the deployment mode of the main video synthesis service based on the available video memory and quantity of the Ascend AI processor;
[0129] A task splitting module 50 is used to split the main video synthesis service task into several subtasks and distribute them to the sub-video synthesis services;
[0130] A service verification module 60, used for the sub-video synthesis service to execute the sub-task and return the result, and the main video synthesis service to verify the integrity of the synthesis frame number;
[0131] The video compression module 70 is used to compress the synthesized video and upload it to the storage server, and update the task status.
[0132] It should be noted that technicians in the relevant field can clearly understand that the specific implementation process of the above-mentioned digital human video generation device based on the Ascend AI processor can refer to the corresponding description in the aforementioned method embodiment. For the convenience and conciseness of the description, it will not be repeated here.
[0133] The above-mentioned digital human video generation device based on the Ascend AI processor can be implemented in the form of a computer program. The computer program can be used in Figure 5 Runs on the computer device shown.
[0134] See also Figure 5 , Figure 5 5 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a terminal or a server, wherein the terminal may be an electronic device with communication functions such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, and a wearable device. The server may be an independent server or a server cluster composed of multiple servers.
[0135] See also Figure 5 The computer device 500 includes a processor 502 , a memory and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0136] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can enable the processor 502 to execute a method for generating a digital human video based on the Ascend AI processor.
[0137] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500 .
[0138] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a digital human video generation method based on the Ascend AI processor.
[0139] The network interface 505 is used to communicate with other devices over the network. Figure 5 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0140] The processor 502 is used to run a computer program 5032 stored in the memory to implement the digital human video generation method based on the Ascend AI processor as described above.
[0141] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0142] It can be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0143] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the digital human video generation method based on the Ascend AI processor as described above.
[0144] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which are computer-readable storage media that can store program codes.
[0145] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0146] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0147] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0148] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0149] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for generating digital human videos based on Ascend AI processor, characterized in that: The following steps are involved: S10, the user terminal submits a digital human synthesis task, including selecting a digital human model, inputting synthesis text or uploading audio; S20, the task forwarding service detects sensitive words in the input content and generates synthesized audio through the speech synthesis service; S30, sending the digital human model parameters and synthesized audio to the main video synthesis service; S40, dynamically adjust the deployment method of the main video synthesis service based on the available video memory and quantity of the Ascend AI processor; S50, the main video synthesis service splits the task into several subtasks and distributes them to the sub-video synthesis services; S60, the sub-video synthesis service executes the sub-task and returns the result, and the main video synthesis service verifies the integrity of the synthesized frame number; S70, the synthesized video is compressed and uploaded to the storage server, and the task status is updated.
2. The method for generating digital human video based on Ascend AI processor according to claim 1, characterized in that: The sensitive word detection in step S20 includes: S21. When a user uploads audio, sensitive words are detected after it is converted into text through automatic speech recognition; S22, when the user inputs the synthetic text, directly perform sensitive word detection; S23, after the detection is passed, a synthetic audio is generated.
3. The method for generating digital human video based on Ascend AI processor according to claim 1, characterized in that: The dynamic adjustment in step S40 includes: The main video synthesis service is dynamically deployed based on the available video memory size and number of processors of the Ascend AI processor in the server, and each main service is associated with several sub-video synthesis services.
4. The method for generating digital human video based on Ascend AI processor according to claim 1, characterized in that: The task splitting in step S50 includes: The main video synthesis service obtains the number of currently available sub-services n, splits the total task into n+1 sub-tasks, assigns the first n sub-tasks to the sub-services, and the last sub-task is executed by the main service.
5. The method for generating digital human video based on Ascend AI processor according to claim 1, characterized in that: The frame number integrity verification in step S60 includes: The main video synthesis service counts the number of synthesized frames returned by the subtask. If the sum is equal to the target total number of frames, the task is considered successful.
6. The method for generating digital human video based on Ascend AI processor according to claim 1, characterized in that: The subtask execution in step S60 includes: The sub-video synthesis service calculates the required number of synthesis frames based on the audio duration and video frame rate, generates video frames or images, and returns them to the main service.
7. The method for generating digital human video based on Ascend AI processor according to claim 1, characterized in that: The synthetic video compression in step S70 includes: The main video synthesis service integrates the synthesis results of all subtasks in frame sequence and uploads them to cloud storage or file server through network protocol.
8. A digital human video generation device based on Ascend AI processor, characterized in that: include: The task submission module is used for the user terminal to submit the digital human synthesis task, including selecting the digital human model, inputting the synthesis text or uploading the audio; The audio synthesis module is used by the task forwarding service to detect sensitive words in the input content and generate synthesized audio through the speech synthesis service; A material sending module, used to send digital human model parameters and synthesized audio to the main video synthesis service; A dynamic adjustment module, which is used to dynamically adjust the deployment mode of the main video synthesis service based on the available video memory and quantity of the Ascend AI processor; The task splitting module is used to split the main video synthesis service task into several subtasks and distribute them to the sub-video synthesis services; The service verification module is used for the sub-video synthesis service to execute sub-tasks and return results, and the main video synthesis service to verify the integrity of the synthesised frames; The video compression module is used to compress the synthesized video and upload it to the storage server, and update the task status.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the digital human video generation method based on the Ascend AI processor as described in any one of claims 1 to 7 when executing the computer program.
10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, can implement the digital human video generation method based on the Ascend AI processor as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Video rendering method and device, computer equipment and storage medium
CN113923519A
Digital human production generation method, device, equipment, medium and program product
CN118587354A
Mounting device for pole-type solar power generation module
KR1020240156718A
Video generation method and apparatus
WO2024198989A1