Abnormal Detection Method and Related Device for Digital Human Video

By detecting and adjusting the movement postures in digital human videos, the image flaws caused by abnormal frames in digital human videos are solved, and the user experience and the performance of digital human image is improved.

CN113888598BActive Publication Date: 2025-06-03SHENZHEN ZHUIYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111212334.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-06-03
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

There may be abnormal frames in digital human videos, resulting in image flaws such as body mosaics during video playback, affecting the image and user experience of digital humans.

Method used

By obtaining the initial action images of the digital character model at multiple time points, detecting whether the action posture of the current frame is abnormal, and adjusting the initial action total trajectory to obtain the target action total trajectory, and generating the target action image to replace the abnormal frame.

Benefits of technology

It effectively avoids generating digital human videos with poor viewing, and improves the broadcast image and user experience of the digital character model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113888598B_ABST
    Figure CN113888598B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose an abnormal detection method and related devices for digital human videos. The method includes: obtaining initial action images of a digital human model at multiple time points; after detecting an abnormal action posture in the current frame, adjusting the initial total action trajectory to obtain a target total action trajectory, and obtaining all target action images of the digital human model according to the target total action trajectory, so as to avoid generating a digital human video with poor visual experience, which affects the broadcast image of the digital human model and the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of Internet technologies, and in particular, to an abnormal detection method for digital human videos and related devices. Background Art

[0002] With the continuous development of information life, human-computer interaction technology penetrates into all aspects of daily life. Among them, the application of digital humans (which can be referred to as robots, virtual humans, or digital character models in the text) is also becoming more and more extensive. A digital human is a virtual three-dimensional human made by using virtual reality technology, human-computer interaction, high-precision three-dimensional portrait simulation, artificial intelligence, motion capture, and facial expression capture technologies. It can replace real people to perform highly realistic business guidance and question answering and other tasks, reducing the cost of artificial customer service.

[0003] In the actual application process, a trained digital human often generates corresponding broadcast lines and / or broadcast postures according to text input and / or action name input to render its image, so that the digital human can simulate a real person and be presented to users in the form of video interaction. Specifically, it can be understood that one frame is a picture and a motion posture, and continuous frames can form a motion trajectory. For example, an animated cartoon is composed of many frames. However, it cannot be ignored that the motion postures generated by the digital human may not appear in the training data, making the video frames prone to abnormalities. For example, due to the excessive or irregular motion amplitude at a certain moment (which is quite different from the existing training data), there are image defects such as mosaics on the limbs of the digital human when the video plays to a certain frame. Therefore, if the generated video frames are not subjected to effect detection and processing, it is easy to affect the image of the digital human and the user's perception, and reduce the user experience. Summary of the Invention

[0004] The embodiments of the present application provide an abnormal detection method for digital human videos and related devices to avoid abnormal video frames played by digital character models.

[0005] The first aspect of the embodiments of the present application provides an abnormal detection method for digital human videos, including:

[0006] Obtain initial action images of a digital character model at multiple time points, where each time point corresponds to one frame of initial action image, and each frame of the initial action image is used to represent the action posture of the digital character model at different time points, and the action postures of all frames of the initial action images form an initial action total trajectory;

[0007] Detect whether the action posture of the current frame is abnormal, where the current frame is any one of the initial action images;

[0008] If so, adjust the initial total action trajectory to obtain a target total action trajectory, and generate a target action image of the digital human model according to the target total action trajectory, where the target action image is used as a video composition frame of the digital human model output to the user.

[0009] Optionally, the detecting whether the action pose of the current frame is abnormal includes:

[0010] Determine the total probability that the pixels corresponding to the action pose in the current frame have defects, where the total probability is used to represent the likelihood of the action pose in the current frame being abnormal;

[0011] Detect whether the total probability exceeds a preset probability.

[0012] Optionally, the determining the total probability that the pixels corresponding to the action pose in the current frame have defects includes:

[0013] Divide the current frame into N pixel sub-regions;

[0014] Detect the sub-probability that the pixels corresponding to the action pose in each pixel sub-region have defects through a classification model;

[0015] Statistically calculate the total probability that the pixels corresponding to the action pose in the current frame have defects according to all the sub-probabilities.

[0016] Optionally, the determining the total probability that the pixels corresponding to the action pose in the current frame have defects includes:

[0017] Divide each initial action image in the current frame, the initial action images of the previous m frames, and the initial action images of the next m frames into N pixel sub-regions;

[0018] Detect the sub-probability that the pixels corresponding to the action pose in each pixel sub-region have defects through a classification model respectively;

[0019] Statistically calculate the total probability that the pixels corresponding to the action pose in the current frame have defects according to all the sub-probabilities.

[0020] Optionally, the adjusting the initial total action trajectory to obtain a target total action trajectory includes:

[0021] Determine the nearest historical normal action pose of the current frame, where the nearest historical normal action pose represents the initial action image of a historical frame before the current frame;

[0022] Generate a target action sub-trajectory in the reverse direction along the initial action sub-trajectory, where the initial action sub-trajectory is the historical action trajectory in the initial total action trajectory that includes the nearest historical normal action pose;

[0023] Update the action trajectory after the initial action sub-trajectory in the initial action total trajectory to the target action sub-trajectory to form a target action total trajectory.

[0024] Optionally, the generating the target action image of the digital human model according to the target action total trajectory includes:

[0025] Generate a corresponding target action image according to the target action sub-trajectory to obtain a digital human video including the initial action image corresponding to the initial action sub-trajectory and the target action image corresponding to the target action sub-trajectory.

[0026] Optionally, the generating the target action image of the digital human model according to the target action total trajectory includes:

[0027] Generate all target action images of the digital human model corresponding to the target action total trajectory.

[0028] Optionally, after detecting whether the action pose of the current frame is abnormal, the method further includes:

[0029] If not, notify the digital human model that the action pose of the current frame is normal, and the current frame is used as a video composition frame of the digital human model output to the user.

[0030] Optionally, after detecting whether the action pose of the current frame is abnormal, the method further includes:

[0031] If so, notify the digital human model that the action pose of the current frame is abnormal, and the current frame is used as a video composition frame of the digital human model output to the user.

[0032] Optionally, the initial action image and the target action image are generated from the action pose key points of the digital human model.

[0033] Optionally, the total probability includes one probability or multiple probabilities among the image black spot probability and the mosaic probability.

[0034] A second aspect of the embodiments of the present application provides an abnormal detection device for a digital human video, including:

[0035] An acquisition module, configured to acquire initial action images of a digital human model at multiple time points, where each time point corresponds to a frame of initial action image, and each frame of the initial action image is used to represent the action pose of the digital human model at different time points, and the action poses of all frames of the initial action images form an initial action total trajectory;

[0036] Anomaly detection module, used to detect whether the action posture of the current frame is abnormal, where the current frame is any one of the initial action images;

[0037] Action processing module, used to adjust the initial action total trajectory to obtain a target action total trajectory;

[0038] Image generation module, used to generate target action images of the digital human model according to the target action total trajectory, and the target action images are used as video composition frames of the digital human model output to the user.

[0039] Optionally, the anomaly detection module is specifically used for:

[0040] Determine the total probability that the pixels corresponding to the action posture in the current frame have defects, and the total probability is used to represent the possibility of the action posture in the current frame being abnormal;

[0041] Detect whether the total probability exceeds a preset probability.

[0042] Optionally, the anomaly detection module is specifically used for:

[0043] Divide the current frame into N pixel sub-regions;

[0044] Detect the sub-probability that the pixels corresponding to the action posture in each pixel sub-region have defects through a classification model;

[0045] Statistically calculate the total probability that the pixels corresponding to the action posture in the current frame have defects according to all the sub-probabilities.

[0046] Optionally, the anomaly detection module is specifically used for:

[0047] Divide each initial action image in the current frame, the initial action images of the previous m frames, and the initial action images of the next m frames into N pixel sub-regions respectively;

[0048] Detect the sub-probability that the pixels corresponding to the action posture in each pixel sub-region have defects through a classification model respectively;

[0049] Statistically calculate the total probability that the pixels corresponding to the action posture in the current frame have defects according to all the sub-probabilities.

[0050] Optionally, the action processing module is specifically used for:

[0051] Determine the nearest historical normal action posture of the current frame, where the nearest historical normal action posture refers to the initial action image of a historical frame before the current frame;

[0052] Generate a target action sub-trajectory in reverse along the initial action sub-trajectory, where the initial action sub-trajectory is the historical action trajectory in the initial total action trajectory that includes the most recent historical normal action posture;

[0053] Update the action trajectory after the initial action sub-trajectory in the initial total action trajectory to the target action sub-trajectory to form a target total action trajectory.

[0054] Optionally, the image generation module is specifically configured to generate a corresponding target action image according to the target action sub-trajectory, so as to obtain a digital human video including the initial action image corresponding to the initial action sub-trajectory and the target action image corresponding to the target action sub-trajectory.

[0055] Optionally, the image generation module is specifically configured to generate all target action images of the digital human model according to the target total action trajectory.

[0056] Optionally, it further includes an output module for notifying the digital human model that the action posture of the current frame is normal, and the current frame is used as a video composition frame output to the user of the digital human model.

[0057] Optionally, it further includes an output module for notifying the digital human model that the action posture of the current frame is abnormal, and the current frame is used as a video composition frame output to the user of the digital human model.

[0058] A third aspect of the embodiments of the present application provides an abnormal detection device for a digital human video, including:

[0059] A central processing unit, a memory, and an input / output interface;

[0060] The memory is a transient storage memory or a persistent storage memory;

[0061] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.

[0062] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, including instructions, which when run on a computer, cause the computer to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.

[0063] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:

[0064] The abnormal detection method for digital human videos in the embodiments of this application includes: obtaining initial action images of a digital human model at multiple time points; after detecting an abnormal action posture in the current frame, adjusting the initial total action trajectory to obtain a target total action trajectory, and obtaining all target action images of the digital human model according to the target total action trajectory, so as to avoid generating digital human videos with poor visual effects and affecting the broadcast image and user experience of the digital human model. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1A FIG. is a schematic diagram of an application environment for an embodiment of this application;

[0066] Figure 1B FIG. is an overall architecture diagram of the abnormal detection method for digital human videos in the embodiments of this application;

[0067] Figure 2 FIG. is a schematic flowchart of the abnormal detection method for digital human videos in the embodiments of this application;

[0068] Figure 3 FIG. is another schematic flowchart of the abnormal detection method for digital human videos in the embodiments of this application;

[0069] Figure 4 FIG. is a schematic structural diagram of the abnormal detection device for digital human videos in the embodiments of this application;

[0070] Figure 5 FIG. is another schematic structural diagram of the abnormal detection device for digital human videos in the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] Please refer to Figure 1A , Figure 1A which shows a schematic diagram of an application environment applicable to the embodiments of this application. The abnormal detection method for digital human videos provided in the embodiments of this application can be applied to an interactive system 100 as shown in Figure 1A . The interactive system 100 includes a terminal device 101 and a server 102, and the server 102 is communicatively connected to the terminal device 101. Among them, the server 102 can be a traditional server or a cloud server, and no specific limitation is made here.

[0072] Among them, the terminal device 101 can be various electronic devices with a display screen and supporting data input, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and wearable electronic devices, etc. Specifically, the data input can be voice input based on the voice module on the terminal device 101, character input based on the character input module, image input based on the image input module, etc., or it can also be based on the gesture recognition module installed on the terminal device 101, enabling the user to implement interaction methods such as gesture input.

[0073] Among them, a client application can be installed on the terminal device 101. A user can communicate with the server 102 based on the client application (such as an APP, a WeChat mini-program, etc.). Specifically, a corresponding server application is installed on the server 102. A user can register a user account with the server 102 based on the client application and communicate with the server 102 based on this user account. For example, the user logs in to the user account in the client application and can input through the client application based on this user account. The input can be text information, voice information, image information, etc. After the client application receives the information input by the user, it can send this information to the server 102, so that the server 102 can receive this information, process it, and store it. The server 102 can also receive this information and return a corresponding output information to the terminal device 101 according to this information.

[0074] In some embodiments, the client application can be used to provide customer service to the user and communicate with the user for customer service. The client application can interact with the user based on a virtual robot. Specifically, the client application can receive the information input by the user and make a response to this information based on the virtual robot. Among them, the virtual robot is a software program based on visual graphics. After this software program is executed, it can present a robot form that simulates biological behaviors or thoughts to the user. The virtual robot can be a robot that simulates a real person, such as a robot similar to a real person established according to the form of the user himself or other people, or a robot with an anime effect, such as a robot in the form of an animal or a cartoon character.

[0075] In some embodiments, after the terminal device 101 obtains a reply message corresponding to the information input by the user, it can display a virtual robot image corresponding to this reply message on the display screen of the terminal device 101 or other image output devices connected thereto. As a way, while playing the virtual robot image, the corresponding audio of the virtual robot image can be played through the speaker of the terminal device 101 or other audio output devices connected thereto, and the text or graphics corresponding to this reply message can also be displayed on the display screen of the terminal device 101, realizing multi-modal interaction with the user in multiple aspects such as image, voice, and text.

[0076] In some embodiments, the device for processing the information input by the user can also be set on the terminal device 101, so that the terminal device 101 can realize interaction with the user without relying on establishing communication with the server 102. At this time, the interaction system 100 can only include the terminal device 101.

[0077] The above application environment is only an example for easy understanding. It can be understood that the embodiments of the present application are not limited to the above application environment.

[0078] Please refer to Figure 1B , after inputting the preset broadcast text and action name (such as waving) that the digital human model needs to respond to, the key points of the digital human model's mouth shape (the key points can be understood as the characteristic coordinate points regarding the mouth shape) can be obtained through the text-to-speech TTS (TTS, text to speech) software and the mouth shape parameter model. Through the action generation module (which can also be understood as the action processing module), the action parameters of the digital human model, that is, the action key points (which can be understood as the characteristic coordinate points regarding the limbs), can be obtained; the full set of key points composed of the mouth shape key points and the action key points can form image frames representing the action postures of the digital human model at different time points through the image generation model (which can be loaded in the image generation module), and multiple frames of images can finally form the audio-visual stream of the digital human conversation system through the video encoder, enabling the digital human to interact with the user in the form of video, such as being applied to the digital human customer service scenario; among them, the use of the anomaly detection module can detect whether each generated frame of image is abnormal, that is, it can be used to detect whether the action posture of the digital human is abnormal. The corresponding manifestations of anomalies include blemish conditions such as black spots or mosaics in the limb movements on the image. It should be noted that the area outside the digital human's body in the image is transparent, and the generation of image blemishes is mainly caused by limb movements, while the mouth shape generally does not produce image blemishes. Therefore, this application mainly takes the anomaly detection and processing of the digital human's limb movements as an example for illustration. In this application, this type of image with normal action postures in the image or simply referred to as a normal frame, and correspondingly, this type of image with abnormal action postures or simply referred to as an abnormal frame.

[0079] Please refer to Figure 1B and Figure 2 , an embodiment of an anomaly detection method for digital human videos provided by the first aspect of this application includes:

[0080] 201. Obtain the initial action images of the digital human model at multiple time points.

[0081] Obtain the initial action images of the digital human model at multiple time points. Among them, since the initial action images are generated from the action posture key points output by the action generation module, obtaining the initial action images here can also be understood as including obtaining the action posture key points. Each time point corresponds to one frame of initial action image, and each frame of initial action image is used to represent the action posture of the digital human model at different time points. The action postures of all frames of initial action images form the initial action total trajectory.

[0082] 202. Detect whether the action posture of the current frame is abnormal.

[0083] To detect whether the action pose of the current frame is abnormal, it can be understood that it is to detect whether the action pose of the digital human in the image shows image defects such as black spots and / or mosaics. Here, the current frame is any initial action image; in practical applications, each initial action image should be subjected to abnormal detection. Since the action poses in the video playback are carried out sequentially in time, therefore, by way of example, the detection results of each frame should be output sequentially, that is, when performing abnormal detection on a certain frame, the detection result of the previous frame is known, so as to perform subsequent abnormal action adjustment processing.

[0084] 203. Perform abnormal action pose adjustment processing.

[0085] When it is detected that the action pose of the current frame is abnormal, abnormal action pose adjustment processing needs to be performed. The specific operations of this processing can include adjusting the total initial action trajectory to obtain the target total action trajectory, so as to avoid the action pose of the digital human finally presented to the user being incoherent or unclear.

[0086] 204. Generate the target action image.

[0087] Generate the target action image of the digital human model according to the target total action trajectory. Here, the target action image will be used as the video composition frame of the digital human model output to the user.

[0088] Please refer to Figure 1B and Figure 3 , another embodiment of an abnormal detection method for a digital human video provided by the present application includes:

[0089] 301. Obtain the initial action images of the digital human model at multiple time points.

[0090] 302. Detect whether the action pose of the current frame is abnormal.

[0091] After determining the total probability that the pixels corresponding to the action pose in the current frame have defects, by detecting whether the total probability exceeds the preset probability, it can be judged whether the action pose of the current frame is abnormal, where the total probability is used to represent the likelihood of the action pose of the current frame being abnormal.

[0092] In a specific embodiment, determining the total probability that the pixels corresponding to the action pose in the current frame have defects can specifically be any one of the following operations:

[0093] (1) Divide the current frame into N pixel sub-regions; detect the sub-probability that the pixels corresponding to the action pose in each pixel sub-region have defects through a classification model; statistically calculate the total probability that the pixels corresponding to the action pose in the current frame have defects. Here, the method of directly detecting the abnormal probability by dividing the region of the current frame is also applicable to images such as the first frame and the last frame that lack adjacent frames on one side.

[0094] (2) Divide each initial action image in the current frame, the initial action images of the previous m frames, and the initial action images of the subsequent m frames into N pixel sub-regions respectively; detect the sub-probability that the pixels corresponding to the action postures in each pixel sub-region have defects through a classification model; and statistically calculate the total probability that the pixels corresponding to the action postures in the current frame have defects. When detecting the abnormal probability by dividing regions for the current frame here, the detection regions of the initial action images of the previous m frames and the subsequent m frames are also detected (the detection window is m + 1 + m frames), which is beneficial to reflecting a more accurate detection effect in the classification model. It can be understood that using the previous m frames and the subsequent m frames as reference frames can make the detection of the current frame more accurate.

[0095] In practical applications, the total probability includes one or more probabilities among the image black spot probability and the mosaic probability. It can be understood that the probability of the existence of black spots and / or mosaics can be detected for each pixel sub-region; the classification model can specifically be an SVM or KNN classification model.

[0096] 303. Perform abnormal action posture adjustment processing.

[0097] The abnormal action posture adjustment processing performed can be any of the following operations:

[0098] (303.1) Notify the action generation module that the current frame has defects, and the action should be immediately returned to the safe posture range and a new safe action path should be re-planned: In a specific embodiment, when it is detected that the action posture of the current frame is abnormal (for example, the total probability exceeds the preset probability), the initial action total trajectory should be adjusted to obtain the target action total trajectory, including: determining the nearest historical normal action posture of the current frame, where the nearest historical normal action posture represents the initial action image of a historical frame before the current frame; generating the target action sub-trajectory in the reverse direction along the initial action sub-trajectory, and the initial action sub-trajectory is the historical action trajectory in the initial action total trajectory that includes the nearest historical normal action posture; updating the action trajectory after the initial action sub-trajectory in the initial action total trajectory to the target action sub-trajectory to form the target action total trajectory.

[0099] Exemplarily, when the generated waving angle exceeds a certain safe posture range (the safe posture can be understood as the action postures or action trajectories that already exist in the training data), correspondingly, in a certain frame of the initial action image, a mosaic will appear in the hand action area of the digital human. (The previous frames are normal frames because they are within the safe posture range, that is, the action postures are normal.) Then, it is possible to move backward along the trajectory of waving to this angle to find the nearest historical frame without a mosaic (the nearest normal historical frame), which can be understood as finding the nearest safe point in the backward trajectory, and re-planning a safe action path - correspondingly manifested as retracting the arm (which can be achieved by notifying the action generation module to regenerate the subsequent action key points, or notifying the action generation module to reverse-adjust the nearest normal historical frame and its previous action key point sequence), so that the action amplitudes of the digital human waving forward and backward are within the safe posture range.

[0100] (303.2) Notify (the digital human conversation system) of the existence of abnormal image frames: In a specific embodiment, when it is detected that the action posture of the current frame is abnormal (for example, the total probability exceeds the preset probability), notify the digital human model that the action posture of the current frame is abnormal, and the current frame can also be used as a video composition frame of the digital human model and output to the user; this is because, generally, a 1-second video corresponds to dozens of frames of images, which makes the appearance time of each frame of image very short, and it is difficult for the human eye to feel such rapid changes, that is, it is difficult to perceive that there are abnormalities in the video frames, so images with abnormalities can also be output; however, in this case, generally, it is also necessary to ensure that subsequent action frames will not have defects before output, otherwise it is likely to affect the visual experience.

[0101] 304. Generate target action images.

[0102] In a specific embodiment, generate the target action images of the digital human model according to the total target action trajectory. Specifically, it can be any of the following operations:

[0103] (1) Generate corresponding target action images according to the target action sub-trajectories to obtain a digital human video including the initial action images corresponding to the initial action sub-trajectories and the target action images corresponding to the target action sub-trajectories. Here, it can be understood that the final output digital human video is composed of: the normal initial action images in the front part plus the normally updated target action images in the back part (simply referred to as normal frames), that is, the initial normal images in the front part are retained for use, only the subsequent action images are generated, and the subsequently generated normal frames will replace those abnormal images corresponding to the back part of the initial action total trajectory and be output as subsequent action video frames.

[0104] (2) Generate all the target action images of the digital human model corresponding to the total target action trajectory. This video frame processing method can be understood as directly regenerating a whole new segment of video frames.

[0105] 305. Perform normal action posture processing.

[0106] When it is detected that the action posture of the current frame is normal (for example, the total probability does not exceed the preset probability), notify the digital human model that the action posture of the current frame is normal, so that the current frame can be used as a video composition frame of the digital human model and output to the user.

[0107] Steps 301 to 304 are respectively similar to steps 201 to 204, and will not be elaborated here specifically.

[0108] Please refer to Figure 4 , the second aspect of the embodiments of the present application provides an abnormal detection device for digital human videos, including:

[0109] An acquisition module 401, configured to acquire initial action images of the digital human model at multiple time points, where each time point corresponds to a frame of initial action image, and each frame of initial action image is used to represent the action posture of the digital human model at different time points, and the action postures of all frames of initial action images form an initial action total trajectory;

[0110] An abnormal detection module 402, configured to detect whether the action posture of the current frame is abnormal, where the current frame is any frame of initial action image;

[0111] An action processing module 403, configured to adjust the initial action total trajectory to obtain a target action total trajectory;

[0112] An image generation module 404, configured to generate a target action image of the digital human model according to the target action total trajectory, and the target action image is used as a video composition frame of the digital human model and output to the user.

[0113] In the embodiments of the present application, the operations performed by each module of the abnormal detection device for digital human videos are similar to those described in the first aspect or any specific method embodiment of the first aspect), and will not be elaborated here specifically.

[0114] Please refer to Figure 5 , the abnormal detection device 500 for digital human videos in the embodiments of the present application may include one or more central processing units CPU (CPU, central processing units) 501 and a memory 505, and one or more application programs or data are stored in the memory 505.

[0115] Among them, the memory 505 can be volatile storage or persistent storage. The programs stored in the memory 505 may include one or more modules, and each module may include a series of instruction operations in the abnormal detection device for digital human videos. Further, the central processing unit 501 can be set to communicate with the memory 505 and execute a series of instruction operations in the memory 505 on the abnormal detection device 500 for digital human videos.

[0116] The abnormal detection device 500 for digital human videos may further include one or more power supplies 502, one or more wired or wireless network interfaces 503, one or more input / output interfaces 504, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0117] The central processing unit 501 can perform the operations executed in the foregoing first aspect or any specific method embodiment of the first aspect, which will not be elaborated herein.

[0118] It can be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the steps do not mean the order of execution. The order of execution of each step should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0119] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.

[0120] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system or device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0121] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0123] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

Claims

1. An abnormal detection method for digital human videos, characterized in that, it includes: Obtain the initial action images of the digital human model at multiple time points, where each time point corresponds to one frame of the initial action image, and each frame of the initial action image is used to represent the action posture of the digital human model at different time points. The action postures of all frames of the initial action images form the initial action total trajectory; Detect whether the action posture of the current frame is abnormal, where the current frame is any one of the initial action images; If so, adjust the initial action total trajectory to obtain the target action total trajectory, and generate the target action images of the digital human model according to the target action total trajectory. The target action images are used as the video composition frames of the digital human model output to the user.

2. The abnormal detection method for digital human videos according to claim 1, characterized in that, The detection of whether the action posture of the current frame is abnormal includes: Determine the total probability that the pixels corresponding to the action posture in the current frame have defects, and the total probability is used to represent the possibility of the action posture in the current frame being abnormal; Detect whether the total probability exceeds the preset probability.

3. The abnormal detection method for digital human videos according to claim 2, characterized in that, The determination of the total probability that the pixels corresponding to the action posture in the current frame have defects includes: Divide the current frame into N pixel sub-regions; Detect the sub-probability that the pixels corresponding to the action posture in each pixel sub-region have defects through a classification model; Statistically calculate the total probability that the pixels corresponding to the action posture in the current frame have defects according to all the sub-probabilities.

4. The abnormal detection method for digital human videos according to claim 2, characterized in that, The determination of the total probability that the pixels corresponding to the action posture in the current frame have defects includes: Divide each frame of the initial action image in the current frame, the initial action images of the previous m frames, and the initial action images of the next m frames into N pixel sub-regions respectively; the detection window composed of the current frame, the previous m frames, and the next m frames is m + 1 + m frames; Detect the sub-probability that the pixels corresponding to the action posture in each pixel sub-region have defects through a classification model respectively; Statistically calculate the total probability that the pixels corresponding to the action posture in the current frame have defects according to all the sub-probabilities.

5. The abnormal detection method for digital human videos according to claim 1, characterized in that, The adjustment of the initial action total trajectory to obtain the target action total trajectory includes: Determine the nearest historical normal action posture of the current frame, where the nearest historical normal action posture represents the initial action image of a historical frame before the current frame; Generate the target action sub-trajectory in the reverse direction along the initial action sub-trajectory, and the initial action sub-trajectory is the historical action trajectory in the initial action total trajectory that includes the nearest historical normal action posture; Update the action trajectory after the initial action sub-trajectory in the initial action total trajectory to the target action sub-trajectory to form the target action total trajectory.

6. The abnormal detection method for digital human videos according to claim 5, characterized in that, the generation of the target action images of the digital human model according to the total target action trajectory includes: generating corresponding target action images according to the target action sub-trajectories, so as to obtain a digital human video including the initial action images corresponding to the initial action sub-trajectories and the target action images corresponding to the target action sub-trajectories.

7. The abnormal detection method for digital human videos according to claim 1 or 5, characterized in that, the generation of the target action images of the digital human model according to the total target action trajectory includes: generating all the target action images of the digital human model corresponding to the total target action trajectory.

8. An abnormal detection device for digital human videos, characterized in that, it includes: an acquisition module, configured to acquire the initial action images of the digital human model at multiple time points, wherein each time point corresponds to one frame of the initial action image, and each frame of the initial action image is used to represent the action posture of the digital human model at different time points, and the action postures of all frames of the initial action images form the total initial action trajectory; an abnormal detection module, configured to detect whether the action posture of the current frame is abnormal, and the current frame is any one of the initial action images; an action processing module, configured to adjust the total initial action trajectory to obtain the total target action trajectory; an image generation module, configured to generate the target action images of the digital human model according to the total target action trajectory, and the target action images are used as the video composition frames output to the user for the digital human model.

9. An abnormal detection device for digital human videos, characterized in that, it includes: a central processing unit, a memory and an input / output interface; the memory is a transient storage memory or a persistent storage memory; the central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, it includes instructions, and when the instructions run on a computer, the computer is made to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform

    US20180204111A1

  • System for interactive organization and browsing of video

    US6278446B1