An artificial intelligence-based video gesture recognition method and system
Through artificial intelligence to identify the overlapping time periods of faces and hands in the video, obtain gesture occlusion feature values and detect the acceleration of pixel changes in real time, it solves the impact of unconscious gestures on user experience and achieves higher video processing accuracy and user experience.
Patent Information
- Application Number
- CN202510427908.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In the prior art, unconscious gestures obscuring the face during video calls or live broadcasts affect the user experience and make video processing difficult.
Through artificial intelligence models, the overlapping time periods of faces and hands in historical videos are identified, the gesture occlusion feature values are obtained, and the acceleration of pixel changes is detected in real time to determine the intention of the gesture. Unobstructed face images are used for coverage to improve the accuracy of video processing.
It improves the accuracy of user intention recognition during video, enhances user experience, and avoids facial occlusion problems caused by unconscious gestures.
Smart Images

Figure CN120299087B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video gesture recognition, and particularly relates to a video gesture recognition method and system based on artificial intelligence. BACKGROUND
[0002] During a video call or a video live broadcast, a user may make gestures or have some unintentional hand waving actions. Some of the gestures are intended by the user, and some are unintentional. These gestures may cause occlusion of the face in the video. Therefore, the intention of the gesture needs to be recognized to facilitate subsequent video processing and improve the user's video experience. SUMMARY
[0003] The present application provides a video gesture recognition method based on artificial intelligence, which is used to solve the problem of unintentional gestures affecting user video experience in the prior art.
[0004] The first aspect of the present application provides a video gesture recognition method based on artificial intelligence, comprising:
[0005] Obtaining a user's historical video, identifying the overlap of a face and a hand in the historical video by an artificial intelligence model, and the time period during which the hand stays on the face;
[0006] Setting a first time interval based on the earliest time point of the overlap time period, obtaining the change relationship of the number of overlapping pixel points of the hand and the face with time in the video image under the first time interval, identifying the pixel change characteristics in the change relationship, and obtaining the gesture occlusion feature value of the user;
[0007] Real-time recognition of the face position in the video, and detection of the hand position, but when the hand in the video exists and the predicted trajectory of the hand movement overlaps the face, the acceleration of the pixel point change of the hand movement is detected; when the pixel point change acceleration does not satisfy the gesture occlusion feature value, the user's gesture is determined to be a non-occlusion gesture.
[0008] Optionally, after the pixel point change acceleration does not satisfy the occlusion gesture occlusion feature value, the user's gesture is determined to be a non-occlusion gesture, the method further comprises:
[0009] Obtaining a user's non-occlusion face image; when there are overlapping pixel points between the user's hand and face in the video, the user's face position is recognized and the user's non-occlusion face image is overlaid.
[0010] Optionally, the method further comprises: setting a second time interval based on the latest time point of the overlap time period, obtaining the change relationship of the number of overlapping pixel points of the hand and the face with time in the video image under the second time interval, identifying the pixel change characteristics in the change relationship, and obtaining the gesture exposure feature value of the user;
[0011] The identifying the user face position and covering with the user non-occluded face image further comprises: when there is a pixel point change acceleration satisfying a gesture exposure characteristic value, removing the user gesture as a non-occluded gesture state, and when there is no overlapping pixel point between the user hand and the face in the video, removing the user non-occluded face image.
[0012] The second aspect of the present application provides a video gesture recognition system based on artificial intelligence, comprising:
[0013] A time period division module is configured to obtain a user historical video, identify a face and a hand overlap in the historical video by using an artificial intelligence model, and obtain a time period during which the hand stays on the face;
[0014] A feature recognition module is configured to set a first time interval based on an earliest time point of the overlap time period, obtain a change relationship of the number of overlapping pixel points between the hand and the face in the video image over time in the first time interval, identify pixel change characteristics in the change relationship, and obtain a gesture occlusion characteristic value of the user;
[0015] A gesture recognition module is configured to identify a face position in a video in real time, and detect a hand position, but when there is a hand in the video and a hand movement predicted trajectory overlaps with the face, detect a pixel point change acceleration of the hand movement; when the pixel point change acceleration does not satisfy the gesture occlusion characteristic value, determine that the user gesture is a non-occluded gesture.
[0016] Optionally, the gesture recognition module further comprises, after determining that the user gesture is a non-occluded gesture when the pixel point change acceleration does not satisfy the gesture occlusion characteristic value:
[0017] Obtain a user non-occluded face image; when there is an overlapping pixel point between the user hand and the face in the video, identify a user face position and cover with the user non-occluded face image.
[0018] Optionally, the feature recognition module further comprises: setting a second time interval based on a latest time point of the overlap time period, obtaining a change relationship of the number of overlapping pixel points between the hand and the face in the video image over time in the second time interval, identifying pixel change characteristics in the change relationship, and obtaining a gesture exposure characteristic value of the user;
[0019] The gesture recognition module further comprises, after identifying the user face position and covering with the user non-occluded face image: when there is a pixel point change acceleration satisfying a gesture exposure characteristic value, removing the user gesture as a non-occluded gesture state, and when there is no overlapping pixel point between the user hand and the face in the video, removing the user non-occluded face image.
[0020] The third aspect of the application provides a video gesture recognition method and device based on artificial intelligence, the device comprising a processor and a memory:
[0021] The memory is used to store program code and transmit the program code to the processor.
[0022] The processor is used to execute the video gesture recognition method based on artificial intelligence according to the instructions in the program code.
[0023] The fourth aspect of the application provides a computer readable storage medium for storing program code, the program code being used to execute the video gesture recognition method based on artificial intelligence according to any one of the first aspect of the application.
[0024] From the above technical solutions, the application has the following advantages: the position and quantity relationship of the hand part and the face part in the user history video in the image pixel points, the pixel point change characteristic value reflecting the user hand shielding face movement characteristic is recognized, and then the user's gesture intention can be judged according to the user's hand movement in the real-time video, so as to facilitate the subsequent processing of the video, improve the accuracy of user intention recognition in the video process, and increase the user video experience. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or the prior art description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0026] Figure 1 It is a flow chart of a video gesture recognition method based on artificial intelligence.
[0027] Figure 2 It is a structure diagram of a video gesture recognition system based on artificial intelligence. DETAILED DESCRIPTION
[0028] In order to make the application purpose, features and advantages of the application more obvious and easy to understand, the technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the embodiments of the application. Obviously, the following described embodiments are only some embodiments of the application, not all embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0029] The application provides an artificial intelligence-based video gesture recognition method, which is used to solve the problem of unconscious gestures affecting user video experience in the prior art.
[0030] Please refer to Figure 1 , Figure 1 The first flowchart of the artificial intelligence-based video gesture recognition method provided by the embodiment of the application.
[0031] S100, acquire user historical video, and identify the face and hand overlap in the historical video and the time period during which the hand stays on the face by using an artificial intelligence model;
[0032] It should be noted that when the user performs video call or live broadcast, the video image content in a period of time after the current video is started can be intercepted as the user historical video, or the video image of a previous completed and ended video call can be selected as the user historical video in the long-term video process of the user; the artificial intelligence model can identify the face and hand in the video image, and position the face and hand in the video to determine the time interval of the hand overlapping the face, the overlapping time period being the interval from the hand starting to block the face to the hand completely not blocking the hand, and the artificial intelligence model also needs to identify the overlapping state of the hand and the face, identify the time period during which the hand and the face are relatively stationary for more than a preset time length, and only keep the data of the hand blocking the face for a period of time, that is, this step identifies the video in which the user consciously blocks the face with the hand, and then obtains the face and hand overlapping time period; the artificial intelligence model in this embodiment can use a deep neural network, but this embodiment does not need to identify the face, only needs to position the face in the video image after identifying the facial features of the human body, and therefore a BP neural network or a CNN neural network model can also be used to identify the features of the face and hand to complete the positioning and overlapping time period identification.
[0033] S200, set a first time interval based on the earliest time point of the overlapping time period, acquire the change relationship between the number of overlapping pixel points of the hand part and the face part in the video image in the first time interval and time, identify the pixel change characteristics in the change relationship, and obtain the gesture blocking feature value of the user;
[0034] It should be noted that the earliest time point of the overlapping time period corresponds to the moment when the pixel points in the area where the hand is located and the pixel points in the area where the face is located start to overlap, and the first time interval is set at the earliest time point, for example, the first time interval is 0.4 seconds long, so that the earliest time point is at 0.2 seconds, and the first time interval is 0.2 seconds before and after the earliest time point. After setting the first time interval, the hand pixel points and face pixel points in the video image corresponding to the time interval are obtained. Since this embodiment is directed to the time period when the hand occludes the face, the change of the occluded face can be reflected according to the value of the number of pixel points corresponding to the face decreasing over time. The image refresh frame rate of the video call is generally 60 frames, and the number of pixel points corresponding to the face at the time point can be obtained 24 times in 0.4 seconds. The value of the number of face pixel points decreasing between adjacent frames corresponds to the area of the hand-face overlap being occluded, that is, the 24 discrete data can be connected to form a change relationship of the number of overlapping pixel points over time, which can be expressed in a function form. The change characteristics of the pixels can be extracted in the change relationship, for example, the earliest time period in the overlapping time period corresponds to the beginning of the hand occluding the face, and the user consciously wants to keep the hand in the face occlusion, so the acceleration of the hand movement should be deceleration. In the change relationship, the change rate value of the overlapping pixel points in the first time interval can be calculated, that is, the function derivative of the change relationship of the number of overlapping pixel points over time is used as the gesture occlusion feature value. The gesture occlusion feature value is negative in the deceleration state, that is, the number of pixel points changing in each frame is getting smaller and smaller. For example, in the case of hand translation without occluding the face, the user may accelerate or move the hand at a constant speed to avoid occlusion, and when the hand occludes the face, the user will decelerate. In this embodiment, the first time interval is set to be short, and the user faces the camera directly, so it can be considered that the user's face does not move.
[0035] S300, the face position in the video is recognized in real time, and the hand position is detected, but when the hand in the video moves and the predicted trajectory of the hand overlaps the face, the acceleration of the change of the pixel points of the hand movement is detected; when the pixel point change acceleration does not satisfy the gesture occlusion feature value, the user's gesture is determined to be a non-occlusion gesture.
[0036] It should be noted that in the video call or live broadcast in real time, the AI artificial intelligence is used to identify the video picture in real time. Generally, the user's face in the video and live broadcast will remain in the picture, and the position of the face area can be determined in real time. The hand is often outside the picture, and the state of entering the video picture is often moving. The position of the hand in the video is also detected in real time. When the hand is recognized and the hand moves, the hand movement prediction trajectory can be obtained based on the preset AI hand movement prediction model. The AI hand movement prediction model can be a hidden Markov model or an autoregressive conditional diffusion model. When the hand movement prediction trajectory overlaps with the real-time face position detected at present, it can be considered that the hand may block the face in the future. At this time, the acceleration of the change of the pixel points in the hand region is detected, for example, the hand part occupies a pixel point a in a certain frame of video image, and the hand part occupies a pixel point b in the next frame of video image. In the b pixel points of the next frame, c pixel points have the same position as the a pixel points of the previous frame. Because the hand moving process is not necessarily a translation facing the camera, it can also have a flip situation or a movement situation facing the camera, resulting in a change in the area of the hand region in the image. Therefore, the number of pixel points a and b is not necessarily the same. However, because the interval time between adjacent frames is short, the acceleration of the change of the area of the hand in the image is slow, and a and b are different, so the total number of pixel points of the earlier frame or the later frame can be selected as the reference of the change of the pixel points, and the change of the pixel points is obtained by subtracting the value of c. The change of the pixel points is calculated according to the time of the adjacent frames.
[0037] The pixel point change acceleration is compared with the gesture blocking feature value calculated in the foregoing step. When the pixel point change acceleration is less than the gesture blocking feature value, it indicates that the user's hand movement does not intend to slow down and stop at the face, and it can be judged that the hand movement at this time is a non-blocking gesture, and the user's purpose is not to block the face with the hand.
[0038] In this embodiment, the pixel point change feature value reflecting the user's hand blocking face movement characteristics is identified through the position and number relationship of the pixel points of the hand and face in the image in the user's historical video. Then, the user's gesture intention can be judged according to the user's hand movement in the real-time video, so as to facilitate the subsequent processing of the video and improve the accuracy of the user's intention recognition in the video process, thereby increasing the user's video experience.
[0039] The above is a detailed description of the first embodiment of the video gesture recognition method based on artificial intelligence provided by the present application. The following is a detailed description of the second embodiment of the video gesture recognition method based on artificial intelligence provided by the present application.
[0040] In the embodiment, a video gesture recognition method based on artificial intelligence is further provided. After the step S300, when the pixel point change acceleration does not satisfy the blocking gesture blocking characteristic value, it is judged that the user gesture is a non-blocking gesture, and the method further comprises the following steps.
[0041] obtaining a user non-blocking face image; when there are overlapping pixel points between the user's hand and face in the video, recognizing the user's face position and covering the user's non-blocking face image;
[0042] It should be noted that the user's non-blocking face image can be obtained in the real-time video image, or an image uploaded by the user in advance. When it is judged that the user gesture is a non-blocking gesture, it can be considered that the user does not want the face to be blocked in the case of subsequent hand blocking of the face, and therefore the user's non-blocking face image can be used to replace the blocked face area, so that the user watching the video or live broadcast sees the non-blocking face, and the best beauty or special effect is retained, and sudden loss of beauty is avoided due to hand blocking.
[0043] In the foregoing step S200, the method further comprises the following steps: setting a second time interval based on the latest time point of the overlapping time period, obtaining the change relationship between the number of overlapping pixel points between the hand pixel points and the face pixel points in the video image and time in the second time interval, recognizing the pixel change characteristics in the change relationship, and obtaining the hand gesture exposure characteristic value of the user;
[0044] After the step of recognizing the user's face position and covering the user's non-blocking face image, the method further comprises the following steps: when there are pixel point change accelerations that satisfy the gesture exposure characteristic value, the non-blocking gesture state of the user's hand gesture is removed, and when there are no overlapping pixel points between the user's hand and face in the video, the user's non-blocking face image is removed.
[0045] It should be noted that the second time interval is set in the same way as the first time interval, and the hand gesture exposure characteristic value of the user is calculated in the same way as the step S200. The pixel point number change acceleration and the user gesture line characteristic value are positive values, which reflect that the user's hand is moving faster and faster. When it is identified that there are pixel point change accelerations that satisfy the gesture exposure characteristic value, it indicates that the user's hand is accelerating to move away from a certain area. When there are no overlapping pixel points between the user's hand and face in the video, it indicates that the non-blocking gesture state of the user ends, and the non-blocking face image does not need to be covered, and the beauty special effect can be completely recognized and covered on the user's face.
[0046] The above is a detailed description of a video gesture recognition method based on artificial intelligence provided in the first aspect of the present application. The following is a detailed description of an embodiment of a video gesture recognition system based on artificial intelligence provided in the second aspect of the present application.
[0047] Referring to Figure 2 , Figure 2 It is a structure diagram of a video gesture recognition system based on artificial intelligence. The embodiment provides a video gesture recognition system based on artificial intelligence, which comprises:
[0048] The time period division module 10 is configured to acquire a user historical video, identify a face and a hand overlap in the historical video by using an artificial intelligence model, and obtain a time period during which the hand stays on the face;
[0049] The feature recognition module 20 is configured to set a first time interval based on an earliest time point of the overlap time period, acquire a change relationship of the number of overlapping pixel points between a hand pixel point and a face pixel point in a video image in the first time interval with time, identify pixel change characteristics in the change relationship, and obtain a gesture occlusion feature value of the user;
[0050] The gesture recognition module 30 is configured to identify a face position in a video in real time, and detect a hand position, but when the hand in the video exists and a hand movement predicted trajectory overlaps the face, detect a pixel point change acceleration of the hand movement; when the pixel point change acceleration does not satisfy the gesture occlusion feature value, determine that a user gesture is a non-occlusion gesture.
[0051] Further, in the gesture recognition module 30, when the pixel point change acceleration does not satisfy the occlusion gesture occlusion feature value, after determining that the user gesture is the non-occlusion gesture, the gesture recognition module 30 further comprises:
[0052] acquiring a user non-occlusion face image; when there are overlapping pixel points between a user hand and a face in the video, identifying a user face position and covering the user non-occlusion face image.
[0053] Further, in the feature recognition module 20, the feature recognition module 20 further comprises: setting a second time interval based on a latest time point of the overlap time period, acquiring a change relationship of the number of overlapping pixel points between a hand pixel point and a face pixel point in a video image in the second time interval with time, identifying pixel change characteristics in the change relationship, and obtaining a gesture exposure feature value of the user;
[0054] In the gesture recognition module 30, after identifying the user face position and covering the user non-occlusion face image, the gesture recognition module 30 further comprises: when there is a pixel point change acceleration satisfying the gesture exposure feature value, removing a non-occlusion gesture state of the user gesture, and when there are no overlapping pixel points between a user hand and a face in the video, removing the user non-occlusion face image.
[0055] The third aspect of the application further provides a video gesture recognition method and device based on artificial intelligence, comprising a processor and a memory: the memory is used for storing program code and transmitting the program code to the processor; and the processor is used for executing the above-mentioned video gesture recognition method based on artificial intelligence according to instructions in the program code.
[0056] The fourth aspect of the application provides a computer readable storage medium, characterized in that the computer readable storage medium is used for storing program code, and the program code is used for executing the above-mentioned video gesture recognition method based on artificial intelligence.
[0057] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-mentioned apparatus and device can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0058] In several embodiments provided in the application, it should be understood that the disclosed system, apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0059] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.
[0060] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0061] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0062] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A video gesture recognition method based on artificial intelligence, characterized in that include: Obtain historical user videos and use an AI model to identify the time periods in which faces and hands overlap and the hands remain on the face. A first time interval is set based on the earliest time point of the overlapping time period, and a change relationship between the number of overlapping pixel points of the hand pixels and the face pixels in the video image of the first time interval is obtained over time, and pixel change characteristics in the change relationship are identified to obtain a user's gesture occlusion feature value; a second time interval is set based on the latest time point of the overlapping time period, and a change relationship between the number of overlapping pixel points of the hand pixels and the face pixels in the video image of the second time interval is obtained over time, and pixel change characteristics in the change relationship are identified to obtain a user's gesture exposure feature value; Real-time recognition of the face position in the video and detection of the hand position. When there are hands in the video and the predicted hand movement trajectory coincides with the face, the pixel change acceleration of the hand movement is detected; when the pixel change acceleration does not meet the gesture occlusion feature value, the user gesture is judged to be a non-occluded gesture; the user's unobstructed face image is obtained; when there are overlapping pixels between the user's hand and face in the video, the user's face position is identified and covered with the user's unobstructed face image; when there are pixel change accelerations that meet the gesture exposure feature value, the non-occluded gesture state of the user gesture is removed, and when there are no overlapping pixels between the user's hand and face in the video, the user's unobstructed face image is removed.
2. A video gesture recognition system based on artificial intelligence, characterized in that: include: The time segmentation module is used to obtain the user's historical videos and use the artificial intelligence model to identify the time periods in which the face and hand overlap and the hand stays on the face in the historical videos; The feature recognition module is configured to set a first time interval based on the earliest time point of the overlapping time period, obtain a temporal relationship between the number of overlapping pixels between the hand pixels and the face pixels in the video image during the first time interval, identify pixel change characteristics in the change relationship, and obtain a user's gesture occlusion feature value; set a second time interval based on the latest time point of the overlapping time period, obtain a temporal relationship between the number of overlapping pixels between the hand pixels and the face pixels in the video image during the second time interval, identify pixel change characteristics in the change relationship, and obtain a user's gesture exposure feature value; The gesture recognition module is used to identify the face position in the video in real time and detect the hand position. When there is a hand in the video and the predicted hand movement trajectory coincides with the face, the pixel change acceleration of the hand movement is detected; when the pixel change acceleration does not meet the gesture occlusion feature value, the user gesture is judged to be a non-occluded gesture; an unobstructed face image of the user is obtained; when there are overlapping pixels between the user's hand and face in the video, the user's face position is identified and covered with the user's unobstructed face image; when there are pixel change accelerations that meet the gesture exposure feature value, the non-occluded gesture state of the user gesture is removed, and when there are no overlapping pixels between the user's hand and face in the video, the user's unobstructed face image is removed.
3. A video gesture recognition device based on artificial intelligence, characterized in that: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the artificial intelligence-based video gesture recognition method of claim 1 according to the instructions in the program code.
4. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the video gesture recognition method based on artificial intelligence according to claim 1.
Citation Information
Patent Citations
Vehicle-mounted intelligent gesture recognition system
CN105334960A
Facial Recognition System
JP3245447U