Mode switching method and device for video playing device, equipment and medium
By collecting and analyzing image frames of user objects and evaluating skeletal and age information, the problem of low accuracy in mode switching in existing technologies is solved, and more accurate mode switching for children is achieved.
Patent Information
- Application Number
- CN202111484622.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2041-12-07
AI Technical Summary
In existing technologies, during the mode switching process of video playback devices, facial recognition technology is limited by factors such as training samples, algorithm models, and occlusion of the user's face, resulting in insufficient accuracy in identifying whether the user is a child and thus low accuracy in mode switching.
By collecting image frames of user objects, performing image recognition and analysis, evaluating the user object's skeletal and age information, and using the skeletal and age information to determine identity attributes and whether to switch to child mode.
It improves the accuracy of user identity attribute judgment and further enhances the accuracy of mode switching for video playback devices, ensuring more accurate switching to children's mode.
Smart Images

Figure CN114222183B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent terminals, and provides a mode switching method and device for a video playing device, an apparatus, and a medium. BACKGROUND
[0002] With the development of intelligent terminal technology, the functions of intelligent terminals are becoming more and more rich. For example, many televisions can be switched to a child mode, and the television can limit and guide children to watch television healthily in the child mode. In the field, the television can automatically switch to the child mode when a child is identified.
[0003] In related technologies, the mode switching process can include that the television can identify the face of a user object through face recognition technology to determine whether the user is a child, and switch to the child mode if the user is a child. However, the face recognition technology is limited by training samples, algorithm models, and factors such as the face of the user object being blocked, and the accuracy of identifying whether the user is a child is not high enough, so that the mode switching process cannot be accurately and timely switched in many cases, resulting in a low accuracy of the above mode switching process. SUMMARY
[0004] The present application provides a mode switching method and device for a video playing device, an apparatus, and a medium, which can solve the problem of low accuracy of the mode switching process in related technologies. The technical solution is as follows:
[0005] In one aspect, a mode switching method for a video playing device is provided, which includes:
[0006] collecting at least one image frame corresponding to a user object currently watching the video playing device;
[0007] performing image recognition and analysis on the at least one image frame to evaluate bone information and age information of the user object;
[0008] judging an identity attribute of the user object according to the bone information and the age information of the user object, and determining whether to switch to a child mode according to the identity attribute.
[0009] In another aspect, a mode switching device for a video playing device is provided, which includes:
[0010] a collecting module configured to collect at least one image frame corresponding to a user object currently watching the video playing device;
[0011] an evaluating module configured to perform image recognition and analysis on the at least one image frame to evaluate bone information and age information of the user object;
[0012] The switching module is configured to determine an identity attribute of the user object according to the skeleton information and the age information of the user object, and determine whether to switch to the child mode according to the identity attribute.
[0013] In one possible implementation, the evaluation module comprises:
[0014] The human body frame determination unit is configured to determine a human body frame of the user object in the at least one image frame.
[0015] The key point determination unit is configured to determine three-dimensional coordinates of human body key points of the user object in a camera coordinate system based on the human body frame of the at least one image frame.
[0016] The prediction unit is configured to evaluate skeleton lengths of multiple skeletons of the user object according to the three-dimensional coordinates of the human body key points, identify a face frame of a face of the user object in the at least one image frame, and predict age information of the user object.
[0017] In one possible implementation, the key point determination unit is configured to, for each image frame, crop the image frame based on the human body frame to obtain a human body frame image, and predict two-dimensional image coordinates of human body key points in the image frame coordinate system of the image frame in the human body frame image; and determine three-dimensional relative coordinates of the human body key points and three-dimensional absolute coordinates of a human body trajectory based on the two-dimensional image coordinates of the human body key points of the at least one image frame through a three-dimensional human body posture model and a human body trajectory model.
[0018] The three-dimensional absolute coordinates of the human body trajectory refer to coordinates of a human body trajectory center point in the camera coordinate system; the three-dimensional relative coordinates of the human body key points refer to coordinates of the human body key points relative to the human body trajectory center point in the camera coordinate system; and the camera coordinate system is a three-dimensional space coordinate system with a camera that collects the image frame as a coordinate origin and an optical axis of the camera as a Z axis.
[0019] In one possible implementation, the key point determination unit is further configured to detect a human body frame of a predetermined human body part of at least one user object in the image frame; adjust a corresponding human body part included in the human body frame of the corresponding user object in the image frame based on two-dimensional image coordinates of human body key points of the at least one user object in a previously collected image frame; and crop the image frame based on the adjusted human body frame to obtain a human body frame image of the at least one user object.
[0020] In a possible implementation, the key point determination unit is further configured to determine three-dimensional absolute coordinates of the human body key points in a camera coordinate system based on the three-dimensional relative coordinates of the human body key points and three-dimensional absolute coordinates of the human body trajectory, and calculate a bone length of each bone based on the three-dimensional absolute coordinates of the human body key points at both ends of the bone.
[0021] In a possible implementation, the prediction unit is further configured to perform any one of the following:
[0022] For each image frame, determine a face frame corresponding to a head of the user object based on respective key points of the head of the user object in the image frame, and determine age information of the user object based on the face frame and a face age recognition model;
[0023] perform face detection on the image frame to obtain a face frame in the image frame, and determine age information of the user object based on the face frame and a face age recognition model.
[0024] In a possible implementation, the switching module is further configured to perform any one of the following:
[0025] predict height information of the user object based on the plurality of bone lengths, and determine an identity attribute of the user object based on the height information and the age information of the user object;
[0026] determine an identity attribute of the user object based on the age information of the user object and the plurality of bone lengths.
[0027] In a possible implementation, the switching module is further configured to perform any one of the following:
[0028] when the height information of the user object does not exceed a first threshold, or when the height information of the user object exceeds the first threshold and the age information of the user object does not exceed a second threshold, determine that the identity attribute of the user object is a child;
[0029] when the age information of the user object does not exceed the second threshold, or when the age information of the user object exceeds the second threshold and the height information of the user object does not exceed the first threshold, determine that the identity attribute of the user object is a child.
[0030] In a possible implementation, the at least one image frame includes at least two user objects; and the switching module is further configured to switch the video playing device to a child mode when an identity attribute of any user object of the at least two user objects is a child, and exit the child mode if an identity attribute of any user object of the at least two user objects is an adult when the video playing device is in the child mode.
[0031] In a possible implementation, the apparatus further includes:
[0032] a distance determination module configured to determine a distance between the user object and the video playing device based on the three-dimensional coordinates of the human body key points of the user object in the camera coordinate system;
[0033] a first reminding module configured to stop playing the video picture and display a first reminding message when the distance is not more than a target distance threshold, the first reminding message being used to remind that at least a target distance threshold is maintained with the video playing device.
[0034] In a possible implementation, the apparatus further includes:
[0035] a sitting posture determination module configured to determine a current sitting posture of the user object based on the three-dimensional coordinates of the human body key points of the user object in the camera coordinate system;
[0036] a second reminding module configured to stop playing the video picture and display a second reminding message when the current sitting posture does not conform to a standard sitting posture threshold, the second reminding message being used to remind to correct the current sitting posture to conform to the standard sitting posture threshold.
[0037] In another aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the above-described mode switching method for a video playing device.
[0038] In another aspect, a computer readable storage medium is provided, having a computer program stored thereon, the computer program being executed by a processor to implement the above-described mode switching method for a video playing device.
[0039] The technical scheme provided in the present application has the beneficial effects that:
[0040] The mode switching method for a video playing device provided in the embodiments of the present application acquires at least one image frame corresponding to a user object currently watching the video playing device; performs image recognition and analysis on the at least one image frame to evaluate bone information and age information of the user object; judges an identity attribute of the user object according to the bone information and the age information of the user object, and determines whether to switch to a child mode according to the identity attribute. The identity attribute is judged by using the age information and the bone information, so as to further determine whether to switch to the child mode, thereby improving the accuracy of judging the identity attribute of the user and further improving the accuracy of mode switching of the video playing device. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced.
[0042] Figure 1 An implementation environment schematic diagram of a mode switching method for a video playing device provided by an embodiment of the present application;
[0043] Figure 2a A flow schematic diagram of a mode switching method for a video playing device provided by an embodiment of the present application;
[0044] Figure 2b A flow schematic diagram of another mode switching method for a video playing device provided by an embodiment of the present application;
[0045] Figure 3 A structure schematic diagram of a mode switching device for a video playing device provided by an embodiment of the present application;
[0046] Figure 4 A structure schematic diagram of another mode switching device for a video playing device provided by an embodiment of the present application;
[0047] Figure 5 A structure schematic diagram of another mode switching device for a video playing device provided by an embodiment of the present application;
[0048] Figure 6 A structure schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0049] The embodiments of the present application will be described below in conjunction with the drawings in the present application. It should be understood that the implementation manners described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitation on the technical solutions of the embodiments of the present application.
[0050] Those skilled in the art can understand that the singular forms "a," "an," and "the" used herein include plural forms unless specifically stated otherwise. It should be further understood that the terms "comprise" and "include" used in the embodiments of the present application refer to the corresponding features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element are connected through an intermediate element. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" indicates implementation as "A", or implementation as "A", or implementation as "A and B".
[0051] For the purpose, technical solutions and advantages of the present application to be clearer, the embodiments of the present application will be described in further detail below with reference to the drawings.
[0052] Many children have developed the habit of watching TV, and even become TV addicts. Long-term watching TV is very harmful to the growth of children, not only causing damage to vision and spine, but also having adverse effects on children's psychology due to bad TV content. In order to limit and guide children to watch TV healthily, many manufacturers set a child mode in their products, which limits the time and content of children watching TV in the child mode.
[0053] The prior art can automatically switch the child mode of the device, simplifying the cumbersome operation of setting the child mode. Specifically, in the prior art, the face recognition technology is used to recognize the face image frame of the user object to identify the age of the user object, but due to the limitation of training samples, algorithm model and user object face shielding, the accuracy of face age recognition is not high enough.
[0054] This application addresses the shortcomings of existing technologies in terms of low recognition accuracy by providing a mode switching method for video playback devices. This method involves acquiring at least one image frame corresponding to a user currently viewing the video playback device; performing image recognition and analysis on the at least one image frame to evaluate and obtain the user's skeletal information and age information; determining the user's identity attribute based on the skeletal and age information; and determining whether to switch to child mode based on the identity attribute. Utilizing age and skeletal information to determine identity attributes, and further determining whether to switch to child mode, improves the accuracy of user identity attribute determination, thereby enhancing the accuracy of mode switching for video playback devices.
[0055] Figure 1 This is a schematic diagram illustrating the implementation environment of a mode switching method for a video playback device provided in an embodiment of this application. For example... Figure 1 As shown, the implementation environment includes a video playback device 101 and a server 102. A communication connection is established between the video playback device 101 and the server 102 via a network. The video playback device 101 can receive video streams sent by the server 102 through this communication connection and play the video streams to user objects. For example, the video playback device 101 is configured with a children's mode. Children's mode refers to a video playback mode specifically designed for children, which is beneficial for children's viewing. For example, children's mode can be configured with video streams specifically for children, and restrictions on children's usage time. In this application, the video playback device 101 can use captured image frames of user objects to identify the user object. When the user object is identified as a child, the video stream playback is switched to children's mode.
[0056] The server 102 can be an independent physical server, a server cluster or a distributed system formed by multiple physical servers, a cloud server or a server cluster providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The network can include but is not limited to a wired network including a local area network, a metropolitan area network, and a wide area network, and a wireless network including Bluetooth, Wi-Fi, and other wireless communication networks. The video playback device can be a smart television, a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a notebook computer, a digital broadcast receiver, a MID (Mobile Internet Device), a PDA (Personal Digital Assistant), a desktop computer, a vehicle terminal (such as a vehicle navigation terminal, a vehicle computer, etc.), a smart speaker, a smart watch, etc. The video playback device 101 and the server 102 can be directly or indirectly connected through wired or wireless communication, but are not limited thereto. The specific connection mode can be determined based on the actual application scenario, and is not limited herein.
[0057] Figure 2a A flowchart of a mode switching method for a video playback device is provided in an embodiment of the present application. The execution subject of the method can be a video playback device, which can be a smart television, a personal computer, or any computer device that can play a video stream. As shown in Figure 2a The method includes the following steps.
[0058] In step S1, the video playback device collects at least one image frame corresponding to a user object currently watching the video playback device.
[0059] The video playback device can collect image frames through a camera, which can be located on the video playback device or be an external camera of the video playback device. The present application does not limit this. In one possible implementation, when the camera is the camera of the video playback device, step S1 can be implemented through the following step 201.
[0060] In step S2, the video playback device performs image recognition and analysis on the at least one image frame to obtain bone information and age information of the user object.
[0061] The video playing device can detect a human body frame in which the user object in the image frame is located, evaluate the bone length of the user object based on the human body frame, and identify a face frame of the user object in the image frame, and predict the age information of the user object. In one possible implementation, step S2 can be implemented through the following steps 202 to 204.
[0062] Step S3, the video playing device determines the identity attribute of the user object according to the bone information and the age information of the user object, and determines whether to switch to the child mode according to the identity attribute.
[0063] The video playing device can further determine the height information according to the bone information, so as to determine the identity attribute based on the height information and the age information. Alternatively, the video playing device can also directly determine the identity attribute based on the bone length and the age information. In one possible implementation, step S3 can be implemented through the following steps 205 to 206.
[0064] The mode switching method for the video playing device provided by the embodiments of the present application comprises the following steps: collecting at least one image frame corresponding to a user object currently watching the video playing device; performing image recognition and analysis on the at least one image frame to obtain bone information and age information of the user object; determining an identity attribute of the user object according to the bone information and the age information of the user object, and determining whether to switch to a child mode according to the identity attribute. The age information and the bone information are used to determine the identity attribute, so as to further determine whether to switch to the child mode, thereby improving the accuracy of determining the identity attribute of the user and further improving the accuracy of mode switching of the video playing device.
[0065] Figure 2b Another flowchart of a mode switching method for a video playing device is provided by the embodiments of the present application. The execution subject of the method can be a video playing device, which can be a smart television, a personal computer or any computer device that can play a video stream. As shown in Figure 2b The method comprises the following steps.
[0066] Step 201, the video playing device collects at least one image frame through a camera of the video playing device.
[0067] Preferably, the video playing device can employ a camera to directly capture the image frames. Each of the captured image frames includes a video of a user object currently watching the video playing device. For example, the image frame can include a human body of the user object, i.e., the image frame can include each human body part of the user object. For example, the image frame can include a head, a neck, a torso and limbs of the user object. In one possible implementation, the plurality of image frames are a plurality of image frames captured by the camera at a target period. Further in this step, the video playing device can periodically capture the human body image frames of the user object at the target period to obtain the plurality of image frames.
[0068] In one possible example, the camera can be located at a central point above a screen of the video playing device. After the video playing device is powered on, the camera can be controlled to capture the full-body image frames of the user object watching the video information.
[0069] Step 202, the video playing device determines a human body bounding box of the user object in the plurality of image frames, and determines three-dimensional coordinates of human body key points of the user object in a camera coordinate system based on the human body bounding box of the at least one image frame.
[0070] It should be noted that the human body bounding box can be a box identifier covering the human body of the user object in the image frame. For example, the human body bounding box can cover the head, the neck, the torso and the limbs of the user object. The video playing device can identify the human body bounding box in each image frame, and determine the three-dimensional coordinates of the human body key points of the user object in the camera coordinate system according to the human body bounding boxes of the plurality of image frames. In one possible implementation, this step can be implemented through steps 2021-2023.
[0071] Step 2021, for each image frame, the video playing device crops the image frame based on the human body bounding box to obtain a human body bounding box image.
[0072] In one possible implementation, for each image frame, if there is no image frame captured before the capture time of the image frame (hereinafter referred to as “previously captured image frame”), and the image frame is the first image frame captured by the video playing device, this step can include that the video playing device can detect a human body bounding box of a predetermined human body part of at least one user object in the image frame, and crop the image frame along the human body bounding box of the at least one user object to obtain a human body bounding box image of the at least one user object.
[0073] In another possible implementation, if the image frame is preceded by a preceding captured image frame, the step can include: for each image frame, the video playing device detecting a human body bounding box of a predetermined human body part of at least one user object in the image frame, the human body bounding box including the predetermined human body part, such as the head, neck, torso and limbs of the user object; the video playing device adjusting a corresponding human body part included in the human body bounding box of the corresponding user in the image frame based on the two-dimensional image coordinates of the human body key points of the at least one user object in the preceding captured image frame; and the video playing device cropping the image frame based on the adjusted human body bounding box to obtain a human body bounding box image of the at least one user object. The human body key points of the user object in the preceding captured image frame can include the cervical joint and the pelvic joint, and the human body part in the human body bounding box in the image frame can be adjusted by the two key points. In one possible example, the process of adjusting the human body bounding box in the image frame based on the preceding captured image frame includes: for each user object, the video playing device identifying the cervical joint and the pelvic joint in the human body bounding box in the image frame according to the cervical joint and the pelvic joint of the user object in the preceding captured image frame, calculating the included angle between the line connecting the cervical joint and the pelvic joint and the vertical line, and rotating the human body in the image frame according to the included angle so that the torso of the rotated user object is aligned in an upward state.
[0074] In one possible example, for each image frame, the video playing device can perform human body detection on the image frame, and when detecting that the image frame includes a human body of a user object, the human body can be labeled according to the position of the human body in the image frame. When the image frame includes human bodies of multiple user objects, the video playing device can label each human body with a corresponding position identifier. The position identifier of each human body is used to identify the position of the human body in the image frame. For example, the image frame includes three user objects A, B and C, the human body of A is located at the left position in the image frame, the human body of B is located at the middle position in the image frame, and the human body of C is located at the right position in the image frame. The position identifier can be in the form of position ID (Identity, number), and the three user objects A, B and C can be distinguished by the corresponding position identifiers "left, middle and right", i.e., the user object at the left position in the previous image frame corresponds to the user object at the left position in the next image frame.
[0075] Exemplarily, the human body frame of the image frame can be adjusted based on the position identifier of the human body frame in the previously collected image frame. The process can include: the video playing device obtains the two-dimensional image coordinates of the neck joint and the pelvic joint of the same user object as the position identifier in the previously collected image frame according to the position identifier of each human body in the image frame, and identifies the neck joint and the pelvic joint of the user object in the image frame according to the two-dimensional image coordinates of the neck joint and the pelvic joint, so as to rotate the human body of the user object in the image frame based on the neck joint and the pelvic joint. Wherein, the two-dimensional image coordinates of the key points of the corresponding user object in the previously collected image frame refer to the two-dimensional image coordinates of the key points in the image coordinate system of the previously collected image frame. Of course, if the image frame is the first image frame collected by the video playing device, the position identifier of the image frame is directly identified, and the human body frame image is obtained by cropping the human body frame.
[0076] In another possible implementation, the image frame can also be detected for a human face first, and the human body frame is estimated based on the detected human face frame. The process can include: for each image frame, the video playing device can detect the human face of the image frame to obtain the human face frame of at least one user object in the image frame; and detect the human body frame of the image frame according to the human face frame; crop the image frame based on the human body frame of the at least one user object to obtain the human body frame image of the at least one user object. Exemplarily, the video playing device can scale the human face frame according to a target magnification factor until the scaled human face frame can include each human body part of the user object, obtain the human body frame, and then crop the image frame based on the human body frame to obtain the human body frame image.
[0077] Step 2022, the video playing device predicts the two-dimensional image coordinates of the human body key points in the image coordinate system of the image frame in the human body frame image.
[0078] For each image frame, the video playing device can detect the human body key points of the image frame through a human body key point prediction model. Wherein, the human body key point prediction model can be a two-dimensional human body key point prediction model constructed based on CNN (Convolutional Neural Networks, convolutional neural network), and the cropped human body frame image of the at least one user object is taken as the input, and the two-dimensional image coordinates of the human body key points of the corresponding user object in the image coordinate system are output. Wherein, the human body key point prediction model can be obtained by pre-training a sample set; for example, the sample set can include a plurality of sample image frames, and each sample image frame is labeled with the true value image coordinates of the human body key points of the user object in the image coordinate system of the sample image frame. The initial model is trained through the sample set to obtain the human body key point prediction model.
[0079] At step 2023, the video playing device determines the three-dimensional relative coordinates of the human body key points and the three-dimensional absolute coordinates of the human body trajectory according to the two-dimensional image coordinates of the human body key points in the at least one image frame, through the three-dimensional human body posture model and the human body trajectory model.
[0080] The plurality of image frames can include a currently captured image frame and a previously captured image frame. That is, for each image frame, the video playing device can use the image frame and the previously captured image frame to perform three-dimensional key point coordinate identification through the three-dimensional human body posture model and the human body trajectory model, to obtain the three-dimensional coordinates of the human body key points in the camera coordinate system.
[0081] The three-dimensional absolute coordinates of the human body trajectory refer to the coordinates of the human body trajectory center point in the camera coordinate system. The three-dimensional relative coordinates of the human body key points refer to the coordinates of the human body key points relative to the human body trajectory center point in the camera coordinate system. The camera coordinate system is a three-dimensional space coordinate system with the camera capturing the image frame as the coordinate origin and the optical axis of the camera as the Z axis. For example, the human body trajectory center point can be the hip joint, and the three-dimensional absolute coordinates of the human body trajectory can be the absolute coordinates of the hip joint in the camera coordinate system. The three-dimensional relative coordinates of the human body key points can be the relative coordinates of other human body key points relative to the hip joint in the camera coordinate system, and the other human body key points can be key points other than the hip joint among the human body key points. For example, the three-dimensional absolute coordinates of the hip joint can be (3, 4, 5), and the three-dimensional relative coordinates of the neck joint relative to the hip joint can be (0, 1, 0), indicating that the relative position coordinates of the neck joint relative to the hip joint in the Y axis direction of the camera coordinate system are 1, and the relative position coordinates in the X axis and Z axis directions of the camera coordinate system are 0, that is, the X axis and Z axis directions of the neck joint have the same coordinate position as the hip joint.
[0082] The three-dimensional human body posture model is used to determine the three-dimensional relative coordinates of the human body key points in the camera coordinate system according to the two-dimensional image coordinates of the human body key points, and the human body trajectory model is used to determine the three-dimensional absolute coordinates of the human body trajectory center point according to the two-dimensional image coordinates of the human body key points. The 3D (3-dimension) human body posture model and the human body trajectory model can be constructed based on a delated casual convolution. In this step, for each image frame, the two-dimensional image coordinates of the human body key points in a plurality of continuous image frames including the image frame can be used as the input of the three-dimensional human body posture model and the human body trajectory model, and the three-dimensional relative coordinates of the human body key points and the three-dimensional absolute coordinates of the human body trajectory are output. Of course, the three-dimensional human body posture model and the human body trajectory model can also be obtained by training based on a large number of sample sets.
[0083] In step 203, the video playback device evaluates the bone lengths of the plurality of bones of the user object according to the three-dimensional coordinates of the human key points.
[0084] The video playback device first determines the three-dimensional absolute coordinates of the human key points, and calculates the bone lengths based on the three-dimensional absolute coordinates of the human key points. In one possible implementation, this step can be implemented through steps 2031-2032.
[0085] In step 2031, the video playback device determines the three-dimensional absolute coordinates of the human key points in the camera coordinate system based on the three-dimensional relative coordinates of the human key points and the three-dimensional absolute coordinates of the human trajectory.
[0086] For each user object, the video playback device can sum the three-dimensional relative coordinates of each human key point of the user object and the three-dimensional absolute coordinates of the human trajectory to obtain the three-dimensional absolute coordinates of the human key point. For example, the three-dimensional absolute coordinates of the hip joint can be (3, 4, 5), and the three-dimensional relative coordinates of the neck joint relative to the hip joint can be (0, 1, 0). Thus, the three-dimensional absolute coordinates of the neck joint can be (3, 5, 5).
[0087] It should be noted that the human key points of the user object can be configured as needed. For example, 21 main key points including the head, neck, limbs, and torso can be obtained. The 21 main key points can include the human key points at both ends of the main bones of the human body, such as the key points at both ends of the lower leg, i.e., the knee joint and the ankle joint.
[0088] In step 2032, the video playback device calculates the bone length of each bone according to the three-dimensional absolute coordinates of the human key points at both ends of the bone.
[0089] The video playback device can take the distance between the three-dimensional absolute coordinates of the human key points at both ends of each bone as the bone length of the bone. For example, the distance between the three-dimensional absolute coordinates of the knee joint and the ankle joint is calculated as the length of the lower leg. Of course, the length of the upper leg can be calculated based on the three-dimensional absolute coordinates of the key points at both ends of the upper leg, the length of the torso can be calculated based on the three-dimensional absolute coordinates of the hip joint and the neck joint, and the length of the head can be calculated based on the three-dimensional absolute coordinates of the neck joint and the key point on the top of the head.
[0090] It should be noted that the three-dimensional coordinates of each key point of the human body are obtained, the three-dimensional spatial positions of each key point are accurately represented, the lengths of each bone of each part of the human body are further obtained based on each key point, and the height of the whole human body can be further calculated. Even if the user object is in any posture such as sitting, standing, half-lying, and lying, the lengths of the bones in the human body can be accurately calculated, and the actual height of the human body can be accurately obtained, thereby improving the accuracy of determining the bone length and height of the user object and further improving the accuracy of the child mode switching. Moreover, the current posture of the user object is not limited, and the child mode switching can be applied to various scenes such as sitting, standing, and lying, thereby further improving the practicability of the child mode switching.
[0091] In step 204, the video playing device identifies the face frame of the face of the user object in the plurality of image frames, and predicts the age information of the user object.
[0092] The video playing device can perform age recognition based on the face part of the user object.
[0093] The video playing device can perform face region recognition based on the two-dimensional image coordinates of the human body key points obtained in step 202 to obtain the face frame. Alternatively, face recognition can be directly performed on the image frames to identify the face frame in the image frames. In the first possible implementation, when the face frame is identified based on the two-dimensional image coordinates of the human body key points in step 202, the present step can include: for each image frame, the video playing device determines the face frame corresponding to the head of the user object based on the key points of the head of the user object in the image frame, and determines the age information of the user object based on the face frame and a face age recognition model. The face age recognition model is used to recognize the age information of the user object according to the face image of the user object. The video playing device can determine the region including the key points of the head of the human body as the face frame according to the two-dimensional image coordinates of the human body key points obtained in step 202; the key points of the head can include the key points of the top of the head, the key points of the chin, the key points of the ears, and the like, and can also include the key points of the facial features, and the like. For example, in the present step, the image frame region in the image frame that can cover the key points of the top of the head, the key points of the ears, and the key points of the chin can be determined as the face frame. In addition, the image frame can be cropped based on the face frame to obtain a face frame image. The face age recognition model can be a face age recognition model constructed based on a CNN network, and at least one face frame image of the user object after cropping is taken as input, and the age of the corresponding user object is output. Of course, the face age recognition model can be obtained by pre-training a sample set; for example, the sample set can include a plurality of sample face images, and each sample face image is labeled with an age true value label of the user object. The initial model is trained through the sample set to obtain the face age recognition model.
[0094] In a second possible implementation, when the face recognition is directly performed on the image frames to obtain the face frame, the step can include: for each image frame, the video playing device performs face detection on the image frame to obtain a face frame in the image frame, and determines age information of the user object based on the face frame and a face age recognition model. Wherein, the way of age recognition based on the face frame and the face age recognition model is the same as the process of recognizing by using the face age recognition model in the first possible implementation, which will not be repeated here.
[0095] Step 205, the video playing device determines the identity attribute of the user object according to the skeleton information and the age information of the user object.
[0096] In a possible implementation, the video playing device can predict the height information of the user object based on the plurality of skeleton lengths, and determine the identity attribute of the user object based on the height information and the age information of the user object. For example, the video playing device determines the sum value accumulated among the plurality of skeleton lengths as the height information of the human body. The video playing device can accumulate the length of the calf, the length of the thigh, and the length of the torso from the pelvis to the neck, and the length of the head from the neck to the top of the head, and determine the accumulated sum value as the height information of the human body.
[0097] In a possible example, the height information can be used to determine first, and then the age information can be used to further correct. The step can include: when the height information of the user object does not exceed a first threshold, or when the height information of the user object exceeds the first threshold and the age information of the user object does not exceed a second threshold, the video playing device determines that the identity attribute of the user object is a child. Wherein, the video playing device can determine whether the height information of the user object exceeds the first threshold. If the height information does not exceed the first threshold, the video playing device directly determines that the identity attribute of the user object is a child. If the height information exceeds the first threshold, the video playing device further determines whether the age information exceeds the second threshold. If the height information exceeds the first threshold but the age information does not exceed the second threshold, the video playing device determines that the identity attribute of the user object is a child. If the height information exceeds the first threshold and the age information exceeds the second threshold, the video playing device determines that the identity attribute of the user object is not a child, for example, the identity attribute of the user object is an adult.
[0098] In another possible example, the judgment can be made based on the age information first, and then further corrected in combination with the height information. The step can include: when the age information of the user object does not exceed a second threshold, or when the age information of the user object exceeds the second threshold and the height information of the user object does not exceed a first threshold, the video playing device determines that the identity attribute of the user object is a child. Wherein, the video playing device can judge whether the age information of the user object exceeds the second threshold, if the age information does not exceed the second threshold, the video playing device directly determines that the identity attribute of the user object is a child; if the age information exceeds the second threshold, the video playing device further judges whether the height information exceeds the first threshold; if the age information exceeds the second threshold but the height information does not exceed the first threshold, the video playing device determines that the identity attribute of the user object is a child; if the age information exceeds the second threshold and the height information exceeds the first threshold, the video playing device determines that the identity attribute of the user object is not a child, for example, the identity attribute of the user object is an adult.
[0099] In yet another possible example, a weighted score can be calculated based on the height information and the age information, and the identity attribute of the user object can be further identified. For example, a first weight corresponding to the height information and a second weight corresponding to the age information can be configured, and a height information score corresponding to each height information range and an age information score corresponding to each age information range can be configured; based on the first weight and the height information score, and the second weight and the age information score, a score of the identity attribute of the user object is determined, and when the score of the identity attribute is within a child threshold range, it is determined that the identity attribute of the user object is a child. For example, a first product value of the height information score and the first weight, and a second product value of the age information score and the second weight are calculated, and a sum value between the first product value and the second product value is taken as the score of the identity attribute. Of course, other combinations of height information and age information can be used to determine the identity attribute of the user object, and the present application only illustrates the above three ways as examples, but does not limit the specific combination of how the height information and the age information are combined to determine the identity attribute.
[0100] In yet another possible implementation, the video playback device can determine the identity attribute of the user object based on the age information of the user object and the plurality of bone lengths. In one possible example, the video playback device can first determine based on the bone lengths and then further correct based on the age information; the process can include: when the plurality of bone lengths meet a target condition, or when the plurality of bone lengths do not meet the target condition and the age information of the user object does not exceed a second threshold, the video playback device determines that the identity attribute of the user object is a child. Wherein the plurality of bone lengths meeting the target condition can include but is not limited to: each bone length does not exceed a third threshold, a target bone length in the plurality of bone lengths does not exceed a fourth threshold, a maximum bone length in the plurality of bone lengths does not exceed a fifth threshold, etc. The target bone length can be set as needed, and the embodiments of the present application do not make specific limitations thereto, for example, the target bone length can be calf bone length, forearm bone length, etc. In another possible example, the video playback device can first determine based on the age information and then further correct based on the bone lengths; the process can include: when the age information of the user object does not exceed the second threshold, or when the age information of the user object exceeds the second threshold and the plurality of bone lengths meet the target condition, the video playback device determines that the identity attribute of the user object is a child.
[0101] Of course, the video playback device can also further identify the identity attribute of the user object in combination with the age information and the bone lengths and respective corresponding weights, and the manner of judging the identity attribute in combination with the weights is the same as the manner of judging the identity attribute based on the height information and the age information as described above. Therefore, details are not repeated here.
[0102] Step 206, the video playback device determines whether to switch to a child mode according to the identity attribute.
[0103] When the identity attribute of the user object is a child, the video playback device switches the video playback device to a child mode.
[0104] The video playback device switches to the child mode and plays a video stream for the user object in the child mode. The content, the playing duration, the playing time period, etc. of the video stream in the child mode are all in the range corresponding to children, for example, the content of the video stream is related to children. Of course, the guardian of the child, for example, the parent of the child, can set the playing duration, for example, set the playing duration to not exceed 2 hours and the playing time period to be from 12:00 noon to 20:00 in the evening. Of course, the video playback device can also establish a communication connection with the mobile phone of the parent and send the child viewing data to the mobile phone of the parent in the child mode to inform the parent in real time through the mobile phone about the child viewing situation, for example, the viewing duration, the viewing content, etc. so that the child viewing process can be supervised by the guardian of the child.
[0105] In a possible implementation, the video playing device can not only switch to the child mode when needed, but also exit the child mode in time under appropriate circumstances. In one possible example, the step of switching the child mode by the video playing device can include: when the identity attribute of any of the at least two user objects is a child, the video playing device switches the video playing device to the child mode; and when the video playing device is in the child mode, if the identity attribute of any of the at least two user objects is an adult, the video playing device exits the child mode. When the at least two user objects only include children, that is, there is no adult accompanying the children to watch, the child mode can be directly switched. When the at least two user objects include an adult, that is, when an adult accompanies the children to watch, the child mode can not be used to restrict the children's watching, and the child mode can be exited at this time. In one possible example, the duration of the adult accompanying the watching can also be used to determine whether to exit the child mode. When it is detected that the identity attribute of any of the plurality of user objects is an adult, the step of the video playing device exiting the child mode can include: when the identity attribute of any of the at least two user objects is an adult, the video playing device detects the continuous watching time of the adult based on the plurality of image frames, and when the continuous watching time exceeds a target time threshold, the video playing device exits the child mode. For example, the target time threshold can be 3 minutes, 20 minutes, etc. When the continuous watching time does not exceed the target time threshold, the video playing device continues to maintain the child mode. The video playing device can calculate the continuous watching time based on the first appearance time of the adult in the plurality of image frames and the time stamps of the plurality of image frames including the adult. For example, an image frame is collected every 3 seconds, the adult first appears in the 15th frame among the currently collected 115 image frames, and the adult is included from the 15th image frame to the 115th image frame, so the continuous watching time is 300 seconds, which exceeds the target time threshold of 3 minutes, and the video playing device can exit the child mode.
[0106] In yet another possible implementation, the video playback device can also monitor some behavior habits of the child when watching and remind, and after the video playback device is switched to the child mode, the process of monitoring the child by the video playback device can include: the video playback device determines a distance between the user object and the video playback device based on the three-dimensional coordinates of the human body key points of the user object in the camera coordinate system; when the distance is less than a target distance threshold, the video playback device stops playing the video picture and displays a first reminding message, and the first reminding message is used to remind to keep at least the target distance threshold from the video playback device. When the distance between the child and the video playback device is close, the video playback device can help the child to watch the video stream at a far distance by displaying the first reminding message, so as to further reduce the influence of watching television on the eyesight of the child. Wherein, the Z-axis direction in the camera coordinate system can be the direction of the camera optical axis, and the distance between the user object and the video playback device can be represented by the coordinate of the Z-axis, for example, the distance between the user object and the video playback device can be the coordinate of the hip joint in the Z-axis direction in the camera coordinate system. For example, the three-dimensional absolute coordinates of the hip joint can be (3, 4, 5), and the distance between the user object and the video playback device is 5 meters. The target distance threshold can be configured as needed, and the embodiments of the present application do not make specific limitations thereto, for example, the target distance threshold can be 3 meters, 5 meters, etc.
[0107] In yet another possible implementation, the video playback device can also monitor the posture of the child when watching and remind, and after the video playback device is switched to the child mode, the process of monitoring the child by the video playback device can include: the video playback device determines a current sitting posture of the user object based on the three-dimensional coordinates of the human body key points of the user object in the camera coordinate system; when the current sitting posture does not conform to a standard sitting posture threshold, the video playback device stops playing the video picture and displays a second reminding message, and the second reminding message is used to remind to correct the current sitting posture to conform to the standard sitting posture threshold. For example, the standard sitting posture threshold can be configured as needed, and the embodiments of the present application do not make specific limitations thereto, for example, the standard sitting posture threshold can be a sitting posture with the upper body upright. For example, when it is detected that the user object leans to the left or right side and lies down, the user object can be reminded to keep the upper body upright.
[0108] The first and second reminder messages can be in any form, such as text, images, animations, or videos; this application embodiment does not impose specific limitations on them. For example, when the current sitting posture does not meet the standard sitting posture threshold, an animation indicating the threshold can be played to visually and concretely correct the user's posture. In one possible example, the video playback device can pause the current screen and display the first or second reminder message. Alternatively, the current screen can be displayed directly without pausing.
[0109] The mode switching method for video playback devices provided in this application determines the three-dimensional coordinates of human key points in the camera coordinate system based on the human body bounding box of a user object in at least two image frames; based on the three-dimensional coordinates of the human body key points in the camera coordinate system, the bone lengths of multiple bones of the user object are determined; furthermore, the height can be accurately predicted based on the lengths of these multiple bones. Since the three-dimensional coordinates of the human body key points in the camera coordinate system can measure the accurate position of the user object in three-dimensional space in units of points, and the lengths of each bone can be accurately located based on the three-dimensional coordinates of the human body key points in the camera coordinate system, and the height information can be further accurately determined, the accuracy of bone length is improved; therefore, the identity attribute of the user object is determined by using the bone length and the age information predicted based on the face bounding box, and if the user is a child, the mode is switched to children's mode; this improves the accuracy of determining the identity attribute of the user object, further improves the accuracy of determining whether to switch to children's mode, and improves the accuracy of children's mode switching.
[0110] Figure 3 This is a schematic diagram of a mode switching device for a video playback device provided in an embodiment of this application. Figure 3 As shown, the device includes:
[0111] Acquisition module 301 is used to acquire at least one image frame corresponding to the user object currently viewing the video playback device;
[0112] Evaluation module 302 is used to perform image recognition and parsing on the at least one image frame to evaluate and obtain the skeletal information and age information of the user object;
[0113] The switching module 303 is used to determine the identity attribute of the user object based on the user object's skeletal information and age information, and determine whether to switch to child mode based on the identity attribute.
[0114] In one possible implementation, the evaluation module 302 includes:
[0115] A human body bounding box determination unit is used to determine the human body bounding box in which the user object is located in the at least one image frame.
[0116] A key point determination unit is used to determine the three-dimensional coordinates of the key points of the user object in the camera coordinate system based on the human body bounding box of at least one image frame.
[0117] The prediction unit is used to evaluate the bone length of multiple bones of the user object based on the three-dimensional coordinates of the human body key points; and to identify the face bounding box of the user object's face in at least one image frame and predict the age information of the user object.
[0118] In one possible implementation, the key point determination unit is used to crop the image frame based on the human body bounding box for each image frame to obtain a human body bounding box image, and predict the two-dimensional image coordinates of the human body key points in the human body bounding box image in the image coordinate system of the image frame; based on the two-dimensional image coordinates of the human body key points in the at least one image frame, the three-dimensional relative coordinates of the human body key points and the three-dimensional absolute coordinates of the human body trajectory are determined by a three-dimensional human body pose model and a human body trajectory model.
[0119] Wherein, the three-dimensional absolute coordinates of the human body trajectory refer to the coordinates of the center point of the human body trajectory in the camera coordinate system; the three-dimensional relative coordinates of the human body key points refer to the coordinates of the human body key points relative to the center point of the human body trajectory in the camera coordinate system; the camera coordinate system is a three-dimensional spatial coordinate system with the camera that captures the image frame as the origin and the optical axis of the camera as the Z-axis.
[0120] In one possible implementation, the key point determination unit is further configured to detect the human body bounding box of at least one predetermined human body part of a user object in the image frame; adjust the corresponding human body part included in the human body bounding box of the corresponding user object in the image frame based on the two-dimensional image coordinates of the human body key points of the at least one user object in the previously acquired image frame; and crop the image frame based on the adjusted human body bounding box to obtain the human body bounding box image of at least one user object.
[0121] In one possible implementation, the key point determination unit is further configured to determine the three-dimensional absolute coordinates of the human key point in the camera coordinate system based on the three-dimensional relative coordinates of the human key point and the three-dimensional absolute coordinates of the human trajectory; and to calculate the bone length of each bone based on the three-dimensional absolute coordinates of the human key points at both ends of each bone.
[0122] In one possible implementation, the prediction unit is also used for any of the following:
[0123] For each image frame, based on the key points of the user's head in the image frame, the face bounding box corresponding to the user's head is determined, and based on the face bounding box and the face age recognition model, the age information of the user is determined.
[0124] Face detection is performed on the image frame to obtain the face bounding box in the image frame. Based on the face bounding box and the face age recognition model, the age information of the user is determined.
[0125] In one possible implementation, the switching module 303 is also used for any of the following:
[0126] Based on the lengths of these multiple bones, the user's height information is predicted, and based on the user's height and age information, the user's identity attributes are determined.
[0127] Based on the user's age information and the lengths of the multiple bones, determine the user's identity attributes.
[0128] In one possible implementation, the switching module 303 is also used for any of the following:
[0129] When the user's height information does not exceed the first threshold, or when the user's height information exceeds the first threshold and the user's age information does not exceed the second threshold, the user's identity attribute is determined to be a child;
[0130] When the user's age information does not exceed the second threshold, or when the user's age information exceeds the second threshold and the user's height information does not exceed the first threshold, the user's identity attribute is determined to be a child.
[0131] In one possible implementation, the at least one image frame includes at least two user objects; the switching module 303 is further configured to switch the video playback device to child mode when the identity attribute of any of the at least two user objects is a child; and if the identity attribute of any of the at least two user objects is an adult when the video playback device is in child mode, exit the child mode.
[0132] In one possible implementation, Figure 4 This is a schematic diagram of another mode switching device for a video playback device provided in an embodiment of this application. Figure 4 As shown, the device also includes:
[0133] The distance determination module 304 is used to determine the distance between the user object and the video playback device based on the three-dimensional coordinates of the user object's human body key points in the camera coordinate system.
[0134] The first reminder module 305 is used to stop playing the video and display a first reminder message when the distance does not exceed the target distance threshold. The first reminder message is used to remind the user to maintain at least the target distance threshold with the video playback device.
[0135] In one possible implementation, Figure 5 This is a schematic diagram of another mode switching device for a video playback device provided in an embodiment of this application. Figure 5 As shown, the device also includes:
[0136] The posture determination module 306 is used to determine the current posture of the user object based on the three-dimensional coordinates of the user object's human body key points in the camera coordinate system.
[0137] The second reminder module 307 is used to stop playing the video and display a second reminder message when the current sitting posture does not meet the standard sitting posture threshold. The second reminder message is used to remind the user to correct the current sitting posture to meet the standard sitting posture threshold.
[0138] The beneficial effects of the technical solution provided in this application are:
[0139] The mode switching device for a video playback device provided in this application acquires at least one image frame corresponding to a user currently viewing the video playback device; performs image recognition and analysis on the at least one image frame to evaluate and obtain the skeletal information and age information of the user; determines the user's identity attribute based on the skeletal information and age information, and determines whether to switch to child mode based on the identity attribute. Using age information and skeletal information to determine identity attributes, and further determining whether to switch to child mode, improves the accuracy of user identity attribute determination, and thus improves the accuracy of mode switching for the video playback device.
[0140] The mode switching device for video playback device in this embodiment can execute the mode switching method for video playback device shown in the above embodiment of this application. The implementation principle is similar, and will not be described again here.
[0141] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. For example... Figure 6 As shown, the computer device includes: a memory and a processor; at least one program, stored in the memory, which, when executed by the processor, can achieve the following compared to the prior art:
[0142] By acquiring at least one image frame corresponding to the user object currently viewing the video playback device; performing image recognition and analysis on the at least one image frame to evaluate and obtain the skeletal information and age information of the user object; determining the user object's identity attribute based on the skeletal information and age information, and determining whether to switch to child mode based on the identity attribute. Using age information and skeletal information to determine identity attributes, and further determining whether to switch to child mode, improves the accuracy of user identity attribute determination and the accuracy of mode switching for the video playback device.
[0143] In one alternative embodiment, a computer device is provided, such as Figure 6 As shown, Figure 6 The computer device 600 shown includes a processor 601 and a memory 603. The processor 601 and the memory 603 are connected, for example, via a bus 602. Optionally, the computer device 600 may further include a transceiver 604, which can be used for data interaction between the computer device and other computer devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 604 is not limited to one type, and the structure of the computer device 600 does not constitute a limitation on the embodiments of this application.
[0144] Processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 601 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0145] Bus 602 may include a pathway for transmitting information between the aforementioned components. Bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 602 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0146] The memory 603 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0147] The memory 603 stores the application code (computer program) for executing the scheme of this application, and its execution is controlled by the processor 601. The processor 601 executes the application code stored in the memory 603 to implement the content shown in the foregoing method embodiments.
[0148] Computer equipment includes, but is not limited to: video playback devices, smart TVs, personal computers, etc.
[0149] This application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content of the mode switching method for a video playback device described in the foregoing method embodiments.
[0150] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned mode switching method for a video playback device.
[0151] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0152] The above description is only a partial embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A mode switching method for a video playing device, characterized in that, The method comprises: collecting at least one image frame corresponding to a user object currently watching a video playing device; performing image recognition and analysis on the at least one image frame to evaluate bone information and age information of the user object; judging an identity attribute of the user object according to the bone information and the age information of the user object, and determining whether to switch to a child mode according to the identity attribute; wherein the performing image recognition and analysis on the at least one image frame to evaluate the bone information and the age information of the user object comprises: determining a human body box in which the user object is located in the at least one image frame; for each image frame, performing cropping on the image frame based on the human body box to obtain a human body box image, and determining three-dimensional coordinates of human body key points of the user object in a camera coordinate system based on the human body box image; evaluating bone lengths of a plurality of bones of the user object according to the three-dimensional coordinates of the human body key points; and recognizing a face box of a face of the user object in the at least one image frame, and predicting age information of the user object based on a face age recognition model and the face box; the judging the identity attribute of the user object according to the bone information and the age information of the user object comprises: if the age information of the user object is not greater than a second threshold value, and the bone lengths of the plurality of bones do not meet a target condition, determining that the identity attribute of the user object is a child; the target condition comprises at least one of each bone length being not greater than a third threshold value, a target bone length in the plurality of bone lengths being not greater than a fourth threshold value, and a maximum bone length in the plurality of bone lengths being not greater than a fifth threshold value; the target bone length is determined based on a length of at least one bone in the plurality of bones; the performing cropping on the image frame based on the human body box to obtain the human body box image for each image frame comprises: for each image frame except a first image frame, performing human body detection on the image frame, and marking a corresponding position identifier for each detected human body; the position identifier of the human body is used to identify a position of the human body in the image frame; human bodies having the same position in adjacent two image frames correspond to the same position identifier; for the position identifiers of the human bodies in the image frame, obtaining two-dimensional image coordinates of a neck joint and a pelvic joint of a user object having the same position identifier as the position identifier in a preceding collected image frame corresponding to the image frame, and identifying the neck joint and the pelvic joint of the user object in the image frame according to the two-dimensional image coordinates of the neck joint and the pelvic joint of the user object in the preceding collected image frame, to rotate a human body of the user object in the image frame based on the neck joint and the pelvic joint of the user object in the image frame, and obtain an adjusted human body box; wherein the preceding collected image frame is an image frame before a collection time of the image frame; performing cropping on the image frame based on the adjusted human body box to obtain a human body box image of at least one user object; The method comprises the following steps: detecting a human body frame of a predetermined human body part of at least one user object in a first image frame, and cropping the first image frame along the human body frame of the at least one user object to obtain a human body frame image of the at least one user object in the first image frame.
2. The method of claim 1, wherein, The method comprises the following steps: determining three-dimensional coordinates of the human body key points of the user object in a camera coordinate system based on the human body frame image. The method comprises the following steps: predicting two-dimensional image coordinates of the human body key points in the human body frame image in an image coordinate system of the image frame; The method comprises the following steps: determining three-dimensional relative coordinates of the human body key points and three-dimensional absolute coordinates of a human body trajectory based on the two-dimensional image coordinates of the human body key points in the at least one image frame, a three-dimensional human body posture model and a human body trajectory model; The three-dimensional absolute coordinates of the human body trajectory refer to coordinates of a human body trajectory center point in the camera coordinate system; the three-dimensional relative coordinates of the human body key points refer to coordinates of the human body key points relative to the human body trajectory center point in the camera coordinate system; and the camera coordinate system is a three-dimensional space coordinate system with a camera for collecting the image frame as a coordinate origin and an optical axis of the camera as a Z axis.
3. The method of claim 2, wherein, The method comprises the following steps: determining three-dimensional absolute coordinates of the human body key points in the camera coordinate system based on the three-dimensional relative coordinates of the human body key points and the three-dimensional absolute coordinates of the human body trajectory; The method comprises the following steps: calculating a bone length of each bone according to the three-dimensional absolute coordinates of the human body key points at both ends of the bone. The method comprises the following steps: identifying a face frame of a face of the user object in the at least one image frame, and predicting age information of the user object.
4. The method of claim 2, wherein, The method comprises the following steps: determining a face frame corresponding to a head of the user object based on each key point of the head of the user object in each image frame, and determining age information of the user object based on the face frame and a face age recognition model; The method comprises the following steps: performing face detection on the image frame to obtain a face frame in the image frame, and determining age information of the user object based on the face frame and a face age recognition model. The at least one image frame comprises at least two user objects; and the method comprises the following steps:
5. The method of claim 1, wherein, When the identity attribute of any user object in the at least two user objects is a child, the video playing device is switched to a child mode; When the video playing device is in the child mode, if the identity attribute of any user object in the at least two user objects is an adult, the child mode is exited. After the step of determining whether to switch to the child mode based on the identity attribute, the method further comprises the following steps:
6. The method of claim 1, wherein, Determining a distance between the user object and the video playing device based on the three-dimensional coordinates of the human body key points of the user object in the camera coordinate system; When the distance is less than a target distance threshold, video pictures are stopped from being played, and a first prompt message is displayed, the first prompt message being used to prompt that at least a target distance threshold is maintained from the video playing device. 7. The method of claim 1, wherein, After determining whether to switch to the child mode according to the identity attribute, the method further comprises: determining a current sitting posture of the user object based on three-dimensional coordinates of the human body key points of the user object in a camera coordinate system; when the current sitting posture does not conform to a standard sitting posture threshold, stopping playing the video picture and displaying a second prompt message for reminding to correct the current sitting posture to conform to the standard sitting posture threshold.
8. A mode switching apparatus for a video playing device, characterized by, The device comprises: a collection module configured to collect at least one image frame corresponding to a user object currently watching a video playing device; an evaluation module configured to perform image recognition and analysis on the at least one image frame to obtain bone information and age information of the user object; a switching module configured to determine an identity attribute of the user object according to the bone information and the age information of the user object, and determine whether to switch to a child mode according to the identity attribute; The evaluation module comprises: a human body frame determination unit configured to determine a human body frame in which the user object is located in the at least one image frame; a key point determination unit configured to, for each image frame, crop the image frame based on the human body frame to obtain a human body frame image, and determine three-dimensional coordinates of human body key points of the user object in a camera coordinate system based on the human body frame image; a prediction unit configured to evaluate bone lengths of a plurality of bones of the user object according to the three-dimensional coordinates of the human body key points, identify a face frame of a face of the user object in the at least one image frame, and predict age information of the user object based on a face age recognition model and the face frame; When determining the identity attribute of the user object according to the bone information and the age information of the user object, the switching module is specifically configured to: if the age information of the user object is not greater than a second threshold value and the bone lengths of the plurality of bones do not conform to a target condition, determine that the identity attribute of the user object is a child; the target condition comprises at least one of each bone length being less than a third threshold value, a target bone length in the plurality of bone lengths being less than a fourth threshold value, and a maximum bone length in the plurality of bone lengths being less than a fifth threshold value; the target bone length is determined based on a length of at least one bone in the plurality of bones; When cropping the image frame based on the human body frame to obtain the human body frame image for each image frame, the key point determination unit is specifically configured to: for each image frame except the first image frame, perform human body detection on the image frame, and mark a corresponding position identifier for each detected human body; the position identifier of the human body is used to identify the position of the human body in the image frame; human bodies having the same position in adjacent two image frames correspond to the same position identifier. For the position identification of each human body in the image frame, a two-dimensional image coordinate of a neck joint and a pelvic joint of a user object same as the position identification in a previous acquisition image frame corresponding to the image frame is acquired, and a two-dimensional image coordinate of the neck joint and the pelvic joint of the user object in the previous acquisition image frame is identified to rotate a human body of the user object in the image frame based on the neck joint and the pelvic joint of the user object in the image frame, so as to obtain an adjusted human body frame; wherein the previous acquisition image frame is an image frame before a collection time of the image frame; Based on the adjusted human body frame, the image frame is cropped to obtain a human body frame image of at least one user object; For the first image frame, a human body frame of a predetermined human body part of at least one user object in the first image frame is detected, and the first image frame is cropped along the human body frame of the at least one user object to obtain a human body frame image of the at least one user object in the first image frame.
9. A computer device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program comprises instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1-8. The processor executes the computer program to implement the mode switching method for the video playing device according to any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the mode switching method for the video playing device according to any one of claims 1 to 7.
Citation Information
Patent Citations
Parental control settings based on body dimensions
CN102221877A
Television watching mode determination method, television and storage medium
CN112770186A
Age identification method and device, electronic equipment and storage medium
CN112990056A