Lip movement detection method, device, readable storage medium and electronic device
By acquiring image sequences for lip key points detection and regional image formation, and combining with neural networks for multi-dimensional information integration, the accuracy of lip motion detection is solved, the accuracy of driver status judgment is improved, and the vehicle driving safety is enhanced.
Patent Information
- Application Number
- CN202210339940.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-01
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-04-01
AI Technical Summary
In the prior art, the method of detecting lip movement based on the distance between the upper and lower lips is inaccurate due to individual differences and face position deviation, which affects the accuracy of the driver's status or behavior judgment.
By acquiring the image sequence, key point detection is performed, key point information is determined, lip area image sequence is formed, and lip action detection is comprehensively used to use multi-dimensional information to comprehensively detect lip action, including the splicing characteristics of the lip area image sequence and key point information, and lip action detection is performed using the trained neural network.
It improves the accuracy of lip motion detection, ensures the accuracy of driver status or behavior judgment, and can promptly remind the driver to pay attention or adjust the driving mode, improving vehicle driving safety.
Smart Images

Figure CN114708578B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to artificial intelligence technology, specifically computer vision technology, deep learning technology, and especially a lip movement detection method, device, readable storage medium and electronic device. Background Art
[0002] During driving, the system can detect whether the driver is yawning to help determine the driver's condition, thereby providing driving assistance. For example, if the driver is detected yawning, it is considered that the driver may be driving fatigued and can be reminded to take a rest.
[0003] In related technologies, the method for detecting whether lip movement occurs is usually to determine the key points of the upper lip and lower lip in the lip image, determine the distance between the upper and lower lips based on the key points of the upper lip and lower lip, compare the distance between the upper and lower lips with a standard value, and when the distance between the upper and lower lips is greater than the standard value, it is determined that lip movement has occurred. Summary of the Invention
[0004] Related technologies detect lip movements based solely on the distance between the upper and lower lips. However, due to differences in lip shape between individuals and facial position shifts in captured images, determining lip movement based solely on the distance between the upper and lower lips can be inaccurate. This, in turn, reduces the accuracy of subsequent lip movement detection results when determining a driver's state or behavior.
[0005] In order to solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a lip movement detection method, device, readable storage medium, and electronic device.
[0006] According to one aspect of an embodiment of the present disclosure, a lip movement detection method is provided, including: acquiring an image sequence including a facial area, the image sequence including multiple frames of images with a time-series relationship; performing key point detection based on each frame image of the image sequence, and determining lip key point information of each frame image in the image sequence; determining lip area images of each frame image based on the lip key point information of each frame image in the image sequence; determining the lip area image sequence based on the lip area images of each frame image; performing lip movement detection based on the lip area image sequence, and determining a lip movement detection result.
[0007] According to another aspect of an embodiment of the present disclosure, a lip movement detection device is provided, including: an image acquisition module for acquiring an image sequence including a facial area, the image sequence including multiple frames of images with a time-series relationship; a first determination module for performing key point detection based on each frame image of the image sequence, and determining the lip key point information of each frame image in the image sequence; a second determination module for determining the lip area image of each frame image according to the lip key point information of each frame image in the image sequence; a third determination module for determining the lip area image sequence according to the lip area image of each frame image; a detection module for performing lip movement detection based on the lip area image sequence, and determining the lip movement detection result.
[0008] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the lip movement detection method.
[0009] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; and the processor for reading the executable instructions from the memory and executing the instructions to implement the lip movement detection method.
[0010] Based on the lip movement detection method, device, readable storage medium and electronic device provided in the above-mentioned embodiments of the present disclosure, an image sequence including a facial area is obtained, key point detection is performed based on each frame image of the image sequence, and lip key point information of each frame image in the image sequence is determined; based on the lip key point information of each frame image in the image sequence, the lip area image of each frame image is determined; based on the lip area image of each frame image, a lip area image sequence is determined; based on the lip area image of each frame image, lip movement detection is performed based on the lip area image sequence to determine the lip movement detection result. The embodiments of the present disclosure determine the lip movement detection result through comprehensive multi-dimensional information, which greatly improves the accuracy of the lip movement detection result, and further improves the accuracy of the subsequent judgment of the driver's state or behavior through the lip movement detection result, so that when the driver exhibits a dangerous state or behavior, the driver can be reminded to pay attention or change the driving mode in a timely manner, thereby effectively improving the safety of vehicle driving.
[0011] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0013] Figure 1 This is a scene diagram applicable to a lip movement detection method according to an embodiment of the present disclosure.
[0014] Figure 2 It is a flowchart of a lip movement detection method provided by an exemplary embodiment of the present disclosure.
[0015] Figure 3 FIG. 1 is a schematic diagram of facial key points of a frame of image provided by an exemplary embodiment of the present disclosure.
[0016] Figure 4 2 is a schematic diagram of a first lip movement detection network structure provided by an exemplary embodiment of the present disclosure.
[0017] Figure 5 It is a flowchart of lip movement detection results provided by another exemplary embodiment of the present disclosure.
[0018] Figure 6 2 is a schematic diagram of a second lip movement detection network structure provided by an exemplary embodiment of the present disclosure.
[0019] Figure 7 1 is a flow chart of lip movement detection results provided by an exemplary embodiment of the present disclosure.
[0020] Figure 8 It is a flowchart of determining splicing features provided by an exemplary embodiment of the present disclosure.
[0021] Figure 9 It is a schematic diagram of a process of determining a lip area image provided by an exemplary embodiment of the present disclosure.
[0022] Figure 10 is a schematic diagram of a lip movement detection device provided by an exemplary embodiment of the present disclosure.
[0023] Figure 11 is a schematic diagram of a second determination module provided by an exemplary embodiment of the present disclosure.
[0024] Figure 12 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0026] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0027] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.
[0028] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.
[0029] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0030] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.
[0031] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.
[0032] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0033] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0034] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0035] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0036] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or specialized computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.
[0037] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.
[0038] Application Overview
[0039] During the implementation of the present disclosure, the inventors discovered that, in related art, the method for detecting whether lip movement has occurred typically involves identifying key points of the upper and lower lips in a lip image, determining the distance between the upper and lower lips based on these key points, comparing the distance with a standard value, and determining that lip movement has occurred when the distance between the upper and lower lips is greater than the standard value. However, due to differences in lip shape between different people or facial position offsets in the captured image, the detection result of whether lip movement has occurred based solely on the distance between the upper and lower lips is inaccurate, resulting in low accuracy in subsequent judgments of the driver's status or behavior based on the lip movement detection results.
[0040] Exemplary Systems
[0041] The present disclosure is applicable to assisted driving scenarios. For example, during the process of driving a vehicle, the driver's lip movements are used to determine the driver's status or behavior, and whether the driver is not conducive to vehicle driving is judged based on the driver's status or behavior. When the driver is not conducive to vehicle driving, the driver can be reminded to pay attention or the vehicle's driving mode can be adjusted in time to improve the safety of vehicle driving. Figure 1 is a scene graph to which this disclosure is applicable. Figure 1As shown, it includes an image acquisition device and a computing platform. The image acquisition device can be set inside the vehicle, and the computing platform can be set inside the vehicle or elsewhere. The image acquisition device is used to acquire multiple frames of images. The computing platform receives the multiple frames of images acquired by the acquisition device, obtains images with facial areas in the multiple frames of images, and forms an image sequence including the facial areas. The image sequence including the facial areas is then processed to obtain lip key point information of each frame image, obtains lip area images of each frame image based on the lip key point information of each frame image, and determines a lip area image sequence based on the lip area images of each frame image. Lip movement detection results are obtained based on the lip area image sequence, and the computing platform outputs the lip movement detection results to the vehicle control device. The embodiment of the present disclosure determines the lip movement detection results through comprehensive determination of multi-dimensional information, greatly improving the accuracy of the lip movement detection results, thereby improving the accuracy of subsequent judgment of the driver's state or behavior through the lip movement detection results. When the driver exhibits dangerous state or behavior, the driver can be reminded to pay attention or change the driving mode in a timely manner, thereby effectively improving the safety of vehicle driving.
[0042] Exemplary Methods
[0043] Figure 2 This is a flow chart of a lip movement detection method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, vehicles, etc. Figure 2 As shown, the following steps are included:
[0044] Step S201: Acquire an image sequence including a facial area.
[0045] The image sequence includes multiple frames of images having a temporal relationship. Multiple frames of images are collected by an image acquisition device, and an image sequence including a facial region is obtained from the collected multiple frames of images using image recognition technology. In the image sequence, each frame of the image includes a facial region. Exemplarily, the collected multiple frames of images are recognized by a trained neural network for facial recognition to determine images including facial regions. The neural network may be Faster-RCNN (Faster Region Convolutional Neural Networks) or YOLO (You Only Look Once). The images including facial regions are arranged in chronological order to obtain an image sequence including facial regions.
[0046] Step S202: performing key point detection based on each frame image in the image sequence to determine lip key point information of each frame image in the image sequence.
[0047] Here, the image recognition technology is used to identify each frame of the image sequence to determine the lip key point information of each frame of the image, and the lip key point information includes the position information of the lip key point. For example, Figure 3 A schematic diagram of facial key points of a frame image is shown. Figure 3 As shown, the face is identified to have 68 facial key points, each of which has a serial number and position information (coordinate value). The serial numbers corresponding to the lip key points are 49 to 68. The lip key points are determined by the correspondence between the serial numbers of the facial key points and the facial positions, and the lip key point information is obtained. The lip key point information includes the position information (coordinate value) of the lip key points. The recognition of facial key points can use a neural network trained for facial key point and lip key point recognition. The neural network can be CNN (Convolutional Neural Networks, convolutional neural network) and the like. When training the neural network, heatmap and regression methods can be used.
[0048] Step S203: determining the lip region image of each frame image according to the lip key point information of each frame image in the image sequence.
[0049] Wherein, according to the position information in the lip key point information in each frame image, the lip area in each frame image is determined, the lip area of each frame image is extracted, and the lip area image of each frame image is obtained.
[0050] Step S204: determining a lip region image sequence according to the lip region image of each frame image.
[0051] The lip region images of each frame image obtained in step S203 are arranged in chronological order to obtain a lip region image sequence.
[0052] Step S205: Perform lip movement detection based on the lip region image sequence to determine a lip movement detection result.
[0053] A neural network trained for lip movement detection can be used to identify a sequence of lip region images to obtain lip movement detection results. The lip movement detection results can include whether the lips are moving and the type of lip movement, such as drinking water.
[0054] It should be noted that, unless otherwise specified in the embodiments of the present disclosure, the number of facial regions in an image may be one or more. When determining that there are multiple facial regions in an image, steps S202 to S205 need to be performed on the lips of each facial region in the image. Furthermore, after identifying the lip key point information of each facial region in the image, a correspondence between each lip key point information and the facial region needs to be established.
[0055] In the disclosed embodiment, the lip key point information of each frame in the image sequence is determined, and then the lip region image of each frame is obtained based on the determined lip key point information. Then, the lip region image sequence formed by the lip region images of each frame is detected to obtain the lip movement detection result. This achieves the comprehensive determination of the lip movement detection result through multi-dimensional information, greatly improving the accuracy of the lip movement detection result, and further improving the accuracy of the subsequent judgment of the driver's state or behavior based on the lip movement detection result. When the driver exhibits dangerous state or behavior, the driver can be promptly reminded to pay attention or change the driving mode, thereby effectively improving the safety of vehicle driving.
[0056] In one embodiment of the present disclosure, the lip movement detection results include a lip movement type. The lip movement type is used to characterize the reason for the lip movement and whether the lip movement occurs. The lip movement type includes any of the following: yawning, talking, eating, drinking, smoking, and no lip movement.
[0057] In one embodiment of the present disclosure, step S205 may include: inputting the lip area image sequence into a first lip movement detection network, and outputting the lip movement detection result through the first lip movement detection network.
[0058] The first lip movement detection network may be a trained neural network for detecting lip movements, such as CNN, GCN (Graph Convolutional Networks, graph convolutional networks), etc. For example, Figure 4 The structure of the first lip movement detection network is shown in FIG. Figure 4As shown, the first lip movement detection network may include: multiple first feature extraction units connected in series, at least one fully connected layer, and a classifier. Each first feature extraction unit includes at least one convolutional layer and at least one pooling layer. After the lip region image sequence is input into the first lip movement detection network, the multiple first feature extraction units extract features from each frame of the image to obtain image features of each lip region image. The image features of each lip region image are then input into the fully connected layer. The fully connected layer can connect the image feature maps of each lip region image and then input them into the classifier. The classifier classifies the lip movement detection results. The convolutional layer performs convolution processing on the image to obtain image features. The convolutional layer can perform 2D or 3D convolution on the image. The convolutional layer outputs the image features of the image to the pooling layer, which compresses the extracted image features. During data transmission between two adjacent first feature extraction units, the convolutional layer in the next first feature extraction unit receives data output by the pooling layer in the previous first feature extraction unit. It should be noted that the number of first feature extraction units, convolution layers, and pooling layers can be set according to actual needs, and the embodiment of the present disclosure does not impose any special restrictions.
[0059] In the disclosed embodiment, a first lip movement detection network trained for lip movement detection is used to detect a sequence of lip region images to obtain lip movement detection results. Because the first lip movement detection network has a strong learning capability, it can perform lip movement detection on different lip region images to obtain accurate lip movement detection results, thereby ensuring the accuracy of subsequent determination of the driver's status or behavior based on the lip movement detection results.
[0060] In one embodiment of the present disclosure, Figure 5 As shown, step S205 may include:
[0061] Step S2501: Input the lip area image sequence into the first sub-network of the second lip movement detection network, and output the image features of each lip area image in the lip area image sequence through the first sub-network.
[0062] Step S2502: Input the image features of each lip area image into the second sub-network of the second lip movement detection network, and output the lip movement detection result through the second sub-network.
[0063] Among them, the second lip movement detection network can be a trained neural network for detecting lip movements. The second lip movement detection network includes a first subnetwork and a second subnetwork. The first subnetwork and the second subnetwork can be CNN, GCN, CNN or TCN (Temporal convolutional network, temporal convolutional neural network) and the like. The first subnetwork is used to extract the features of the lip area image to obtain the image features of each lip area image. The second subnetwork is used to process the image features of each lip area image to obtain the lip movement detection result. For example, Figure 6 The structure of the second lip movement detection network is shown in FIG. Figure 6 As shown, the first subnetwork may include multiple second feature extraction units arranged in parallel. Each second extraction unit includes multiple feature extraction subunits arranged in series, and each feature extraction subunit includes a convolution layer and a pooling layer. In two adjacent feature extraction subunits, the convolution layer in the next feature extraction subunit receives the data output by the pooling layer in the previous feature extraction subunit. The second subnetwork includes a convolution layer, a fully connected layer, and a classifier. The lip area image sequence is input into the first subnetwork, and after feature extraction by the second feature extraction unit, the image features of each lip area image are obtained. The image features of each lip area image are input into the second subnetwork, and after convolution processing by the convolution layer of the second subnetwork, the features of the image features of each lip area image are extracted. The features of the image features of each lip area image are then output to the fully connected layer for compression processing, and then enter the classifier. After separation processing by the classifier, the lip movement detection result is obtained. When the second subnetwork is a TCN, the convolution layer of the second subnetwork performs temporal convolution processing on the image features. It should be noted that the number of the second feature extraction unit, convolution layer and pooling layer can be set according to actual needs, and the embodiment of the present disclosure does not impose any special restrictions.
[0064] In an embodiment of the present disclosure, the first subnetwork in the second lip movement detection network recognizes a lip image sequence to obtain image features of each lip region image. The second subnetwork then detects the image features of each lip region image to obtain lip movement detection results, achieving efficient and accurate lip movement detection results. Based on the powerful learning capabilities of the second lip movement detection network, the first subnetwork can extract image features from different lip region images to obtain image features of different lip region images, thereby improving the accuracy of the lip movement detection results obtained by the subsequent second subnetwork using the image features of different lip region images. The second subnetwork can detect different lip region images to obtain accurate lip movement detection results, thereby ensuring the accuracy of the subsequent determination of the driver's status or behavior based on the lip movement detection results. In one embodiment of the present disclosure, the present disclosure further includes: obtaining a lip key point information sequence based on the lip key point information of each frame image in the image sequence.
[0065] In one embodiment of the present disclosure, Figure 7 As shown, step 205 may include the following steps:
[0066] Step S2053: determining the splicing features based on the lip region image sequence and the lip key point information sequence.
[0067] Step S2054: input the splicing features into the third lip movement detection network, and output the lip movement detection results through the third lip movement detection network.
[0068] In which, the lip key point information of each frame image in the image sequence is extracted, and the lip key point information of each frame image is arranged in time sequence to obtain a lip key point information sequence. The image features of each lip area image in the lip area image sequence are determined. The features of each lip key point information in the lip key point information sequence are determined. The image features of each lip area image and the features of each lip key point information are spliced to obtain spliced features. The spliced features are input into the third lip movement detection network, and the lip movement detection results are output through the third lip movement detection network. In which, the third lip movement detection network can be a trained neural network for detecting lip movements, and the third lip movement detection network can be CNN, GCN or RNN (Rerrent Neural Network, recurrent neural network), etc.
[0069] In the disclosed embodiment, the lip region image sequence and the lip key point information sequence are first preprocessed to obtain the lip region features of each lip region image and the lip key point features of each lip key point information. The lip region features of each lip region image and the lip key point features of each lip key point information are used to obtain the splicing features, so that the information included in the splicing features is rich and accurate, which greatly improves the accuracy of the lip movement detection results obtained by the subsequent third lip movement detection network through the splicing features. In addition, because the lip movement detection network has powerful learning and detection capabilities, it can detect various splicing features and obtain accurate lip movement detection results, thereby ensuring the accuracy of the subsequent determination of the driver's status or behavior through the lip movement detection results.
[0070] In one embodiment of the present disclosure, Figure 8 As shown, step 2054 may include the following steps:
[0071] Step S20541: extract features from each lip region image in the lip region image sequence to obtain a lip region feature sequence.
[0072] Step S20542: extract features from the lip key point information sequence to obtain a lip key point feature sequence.
[0073] Step S20543: splice the lip region feature sequence and the lip key point feature sequence to obtain a spliced feature.
[0074] In this process, a neural network trained for image feature extraction is used to extract features from a lip region image sequence to obtain image features of each lip region image. The image features of each lip region image are used as lip region features of each lip region image. The lip region features of each lip region image are arranged in time sequence to obtain a lip region feature sequence. A neural network trained for key point feature extraction is used to extract features from a lip key point information sequence to obtain features of each lip key point information in the lip key point information sequence. The features of each lip key point information are used as lip key point features of each lip key point information. The lip key point features of each lip key point information are arranged in time sequence to obtain a lip key point feature sequence. The lip region feature sequence and the lip key point feature sequence are spliced together to obtain spliced features. In this process, the neural network for image feature extraction and the neural network for key point feature extraction can be CNN, Faster-RCNN, GCN, etc.
[0075] In the embodiment of the present disclosure, by obtaining the lip key point feature sequence and the lip area feature sequence, and splicing the lip key point feature sequence and the lip area feature sequence to obtain accurate splicing features, the accuracy of the subsequent lip movement detection results obtained by splicing features is improved.
[0076] In one embodiment of the present disclosure, Figure 9 As shown, step 203 may include the following steps:
[0077] Step S2031: Determine a first detection frame of the lip region in each frame image in the image sequence based on the lip key point information of each frame image in the image sequence.
[0078] Step S2032, taking the center position of the first detection frame of the lip area in each frame image as the center, according to a preset magnification factor, the length and / or width of the first detection frame of the lip area in each frame image is magnified to obtain the second detection frame of the lip area in each frame image in the image sequence.
[0079] Step S2033: Determine the lip region image of each frame image according to the second detection frame of the lip region in each frame image.
[0080] In step S201, the coordinate values of the facial region in each image frame are determined, for example, the coordinate values of the vertices of the facial region in each frame of image; step S202 determines the facial key points and the coordinate values of the facial key points in the facial region in each frame of image, and determines the lip key points and the coordinate values of the lip key points from the facial key points, that is, obtains the lip key point information. The lip key point information includes the two-dimensional coordinate values (x, y) of the lip key points, determines the maximum horizontal coordinate value and the vertical coordinate value, and the minimum horizontal coordinate value and the vertical coordinate value of the lip key points, uses the maximum horizontal coordinate value and the vertical coordinate value as the first endpoint, uses the minimum horizontal coordinate value and the vertical coordinate value as the second endpoint, and uses the line connecting the first endpoint and the second endpoint as the diagonal line to form a rectangular frame, that is, obtains the first detection frame of the lip region. The center position of the first detection frame can be determined based on the coordinate values of the vertices of the first detection frame. With the center position of the first detection frame as the center, the length and / or width of the first detection frame are magnified according to a preset magnification factor to obtain a second detection frame for the lip area in each frame image in the image sequence, thereby avoiding lip key points or lip areas not being in the detection frame. The preset magnification factor can be set according to actual needs and is not particularly limited in the embodiment of the present disclosure. The image in the second detection frame of the lip area in each frame image is extracted to obtain the lip area image of each frame image.
[0081] Among them, before determining the first detection frame of the lip area in each frame image in the image sequence based on the lip key point information of each frame image in the image sequence, the lip key point information of each frame image can be normalized to further improve the accuracy of lip area image determination. Specifically, according to the preset rules, a frame image in the image sequence is used as a reference image. For example, the time sequence window is 10, and the first frame image in the 10 frames is determined as the reference image. According to the coordinate values of the lip key points in the reference image, the center point of the lips in the reference image is determined, and the center point is used as the reference point. According to the coordinate values of the lip key points in each frame image, the horizontal coordinate value and the vertical coordinate value with the largest absolute value among the lip key points in each frame image are determined. The absolute values of the horizontal and vertical coordinates with the largest absolute values are used as the horizontal and vertical coordinate reference values. The horizontal coordinate value of the lip key point of each frame image is subtracted from the horizontal coordinate value of the reference point, and then divided by the horizontal coordinate reference value, and the vertical coordinate value is subtracted from the vertical coordinate value of the reference point, and then divided by the vertical coordinate reference value, thereby completing the normalization processing of the lip key point information.
[0082] In the disclosed embodiment, a first detection frame is determined by using lip key point information, and then the length and / or width of the first detection frame are enlarged with the center position of the first detection frame as the center to obtain a second detection frame, and a lip area image is obtained based on the second detection frame, thereby achieving accurate acquisition of the lip area image.
[0083] Any lip movement detection method provided in the embodiments of the present disclosure can be executed by any appropriate device with data processing capabilities, including but not limited to a terminal device and a server. Alternatively, any lip movement detection method provided in the embodiments of the present disclosure can be executed by a processor, such as a processor that executes any lip movement detection method mentioned in the embodiments of the present disclosure by invoking corresponding instructions stored in a memory. This will not be further described below.
[0084] Exemplary devices
[0085] Figure 10 FIG. 1 is a block diagram of a lip movement detection device according to an embodiment of the present disclosure. Figure 10 As shown, the lip movement inspection device includes: an image acquisition module 100 , a first determination module 101 , a second determination module 102 , a third determination module 103 , and a detection module 104 .
[0086] An image acquisition module 100 is configured to acquire an image sequence including a facial region, wherein the image sequence includes multiple frames of images having a time-sequential relationship;
[0087] A first determining module 101 is configured to perform key point detection based on each frame of the image sequence to determine lip key point information of each frame of the image sequence;
[0088] A second determining module 102 is configured to determine a lip region image of each frame of the image sequence based on lip key point information of each frame of the image sequence;
[0089] A third determining module 103 is configured to determine the lip region image sequence according to the lip region images of the respective frames of image;
[0090] The detection module 104 is configured to perform lip movement detection based on the lip region image sequence and determine a lip movement detection result.
[0091] In one embodiment of the present disclosure, the detection module 104 includes:
[0092] The first detection submodule 1041 is configured to input the lip region image sequence into a first lip movement detection network, and output the lip movement detection result via the first lip movement detection network.
[0093] In one embodiment of the present disclosure, the detection module 104 further includes:
[0094] A second detection submodule 1042 is configured to input the lip region image sequence into a first subnetwork of a second lip movement detection network, and output image features of each lip region image in the lip region image sequence via the first subnetwork;
[0095] The third detection submodule 1043 is used to input the image features of each lip area image into the second subnetwork of the second lip movement detection network, and output the lip movement detection result through the second subnetwork.
[0096] In one embodiment of the present disclosure, the lip motion device further comprises:
[0097] A key point sequence acquisition module, configured to acquire a lip key point information sequence based on the lip key point information of each frame image in the image sequence;
[0098] The detection module 104 further includes:
[0099] A first stitching submodule 1044 is configured to determine stitching features based on the lip region image sequence and the lip key point information sequence;
[0100] The fourth detection submodule 1045 is configured to input the splicing features into a third lip movement detection network, and output the lip movement detection result via the third lip movement detection network.
[0101] In one embodiment of the present disclosure, the first splicing submodule 1044 includes:
[0102] a first feature acquisition unit 10441, configured to extract features from each lip region image in the lip region image sequence to obtain a lip region feature sequence;
[0103] The second feature acquisition unit 10442 performs feature extraction on the lip key point information sequence to obtain a lip key point feature sequence;
[0104] The splicing unit 10443 is used to splice the lip region feature sequence and the lip key point feature sequence to obtain a spliced feature.
[0105] Figure 11 FIG. 1 is a structural block diagram of the second determination module 102 in one embodiment of the present disclosure. Figure 11 As shown, the second determining module 102 includes:
[0106] A first sub-determination module 1021 is configured to determine a first detection frame of a lip region in each frame of the image sequence based on lip key point information of each frame of the image sequence;
[0107] A second sub-determination module 1022 is configured to enlarge the length and / or width of the first detection frame of the lip region in each frame image according to a preset magnification factor, with the center position of the first detection frame of the lip region in each frame image as the center, to obtain a second detection frame of the lip region in each frame image in the image sequence;
[0108] The third sub-determination module 1023 determines the lip region image of each frame image according to the second detection frame of the lip region in each frame image.
[0109] In one embodiment of the present disclosure, the lip movement detection result includes: lip movement type, wherein the lip movement type includes any one of the following: yawning, talking, eating, drinking, smoking, and no lip movement.
[0110] Exemplary electronic devices
[0111] Below, reference Figure 12 To describe the electronic device according to the embodiment of the present disclosure. Figure 12 As shown, the electronic device includes one or more processors 11 and a memory 12 .
[0112] The processor 11 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0113] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute the program instructions to implement the lip movement detection method of the various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage medium.
[0114] In one example, the electronic device may further include an input device 13 and an output device 14 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0115] For example, the input device 13 may be the microphone or microphone array mentioned above, which is used to capture the input signal of the sound source. In addition, the input device 13 may also include, for example, a keyboard, a mouse, etc.
[0116] The output device 14 can output various information to the outside, including determined distance information, direction information, etc. The output device 14 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output device, etc.
[0117] Of course, to simplify, Figure 12 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.
[0118] Exemplary computer program products and computer-readable storage media
[0119] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the lip movement detection method according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.
[0120] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0121] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the lip movement detection method according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0122] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0123] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0124] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.
[0125] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0126] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0127] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0128] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0129] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A lip movement detection method, comprising: Acquire an image sequence including a facial region, wherein the image sequence includes a plurality of image frames having a time sequence relationship; Performing key point detection based on each frame image of the image sequence to determine lip key point information of each frame image in the image sequence; Determining a lip region image of each frame image in the image sequence according to lip key point information of each frame image; determining the lip region image sequence according to the lip region images of the respective frames of image; The lip movement detection is performed on the lip area image sequence using the first lip movement detection network, the second lip movement detection network or the third lip movement detection network to obtain a lip movement detection result, including: inputting the lip area image sequence into the first subnetwork of the second lip movement detection network, and outputting the image features of each lip area image in the lip area image sequence through the first subnetwork, the first subnetwork including a plurality of second feature extraction units arranged in parallel, each second extraction unit including a plurality of feature extraction subunits arranged in series, and each feature extraction subunit including a convolution layer and a pooling layer; inputting the image features of each lip area image into the second subnetwork of the second lip movement detection network, and outputting the lip movement detection result through the second subnetwork, the second subnetwork including a convolution layer, a fully connected layer and a classifier; the lip movement detection result includes: lip movement type, and the lip movement type includes any one of the following: yawning, talking, eating, drinking, smoking, and no lip movement.
2. The method according to claim 1, wherein The performing lip movement detection based on the lip region image sequence and determining the lip movement detection result includes: The lip area image sequence is input into a first lip movement detection network, and the lip movement detection result is output through the first lip movement detection network.
3. The method according to claim 1, wherein Before determining the lip region image sequence according to the lip region images of the respective frames of images, the method further includes: Acquire a lip key point information sequence according to the lip key point information of each frame image in the image sequence; The performing lip movement detection based on the lip region image sequence and determining the lip movement detection result includes: Determining a splicing feature based on the lip region image sequence and the lip key point information sequence; The splicing features are input into a third lip movement detection network, and the lip movement detection result is output through the third lip movement detection network.
4. The method according to claim 3, wherein: The determining of the splicing features based on the lip region image sequence and the lip key point information sequence includes: performing feature extraction on each lip region image in the lip region image sequence to obtain a lip region feature sequence; Performing feature extraction on the lip key point information sequence to obtain a lip key point feature sequence; The lip region feature sequence and the lip key point feature sequence are spliced to obtain a spliced feature.
5. The method according to any one of claims 1 to 4, wherein The step of determining the lip region image of each frame image in the image sequence according to the lip key point information of each frame image includes: determining a first detection frame of a lip region in each frame image in the image sequence according to lip key point information of each frame image in the image sequence; Taking the center position of the first detection frame of the lip region in each frame image as the center, according to a preset magnification factor, the length and / or width of the first detection frame of the lip region in each frame image is magnified to obtain a second detection frame of the lip region in each frame image in the image sequence; The lip region image of each frame image is determined according to the second detection frame of the lip region in each frame image.
6. A lip movement detection device comprising: An image acquisition module, configured to acquire an image sequence including a facial area, wherein the image sequence includes multiple frames of images having a time sequence relationship; A first determining module is configured to perform key point detection based on each frame image of the image sequence to determine lip key point information of each frame image in the image sequence; A second determining module is configured to determine a lip region image of each frame image in the image sequence based on lip key point information of each frame image; a third determining module, configured to determine the lip region image sequence according to the lip region images of the respective frame images; a detection module, configured to perform lip movement detection on the lip region image sequence using a first lip movement detection network, a second lip movement detection network, or a third lip movement detection network to obtain a lip movement detection result, wherein the lip movement detection result includes: a lip movement type, wherein the lip movement type includes any one of the following: yawning, talking, eating, drinking, smoking, and no lip movement; The detection module includes: a second detection submodule, used to input the lip area image sequence into the first subnetwork of the second lip movement detection network, and output the image features of each lip area image in the lip area image sequence through the first subnetwork, the first subnetwork includes a plurality of second feature extraction units arranged in parallel, each second extraction unit includes a plurality of feature extraction subunits arranged in series, and each feature extraction subunit includes a convolution layer and a pooling layer; a third detection submodule, used to input the image features of each lip area image into the second subnetwork of the second lip movement detection network, and output the lip movement detection result through the second subnetwork, and the second subnetwork includes a convolution layer, a fully connected layer and a classifier.
7. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the lip movement detection method according to any one of claims 1 to 5.
8. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the lip movement detection method described in any one of claims 1-5 above.
Citation Information
Patent Citations
Lip action detection method and device, equipment and storage medium
CN113642469A
Dangerous expression detection method and system based on deep learning for person to be detected
CN113723165A