Screen operation method, system and equipment based on air blowing and medium

By using facial contour, mouth shape, voiceprint and eye focus recognition algorithms, combined with blowing audio and facial images, the problem of inaccurate recognition in blowing interaction is solved, precise user interaction is achieved, and the user experience and natural interaction are improved.

CN120595935APending Publication Date: 2025-09-05ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410243348.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing air-blowing interaction technology cannot accurately identify the user's blowing action or intention, resulting in poor human-vehicle interaction and affecting the user's driving experience.

Method used

Through facial contour detection, blowing mouth shape recognition, blowing voiceprint recognition and eye focus recognition algorithms, combined with obtaining the user's blowing audio and facial image, the user's identity and line of sight focus are identified to achieve precise blowing interaction.

Benefits of technology

The accuracy of blowing behavior detection is improved, the naturalness and intuitiveness of user interaction are enhanced, personalized services are provided, and the user's interactive experience is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595935A_ABST
    Figure CN120595935A_ABST
Patent Text Reader

Abstract

The invention relates to a blowing-based screen operation method, system and equipment and a medium, and belongs to the field of blowing detection. The method comprises the following steps: acquiring a blowing audio and a face image of a user, and performing face contour detection on the face image based on a face contour recognition algorithm; when the face contour of the user meets a preset detection condition, whether the mouth shape of the user in the face image is a blowing mouth shape or not is judged based on a blowing mouth shape recognition algorithm; when the user is in the blowing nozzle type, the identity of the user is recognized according to the blowing audio of the user based on a blowing voiceprint recognition algorithm; wherein the identity of the user comprises a registered user and a non-registered user; when the identity of the user is a registered user, acquiring a sight focus from the face image based on an eyeball focus recognition algorithm; and executing an operation of the corresponding screen area according to the sight focus. The accuracy of blowing behavior detection is greatly improved, and effective blowing interaction between the screen and the user is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of air blowing detection, and in particular to an air blowing-based screen operation method, system, device and medium. Background Art

[0002] With the continuous advancement of intelligent technology, air-blowing interaction technology has gradually been widely used. Air-blowing interaction means that after the user blows on the screen, the screen detects the user's blowing and triggers a specific function, thereby realizing interaction with the user. This air-blowing interaction technology greatly frees up the user's hands, allowing the user to interact with the screen without directly touching the screen. However, existing air-blowing interaction solutions are often unable to accurately identify the user's blowing action or blowing intention, resulting in poor human-vehicle interaction, thereby affecting the user's driving experience. Therefore, it is necessary to provide a screen operation method, system, device and medium based on air blowing. Summary of the Invention

[0003] The present invention provides a screen operation method and system based on blowing, so as to solve the problem in the prior art that the user's blowing behavior cannot be accurately detected when the user's blowing behavior is detected.

[0004] The present invention provides a screen operation method based on blowing, comprising: obtaining a user's blowing audio and facial image, and performing facial contour detection on the facial image based on a facial contour recognition algorithm; when the user's facial contour meets preset detection conditions, judging whether the user's mouth shape in the facial image is a blowing mouth shape based on a blowing mouth shape recognition algorithm; when the user's mouth shape is a blowing mouth shape, identifying the user's identity based on the user's blowing audio based on a blowing voiceprint recognition algorithm; wherein the user's identity includes a registered user and an unregistered user; when the user's identity is a registered user, obtaining a sight focus from the facial image based on an eye focus recognition algorithm; and performing an operation on the corresponding screen area based on the sight focus.

[0005] In one embodiment of the present invention, the method of obtaining the sight focus from a facial image based on an eye focus recognition algorithm includes: extracting an eye area image from the facial image; performing feature extraction on the eye area image based on a stacked hourglass network to obtain an image in the center of the eye; and performing regression processing on the image in the center of the eye based on a densely connected network to obtain the sight focus.

[0006] In one embodiment of the present invention, before extracting the eye region image from the facial image, the method further includes: performing grayscale processing on the facial image.

[0007] In one embodiment of the present invention, the operation of the corresponding screen area is also affected by the position of the sound source, and obtaining the position of the sound source includes: locating the position of the sound source of the blowing audio based on a sound field localization algorithm.

[0008] In one embodiment of the present invention, the method of determining whether the user's mouth shape in a facial image is a blowing mouth shape based on a blowing mouth shape recognition algorithm includes: obtaining multiple key point information of the mouth in the facial image based on a facial key point detection algorithm; obtaining the height and width of the mouth according to each key point information, and calculating the ratio of the height to the width; determining whether the ratio is greater than a preset ratio threshold; if so, the mouth shape is a blowing mouth shape; if not, the mouth shape is not a blowing mouth shape.

[0009] In one embodiment of the present invention, the blow voiceprint recognition algorithm is based on the user's blow audio to identify the user's identity, including: resampling the blow audio to obtain monophonic fixed-frequency audio; performing a short-time Fourier transform on the monophonic fixed-frequency audio to obtain a spectrogram; extracting feature data from the spectrogram to obtain multiple blow voiceprint features; feature fusing the multiple blow voiceprint features, and determining the user's identity based on the fused blow voiceprint features.

[0010] In one embodiment of the present invention, before resampling the blowing audio to obtain the monophonic fixed-frequency audio, the method further includes normalizing the blowing audio.

[0011] In another aspect of the present invention, a screen operating system based on blowing is also provided, which includes: a facial contour recognition module, which is used to obtain a user's blowing audio and facial image, and perform facial contour detection on the facial image based on a facial contour recognition algorithm; a mouth shape recognition module, which is used to determine whether the user's mouth shape in the facial image is a blowing mouth shape based on a blowing mouth shape recognition algorithm when the user's facial contour meets preset detection conditions; an identity recognition module, which is used to identify the user's identity based on the user's blowing audio based on a blowing voiceprint recognition algorithm when the user's mouth shape is a blowing mouth shape; wherein the user's identity includes registered users and unregistered users; an eye tracking module, which is used to obtain the line of sight focus from the facial image based on the eye focus recognition algorithm when the user's identity is a registered user; and an operation module, which is used to perform operations on the corresponding screen area based on the line of sight focus.

[0012] In one embodiment of the present invention, an electronic device is also provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device implements any of the above-mentioned screen operation methods based on blowing.

[0013] In one embodiment of the present invention, a computer-readable storage medium is further provided, characterized in that a computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute any of the above-mentioned screen operation methods based on blowing.

[0014] The present invention proposes a screen operation method, system, device, and medium based on blowing. When a user blows toward the screen, a facial contour detection algorithm is used to detect the user's facial contour. If the detection conditions are met, a blow mouth shape recognition algorithm is used to determine whether the user has a blow mouth shape. If the user has a blow mouth shape, a blow voiceprint recognition algorithm is used to identify the user's identity based on the blowing audio. If the user is a registered user, an eye focus recognition algorithm is used to determine the user's gaze focus. Based on this gaze focus, the corresponding operation is performed in the corresponding area of ​​the screen. On one hand, the present invention can accurately capture and understand the user's natural behavior patterns, greatly improving the naturalness and intuitiveness of the interaction process. It also enhances the user's interactive experience through personalized user identification and response. On the other hand, the present invention combines multiple algorithms, including a facial contour detection algorithm, a blow mouth shape recognition algorithm, a blow voiceprint recognition algorithm, and an eye focus recognition algorithm, to comprehensively identify the interrelated facial contours, blow mouth shape, user identity, and gaze focus, thereby significantly improving the accuracy of blowing behavior detection. This allows the screen to perform corresponding operations based on the user's blowing behavior, realizing effective blowing interaction between the screen and the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A schematic flow chart of a screen operation method based on air blowing provided in an embodiment of the present invention;

[0016] Figure 2 A schematic diagram of the user identity detection process provided by an embodiment of the present invention;

[0017] Figure 3 Shown is a schematic diagram of the flow of sight focus detection provided by an embodiment of the present invention;

[0018] Figure 4 Shown is a schematic diagram of key points of a face provided by an embodiment of the present invention;

[0019] Figure 5 A schematic diagram showing the information entry process provided by an embodiment of the present invention;

[0020] Figure 6 Shown is a structural block diagram of a screen operation based on blowing provided by an embodiment of the present invention;

[0021] Figure 7 Shown is a structural schematic diagram of an electronic device using a screen operation method based on air blowing. DETAILED DESCRIPTION

[0022] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0023] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0024] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0025] The present invention provides a screen operation method based on blowing. When a user blows toward the screen, a facial contour detection algorithm is used to detect the user's facial contour. If the detection conditions are met, a blow mouth shape recognition algorithm is used to determine whether the user has a blow mouth shape. If the user has a blow mouth shape, a blow voiceprint recognition algorithm is used to identify the user's identity based on the blowing audio. If the user is a registered user, an eye focus recognition algorithm is used to obtain the user's gaze focus. Based on this gaze focus, corresponding operations are performed in the corresponding area of ​​the screen. On one hand, the present invention can accurately capture and understand the user's natural behavioral patterns, not only ensuring the purposefulness and intention of user interaction, but also providing personalized services by identifying specific users, thereby enhancing the user's interactive experience. On the other hand, the present invention combines multiple algorithms, such as a facial contour detection algorithm, a blow mouth shape recognition algorithm, a blow voiceprint recognition algorithm, and an eye focus recognition algorithm, to comprehensively identify the interrelated facial contours, blow mouth shape, user identity, and gaze focus. This improves the accuracy of identifying user blowing interactions, reduces the occurrence of misidentification, and greatly improves the accuracy of blowing behavior detection. This allows the screen to respond to the user's blowing behavior, achieving effective interaction between the screen and the user. Users only need to perform simple blowing movements to interact with the device, eliminating the need for complex operations. This is simple and convenient, greatly improving the user experience.

[0026] See Figures 1 to 3 , the screen operation method based on blowing includes the following steps:

[0027] S1. Obtain the user's blowing audio and facial image, and perform facial contour detection on the facial image based on a facial contour recognition algorithm.

[0028] After the user blows towards the screen, the recording device (microphone, etc.) on the screen will collect the user's blowing audio, and the camera installed on the screen will automatically capture the user's facial image at this time, and detect it through the facial contour recognition algorithm. Specifically, the facial contour can be detected based on the key point algorithm, and the offset of the user's face relative to the screen can be obtained according to the number of detected key points on the user's face and the corresponding positions of the key points on the face, thereby determining the angular relationship between the face and the screen during the interaction. For example, the center of the screen can be used as a reference point, and the angular relationship between the face and the screen can be obtained according to the degree of deviation of the key points on the face relative to the reference point. It should be noted that the calculation method of the angular relationship between the face and the screen is not limited to the above method, and those skilled in the art can adaptably select an appropriate method based on actual needs. It can be understood that the screen mentioned in the present invention can be a car screen or a screen of other electronic equipment, which is not limited here.

[0029] S2. When the user's facial contour meets the preset detection conditions, based on the blowing mouth shape recognition algorithm, determine whether the user's mouth shape in the facial image is a blowing mouth shape.

[0030] The facial contour recognition algorithm can be used to determine the angular relationship between the user's face and the screen. The detection condition can be to determine whether the angle is less than a preset angle threshold. If so, it is considered that the user can observe the operations presented on the screen at this time, and therefore the user's facial contour is determined to meet the detection condition. If not, it is considered that the user is far away from the screen at this time and cannot clearly observe the operations presented on the screen, and therefore the user's facial contour is determined to not meet the detection condition. When the detection condition is met, the user's blowing mouth shape can be identified to distinguish whether the user is in a blowing state. In one embodiment of the present invention, the method of determining whether the user's mouth shape in the facial image is a blowing mouth shape based on the blowing mouth shape recognition algorithm includes:

[0031] Based on the facial key point detection algorithm, multiple key point information of the mouth in the facial image is obtained;

[0032] Get the height and width of the mouth based on the information of each key point, and calculate the ratio of height to width;

[0033] Determining whether the ratio is greater than a preset ratio threshold;

[0034] If so, the mouth shape is a blowing mouth shape;

[0035] If not, the mouth shape is not a blowing mouth shape.

[0036] See also Figure 4 Considering that people tend to pout their lips when blowing, a facial key point detection algorithm can be used to obtain 68 key points from the facial image. Key point information for points 52, 58, 49, and 55 is extracted from these 68 key points to obtain mouth position information, thereby identifying the lip blowing action. Specifically, the mouth height h is calculated based on the distance between points 52 and 58, and the mouth width w is calculated based on the distance between points 49 and 55. By calculating the ratio of height to width R = w / h, if the ratio R is greater than the ratio threshold, it is determined to be a blowing mouth shape. If R is less than the ratio threshold, it is determined not to be a blowing mouth shape.

[0037] S3. When the user's mouth is in the blowing position, the user's identity is identified based on the blowing voiceprint recognition algorithm according to the user's blowing audio; wherein the user's identity includes registered users and unregistered users.

[0038] Taking into account the correspondence between the screen and the user, for example, for the car screen, there are only a few fixed users who frequently use the car. Since these fixed users can enter their identity information before interacting, they become registered users. Users who use the screen temporarily are called unregistered users. For registered users, more diverse operations can be performed. Unregistered users can only perform some simple interactions with the screen because there is no information about this user in the system. In order to distinguish the identity of the user, when the user has a blowing mouth shape, the blowing audio is input into a pre-trained audio recognition model. Through the feature extraction and matching of multiple blowing sounds, the user's blowing voiceprint is identified to determine the user's identity. Specifically, the audio recognition model is the YAMNet model. YAMNet is an open source pre-trained model of the TensorFlow Center. It has been trained for audio event prediction on millions of YouTube videos. The network is based on the MobileNet architecture and is very suitable for embedded applications. It can provide a good benchmark for application developers. Specifically, please refer to Figure 2 In one embodiment of the present invention, the method of identifying the user's identity based on the user's blowing audio based on the blowing voiceprint recognition algorithm includes:

[0039] S31, resampling the blowing audio to obtain a monophonic fixed-frequency audio;

[0040] S32, performing short-time Fourier transform on the monophonic fixed-frequency audio to obtain a spectrogram;

[0041] S33, extracting characteristic data of the spectrogram to obtain multiple blowing voiceprint features;

[0042] S34. Feature fusion of multiple blowing voiceprint features, and determining the user's identity based on the fused blowing voiceprint features.

[0043] The blowing audio is resampled to a 16kHz monophonic fixed audio. Since the sound signal is a one-dimensional signal, only the time domain information can be seen intuitively, and the frequency domain information cannot be seen. The frequency domain can be transformed through Fourier transform (FT), but the time domain information is lost and the time-frequency relationship cannot be seen. Therefore, in this application, a time window of 25 milliseconds in length and 10 milliseconds in step size is added to the monophonic fixed audio, and a short-time Fourier transform is performed to calculate the spectrogram. Since the obtained spectrogram is large, in order to obtain sound features of appropriate size, the spectrogram needs to be converted into a Mel spectrum through a Mel-scale filter bank. The Mel spectrum is fed into the Mobilenet_v1 model to extract multiple blowing voiceprint features. Multiple blowing voiceprint features are feature fused, and based on the fused blowing voiceprint features, it can be determined whether the user is a registered user. By identifying the blowing voiceprint and judging the user's identity, personalized identification and services for different users can be achieved, thereby improving user satisfaction.

[0044] In one embodiment of the present invention, before resampling the blowing audio to obtain the monophonic fixed-frequency audio, the method further includes normalizing the blowing audio.

[0045] S4. When the user is a registered user, the eye focus is obtained from the facial image based on the eye focus recognition algorithm.

[0046] When the user is a registered user, the eye focus recognition algorithm can be used to determine the screen position where the user's eyes are focused when using the interaction. For details, please refer to Figure 3 In one embodiment of the present invention, obtaining the sight focus from the facial image based on the eye focus recognition algorithm includes:

[0047] S41, extracting an eye region image from the facial image;

[0048] S42. Perform feature extraction on the eye region image according to the stacked hourglass network to obtain an image of the center of the eye;

[0049] S43. Perform regression processing on the image in the middle of the eye based on a densely connected network to obtain the visual focus.

[0050] After extracting the eye region image from the facial image, the extracted eye region image is input into the gaze tracking model. The gaze tracking model of the present invention is a deep pictorial gaze estimation model. Assuming that the eyeball is a standard sphere and the pupil is a standard circle on the sphere, the projection of the eyeball and pupil onto the image is a standard circle and a small ellipse within the circle (when the eyeball does not rotate, the pupil projection becomes a circle. The greater the eyeball rotates, the more elliptical the ellipse). The resulting image is called a gazemap. Through this three-dimensional to two-dimensional mapping relationship, the image is converted from three dimensions to two dimensions. Since the gaze tracking model includes a stacked hourglass network (stacked hourglass network) and a densely connected network (DenseNet), features can be extracted from the eye image through the stacked hourglass network. According to the 3D to 2D mapping relationship, the eyeball and pupil are mapped into a two-dimensional image to generate an image in the middle of the eye (gazemap). The image in the middle of the eye is input into the densely connected network, and through regression processing, the gaze focus output by the gaze tracking model is obtained. By identifying the focus of the eyeball, the user's line of sight can be determined, so that the device can respond promptly to changes in the user's line of sight, thereby improving the device's interactive sensitivity.

[0051] In one embodiment of the present invention, before extracting the eye region image from the facial image, the process further includes grayscaling the facial image. Grayscaling the color image reduces the data size from the original 0-256*256*256 to 256 values ​​of 0-255, significantly reducing computational complexity and improving the model's recognition rate.

[0052] S5. Execute operations on the corresponding screen area according to the visual focus.

[0053] See also Figure 5 , according to the user's gaze focus and the preset scene, certain operations are performed in the screen area corresponding to the gaze, and displayed on the screen area, thereby realizing interaction with the user. For example, when the gaze focus is on point A on the screen, for the candle blowing scene, after detecting the user's blowing voiceprint, the candle at point A is extinguished, and the extinguishing process is displayed in area A. It should be noted that before the user performs this blowing interaction, he will first select a certain preset scene on the screen, so that when the user blows, the screen will operate on the corresponding area on the screen according to the user's blowing and the preset scene.

[0054] Furthermore, in one embodiment of the present invention, the operation of the corresponding screen area is also affected by the location of the sound source. Acquiring the location of the sound source includes: locating the location of the sound source of the blowing audio based on a sound field localization algorithm. Sound field localization algorithms include, but are not limited to, MUSIC algorithms, MVDR algorithms, etc. For example, for a book page turning scenario, if the sound source is detected on the left side of the screen, the page is turned from the left to the right, and vice versa, the page is turned from the right to the left. The left side refers to the user's left side when the user is facing the display interface of the screen; the right side refers to the user's right side when the user is facing the display interface of the screen.

[0055] The process of entering relevant information of a registered user of the present invention is as follows: using a camera to capture the user's usual mouth shape and the mouth shape when blowing, and entering the mouth images of different shapes into the system. By entering mouth images of different shapes, it is possible to effectively avoid the misidentification of users with naturally upturned lips. When entering a facial image, the user is required to expose both eyes and part of the mouth so that the user's facial features can be fully extracted, which facilitates facial contour recognition of the user during subsequent use. The user blows air while looking at the corresponding letters on the screen according to the prompts, which enables eye tracking and voiceprint recognition of the user. Through the above process, the user's information can be saved in the system, so that the next time the user interacts with the screen, the user can be identified as a registered user, thereby achieving more accurate interactive control. It is understandable that when the user enters a facial image, various angles can be adapted, including but not limited to looking straight at the screen, looking at the screen sideways, looking down at the screen, and looking up at the screen.

[0056] See Figure 6 The air-blowing screen operating system 100 includes a facial contour recognition module 110, a mouth shape recognition module 120, an identity recognition module 130, an eye tracking module 140, and an operation module 150. The facial contour recognition module 110 is used to obtain the user's air-blowing audio and facial image, and perform facial contour detection on the facial image based on a facial contour recognition algorithm. The mouth shape recognition module 120 is used to determine whether the user's mouth shape in the facial image is an air-blowing mouth shape based on the air-blowing mouth shape recognition algorithm when the user's facial contour meets preset detection conditions. The identity recognition module 130 is used to identify the user's identity based on the air-blowing audio using the air-blowing voiceprint recognition algorithm if the user's mouth shape is an air-blowing mouth shape. User identities include registered and unregistered users. The eye tracking module 140 is used to obtain the gaze focus from the facial image based on the eye focus recognition algorithm if the user's identity is a registered user. The operation module 150 is used to perform operations on the corresponding screen area based on the gaze focus.

[0057] It should be noted that, in order to highlight the innovative part of the present invention, this embodiment does not introduce modules that are not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other modules in this embodiment.

[0058] See Figure 7 The electronic device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a screen operation program based on blowing.

[0059] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 12 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 1. Furthermore, the memory 12 can also include both an internal storage unit of the electronic device 1 and an external storage device. The memory 12 can not only be used to store application software and various types of data installed in the electronic device 1, such as codes for screen operations based on blowing, etc., but can also be used to temporarily store data that has been output or is to be output.

[0060] In some embodiments, the processor 13 may be composed of an integrated circuit, such as a single packaged integrated circuit or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting the various components of the entire electronic device 1 using various interfaces and circuits. It executes or runs programs or modules stored in the memory 12 (such as a charging solution recommendation program) and calls data stored in the memory 12 to perform various functions of the electronic device 1 and process data.

[0061] The processor 13 executes the operating system and various installed applications of the electronic device 1. The processor 13 executes the applications to implement the steps in the above-mentioned screen operation method based on blowing.

[0062] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into a facial contour recognition module 110, a mouth shape recognition module 120, an identity recognition module 130, an eye tracking module 140, and an operation module 150.

[0063] The above-mentioned integrated unit implemented in the form of a software function module can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above-mentioned software function module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, computer device, or network device, etc.) or a processor to perform part of the functions of the screen operation method based on blowing described in various embodiments of the present application.

[0064] In summary, the present invention discloses a screen operation method, system, device, and medium based on blowing. When a user blows toward the screen, a facial contour detection algorithm is used to detect the user's facial contour. If the detection conditions are met, a blowing mouth shape recognition algorithm is used to determine whether the user has a blowing mouth shape. If the user has a blowing mouth shape, a blowing voiceprint recognition algorithm is used to identify the user's identity based on the blowing audio. If the user is a registered user, an eye focus recognition algorithm is used to obtain the user's gaze focus. Based on the gaze focus, the corresponding operation is performed in the corresponding area of ​​the screen. On the one hand, the present invention can accurately capture and understand the user's natural behavior patterns, not only ensuring the purposefulness and intention of user interaction, but also providing personalized services by identifying specific users, thereby enhancing the user's interactive experience. On the other hand, the present invention combines multiple algorithms, such as a facial contour detection algorithm, a blowing mouth shape recognition algorithm, a blowing voiceprint recognition algorithm, and an eye focus recognition algorithm, to comprehensively identify the interrelated facial contours, blowing mouth shape, user identity, and gaze focus, thereby greatly improving the accuracy of blowing behavior detection. This allows the screen to respond to the user's blowing behavior, achieving effective interaction between the screen and the user. Furthermore, compared to contact-based blowing, this method of screen interaction through air blowing significantly increases user freedom, making it easier for people with disabilities and other difficulties operating screens to interact with the screen. It also eliminates contact with the screen, resulting in higher hygiene standards. Therefore, this invention effectively overcomes the shortcomings of existing technologies and has high industrial application value.

[0065] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. A screen operation method based on blowing, characterized in that: The method comprises: Obtain the user's blowing audio and facial image, and perform facial contour detection on the facial image based on the facial contour recognition algorithm; When the user's facial contour meets the preset detection conditions, the user's mouth shape in the facial image is judged to be a blowing mouth shape based on the blowing mouth shape recognition algorithm; When the user's mouth is in the blowing position, the user's identity is identified based on the blowing audio based on the blowing voiceprint recognition algorithm; the user's identity includes registered users and unregistered users; When the user is a registered user, the eye focus is obtained from the facial image based on the eye focus recognition algorithm; Execute operations on the corresponding screen area based on the line of sight.

2. The screen operation method based on blowing according to claim 1, characterized in that: The step of obtaining the sight focus from the facial image based on the eye focus recognition algorithm includes: Extracting eye region images from facial images; Based on the stacked hourglass network, feature extraction is performed on the eye area image to obtain the image in the middle of the eye; The image in the middle of the eye is regressed based on a densely connected network to obtain the visual focus.

3. The screen operation method based on blowing according to claim 2, characterized in that: Before extracting the eye region image from the facial image, the method further includes: performing grayscale processing on the facial image.

4. The screen operation method based on blowing according to claim 1, characterized in that: The operation of the corresponding screen area is also affected by the position of the sound source. The acquisition of the sound source position includes: locating the sound source position of the blowing audio based on a sound field localization algorithm.

5. The screen operation method based on blowing according to claim 1, characterized in that: The method of determining whether the user's mouth shape in the facial image is a blowing mouth shape based on the blowing mouth shape recognition algorithm includes: Based on the facial key point detection algorithm, multiple key point information of the mouth in the facial image is obtained; Get the height and width of the mouth based on the information of each key point, and calculate the ratio of height to width; Determining whether the ratio is greater than a preset ratio threshold; If so, the mouth shape is a blowing mouth shape; If not, the mouth shape is not a blowing mouth shape.

6. The screen operation method based on blowing according to claim 1, characterized in that: The method of identifying the user's identity based on the user's blowing audio based on the blowing voiceprint recognition algorithm includes: Resample the blowing audio to obtain mono fixed-frequency audio; Perform short-time Fourier transform on the monophonic fixed-frequency audio to obtain a spectrogram; Extracting characteristic data of the spectrogram to obtain multiple blowing voiceprint features; The feature fuses multiple blowing voiceprint features and determines the user's identity based on the fused blowing voiceprint features.

7. The screen operation method based on blowing according to claim 6, characterized in that: Before resampling the blowing audio to obtain the monophonic fixed-frequency audio, the method further includes normalizing the blowing audio.

8. A screen operating system based on air blowing, characterized in that: The system comprises: A facial contour recognition module is used to obtain the user's blowing audio and facial image, and perform facial contour detection on the facial image based on the facial contour recognition algorithm; A mouth shape recognition module is used to determine whether the user's mouth shape in the facial image is a blowing mouth shape based on the blowing mouth shape recognition algorithm when the user's facial contour meets the preset detection conditions; An identity recognition module is used to identify the user's identity based on the user's blowing audio when the user is in the blowing mouth shape based on the blowing voiceprint recognition algorithm; the user identity includes registered users and unregistered users; An eye tracking module, which is used to obtain the gaze focus from the facial image based on the eye focus recognition algorithm when the user is a registered user; The operation module is used to perform operations on the corresponding screen area according to the visual focus.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the screen operation method based on blowing as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the screen operation method based on blowing according to any one of claims 1 to 7.

Citation Information

Cited By

  • Eye movement tracking method and system based on stacked hourglass network and homography transformation

    CN121635683A