Sight Detection Method, Device, Electronic Device and Storage Medium
By determining the classification results of lateral eye in line of sight detection and performing eye area correction processing, the problem of low line of sight detection accuracy for people with abnormal vision is solved, and effective line of sight detection for people with transverse eyes is achieved.
Patent Information
- Application Number
- CN202210096231.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The prior art cannot effectively detect the line of sight direction of people with abnormal vision, resulting in low or unachievable visual detection accuracy.
By acquiring the image to be detected, the squinting eyes classification results in the face image are determined, and the corresponding line of sight direction detection method is adopted, including correcting the eye area in the strabismus state, and using the line of sight direction detection model to detect the line of sight direction.
The line of sight detection for people with squinting eyes is realized, and the applicability and accuracy of line of sight detection is improved.
Smart Images

Figure CN114495252B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method, apparatus, electronic device, and storage medium for gaze detection. Background Art
[0002] Through research, it is found that in some scenarios (such as vehicle driving scenarios, VR game scenarios), it is necessary to perform gaze detection on a target person to determine the gaze direction of the target person. However, currently, gaze detection can only be performed on people with normal vision, and for some people with abnormal vision (such as strabismus), since the human eye will deviate from the fixation target, gaze detection cannot be achieved or the gaze detection accuracy is low. Summary of the Invention
[0003] Embodiments of the present disclosure at least provide a method, apparatus, electronic device, and storage medium for gaze detection.
[0004] Embodiments of the present disclosure provide a method for gaze detection, including:
[0005] Obtain an image to be detected;
[0006] When it is determined that the image to be detected contains a face image, determine the strabismus classification result of the target person indicated by the face image, where the strabismus classification result represents that the eye state of the target person is a strabismus state with inconsistent binocular gaze or a normal state with consistent binocular gaze;
[0007] Detect the gaze direction of the target person by using a gaze direction detection method corresponding to the strabismus classification result.
[0008] In a possible implementation manner, the detecting the gaze direction of the target person by using a gaze direction detection method corresponding to the strabismus classification result includes:
[0009] When the strabismus classification result represents that the eye state of the target person is the normal state, input the face image into a gaze direction detection model to obtain the gaze direction of the target person; or,
[0010] When the strabismus classification result represents that the eye state of the target person is the strabismus state, perform correction processing on the eye region of the target person in the face image to obtain a face image or eye image with the corrected eye state, and input the face image or eye image with the corrected eye state into the gaze direction detection model to obtain the gaze direction of the target person.
[0011] In a possible implementation manner, the performing correction processing on the eye region of the target person in the face image includes:
[0012] Perform correction processing on the area corresponding to the single eye in the esotropia state of both eyes of the target person.
[0013] In a possible implementation manner, the esotropia state includes a left-eye esotropia state and a right-eye esotropia state, and the esotropic eye classification result includes a result indicating that the eye state of the target person is a left-eye esotropia state, or a result indicating that the eye state of the target person is a right-eye esotropia state;
[0014] When the esotropic eye classification result indicates that the eye state of the target person is the esotropia state, performing correction processing on the eye area of the target person in the face image includes:
[0015] When the esotropic eye classification result indicates that the eye state of the target person is a left-eye esotropia state, performing masking processing on the left-eye area of the target person in the face image to remove the left-eye area of the target person in the face image; or,
[0016] When the esotropic eye classification result indicates that the eye state of the target person is a right-eye esotropia state, performing masking processing on the right-eye area of the target person in the face image to remove the right-eye area of the target person in the face image.
[0017] In a possible implementation manner, determining the esotropic eye classification result of the target person indicated by the face image includes:
[0018] Identifying the identity information of the target person indicated by the face image;
[0019] Based on the preset correspondence between the personnel identity information and the esotropic eye category, determining the esotropic eye category corresponding to the identity information of the identified target person as the esotropic eye classification result of the target person.
[0020] In a possible implementation manner, determining the esotropic eye classification result of the target person indicated by the face image includes:
[0021] Based on the face image, determining the human eye area image;
[0022] Based on the trained esotropia detection model, performing esotropia detection on the human eye area image to obtain the esotropic eye classification result of the target person.
[0023] In a possible implementation manner, the method further includes:
[0024] Obtaining an image sample set, where the image sample set includes human eye images or face images in a single-eye esotropia state;
[0025] Based on the set of image samples, train the strabismus detection model to be trained to obtain the trained strabismus detection model.
[0026] In a possible implementation manner, the images in the set of image samples are obtained by the following method:
[0027] Obtain a target face image or a target eye image with both eyes looking straight ahead;
[0028] Perform a line-of-sight redirection process on a single eye in the target face image or the target eye image to obtain a processed target face image or a processed target eye image;
[0029] Generate the images in the set of image samples based on the processed target face image or the processed target eye image.
[0030] In a possible implementation manner, the images in the set of image samples are obtained by the following method:
[0031] Respectively obtain a first target face image with both eyes looking straight ahead and a second target face image with both eyes squinting, or respectively obtain a first target eye image with both eyes looking straight ahead and a second target eye image with both eyes squinting; and
[0032] Generate the images in the set of image samples by at least one of the following methods:
[0033] Stitch the left eye image region in the first target face image with the right eye image region in the second target face image;
[0034] Stitch the right eye image region in the first target face image with the left eye image region in the second target face image;
[0035] Stitch the left eye image region in the first target eye image with the right eye image region in the second target eye image;
[0036] Stitch the right eye image region in the first target eye image with the left eye image region in the second target eye image.
[0037] In a possible implementation manner, the image to be detected includes multiple frames of images. When it is determined that the image to be detected contains a face image, determining the strabismus classification result of the target person indicated by the face image includes:
[0038] When it is determined that at least one frame of the multiple frames of images to be detected contains a face image, determine the strabismus classification result of each frame of the image to be detected of the target person under the at least one frame of the image to be detected;
[0039] Determine the squint classification result of the target person based on the squint classification results of each frame of the image to be detected among the at least one frame of the image to be detected.
[0040] In a possible implementation manner, the image to be detected includes an image of the driving area of a vehicle in a driving state, and the method further includes:
[0041] Based on the detection result of the line of sight direction of the target person and the driving direction of the vehicle, determine whether the line of sight direction of the target person deviates from a preset direction;
[0042] When the line of sight direction of the target person deviates from the preset direction, send a prompt message.
[0043] An embodiment of the present disclosure provides a line of sight detection device, including:
[0044] An image acquisition module, configured to acquire an image to be detected;
[0045] A squint detection module, configured to determine the squint classification result of the target person indicated by the face image when it is determined that the image to be detected contains a face image, where the squint classification result characterizes that the eye state of the target person is a strabismus state with inconsistent binocular lines of sight or a normal state with consistent binocular lines of sight;
[0046] A line of sight detection module, configured to detect the line of sight direction of the target person by using a line of sight direction detection method corresponding to the squint classification result.
[0047] In a possible implementation manner, the line of sight detection module is specifically configured to:
[0048] When the squint classification result characterizes that the eye state of the target person is the normal state, input the face image into a line of sight direction detection model to obtain the line of sight direction of the target person; or,
[0049] When the squint classification result characterizes that the eye state of the target person is the strabismus state, perform correction processing on the eye area of the target person in the face image to obtain a face image or an eye image after correcting the eye state, and input the face image or the eye image after correcting the eye state into the line of sight direction detection model to obtain the line of sight direction of the target person.
[0050] In a possible implementation manner, the line of sight detection module is specifically configured to:
[0051] Perform correction processing on the area corresponding to the monocular with a strabismus state in the two eyes of the target person.
[0052] In a possible implementation, the strabismus state includes a left-eye strabismus state and a right-eye strabismus state, and the strabismus classification result includes a result indicating that the eye state of the target person is the left-eye strabismus state, or a result indicating that the eye state of the target person is the right-eye strabismus state;
[0053] The line-of-sight detection module is specifically configured to:
[0054] When the strabismus classification result indicates that the eye state of the target person is the left-eye strabismus state, perform masking processing on the left-eye area of the target person in the face image to remove the left-eye area of the target person in the face image; or,
[0055] When the strabismus classification result indicates that the eye state of the target person is the right-eye strabismus state, perform masking processing on the right-eye area of the target person in the face image to remove the right-eye area of the target person in the face image.
[0056] In a possible implementation, the strabismus detection module is specifically configured to:
[0057] Identify the identity information of the target person indicated by the face image;
[0058] Based on the preset correspondence between the personnel identity information and the strabismus category, determine the strabismus category corresponding to the identity information of the identified target person as the strabismus classification result of the target person.
[0059] In a possible implementation, the strabismus detection module is specifically configured to:
[0060] Based on the face image, determine the human eye area image;
[0061] Based on the trained strabismus detection model, perform strabismus detection on the human eye area image to obtain the strabismus classification result of the target person.
[0062] In a possible implementation, the device further includes a model training module, and the model training module is used to:
[0063] Obtain an image sample set, where the image sample set includes human eye images or face images with monocular strabismus states;
[0064] Based on the image sample set, train the to-be-trained strabismus detection model to obtain the trained strabismus detection model.
[0065] In a possible implementation, the model training module is specifically configured to:
[0066] Obtain a target face image or a target human eye image with both eyes looking straight ahead;
[0067] Perform gaze redirection processing on a single eye in the target face image or the target eye image to obtain the processed target face image or target eye image;
[0068] Generate an image in the image sample set based on the processed target face image or target eye image.
[0069] In a possible implementation manner, the model training module is specifically configured to:
[0070] Obtain a first target face image with both eyes looking straight ahead and a second target face image with both eyes squinting respectively, or obtain a first target eye image with both eyes looking straight ahead and a second target eye image with both eyes squinting respectively; and
[0071] Generate an image in the image sample set by at least one of the following methods:
[0072] Stitch the left eye image region in the first target face image with the right eye image region in the second target face image;
[0073] Stitch the right eye image region in the first target face image with the left eye image region in the second target face image;
[0074] Stitch the left eye image region in the first target eye image with the right eye image region in the second target eye image;
[0075] Stitch the right eye image region in the first target eye image with the left eye image region in the second target eye image.
[0076] In a possible implementation manner, the image to be detected includes multiple frames of images, and the squint detection module is specifically configured to:
[0077] When it is determined that at least one frame of the images to be detected in the multiple frames of images to be detected contains a face image, determine the squint classification result of each frame of the images to be detected of the target person under the at least one frame of the images to be detected;
[0078] Based on the squint classification results of each frame of the images to be detected under the at least one frame of the images to be detected, determine the squint classification result of the target person.
[0079] In a possible implementation manner, the image to be detected includes an image of the driving area of a vehicle in a driving state, and the device further includes a gaze judgment module, and the gaze judgment module is used for:
[0080] Determine whether the line-of-sight direction of the target person deviates from the preset direction based on the detection result of the line-of-sight direction of the target person and the driving direction of the vehicle;
[0081] When the line-of-sight direction of the target person deviates from the preset direction, a prompt message is sent.
[0082] An embodiment of the present disclosure provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the line-of-sight detection method described in any of the above embodiments are executed.
[0083] An embodiment of the present disclosure provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the line-of-sight detection method described in any of the above embodiments are executed.
[0084] In the line-of-sight detection method, device, electronic device, and storage medium provided in the embodiments of the present disclosure, when it is determined that the to-be-detected image contains a face image, first determine the squint classification result of the target person indicated by the face image, and then based on the squint classification result of the target person, adopt a line-of-sight direction detection method corresponding to the squint classification result to detect the line-of-sight direction of the target person. In this way, not only can the line-of-sight detection of squint people be realized, the applicability of the line-of-sight detection method is improved, but also the accuracy of the line-of-sight detection can be improved.
[0085] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for the embodiments. The accompanying drawings are incorporated into the specification and form a part of the specification. These drawings show embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0087] Figure 1 Shows a flowchart of a line-of-sight detection method provided by an embodiment of the present disclosure;
[0088] Figure 2Shows a flowchart of a method for determining the squint classification result of a target person provided by an embodiment of the present disclosure;
[0089] Figure 3 Shows a flowchart of a method for training a strabismus detection model provided by an embodiment of the present disclosure;
[0090] Figure 4 Shows a flowchart of a method for obtaining an image sample provided by an embodiment of the present disclosure;
[0091] Figure 5 Shows a schematic diagram of the first type of face image sample provided by an embodiment of the present disclosure;
[0092] Figure 6 Shows Figure 5 Schematic diagram after redirecting the face image in;
[0093] Figure 7 Shows a flowchart of another method for obtaining an image sample provided by an embodiment of the present disclosure;
[0094] Figure 8 Shows a schematic diagram of the second type of face image sample provided by an embodiment of the present disclosure;
[0095] Figure 9 Shows a schematic diagram of the third type of face image sample provided by an embodiment of the present disclosure;
[0096] Figure 10 Shows a schematic diagram of a human eye image sample provided by an embodiment of the present disclosure;
[0097] Figure 11 Shows a flowchart of another line-of-sight detection method provided by an embodiment of the present disclosure;
[0098] Figure 12 Shows a schematic structural diagram of a line-of-sight detection device provided by an embodiment of the present disclosure;
[0099] Figure 13 Shows a schematic structural diagram of another line-of-sight detection device provided by an embodiment of the present disclosure;
[0100] Figure 14 Shows a schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0101] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Components of the embodiments of the present disclosure described and illustrated in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0102] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.
[0103] As used herein, the term "and / or" merely describes an associated relationship and indicates that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, both A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may mean including any one or more elements selected from the set composed of A, B, and C.
[0104] Through research, it has been found that in some scenarios (such as vehicle driving scenarios, VR game scenarios), it is necessary to perform gaze detection on a target person to determine the gaze direction of the target person. However, currently, gaze detection can only be achieved for people with normal vision. For some people with abnormal vision (such as strabismus), since the human eye will deviate from the fixation target, gaze detection cannot be achieved or the gaze detection accuracy is low.
[0105] Based on the above research, the present disclosure provides a gaze detection method, including: obtaining an image to be detected; in the case of determining that the image to be detected contains a face image, determining the strabismus classification result of the target person indicated by the face image, where the strabismus classification result characterizes that the eye state of the target person is a strabismus state with inconsistent binocular gaze or a normal state with consistent binocular gaze; and detecting the gaze direction of the target person by using a gaze direction detection method corresponding to the strabismus classification result.
[0106] In an embodiment of the present disclosure, when it is determined that the image to be detected contains a face image, first determine the squint classification result of the target person indicated by the face image, and then based on the squint classification result of the target person, adopt a line-of-sight direction detection method corresponding to the squint classification result to detect the line-of-sight direction of the target person. In this way, not only can the line-of-sight detection of squint people be realized, the applicability of the line-of-sight detection method can be improved, but also the accuracy of the line-of-sight detection can be improved.
[0107] Next, with reference to the accompanying drawings, the line-of-sight detection method provided in the embodiments of the present disclosure will be introduced in detail. Refer to Figure 1 As shown, it is a flowchart of the line-of-sight detection method provided in the embodiments of the present disclosure. The line-of-sight detection method includes the following S101 to S103:
[0108] S101, obtain the image to be detected.
[0109] Exemplarily, video data of a target scene or a target area can be obtained through a camera device, and then the image to be detected can be obtained by decoding the video data. The target scene and the target area can be different according to different scenarios. For example, in a vehicle driving scene, the target area can be the driving area inside the vehicle. At this time, it is necessary to detect the line-of-sight direction of the driver to determine whether the line-of-sight direction of the driver deviates from a preset direction, thereby improving driving safety.
[0110] Among them, video data refers to a continuous sequence of images. In essence, it is composed of a group of continuous images. Among them, an image frame is the smallest visual unit that makes up a video and is a static image. Synthesizing a sequence of image frames that are continuous in time together forms a dynamic video. Exemplarily, in order to facilitate subsequent detection and recognition, the video data can be decoded by means of FFmpeg technology or the like, and then the image to be detected can be obtained.
[0111] Among them, the FFmpeg technology is an open-source computer program that can be used to record, convert digital audio and video, and can convert them into streams. It adopts the LGPL or GPL license, provides a complete solution for recording, converting and streaming audio and video, and includes a very advanced audio / video codec library libavcodec. In order to ensure high portability and codec quality, many codes in libavcodec can be developed from scratch, so as to simplify the data usage and adaptation. Correspondingly, in actual use, the video data of the driver's driving behavior captured can also be saved and transmitted using the FFmpeg technology.
[0112] It can be understood that since video data usually includes many frame images per second (for example, 24 frame images per second), when decoding video data, image frame extraction can be performed on the video data. Among them, frame extraction refers to extracting frames according to a preset number of interval frames. For example, one frame of the image to be detected is extracted every 20 frames; frame extraction can also be performed according to a preset time interval. For example, one frame of the image to be detected is extracted every 10 ms.
[0113] It should be noted that the specific number of interval frames and the interval time can be set according to actual needs and are not limited here.
[0114] In addition, the execution subject of this gaze detection method can be a terminal device, where the terminal device includes but is not limited to in-vehicle devices, wearable devices, user terminals, handheld devices, etc. In other embodiments, the execution subject of this gaze detection method can also be a server, where the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. In some possible implementation manners, this gaze detection method can also be implemented by a processor calling computer-readable instructions stored in a memory.
[0115] S102. When it is determined that the image to be detected contains a face image, determine the squint classification result of the target person indicated by the face image, where the squint classification result represents that the eye state of the target person is a squint state with inconsistent binocular gaze or a normal state with consistent binocular gaze.
[0116] The squint state usually refers to the state where the gaze directions of the left and right eyes are inconsistent when both eyes observe the same direction. For example, when both eyes observe directly ahead, the gaze direction of one of the left and right eyes is directly ahead, and the gaze direction of the other is obliquely ahead. The normal state usually refers to the state where the binocular directions are consistent when both eyes look at the same direction.
[0117] Exemplarily, after obtaining the image to be detected, face detection can be performed on the image to be detected to determine whether the image to be detected contains a face image. In some embodiments, the face detection model trained in advance can be used to perform face detection on the image to be detected. In other embodiments, other methods can also be used to perform face detection on the image to be detected, such as using a preset face recognition algorithm to determine whether the image to be detected contains a face image. The specific face detection method is not limited, as long as the result of face detection can be achieved.
[0118] It can be understood that if it is determined that the image to be detected contains a face image, it indicates that there is a target person in the target area or target place, and subsequent steps need to be executed to further detect the line-of-sight direction of the target person; if it is determined that the image to be detected does not contain a face image, it indicates that there is no target person in the target area or target place currently, and the process ends, without the need to perform subsequent further detection processes, thereby avoiding waste of resources.
[0119] Exemplarily, in the case of determining that the image to be detected contains a face image, the identity information of the target person indicated by the face image can be recognized, and then based on the preset correspondence between the personnel identity information and the squint category, the squint category corresponding to the recognized identity information of the target person is determined as the squint classification result of the target person.
[0120] Specifically, in some embodiments, the correspondence between the personnel identity information and the squint category can be established in advance. The correspondence between the personnel identity information and the squint category includes different target persons and the squint classification results corresponding to each driver. For example, for the 1st target person (such as Zhang San), the squint classification result is right-eye strabismus; for the 2nd target person (such as Li Si), the squint classification result is left-eye strabismus; for the 3rd target person (such as Wang Wu), the squint classification result is the normal state.
[0121] Therefore, after determining the target person corresponding to the face image, according to the correspondence between the personnel identity information and the squint category, the squint classification result of the target person can be determined. For example, if the target person indicated by the face image is the 1st driver, the squint classification result of the target person can be determined as right-eye strabismus according to the pre-stored information.
[0122] Among them, when determining that the image to be detected contains a face image, the face image can be compared with a preset face image set. After the face image is successfully compared with any face image in the face image set, the target person indicated by the face image can be determined. Among them, the successful comparison of the face image with any face image in the face image set means that the similarity between the face image and any face image in the face image set is greater than a preset threshold (such as 90%).
[0123] It should be noted that in this embodiment, since the correspondence between the personnel identity information and the squint category is established in advance, the squint classification results of different personnel need to be obtained in advance. Among them, the method for obtaining the squint classification results is not limited. For example, the squint classification results can be determined by a professional squint detection device (such as a vision detection device in a hospital).
[0124] In the embodiments of the present disclosure, since the squint classification result of the target person can be determined based on the correspondence between the preset personal identity information and the squint category, the detection efficiency of the line-of-sight direction can be improved accordingly.
[0125] In some other embodiments, the eye region image of the person can also be determined based on the face image, and then the squint detection is performed on the eye region image based on the trained squint detection model to obtain the squint classification result of the target person. Among them, the eye region image can be determined by detecting the key points of the eyes in the face image. In addition, the method for obtaining the squint detection model will be described in detail later.
[0126] It should be understood that in this embodiment, when the squint classification result of the target person indicated by the face image is not known in advance, the eye region image can be determined based on the face image, and then the squint classification result is obtained by detecting the eye region image. In this way, even when the squint classification result of the target person is not known in advance, the squint classification result of the target person can still be determined, and then the line-of-sight direction can be detected based on the squint classification result, thereby improving the applicability of the line-of-sight detection method.
[0127] Exemplarily, for the vehicle driving scenario, the driver can be subjected to squint detection when the driver first registers to obtain the squint classification result, and the squint eye classification result is bound to the driver identity obtained through face recognition. In this way, when the driver logs in later, the driver's identity can be determined through face recognition, and then the corresponding squint classification result of the driver can be determined. In this way, it is possible to avoid performing squint detection on the same driver every time, avoiding waste of resources and improving the efficiency of line-of-sight detection at the same time.
[0128] In some embodiments, in order to improve the accuracy of squint detection, squint detection can be performed on multiple frames of images to be detected respectively, and then the final squint classification result is obtained by combining the detection results of multiple frames. Therefore, as shown in Figure 2 In this embodiment, when determining the squint classification result of the target person indicated by the face image, the following S1021-S1022 may be included:
[0129] S1021, when it is determined that at least one frame of the multiple frames of images to be detected contains a face image, determining the squint classification result of each frame of the target person in the at least one frame of the images to be detected;
[0130] S1022, determining the squint classification result of the target person based on the squint classification result of each frame of the at least one frame of the images to be detected.
[0131] Exemplarily, multiple frames of images to be detected can be obtained, and then squint detection can be performed on at least one frame of the images to be detected with face images respectively to obtain the squint classification results of each frame of the images to be detected of the target person under the at least one frame of the images to be detected. Then, based on the detection results of multiple frames, the squint classification result of the target person can be determined. For example, if a total of 5 frames of images to be detected are obtained, and 4 of them have face images, and the detection results of 3 of the 4 frames show right eye squint, and the detection result of 1 frame shows normal both eyes, therefore, the squint classification result of the target person can be determined as right eye squint according to the quantity, so that the situation where the detection result is inaccurate due to abnormal single-frame detection results can be reduced, and thus the accuracy of squint detection is improved.
[0132] S103, adopt a line-of-sight direction detection method corresponding to the squint classification result to detect the line-of-sight direction of the target person.
[0133] It can be understood that in order to implement line-of-sight detection for different squint classification results and improve the line-of-sight detection accuracy of different squint classification results, different line-of-sight detection methods can be adopted for different squint classification results.
[0134] Exemplarily, the line-of-sight direction of the target person can be predicted by using a neural network model. For example, the face image and / or eye image of the target person can be input into a pre-trained line-of-sight direction detection model to obtain the line-of-sight direction of the target person.
[0135] Specifically, when the squint classification result indicates that the eye state of the target person is the normal state, the face image is input into the line-of-sight direction detection model to obtain the line-of-sight direction of the target person; while when the squint classification result indicates that the eye state of the target person is the squint state, correction processing needs to be performed on the eye region of the target person in the face image to obtain the face image or eye image after correcting the eye state, and then the face image or eye image after correcting the eye state is input into the line-of-sight direction detection model to obtain the line-of-sight direction of the target person.
[0136] Specifically, when the squint classification result indicates that the eye state of the target person is the squint state, correction processing can be performed only on the region corresponding to the single eye with the squint state in the two eyes of the target person, so that unnecessary processing can be reduced, and thus computing resources can be saved and the efficiency of the correction processing can be improved.
[0137] In some embodiments, the squint state includes the left eye squint state and the right eye squint state, and the squint classification result includes the result indicating that the eye state of the target person is the left eye squint state or the result indicating that the eye state of the target person is the right eye squint state.
[0138] In some embodiments, when the cross-eye classification result indicates that the eye state of the target person is the left-eye strabismus state, the left-eye area of the target person in the face image may be masked to remove the left-eye area of the target person in the face image; when the cross-eye classification result indicates that the eye state of the target person is the right-eye strabismus state, the right-eye area of the target person in the face image may be masked to remove the right-eye area of the target person in the face image.
[0139] Specifically, the specific manner of the masking process can be achieved by adjusting the gray value of the eye area. For example, when performing gray processing on the left-eye area in the face image, the gray adjustment value of the left-eye area can be preset to a threshold value (such as 100), thereby achieving the masking process of the left-eye area.
[0140] In the embodiments of the present disclosure, when it is determined that the image to be detected contains a face image, first determine the cross-eye classification result of the target person indicated by the face image, and then based on the cross-eye classification result of the target person, adopt a line-of-sight direction detection method corresponding to the cross-eye classification result to detect the line-of-sight direction of the target person. In this way, not only can the line-of-sight detection of cross-eye people be realized, improving the applicability of the line-of-sight detection method, but also the accuracy of the line-of-sight detection can be improved.
[0141] The following details the specific manner of obtaining the above cross-eye detection model. Refer to Figure 3 As shown, in some embodiments, before obtaining the image to be detected in the driving area of the vehicle, the cross-eye detection model is obtained through the following steps S201 to S202.
[0142] S201, obtain an image sample set, where the image sample set includes eye images or face images of people in a monocular strabismus state.
[0143] S202, based on the image sample set, train the cross-eye detection model to be trained to obtain the trained cross-eye detection model.
[0144] Exemplarily, a large number of face image or human eye image samples can be obtained. Among them, in order to enable the trained strabismus detection model to detect strabismus, the image sample set needs to include human eye images or face images in a monocular strabismus state. Herein, the monocular strabismus state means that both eyes are visible (open) and only one eye is in a strabismus state, or both eyes are visible and one eye is looking straight ahead while the other eye is in a strabismus state. That is to say, the image sample can be an image sample in a right-eye strabismus state or an image sample in a left-eye strabismus state, and the strabismus degree of different image samples can be different. In this way, by training the strabismus detection model to be trained with samples of different strabismus types, the detection accuracy of the obtained strabismus detection model can be improved.
[0145] See Figure 4 As shown, in some embodiments, the acquisition method of the images in the image sample set can be implemented through the following steps S301 to S303.
[0146] S301, obtain a target face image or a target human eye image with both eyes looking straight ahead.
[0147] Exemplarily, see Figure 5 As shown, first obtain a target face image with both eyes looking straight ahead. Among them, the processing method of human eye images is similar to that of face images. Here, the face image is taken as an example for illustration.
[0148] S302, perform a line-of-sight redirection process on a single eye in the target face image or target human eye image to obtain a processed target face image or target human eye image.
[0149] Exemplarily, a line-of-sight adjustment model can be used to perform a redirection process on a single eye in the target face image or target human eye image to obtain the performance of the human eye under different lines of sight. See Figure 6 As shown, it is a schematic diagram of the right eye (region C in the figure) in the face image in Figure 5 after being redirected. Among them, the redirection result can be set according to specific requirements and is not limited here. For example, the left eye of the two eyes can be redirected, or the right eye of the two eyes can be redirected, and the redirection amplitude can be set according to the actual situation.
[0150] S303, generate the images in the image sample set based on the processed target face image or target human eye image.
[0151] In the embodiments of the present disclosure, by redirecting the images in the image sample set, not only can the problem of difficult sampling of strabismus image samples in the real situation be avoided, but also the richness of human eye image samples can be improved, which further helps to improve the training accuracy of the model.
[0152] SeeFigure 7 As shown, in some embodiments, the acquisition method of the human eye image samples can also be implemented through the following steps S401 to S402.
[0153] S401: Obtain a first target face image with both eyes looking straight ahead and a second target face image with both eyes looking obliquely, or obtain a first target human eye image with both eyes looking straight ahead and a second target human eye image with both eyes looking obliquely, respectively.
[0154] Among them, binocular squint means that both eyes look at a certain position diagonally forward at the same time, and the line-of-sight directions of the two eyes are the same and both are diagonally forward.
[0155] Please refer to Figure 8 and Figure 9 at the same time. In some embodiments, a first target face image with both eyes looking straight ahead (as shown in Figure 8 ) and a second target face image with both eyes looking obliquely (as shown in Figure 9 ) can be obtained respectively. Among them, the head of the target person can be kept upright and facing forward, and both eyes looking straight ahead, then the first target face image can be obtained; the second target face image can be obtained based on the first target face image. For example, on the basis of the first target face image, keeping the head upright and facing forward, and then rotating the eyeballs of both eyes 60 degrees to the right, the second target face image can be obtained.
[0156] Among them, the human eye image can be segmented by detecting key eye points in the face image, etc. After obtaining the human eye image, the processing method of the human eye image is similar to that of the following face image. Here, the face image is taken as an example for illustration, and the process of the human eye image will not be elaborated.
[0157] S402: Generate the image sample set based on the first target face image and the second target face image, or generate the image sample set based on the first target human eye image and the second target human eye image.
[0158] Exemplarily, the left eye image area in the first target face image (as shown by B2 in Figure 8 ) can be spliced with the right eye image area in the second target face image (as shown by B1 in Figure 9 ); or, the right eye image area in the first target face image can be spliced with the left eye image area in the second target face image, and then an image sample set is generated based on the spliced image.
[0159] Similarly, the left-eye image region in the first target eye image can be spliced with the right-eye image region in the second target eye image; or, the right-eye image region in the first target eye image can be spliced with the left-eye image region in the second target eye image, and then an image sample set is generated based on the spliced image.
[0160] See Figure 10 As shown, it is a schematic diagram of a spliced human eye image sample provided by an embodiment of the present disclosure. It can be understood that if the spliced image is a face image, human eye key point detection and recognition can be performed on the face image to obtain a human eye image sample.
[0161] In the embodiments of the present disclosure, it is not necessary to obtain a human eye image sample based on the actually existing cross-eyed population. A human eye image sample can be obtained based on the population with normal vision. In this way, the convenience of sample acquisition is improved.
[0162] See Figure 11 As shown, it is a flowchart of another line-of-sight detection method provided by an embodiment of the present disclosure. Different from the line-of-sight detection method in Figure 1 , this line-of-sight detection method further includes the following S104 - S105 after step S103:
[0163] S104, based on the detection result of the line-of-sight direction of the target person and the driving direction of the vehicle, determine whether the line-of-sight direction of the target person deviates from the preset direction.
[0164] S105, when the line-of-sight direction of the target person deviates from the preset direction, send a prompt message.
[0165] In some embodiments, taking the vehicle driving scenario as an example, the image to be detected includes an image of the driving area of the vehicle in the driving state. By performing line-of-sight detection on the image of the driving area of the vehicle, after detecting the line-of-sight direction of the target person (driver), based on the driving direction of the vehicle, it is determined whether the line-of-sight direction of the target person deviates from the preset direction (such as the direction in which the vehicle travels). When the line-of-sight direction of the target person deviates from the preset direction, a prompt message is sent to warn the driver, prompting the driver to improve the attention during driving, thereby improving driving safety.
[0166] Among them, when the angle between the line-of-sight direction of the driver and the driving direction of the vehicle is greater than a preset angle (such as 45 degrees), it is determined that the line-of-sight direction of the driver deviates from the preset direction.
[0167] Among them, the driving area refers to the area inside the vehicle where the driver performs vehicle driving operations. Among them, vehicle driving operations include, but are not limited to, steering wheel control operations, accelerator pedal control operations, etc. Optionally, video data of the driver during the vehicle driving process can be collected by a camera device installed in the vehicle cabin, and then the video data collected by the camera device can be obtained through a terminal device.
[0168] This camera device is an essential hardware of the driver monitoring system (DMS). The driver monitoring system uses a camera to obtain images, and through technologies such as visual tracking, object detection, and action recognition, it conducts real-time intelligent detection and reminder of situations such as the driver's fatigue driving, driving distraction, and dangerous actions, so as to reduce the probability of traffic accidents.
[0169] Specifically, the camera device can be installed on the A-pillar of the vehicle or the interior rearview mirror, etc., and be oriented towards the driving area of the vehicle cabin. Of course, the camera device can also be installed in other parts of the vehicle interior, as long as it can achieve video acquisition of the vehicle driving area. In addition, the number of camera devices is not limited here. For example, it can be one, two, or multiple.
[0170] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0171] Based on the same inventive concept, the present disclosure embodiments also provide a gaze detection device corresponding to the gaze detection method. Since the principle of solving problems by the device in the present disclosure embodiments is similar to the above gaze detection method in the present disclosure embodiments, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0172] Refer to Figure 12 As shown, it is a schematic diagram of a gaze detection device 500 provided by the present disclosure embodiments. The gaze detection device includes:
[0173] An image acquisition module 501, configured to acquire an image to be detected;
[0174] A squint detection module 502, configured to determine the squint classification result of the target person indicated by the face image in the case where the image to be detected contains a face image, where the squint classification result represents that the eye state of the target person is a squint state with inconsistent binocular gaze or a normal state with consistent binocular gaze;
[0175] A gaze detection module 503, configured to detect the gaze direction of the target person by using a gaze direction detection method corresponding to the squint classification result.
[0176] In a possible implementation manner, the line-of-sight detection module 503 is specifically configured to:
[0177] When the squint classification result indicates that the eye state of the target person is the normal state, input the face image into the line-of-sight direction detection model to obtain the line-of-sight direction of the target person; or,
[0178] When the squint classification result indicates that the eye state of the target person is the squint state, perform correction processing on the eye region of the target person in the face image to obtain a face image or an eye image with the corrected eye state, and input the face image or the eye image with the corrected eye state into the line-of-sight direction detection model to obtain the line-of-sight direction of the target person.
[0179] In a possible implementation manner, the line-of-sight detection module 503 is specifically configured to:
[0180] Perform correction processing on the area corresponding to the monocular with the squint state in the two eyes of the target person.
[0181] In a possible implementation manner, the squint state includes a left-eye squint state and a right-eye squint state, and the squint classification result includes a result indicating that the eye state of the target person is the left-eye squint state, or a result indicating that the eye state of the target person is the right-eye squint state;
[0182] The line-of-sight detection module 503 is specifically configured to:
[0183] When the squint classification result indicates that the eye state of the target person is the left-eye squint state, perform masking processing on the left-eye area of the target person in the face image to remove the left-eye area of the target person in the face image; or,
[0184] When the squint classification result indicates that the eye state of the target person is the right-eye squint state, perform masking processing on the right-eye area of the target person in the face image to remove the right-eye area of the target person in the face image.
[0185] In a possible implementation manner, the squint detection module 502 is specifically configured to:
[0186] Identify the identity information of the target person indicated by the face image;
[0187] Based on the preset correspondence between the personnel identity information and the squint category, determine the squint category corresponding to the identified identity information of the target person as the squint classification result of the target person.
[0188] In a possible implementation manner, the cross-eye detection module 502 is specifically configured to:
[0189] Determine a human eye region image based on the face image;
[0190] Perform cross-eye detection on the human eye region image based on a trained cross-eye detection model to obtain a cross-eye classification result of the target person.
[0191] See Figure 13 As shown, in a possible implementation manner, the device further includes a model training module 504, and the model training module 504 is configured to:
[0192] Obtain an image sample set, where the image sample set includes human eye images or face images in a monocular cross-eye state;
[0193] Train a cross-eye detection model to be trained based on the image sample set to obtain the trained cross-eye detection model.
[0194] In a possible implementation manner, the model training module 504 is specifically configured to:
[0195] Obtain a target face image or a target human eye image with both eyes looking straight ahead;
[0196] Perform line-of-sight redirection processing on one eye in the target face image or the target human eye image to obtain a processed target face image or a processed target human eye image;
[0197] Generate an image in the image sample set based on the processed target face image or the processed target human eye image.
[0198] In a possible implementation manner, the model training module 504 is specifically configured to:
[0199] Respectively obtain a first target face image with both eyes looking straight ahead and a second target face image with both eyes cross-eyed, or respectively obtain a first target human eye image with both eyes looking straight ahead and a second target human eye image with both eyes cross-eyed; and
[0200] Generate an image in the image sample set by at least one of the following methods:
[0201] Stitch the left eye image region in the first target face image with the right eye image region in the second target face image;
[0202] Stitch the right eye image region in the first target face image with the left eye image region in the second target face image;
[0203] Stitch the left eye image region in the first target eye image with the right eye image region in the second target eye image;
[0204] Stitch the right eye image region in the first target eye image with the left eye image region in the second target eye image.
[0205] In a possible implementation, the image to be detected includes multiple frames of images, and the cross-eye detection module 502 is specifically configured to:
[0206] When it is determined that at least one frame of the images to be detected among the multiple frames of images to be detected contains a face image, determine the cross-eye classification result of each frame of the images to be detected of the target person under the at least one frame of images to be detected;
[0207] Based on the cross-eye classification results of each frame of the images to be detected under the at least one frame of images to be detected, determine the cross-eye classification result of the target person.
[0208] In a possible implementation, the image to be detected includes an image of the driving area of a vehicle in a driving state, and the device further includes a line-of-sight judgment module 505, and the line-of-sight judgment module 505 is configured to:
[0209] Based on the detection result of the line-of-sight direction of the target person and the driving direction of the vehicle, determine whether the line-of-sight direction of the target person deviates from a preset direction;
[0210] When the line-of-sight direction of the target person deviates from the preset direction, send a prompt message.
[0211] The description of the processing flow of each module in the device and the interaction flow between the modules can refer to the relevant descriptions in the above method embodiments and will not be elaborated here.
[0212] Based on the same inventive concept, the embodiments of the present disclosure also provide an electronic device. Refer to Figure 14 As shown, it is a schematic structural diagram of an electronic device 700 provided by the embodiments of the present disclosure, including a processor 701, a memory 702, and a bus 703. Among them, the memory 702 is used to store execution instructions, including an internal memory 7021 and an external memory 7022; here, the internal memory 7021 is also called the main memory, which is used to temporarily store the operation data in the processor 701 and the data exchanged with the external memory 7022 such as a hard disk, and the processor 701 exchanges data with the external memory 7022 through the internal memory 7021.
[0213] In the embodiments of the present application, the memory 702 is specifically configured to store the application program code for executing the solutions of the present application, and is controlled by the processor 701 to execute. That is, when the electronic device 700 runs, the processor 701 communicates with the memory 702 through the bus 703, so that the processor 701 executes the application program code stored in the memory 702, and further executes the methods described in any of the foregoing embodiments.
[0214] Among them, the memory 702 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0215] The processor 701 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0216] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 700. In other embodiments of the present application, the electronic device 700 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0217] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the line-of-sight detection method in the above method embodiments. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.
[0218] Embodiments of the present disclosure also provide a computer program product. The computer program product carries program codes, and the instructions included in the program codes can be used to execute the steps of the line-of-sight detection method in the above method embodiments. For details, reference can be made to the above method embodiments and will not be elaborated here.
[0219] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0220] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0221] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0222] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0223] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0224] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A line-of-sight detection method, characterized in that, Including: Obtain an image to be detected; When it is determined that the image to be detected contains a face image, determine the cross-eye classification result of the target person indicated by the face image, where the cross-eye classification result characterizes that the eye state of the target person is a strabismus state with inconsistent binocular line of sight or a normal state with consistent binocular line of sight; Adopt a line-of-sight direction detection method corresponding to the cross-eye classification result to detect the line-of-sight direction of the target person; The adopting a line-of-sight direction detection method corresponding to the cross-eye classification result to detect the line-of-sight direction of the target person includes: When the cross-eye classification result characterizes that the eye state of the target person is the strabismus state, perform correction processing on the eye region of the target person in the face image to obtain a face image or an eye image after correcting the eye state, and input the face image or the eye image after correcting the eye state into the line-of-sight direction detection model to obtain the line-of-sight direction of the target person; The strabismus state includes a left-eye strabismus state and a right-eye strabismus state, and the cross-eye classification result includes a result characterizing that the eye state of the target person is the left-eye strabismus state or a result characterizing that the eye state of the target person is the right-eye strabismus state; The performing correction processing on the eye region of the target person in the face image when the cross-eye classification result characterizes that the eye state of the target person is the strabismus state includes: When the cross-eye classification result characterizes that the eye state of the target person is the left-eye strabismus state, perform masking processing on the left-eye region of the target person in the face image to remove the left-eye region of the target person in the face image; or, When the cross-eye classification result characterizes that the eye state of the target person is the right-eye strabismus state, perform masking processing on the right-eye region of the target person in the face image to remove the right-eye region of the target person in the face image.
2. The method according to claim 1, wherein The adopting a line-of-sight direction detection method corresponding to the cross-eye classification result to detect the line-of-sight direction of the target person includes: When the cross-eye classification result characterizes that the eye state of the target person is the normal state, input the face image into the line-of-sight direction detection model to obtain the line-of-sight direction of the target person.
3. The method according to claim 2, wherein The performing correction processing on the eye region of the target person in the face image includes: Performing correction processing on the region corresponding to the single eye in the strabismus state among the two eyes of the target person.
4. The method according to claim 1, wherein The determining the cross-eye classification result of the target person indicated by the face image includes: Identify the identity information of the target person indicated by the face image; Based on the preset correspondence between the personnel identity information and the cross-eye category, determine the cross-eye category corresponding to the identity information of the identified target person as the cross-eye classification result of the target person.
5. The method according to claim 1, wherein The determining the cross-eye classification result of the target person indicated by the face image includes: Based on the face image, determine the eye region image; Perform strabismus detection on the eye region image based on a trained strabismus detection model to obtain the cross-eye classification result of the target person.
6. The method according to claim 5, wherein The method further includes: Obtaining an image sample set, where the image sample set includes eye images or face images of a person with monocular strabismus; Training a to-be-trained strabismus detection model based on the image sample set to obtain the trained strabismus detection model.
7. The method according to claim 6, wherein The images in the image sample set are obtained by the following method: Obtaining a target face image or a target eye image with both eyes looking straight ahead; Performing a line-of-sight redirection process on one eye in the target face image or the target eye image to obtain a processed target face image or a processed target eye image; Generating the images in the image sample set based on the processed target face image or the processed target eye image.
8. The method according to claim 6, wherein The images in the image sample set are obtained by the following method: Respectively obtaining a first target face image with both eyes looking straight ahead and a second target face image with both eyes squinting, or respectively obtaining a first target eye image with both eyes looking straight ahead and a second target eye image with both eyes squinting; and Generating the images in the image sample set by at least one of the following methods: Stitching the left-eye image region in the first target face image with the right-eye image region in the second target face image; Stitching the right-eye image region in the first target face image with the left-eye image region in the second target face image; Stitching the left-eye image region in the first target eye image with the right-eye image region in the second target eye image; Stitching the right-eye image region in the first target eye image with the left-eye image region in the second target eye image.
9. The method according to any one of claims 1 to 8, characterized in that, The image to be detected includes multiple frames of images. When it is determined that the image to be detected contains a face image, determining the strabismus classification result of the target person indicated by the face image includes: When it is determined that at least one frame of the multiple frames of images to be detected contains a face image, determining the strabismus classification result of the target person in each frame of the images to be detected; Based on the strabismus classification results of each frame of the images to be detected, determining the strabismus classification result of the target person.
10. The method according to any one of claims 1 to 8, characterized in that, The image to be detected includes an image of the driving area of a vehicle in a driving state. The method further includes: Based on the detection result of the line-of-sight direction of the target person and the driving direction of the vehicle, determining whether the line-of-sight direction of the target person deviates from a preset direction; When the line-of-sight direction of the target person deviates from the preset direction, sending a prompt message.
11. A line-of-sight detection device, characterized in that, Including: An image acquisition module for acquiring an image to be detected; A strabismus detection module for determining the strabismus classification result of the target person indicated by the face image when it is determined that the image to be detected contains a face image, where the strabismus classification result represents that the eye state of the target person is a strabismus state with inconsistent binocular line of sight or a normal state with consistent binocular line of sight; A line-of-sight detection module for detecting the line-of-sight direction of the target person by using a line-of-sight direction detection method corresponding to the strabismus classification result; Specifically, the line-of-sight detection module is used for: When the cross-eye classification result indicates that the eye state of the target person is the strabismus state, perform a correction process on the eye region of the target person in the face image to obtain a face image or an eye image with the corrected eye state, and input the face image or the eye image with the corrected eye state into the line-of-sight direction detection model to obtain the line-of-sight direction of the target person; The strabismus state includes a left-eye strabismus state and a right-eye strabismus state, and the cross-eye classification result includes a result indicating that the eye state of the target person is the left-eye strabismus state, or a result indicating that the eye state of the target person is the right-eye strabismus state; The performing a correction process on the eye region of the target person in the face image when the cross-eye classification result indicates that the eye state of the target person is the strabismus state includes: When the cross-eye classification result indicates that the eye state of the target person is the left-eye strabismus state, perform a masking process on the left-eye region of the target person in the face image to remove the left-eye region of the target person in the face image; Or, When the cross-eye classification result indicates that the eye state of the target person is the right-eye strabismus state, perform a masking process on the right-eye region of the target person in the face image to remove the right-eye region of the target person in the face image.
12. An electronic device, characterized in that, Includes: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the line-of-sight detection method according to any one of claims 1-10 are executed.
13. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the line-of-sight detection method according to any one of claims 1-10 are executed.
Citation Information
Patent Citations
Line of sight detection device and line of sight detection method
CN108027977A
An image analysis method for eyes
CN110288567A
Personnel watching position detection method and device
CN113807119A