Deep learning-based eye exercise action recognition method, device and related equipment
By building a pre-trained lightweight model, identifying the eye direction value of the user's eyes, the problem of low accuracy in eyeball movement recognition in the prior art is solved, and efficient and accurate recognition effect is achieved.
Patent Information
- Application Number
- CN202210544220.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The prior art has large errors in identifying eyeball maneuvers and is sensitive to the setting of thresholds, resulting in low recognition accuracy.
Using a deep learning-based method, the pre-trained lightweight model is constructed to identify the visual direction value of the user's eyes, and improve the recognition efficiency and accuracy of eyeball maneuvers.
It realizes the rapid and accurate identification of eyeball movements, avoids misjudgment and missed counting, and improves the recognition efficiency and accuracy.
Smart Images

Figure CN114898452B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device and related equipment for identifying eye exercise actions based on deep learning. Background Art
[0002] At present, online teaching and working from home are gradually increasing, and it is particularly important to accurately perform eye exercise actions. Existing algorithms for identifying and detecting eyes and eyeballs are relatively complex. Most of them first detect the position of the eyes, then extract the images of the eye regions, and further calibrate the position of the black eyeballs.
[0003] However, when calibrating the position of the black eyeballs in the prior art, there are still large errors, and it is very sensitive to the setting of thresholds. During the recognition process of eye exercises, misjudgment and missed counting are likely to occur, resulting in low accuracy of eye exercise recognition.
[0004] Therefore, it is necessary to propose a method that can quickly and accurately identify the direction of eye exercises. Summary of the Invention
[0005] In view of the above, it is necessary to propose a method, device and related equipment for identifying eye exercise actions based on deep learning. By constructing a pre-trained lightweight model, the visual direction value of the current user's eyes is recognized, improving the recognition efficiency and accuracy of eye exercise actions.
[0006] The first aspect of the present invention provides a method for identifying eye exercise actions based on deep learning, the method comprising:
[0007] Obtain a first face image set, and perform visual direction value annotation on the first face image set to obtain a second face image set;
[0008] Construct a pre-trained lightweight model, and train the pre-trained lightweight model based on the second face image set to obtain a target lightweight model;
[0009] In response to receiving an eye exercise recognition request, obtain a to-be-recognized eye exercise video;
[0010] Determine a plurality of target video frames based on the to-be-recognized eye exercise video, and sequentially input each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame;
[0011] Identify the eye exercise actions of the current user's eyes based on the multiple visual direction values of the current user's eyes in the multiple target video frames to obtain the gaze path of the current user's eyes;
[0012] Identify the gaze path of the current user's eyes and record the number of times the eye exercise actions are up to standard.
[0013] Optionally, the construction of the pre-trained lightweight model includes:
[0014] Obtain an initial lightweight model, and replace the last classification layer in the initial lightweight model with a regression layer with two nodes to obtain a pre-trained lightweight model.
[0015] Optionally, the step of sequentially inputting each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame includes:
[0016] Input each target video frame into the target lightweight model in the chronological order of the multiple target video frames to obtain the visual direction value of the current user's eyes in each target video frame.
[0017] Optionally, the step of identifying the eye exercise actions of the current user's eyes based on the multiple visual direction values of the current user's eyes in the multiple target video frames to obtain the gaze path of the current user's eyes includes:
[0018] Obtain a preset two-dimensional spatial map of visual directions, where the two-dimensional spatial map of visual directions contains multiple regions and the label of each region;
[0019] Sort the multiple target video frames in chronological order to obtain a target queue;
[0020] Starting from the head of the target queue, map the visual direction value of the current user's eyes in each target video frame to the two-dimensional spatial map of visual directions one by one, and obtain the target label of the target region where the current user's eyes are located in the two-dimensional spatial map of visual directions for each target video frame;
[0021] Connect the multiple target labels in the order of obtaining the target label of the current user's eyes in each target video frame to generate the gaze path of the current user's eyes.
[0022] Optionally, the step of identifying the gaze path of the current user's eyes and recording the number of times the eye exercise actions are up to standard includes:
[0023] When the eye exercise type of the eye exercise video to be recognized is the first eye operation type, obtain the valid region, invalid region, and ignored region corresponding to the first eye operation type;
[0024] Start identifying from the first node of the gaze path. When it is recognized that the node of the gaze path first falls into the valid region, determine the first fallen node as the starting position;
[0025] When it is recognized that the next node at the starting position falls within the valid region and the target label of the next node at the starting position is different from the target label at the starting position, the next node at the starting position is determined as the ending position; or when it is recognized that the next node at the starting position falls within the ignored region, the next node at the starting position is ignored, and the next node of the initial node is re-determined and recognized in the order of the nodes in the gaze path; or when it is recognized that the next node at the starting position falls within the invalid region, the starting position is re-determined and recognized in the order of the nodes in the gaze path;
[0026] When it is recognized that the next node at the ending position falls within the valid region and is the same as the target label at the starting position, it is determined as an eye exercise action that meets the standard once, and the count is incremented by one. By analogy, the number of times the eye exercise action meets the standard is recorded.
[0027] Optionally, the recognizing the gaze path of the current user's eyes and recording the number of times the eye exercise action meets the standard includes:
[0028] When the eye exercise type of the eye exercise video to be recognized is the second eye operation type, obtain the first target value of the valid region corresponding to the second eye operation type;
[0029] Obtain the target label of each node and connect them in the order of the obtained target labels to obtain the first path;
[0030] Determine the first node of the first path as the current node and sequentially recognize each node in the gaze path;
[0031] Determine whether the target label of the next node of the current node in the first path is greater than the target label of the current node;
[0032] When the target label of the next node of the current node in the first path is greater than the target label of the current node, calculate the difference between the target label of the next node of the current node and the first target value to obtain the new target label of the next node of the current node;
[0033] Update the target label in the first path based on the next node of the current node to obtain the second path;
[0034] Determine the first node of the second path as the current node and start sequentially recognizing from the current node of the second path;
[0035] Calculate the difference between the target label of the next node of the current node and the target label of the current node;
[0036] When the calculated difference is greater than or equal to a preset first threshold, determine the current node as the starting position;
[0037] Re - determine the next node of the current node as the current node in the order of the nodes of the second path, and repeat the operation of calculating the difference between the target label of the next node of the current node and the target label of the current node. When the calculated difference is greater than or equal to the preset first threshold, determine the current node as the target node of the starting position. Until the target label of the target node is consistent with the target label of the starting position, determine it as an eye exercise action that meets the standard once, and count it once. And so on, record the number of times the eye exercise action meets the standard.
[0038] Optionally, the identifying the gaze path of the current user's eyes and recording the number of times the eye exercise action meets the standard includes:
[0039] When the eye exercise type of the eye exercise video to be identified is the third eye operation type, obtain the second target value of the effective area corresponding to the third eye operation type;
[0040] Obtain the target label of each node, and connect them in the order of the obtained target labels to get the third path;
[0041] Determine the first node of the third path as the current node, and sequentially identify each node of the gaze path;
[0042] Judge whether the target label of the next node of the current node in the third path is less than the target label of the current node;
[0043] When the target label of the next node of the current node in the third path is less than the target label of the current node, calculate the sum of the target label of the next node of the current node and the target value to obtain the new target label of the next node of the current node;
[0044] Update the target label in the third path based on the new target label of the next node of the current node to obtain the fourth path;
[0045] Determine the first node of the fourth path as the current node, and start to sequentially identify from the current node of the fourth path;
[0046] Calculate the difference between the target label of the next node of the current node and the target label of the current node;
[0047] When the calculated difference is less than or equal to a preset second threshold, determine the current node as the starting position;
[0048] Redetermine the next node of the current node in the order of the nodes of the fourth path as the current node, and repeatedly execute calculating the difference between the target label of the next node of the current node and the target label of the current node. When the calculated difference is less than or equal to a preset second threshold, determine the current node as the target node at the starting position. Until the target label of the target node is the same as the target label at the starting position, determine it as a qualified eye exercise action once, and count once. And so on, record the number of qualified times of the eye exercise action.
[0049] The second aspect of the present invention provides an eye exercise action recognition device based on deep learning, and the device includes:
[0050] A first acquisition module, configured to acquire a first face image set, and perform visual direction value annotation on the first face image set to obtain a second face image set;
[0051] A construction module, configured to construct a pre-trained lightweight model, and train the pre-trained lightweight model based on the second face image set to obtain a target lightweight model;
[0052] A second acquisition module, configured to acquire an eye exercise video to be recognized in response to an received eye exercise recognition request;
[0053] A determination module, configured to determine a plurality of target video frames based on the eye exercise video to be recognized, and sequentially input each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame;
[0054] An identification module, configured to identify the eye exercise action of the current user's eyes based on the plurality of visual direction values of the current user's eyes in the plurality of target video frames to obtain the gaze path of the current user's eyes;
[0055] An identification and recording module, configured to identify the gaze path of the current user's eyes and record the number of qualified times of the eye exercise action.
[0056] The third aspect of the present invention provides an electronic device, and the electronic device includes a processor and a memory. When the processor executes a computer program stored in the memory, it implements the above-mentioned eye exercise action recognition method based on deep learning.
[0057] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned eye exercise action recognition method based on deep learning.
[0058] In summary, for the method, device and related equipment for identifying eye exercise actions based on deep learning according to the present invention, by constructing a pre-trained lightweight model, training the pre-trained lightweight model based on the second face image set to obtain a target lightweight model, and relying on the high performance and strong robustness of the target lightweight model, the visual direction value of the current user's eyes is accurately identified. Based on the eye exercise video to be identified, a plurality of target video frames are determined, and each target video frame is sequentially input into the target lightweight model, and the eye exercise actions of the current user's eyes are identified according to the visual direction values of the current user's eyes in each identified target video frame, avoiding the phenomena of misjudgment and missed counting, and improving the identification efficiency and accuracy of eye exercise actions. Description of the Drawings
[0059] Figure 1 It is a flowchart of the method for identifying eye exercise actions based on deep learning provided in Embodiment 1 of the present invention.
[0060] Figure 2 It is a schematic diagram of the two-dimensional space map of the visual direction provided in Embodiment 1 of the present invention.
[0061] Figure 3 It is a schematic diagram of the left-right eye rotation eye exercise provided in Embodiment 1 of the present invention.
[0062] Figure 4 It is a schematic diagram of the up-down eye rotation eye exercise provided in Embodiment 1 of the present invention.
[0063] Figure 5 It is a schematic diagram of the upper-left to lower-right eye rotation eye exercise provided in Embodiment 1 of the present invention.
[0064] Figure 6 It is a schematic diagram of the upper-right to lower-left eye rotation eye exercise provided in Embodiment 1 of the present invention.
[0065] Figure 7 It is a schematic diagram of the clockwise eye rotation eye exercise provided in Embodiment 1 of the present invention.
[0066] Figure 8 It is a schematic diagram of the counterclockwise eye rotation eye exercise provided in Embodiment 1 of the present invention.
[0067] Figure 9 It is a structural diagram of the device for identifying eye exercise actions based on deep learning provided in Embodiment 2 of the present invention.
[0068] Figure 10 It is a schematic structural diagram of the electronic device provided in Embodiment 3 of the present invention. Detailed Embodiments
[0069] To better understand the above objects, features, and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0071] Embodiment 1
[0072] Figure 1 is a flowchart of the eye exercise action recognition method based on deep learning provided by Embodiment 1 of the present invention.
[0073] In this embodiment, the eye exercise action recognition method based on deep learning can be applied to an electronic device. For an electronic device that needs to perform eye exercise action recognition based on deep learning, the function of the eye exercise action recognition provided by the method of the present invention can be directly integrated on the electronic device, or run on the electronic device in the form of a Software Development Kit (SDK).
[0074] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0075] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning, deep learning.
[0076] As Figure 1 shown, the eye exercise action recognition method based on deep learning specifically includes the following steps. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0077] S11, obtain a first face image set, and perform visual direction value annotation on the first face image set to obtain a second face image set.
[0078] In this embodiment, the first face image set can be obtained from a public network platform. After obtaining the first face image set, each first face image is labeled with a visual direction value to obtain a second face image set. Specifically, when labeling the visual direction value, Row represents the deflection angle in the left-right direction, with the left being positive (+) and the right being negative (-); Pitch represents the deflection angle in the up-down direction, with the up being positive (+) and the down being negative (-). The unit of the angle here is degree (°), and it is 90° when facing the leftmost, rightmost, uppermost, and lowermost directions, and both Row and Pitch are 0° when facing directly forward.
[0079] S12. Construct a pre-trained lightweight model, and train the pre-trained lightweight model based on the second face image set to obtain a target lightweight model.
[0080] In this embodiment, in order to obtain an algorithm logic that is as simple as possible, rely on as little computing power resources as possible, and ensure the smooth operation of the algorithm on the mobile phone side, this embodiment uses a very famous lightweight model in the field of deep learning, that is, the lightweight basic convolutional neural network MobileNetV2.
[0081] In an optional embodiment, the constructing the pre-trained lightweight model includes:
[0082] Obtain an initial lightweight model, and replace the last classification layer in the initial lightweight model with a regression layer with two nodes to obtain a pre-trained lightweight model.
[0083] In this embodiment, the network of the pre-trained lightweight model is different from that of the initial lightweight model. The pre-trained lightweight model replaces the last classification layer of MobileNetV2 with a regression layer that can output two nodes, that is, a regression layer that can output the left-right direction value and the up-down direction value in each face image.
[0084] In an optional embodiment, the training the pre-trained lightweight model based on the second face image set to obtain a target lightweight model includes:
[0085] Obtain each face image in the second face image set and the visual direction value of the user's eyes labeled in each face image to form a data set, where the visual direction value includes the left-right direction value and the up-down direction value;
[0086] Randomly divide the data set into a first number of training sets and a second number of test sets;
[0087] Input the training set into the pre-trained lightweight model for training to obtain a lightweight model;
[0088] Input the test set into the lightweight model for testing to obtain a test pass rate;
[0089] Determine whether the passing rate of the test is greater than a preset passing rate threshold;
[0090] When the passing rate of the test is greater than or equal to the preset passing rate threshold, end the training of the lightweight model to obtain a target lightweight model.
[0091] Further, the determining whether the passing rate of the test is greater than a preset passing rate threshold includes:
[0092] When the passing rate of the test is less than the preset passing rate threshold, increase the number of the training set and retrain the lightweight model.
[0093] In this embodiment, during the training of the target lightweight model, the input parameter is the second face image set, and the target is the annotation data corresponding to each face image: the left - right direction value and the up - down direction value. Then, the left - right direction value and the up - down direction value are used as the training set and input into the constructed pre - trained lightweight model to obtain the target lightweight model. Subsequently, only by acquiring the face image of the user, the visual direction value of the user can be recognized through the target lightweight model, with high accuracy.
[0094] S13, in response to the received eye exercise recognition request, obtain the eye exercise video to be recognized.
[0095] In this embodiment, the eye exercise recognition request is used to request to recognize whether the current user's eye exercise is up to standard. The eye exercise recognition request is initiated from the client to the server. The server can be an eye exercise recognition subsystem, receive the eye exercise recognition request sent by the client, and parse the eye exercise recognition request to obtain the eye exercise video to be recognized.
[0096] S14, determine a plurality of target video frames based on the eye exercise video to be recognized, and sequentially input each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame.
[0097] In this embodiment, the visual direction value includes a left - right direction value and an up - down direction value. By inputting each target video frame, the visual direction value of the current user's eyes in each video frame can be accurately recognized. The steps are simple, improving the recognition efficiency of the visual direction value.
[0098] In an alternative embodiment, the determining a plurality of target video frames based on the eye exercise video to be recognized includes:
[0099] Decode the eye exercise video to be recognized using a preset decoder to obtain a plurality of video frames;
[0100] Screen the plurality of video frames according to a preset screening rule to obtain a plurality of target video frames.
[0101] In this embodiment, a preset video decoder may be used to decode the video frame. Specifically, when the video to be recognized is obtained, a preset decoder is obtained according to the video to be recognized, and the video frame is decompressed based on the preset video decoder to obtain a plurality of video frames.
[0102] In this embodiment, a screening rule may be preset. Specifically, the screening rule may be preset in advance according to the standard time of each set of eye exercises. For example, it may be set to select one video frame every 2 frames or select one video frame every 3 frames as the target video frame. This embodiment does not limit this here.
[0103] In an alternative embodiment, the step of sequentially inputting each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame includes:
[0104] Sequentially input each target video frame into the target lightweight model according to the time sequence of the plurality of target video frames to obtain the visual direction value of the current user's eyes in each target video frame.
[0105] In this embodiment, by presetting the screening rule in advance according to the standard time of each set of eye exercises, if the operation time of the eye exercise video to be recognized is long, it is determined that the current user's eye exercise action is slow. If one target video frame is selected every 1 frame, a large number of target video frames are obtained, resulting in low recognition efficiency of the target lightweight model; if the operation time of the eye exercise video to be recognized is short, it is determined that the current user's eye movement is fast. If one target video frame is selected every 3 frames, a small number of target video frames are obtained, and there may be a problem of missing the visual direction, resulting in low recognition accuracy of the target lightweight model.
[0106] S15. Recognize the eye exercise action of the current user's eyes based on the multiple visual direction values of the current user's eyes in the plurality of target video frames to obtain the gaze path of the current user's eyes.
[0107] In this embodiment, the gaze path of the current user's eyes is determined by the visual direction values of the current user's eyes in a plurality of target video frames.
[0108] In an alternative embodiment, the step of recognizing the eye exercise action of the current user's eyes based on the multiple visual direction values of the current user's eyes in the plurality of target video frames to obtain the gaze path of the current user's eyes includes:
[0109] Obtain a preset two-dimensional spatial map of visual directions, where the two-dimensional spatial map of visual directions contains a plurality of regions and the label of each region;
[0110] Sort the multiple target video frames in chronological order to obtain a target queue;
[0111] Starting from the head of the target queue, map the visual direction value of the current user's eyes in each target video frame to the visual direction two-dimensional space map one by one, and obtain the target label of the target area where the current user's eyes are located in the visual direction two-dimensional space map for each target video frame;
[0112] Connect multiple target labels in the order of obtaining the target labels of the current user's eyes in each target video frame to generate the gaze path of the current user's eyes.
[0113] Refer to Figure 2 The visual direction two-dimensional space map shown. The visual direction two-dimensional map is calibrated through preset quantization. For example, two flag bits are set: 20° and -20°. Based on the two calibration positions, the left-right direction is divided into 3 segments, and the up-down direction is divided into 3 segments. Among them, the middle area 0 represents that the current user is looking straight ahead, and the remaining 1 to 8 areas respectively represent that the current user's visual direction is: left, upper left, up, upper right, right, lower right, down, lower left. By identifying the visual direction determination results of a series of video frames corresponding to the eye exercise video to be identified, it is possible to determine whether the actions of the eye exercise are up to standard and perform counting.
[0114] S16. Identify the gaze path of the current user's eyes and record the number of times the eye exercise actions meet the standard.
[0115] In this embodiment, the gaze path of the current user's eyes refers to being determined by the rotation direction of the current user's eyes in a series of video frames in the eye exercise video to be identified. By identifying the gaze path of the current user's eyes, it is possible to quickly identify whether the eye exercise actions are correct and record the number of times the eye exercise actions meet the standard.
[0116] In an alternative embodiment, the identifying the gaze path of the current user's eyes and recording the number of times the eye exercise actions meet the standard includes:
[0117] When the eye exercise type of the eye exercise video to be identified is the first eye operation type, obtain the valid area, invalid area, and ignored area corresponding to the first eye operation type;
[0118] Identify sequentially starting from the first node of the gaze path. When it is identified that the node of the gaze path first falls into the valid area, determine the first falling node as the starting position;
[0119] When it is recognized that the next node at the starting position falls into the valid region and the target label of the next node at the starting position is inconsistent with the target label of the starting position, the next node at the starting position is determined as the ending position; or when it is recognized that the next node at the starting position falls into the ignored region, the next node at the starting position is ignored, and the next node of the initial node is re-determined in the order of the nodes of the gaze path for recognition; or when it is recognized that the next node at the starting position falls into the invalid region, the starting position is re-determined in the order of the nodes of the gaze path for recognition;
[0120] When it is recognized that the next node at the ending position falls into the valid region and is consistent with the target label of the starting position, it is determined as an eye exercise action that meets the standard once, and the count is incremented by one. By analogy, the number of times the eye exercise action meets the standard is recorded.
[0121] In this embodiment, the types of eye exercises include a first type of eye operation, a second type of eye operation, and a third type of eye operation.
[0122] Specifically, the first type of eye operation may include: left - right rotation, up - down rotation, upper - left to lower - right rotation, and upper - right to lower - left rotation.
[0123] Refer to Figure 3 the schematic diagram of the left - right rotation eye exercise shown. Among them, the double - arrow line represents the rotation path. The target labels of the valid region are 1 and 5, the target labels of the invalid region are 2, 3, 4, 6, 7, and 8, the target label of the ignored region is 0. The starting position can start from 1 or 5. If the node of the gaze path first falls into the valid region 1, 1 is determined as the starting position. The ignored region 0 can appear in the middle of the path, but the invalid region cannot appear. If the invalid region appears, the starting position is re - calibrated. A complete and valid 1 - 5 - 1 is considered that the action meets the standard once, and the count is incremented by one. By analogy.
[0124] Refer to Figure 4 the schematic diagram of the up - down rotation eye exercise shown. Among them, the double - arrow line represents the rotation path. The target labels of the valid region are 3 and 7, the target labels of the invalid region are 4, 5, 6, 2, 1, and 8, the target label of the ignored region is 0. The starting position can start from 3 or 7. If the node of the gaze path first falls into the valid region 3, 3 is determined as the starting position. The ignored region 0 can appear in the middle of the path, but the invalid region cannot appear. If the invalid region appears, the starting position is re - calibrated. A complete and valid 3 - 7 - 3 is considered that the action meets the standard once, and the count is incremented by one. By analogy.
[0125] Refer to Figure 5Schematic diagram of the eye movement exercise of turning the eyes from the upper left to the lower right. Among them, the double-arrow line represents the movement path. The target labels of the effective area are 2 and 6, the target labels of the invalid area are 3, 4, 5, 1, 7, and 8, and the target label of the ignored area is 0. The starting position can start from 2 or 6. If the node of the gaze path first falls into the effective area 2, 2 is determined as the starting position. The ignored area 0 can appear in the middle of the path, but the invalid area cannot appear. If the invalid area appears, the starting position is re-calibrated. A complete and effective 2-6-2 is considered to meet the standard for one action, and the count is incremented by 1, and so on.
[0126] Refer to Figure 6 Schematic diagram of the eye movement exercise of turning the eyes from the upper right to the lower left. Among them, the double-arrow line represents the movement path. The target labels of the effective area are 4 and 8, the target labels of the invalid area are 1, 2, 3, 5, 6, and 7, and the target label of the ignored area is 0. The starting position can start from 4 or 8. If the node of the gaze path first falls into the effective area 4, 4 is determined as the starting position. The ignored area 0 can appear in the middle of the path, but the invalid area cannot appear. If the invalid area appears, the starting position is re-calibrated. A complete and effective 4-8-4 is considered to meet the standard for one action, and the count is incremented by 1, and so on.
[0127] In an optional embodiment, the identifying the gaze path of the current user's eyes and recording the number of times the eye movement exercise actions meet the standard includes:
[0128] When the eye movement exercise type of the eye movement exercise video to be identified is the second eye movement operation type, obtain the first target value of the effective area corresponding to the second eye movement operation type;
[0129] Obtain the target label of each node, and connect them in the order of the obtained target labels to obtain the first path;
[0130] Determine the first node of the first path as the current node, and sequentially identify each node of the gaze path;
[0131] Judge whether the target label of the next node of the current node in the first path is greater than the target label of the current node;
[0132] When the target label of the next node of the current node in the first path is greater than the target label of the current node, calculate the difference between the target label of the next node of the current node and the first target value to obtain the new target label of the next node of the current node;
[0133] Update the target label in the first path based on the next node of the current node to obtain the second path;
[0134] Determine the first node of the second path as the current node, and sequentially identify starting from the current node of the second path;
[0135] Calculate the difference between the target label of the next node of the current node and the target label of the current node;
[0136] When the calculated difference is greater than or equal to a preset first threshold, determine the current node as the starting position;
[0137] Re - determine the next node of the current node as the current node in the order of the nodes of the second path, and repeat the operation of calculating the difference between the target label of the next node of the current node and the target label of the current node. When the calculated difference is greater than or equal to the preset first threshold, determine the current node as the target node of the starting position. Until the target label of the target node is consistent with the target label of the starting position, determine it as a qualified eye exercise action and count once. And so on, record the qualified times of the eye exercise actions.
[0138] Specifically, the second type of eye operation is clockwise rotation.
[0139] In this embodiment, refer to Figure 7 As shown in the schematic diagram of the clockwise eye rotation exercise, where the circular arrow represents the rotation path, the first target value is 8, the first threshold is - 3, and a standard eye - fixing path should be: 1 - 8 - 7 - 6 - 5 - 4 - 3 - 2 - 1. It can start from any node. Since the clockwise rotation path is a cycle, the nodes need to be processed first. The target nodes in the area passed by the clockwise rotation are continuously decreasing. When identifying the eye gaze path of the current user, if the first path obtained by connecting the target labels in the eye gaze path of the current user is: 1 - 8 - 7 - 6 - 5 - 4 - 3 - 2 - 1, when it is judged that the target label of the next node of the current node is greater than the target label of the current node, subtract 8 from the target label of the next node of the current node to get the second path: 1 - 0 - (-1) - (-2) - (-3) - (-4) - (-5) - (-6) - (-7). After obtaining the second path, if the difference between the target label of the next node of the current node of the second path and the target label of the current node is greater than or equal to - 3, it is considered valid, otherwise it is considered that the eye exercise action is incorrect, and the starting position is re - determined. When the cycle returns to the starting position, it is considered that one action is qualified and the count is incremented by 1. And so on.
[0140] In this embodiment, the first threshold is set to - 3, which determines that the area passed between the current node and the next node during the clockwise rotation of the eye exercise does not exceed 3. If it exceeds 3, it is determined that the eye exercise action is incorrect.
[0141] Further, the step of determining whether the target label of the next node of the current node in the first path is greater than the target label of the current node further includes:
[0142] When the target label of the next node of the current node in the first path is less than or equal to the target label of the current node, it is determined that the eye exercise is correct, and the recognition continues according to the node order of the first path.
[0143] Further, the step of calculating the difference between the target label of the next node of the current node and the target label of the current node further includes:
[0144] When the calculated difference is less than the preset first threshold, it is determined that the eye exercise action of the current node is incorrect, and the starting position is re-determined.
[0145] In an alternative embodiment, the step of recognizing the gaze path of the current user's eyes and recording the number of times the eye exercise actions are up to standard includes:
[0146] When the eye exercise type of the eye exercise video to be recognized is the third eye operation type, obtain the second target value of the effective area corresponding to the third eye operation type;
[0147] Obtain the target label of each node, and connect them in the order of the obtained target labels to obtain a third path;
[0148] Determine the first node of the third path as the current node, and sequentially recognize each node of the gaze path;
[0149] Determine whether the target label of the next node of the current node in the third path is less than the target label of the current node;
[0150] When the target label of the next node of the current node in the third path is less than the target label of the current node, calculate the sum of the target label of the next node of the current node and the target value to obtain the new target label of the next node of the current node;
[0151] Update the target label in the third path based on the new target label of the next node of the current node to obtain a fourth path;
[0152] Determine the first node of the fourth path as the current node, and start sequential recognition from the current node of the fourth path;
[0153] Calculate the difference between the target label of the next node of the current node and the target label of the current node;
[0154] When the calculated difference is less than or equal to the preset second threshold, determine the current node as the starting position;
[0155] Redetermine the next node of the current node in the order of the nodes of the fourth path as the current node, and repeatedly execute the calculation of the difference between the target label of the next node of the current node and the target label of the current node. When the calculated difference is less than or equal to a preset second threshold, determine the current node as the target node at the starting position. Until the target label of the target node is the same as the target label at the starting position, determine it as a qualified eye exercise action and count once. And so on, record the qualified times of the eye exercise actions.
[0156] Specifically, the third type of eye operation is counterclockwise rotation.
[0157] In this embodiment, refer to Figure 8 the schematic diagram of the counterclockwise eye rotation exercise shown in the figure. Among them, the circular arrow represents the rotation path, the second target value is 8, and the second threshold is 3. A standard gaze calibration path should be: 1-2-3-4-5-6-7-8-1. It can start from any node. At the same time, since the counterclockwise rotation path is a cycle, the nodes need to be processed first. The target nodes in the area of the counterclockwise rotation path are constantly increasing. When identifying the gaze path of the current user's eyes, if the third path obtained by connecting the target labels in the gaze path of the current user's eyes is: 1-2-3-4-5-6-7-8-1, when it is judged that the target label of the next node of the current node is less than the target label of the current node, add 8 to the target label of the next node of the current node to obtain the fourth path: 1-2-3-4-5-6-7-8-9. After obtaining the fourth path, if the difference between the target label of the next node of the current node in the fourth path and the target label of the current node is less than or equal to 3, it is considered that the eye exercise action is correct, otherwise it is considered that the eye exercise action is incorrect, and the starting position is redetermined. When the loop returns to the starting position, it is considered that one action is qualified and the count is incremented by 1. And so on.
[0158] Further, the judgment of whether the target label of the next node of the current node in the third path is less than the target label of the current node further includes:
[0159] When the target label of the next node of the current node in the third path is greater than or equal to the target label of the current node, determine that the eye exercise is correct and continue to identify according to the order of the nodes of the third path.
[0160] Further, the calculation of the difference between the target label of the next node of the current node and the target label of the current node includes:
[0161] When the calculated difference is greater than the preset second threshold, determine that the eye exercise action of the current node is incorrect and redetermine the starting position.
[0162] In this embodiment, by leveraging the high performance and strong robustness of the target lightweight model, the visual direction value of the current user's eyes is accurately identified. Then, the gaze path of the current user's eyes is determined using the visual direction values of the sequential target video frames. Based on the gaze path of the current user's eyes, logical judgments are made to further identify and count the eye exercise actions, avoiding misjudgment and missed counting, and improving the recognition efficiency and accuracy of the eye exercise actions.
[0163] In summary, for the method for identifying eye exercise actions based on deep learning described in this embodiment, a pre-trained lightweight model is constructed, and the pre-trained lightweight model is trained based on the second face image set to obtain a target lightweight model. By leveraging the high performance and strong robustness of the target lightweight model, the visual direction value of the current user's eyes is accurately identified. Based on the eye exercise video to be recognized, multiple target video frames are determined, and each target video frame is sequentially input into the target lightweight model. The eye exercise actions of the current user's eyes are recognized according to the visual direction values of the current user's eyes in each recognized target video frame, avoiding misjudgment and missed counting, and improving the recognition efficiency and accuracy of the eye exercise actions.
[0164] Embodiment 2
[0165] Figure 9 It is a structural diagram of an apparatus for identifying eye exercise actions based on deep learning provided in Embodiment 2 of the present invention.
[0166] In some embodiments, the apparatus 20 for identifying eye exercise actions based on deep learning may include multiple functional modules composed of program code segments. The program code of each program segment in the apparatus 20 for identifying eye exercise actions based on deep learning can be stored in the memory of the electronic device and executed by the at least one processor to perform (see details in Figures 1 to 8 the description) the functions of identifying eye exercise actions based on deep learning.
[0167] In this embodiment, the apparatus 20 for identifying eye exercise actions based on deep learning can be divided into multiple functional modules according to the functions it performs. The functional modules may include: a first acquisition module 201, a construction module 202, a second acquisition module 203, a determination module 204, an identification module 205, and an identification and recording module 206. What is referred to as a module in the present invention means a series of computer-readable instruction segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0168] The first acquisition module 201 is configured to acquire a first face image set and perform visual direction value annotation on the first face image set to obtain a second face image set.
[0169] In this embodiment, the first face image set can be obtained from a public network platform. After obtaining the first face image set, each first face image is labeled with a visual direction value to obtain a second face image set. Specifically, when labeling the visual direction value, Row represents the deflection angle in the left-right direction, with the left being positive (+) and the right being negative (-); Pitch represents the deflection angle in the up-down direction, with the up being positive (+) and the down being negative (-). The unit of the angle here is degree (°), and it is 90° when facing the leftmost, rightmost, uppermost, and lowermost directions, and both Row and Pitch are 0° when facing directly forward.
[0170] A construction module 202 is configured to construct a pre-trained lightweight model and train the pre-trained lightweight model based on the second face image set to obtain a target lightweight model.
[0171] In this embodiment, in order to obtain an algorithm logic as simple as possible, rely on as little computing power resources as possible, and ensure the smooth operation of the algorithm on the mobile phone side, this embodiment uses a very famous lightweight model in the field of deep learning, that is, the lightweight basic convolutional neural network MobileNetV2.
[0172] In an optional embodiment, the construction of the pre-trained lightweight model by the construction module 202 includes:
[0173] Obtain an initial lightweight model, and replace the last classification layer in the initial lightweight model with a regression layer with two nodes to obtain a pre-trained lightweight model.
[0174] In this embodiment, the network of the pre-trained lightweight model is different from that of the initial lightweight model. The pre-trained lightweight model replaces the last classification layer of MobileNetV2 with a regression layer that can output two nodes, that is, a regression layer that can output the left-right direction value and the up-down direction value in each face image.
[0175] In an optional embodiment, the construction module 202 training the pre-trained lightweight model based on the second face image set to obtain a target lightweight model includes:
[0176] Obtain each face image in the second face image set and the visual direction value of the user's eyes labeled in each face image to form a data set, where the visual direction value includes the left-right direction value and the up-down direction value;
[0177] Randomly divide the data set into a first number of training sets and a second number of test sets;
[0178] Input the training set into the pre-trained lightweight model for training to obtain a lightweight model;
[0179] Input the test set into the lightweight model for testing to obtain a test pass rate.
[0180] Determine whether the test pass rate is greater than a preset pass rate threshold.
[0181] When the test pass rate is greater than or equal to the preset pass rate threshold, end the training of the lightweight model to obtain a target lightweight model.
[0182] Further, the determination of whether the test pass rate is greater than the preset pass rate threshold includes:
[0183] When the test pass rate is less than the preset pass rate threshold, increase the number of the training set and retrain the lightweight model.
[0184] In this embodiment, during the training process of the target lightweight model, the input parameter is the second face image set, and the target is the annotation data corresponding to each face image: the left - right direction value and the up - down direction value. The left - right direction value and the up - down direction value are used as the training set and input into the pre - trained lightweight model that has been constructed to obtain the target lightweight model. Subsequently, only by obtaining the user's face image, the visual direction value of the user can be identified through the target lightweight model, with high accuracy.
[0185] A second acquisition module 203, configured to acquire a to - be - recognized eye exercise video in response to the received eye exercise recognition request.
[0186] In this embodiment, the eye exercise recognition request is used to request to recognize whether the current user's eye exercise meets the standard. The eye exercise recognition request is initiated by the client to the server, and the server can be an eye exercise recognition subsystem. The server receives the eye exercise recognition request sent by the client and parses the eye exercise recognition request to acquire the to - be - recognized eye exercise video.
[0187] A determination module 204, configured to determine a plurality of target video frames based on the to - be - recognized eye exercise video, and sequentially input each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame.
[0188] In this embodiment, the visual direction value includes a left - right direction value and an up - down direction value. By inputting each target video frame, the visual direction value of the current user's eyes in each video frame can be accurately recognized. The steps are simple, improving the recognition efficiency of the visual direction value.
[0189] In an optional embodiment, the determination module 204 determines a plurality of target video frames based on the to - be - recognized eye exercise video, including:
[0190] Decode the to-be-recognized eye exercise video using a preset decoder to obtain multiple video frames;
[0191] Filter the multiple video frames according to a preset filtering rule to obtain multiple target video frames.
[0192] In this embodiment, a preset video decoder can be used to decode the video frames. Specifically, when the to-be-recognized video is obtained, a preset decoder is obtained according to the to-be-recognized video, and the video frames are decompressed based on the preset video decoder to obtain multiple video frames.
[0193] In this embodiment, the filtering rule can be preset. Specifically, the filtering rule can be preset in advance according to the standard time of each set of eye exercises. For example, it can be set to select one video frame every 2 frames or select one video frame every 3 frames as the target video frame. This embodiment does not limit this here.
[0194] In an alternative embodiment, the determining module 204 sequentially inputs each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame, including:
[0195] According to the chronological order of the multiple target video frames, sequentially input each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame.
[0196] In this embodiment, by presetting the filtering rule in advance according to the standard time of each set of eye exercises, if the operation time of the to-be-recognized eye exercise video is long, it is determined that the current user's eye exercise actions are slow. If one target video frame is selected every 1 frame, a large number of target video frames are obtained, resulting in low recognition efficiency of the target lightweight model; if the operation time of the to-be-recognized eye exercise video is short, it is determined that the current user's eye movements are fast. If one target video frame is selected every 3 frames, a small number of target video frames are obtained, and there may be a problem of missing visual directions, resulting in low recognition accuracy of the target lightweight model.
[0197] The recognition module 205 is configured to recognize the eye exercise actions of the current user's eyes based on the multiple visual direction values of the current user's eyes in the multiple target video frames to obtain the gaze path of the current user's eyes.
[0198] In this embodiment, the gaze path of the current user's eyes is determined by the visual direction values of the current user's eyes in multiple target video frames.
[0199] In an optional embodiment, the recognition module 205 recognizes the eye exercise actions of the current user's eyes based on multiple visual direction values of the current user's eyes in the multiple target video frames, and obtaining the visual path of the current user's eyes includes:
[0200] Obtain a preset two-dimensional spatial map of visual directions, where the two-dimensional spatial map of visual directions contains multiple regions and the label of each region;
[0201] Sort the multiple target video frames in chronological order to obtain a target queue;
[0202] Starting from the head of the target queue, map the visual direction value of the current user's eyes in each target video frame to the two-dimensional spatial map of visual directions one by one, and obtain the target label of the target region where the current user's eyes are located in the two-dimensional spatial map of visual directions for each target video frame;
[0203] Connect multiple target labels in the order of obtaining the target label of the current user's eyes in each target video frame to generate the visual path of the current user's eyes.
[0204] Refer to Figure 2 The shown two-dimensional spatial map of visual directions is calibrated through preset quantization. For example, set two flag bits: 20° and -20°. Based on the two calibration positions, the left-right direction is divided into 3 segments, and the up-down direction is divided into 3 segments. Among them, the middle region 0 represents that the current user is looking straight ahead, and the remaining 1 to 8 regions respectively represent that the current user's visual direction is: left, upper left, up, upper right, right, lower right, down, lower left. By identifying the visual direction determination results of a series of video frames corresponding to the eye exercise video to be recognized, it is possible to determine whether the actions of the eye exercise are up to standard and perform counting.
[0205] The recognition and recording module 206 is used to recognize the visual path of the current user's eyes and record the number of times the eye exercise actions meet the standard.
[0206] In this embodiment, the visual path of the current user's eyes refers to being determined by the rotation direction of the current user's eyes in a series of video frames in the eye exercise video to be recognized. By recognizing the visual path of the current user's eyes, it is possible to quickly identify whether the eye exercise actions are correct and record the number of times the eye exercise actions meet the standard.
[0207] In an optional embodiment, the recognition and recording module 206 recognizes the visual path of the current user's eyes and records the number of times the eye exercise actions meet the standard, including:
[0208] When the eye exercise type of the eye exercise video to be recognized is the first eye operation type, obtain the valid region, invalid region, and ignored region corresponding to the first eye operation type;
[0209] Identify them in sequence starting from the first node of the gaze path. When it is identified that a node of the gaze path first falls into the valid area, determine the node that first falls in as the starting position;
[0210] When it is identified that the next node of the starting position falls into the valid area and the target label of the next node of the starting position is inconsistent with the target label of the starting position, determine the next node of the starting position as the ending position; or when it is identified that the next node of the starting position falls into the ignored area, ignore the next node of the starting position, re-determine the next node of the initial node in the order of the nodes of the gaze path and conduct identification; or when it is identified that the next node of the starting position falls into the invalid area, re-determine the starting position in the order of the nodes of the gaze path and conduct identification;
[0211] When it is identified that the next node of the ending position falls into the valid area and is consistent with the target label of the starting position, it is determined as an eye exercise action that meets the standard once, and count it once. And so on, record the number of times the eye exercise action meets the standard.
[0212] In this embodiment, the types of eye exercises include the first type of eye operation, the second type of eye operation, and the third type of eye operation.
[0213] Specifically, the first type of eye operation may include: left - right rotation, up - down rotation, upper - left to lower - right rotation, and upper - right to lower - left rotation.
[0214] Refer to Figure 3 The schematic diagram of the left - right rotation eye exercise as shown. Among them, the double - arrow line represents the rotation path. The target labels of the valid area are 1 and 5, the target labels of the invalid area are 2, 3, 4, 6, 7, and 8, the target label of the ignored area is 0. The starting position can start from 1 or 5. If a node of the gaze path first falls into the valid area 1, determine 1 as the starting position. The ignored area 0 can appear in the middle of the path, but the invalid area cannot appear. If the invalid area appears, re - calibrate the starting position. A complete and valid 1 - 5 - 1 is considered as an action meeting the standard once, and the count +1. And so on.
[0215] Refer to Figure 4Schematic diagram of the up-and-down eye-rolling exercise shown, where the double-arrow line represents the rotation path. The target labels for the effective area are 3 and 7, the target labels for the invalid area are 4, 5, 6, 2, 1, and 8, and the target label for the ignored area is 0. The starting position can be 3 or 7. If the node of the gaze path first falls within the effective area 3, 3 is determined as the starting position. The ignored area 0 can appear in the middle of the path, but the invalid area cannot. If the invalid area appears, the starting position is re-determined. A complete and valid 3-7-3 is considered one successful action, and the count is incremented by 1, and so on.
[0216] Refer to Figure 5 Schematic diagram of the left-up and right-down eye-rolling exercise shown, where the double-arrow line represents the rotation path. The target labels for the effective area are 2 and 6, the target labels for the invalid area are 3, 4, 5, 1, 7, and 8, and the target label for the ignored area is 0. The starting position can be 2 or 6. If the node of the gaze path first falls within the effective area 2, 2 is determined as the starting position. The ignored area 0 can appear in the middle of the path, but the invalid area cannot. If the invalid area appears, the starting position is re-determined. A complete and valid 2-6-2 is considered one successful action, and the count is incremented by 1, and so on.
[0217] Refer to Figure 6 Schematic diagram of the right-up and left-down eye-rolling exercise shown, where the double-arrow line represents the rotation path. The target labels for the effective area are 4 and 8, the target labels for the invalid area are 1, 2, 3, 5, 6, and 7, and the target label for the ignored area is 0. The starting position can be 4 or 8. If the node of the gaze path first falls within the effective area 4, 4 is determined as the starting position. The ignored area 0 can appear in the middle of the path, but the invalid area cannot. If the invalid area appears, the starting position is re-determined. A complete and valid 4-8-4 is considered one successful action, and the count is incremented by 1, and so on.
[0218] In an optional embodiment, the recognition and recording module 206 recognizes the gaze path of the current user's eyes and records the number of successful eye-rolling exercise actions, including:
[0219] When the eye-rolling exercise type of the eye-rolling exercise video to be recognized is the second eye operation type, obtain the first target value of the effective area corresponding to the second eye operation type;
[0220] Obtain the target label of each node, and connect them in the order of the obtained target labels to obtain the first path;
[0221] Determine the first node of the first path as the current node, and sequentially recognize each node of the gaze path;
[0222] Determine whether the target label of the next node of the current node in the first path is greater than the target label of the current node;
[0223] When the target label of the next node of the current node in the first path is greater than the target label of the current node, calculate the difference between the target label of the next node of the current node and the first target value to obtain the new target label of the next node of the current node;
[0224] Update the target label in the first path based on the next node of the current node to obtain a second path;
[0225] Determine the first node of the second path as the current node and sequentially identify from the current node of the second path;
[0226] Calculate the difference between the target label of the next node of the current node and the target label of the current node;
[0227] When the calculated difference is greater than or equal to a preset first threshold, determine the current node as the starting position;
[0228] Re - determine the next node of the current node as the current node in the order of the nodes of the second path, repeat the operation of calculating the difference between the target label of the next node of the current node and the target label of the current node, and when the calculated difference is greater than or equal to the preset first threshold, determine the current node as the target node of the starting position. Until the target label of the target node is consistent with the target label of the starting position, determine it as a qualified eye exercise action and count once. And so on, record the qualified times of the eye exercise actions.
[0229] Specifically, the second eye operation type is clockwise rotation.
[0230] In this embodiment, refer to Figure 7Schematic diagram of the clockwise eye rotation exercise. Among them, the circular arrow indicates the rotation path. The first target value is 8, and the first threshold is -3. A standard eye fixation path should be: 1-8-7-6-5-4-3-2-1. It can start from any node. At the same time, since the clockwise rotation path is a cycle, the nodes need to be processed first. The target nodes in the area passed by the clockwise rotation are constantly decreasing. When identifying the eye gaze path of the current user, if the first path obtained by connecting the target labels in the eye gaze path of the current user is: 1-8-7-6-5-4-3-2-1, when it is determined that the target label of the next node of the current node is greater than the target label of the current node, subtract 8 from the target label of the next node of the current node to obtain the second path: 1-0-(-1)-(-2)-(-3)-(-4)-(-5)-(-6)-(-7). After obtaining the second path, if the difference between the target label of the next node of the current node in the second path and the target label of the current node is greater than or equal to -3, it is considered valid; otherwise, it is considered that the eye exercise action is incorrect, and the starting position is re-determined. When the cycle returns to the starting position, it is considered that one action is up to standard, and the count is incremented by 1, and so on.
[0231] In this embodiment, the first threshold is set to -3 to determine that the area passed between the current node and the next node of the current node during the clockwise rotation of the eye exercise does not exceed 3. If it exceeds 3, it is determined that the eye exercise action is incorrect.
[0232] Further, the determination of whether the target label of the next node of the current node in the first path is greater than the target label of the current node further includes:
[0233] When the target label of the next node of the current node in the first path is less than or equal to the target label of the current node, it is determined that the eye exercise is correct, and the recognition continues in the order of the nodes in the first path.
[0234] Further, the calculation of the difference between the target label of the next node of the current node and the target label of the current node further includes:
[0235] When the calculated difference is less than the preset first threshold, it is determined that the eye exercise action of the current node is incorrect, and the starting position is re-determined.
[0236] In an alternative embodiment, the recognition and recording module 206 recognizes the eye gaze path of the current user and records the number of times the eye exercise action reaches the standard, including:
[0237] When the eye exercise type of the eye exercise video to be recognized is the third eye operation type, obtain the second target value of the effective area corresponding to the third eye operation type;
[0238] Obtain the target label of each node, and connect them in the order of the obtained target labels to obtain the third path;
[0239] Determine the first node of the third path as the current node, and sequentially identify each node of the gaze path;
[0240] Judge whether the target label of the next node of the current node in the third path is less than the target label of the current node;
[0241] When the target label of the next node of the current node in the third path is less than the target label of the current node, calculate the sum of the target label of the next node of the current node and the target value to obtain the new target label of the next node of the current node;
[0242] Update the target label in the third path based on the new target label of the next node of the current node to obtain the fourth path;
[0243] Determine the first node of the fourth path as the current node, and start to sequentially identify from the current node of the fourth path;
[0244] Calculate the difference between the target label of the next node of the current node and the target label of the current node;
[0245] When the calculated difference is less than or equal to the preset second threshold, determine the current node as the starting position;
[0246] Re-determine the next node of the current node as the current node in the order of the nodes of the fourth path, and repeat the operation of calculating the difference between the target label of the next node of the current node and the target label of the current node. When the calculated difference is less than or equal to the preset second threshold, determine the current node as the target node of the starting position. Until the target label of the target node is consistent with the target label of the starting position, determine it as a qualified eye exercise action and count once. By analogy, record the qualified times of the eye exercise actions.
[0247] Specifically, the third type of eye movement operation is counterclockwise rotation.
[0248] In this embodiment, refer to Figure 8Schematic diagram of the counterclockwise eye rotation exercise. Among them, the circular arrow indicates the rotation path, the second target value is 8, the second threshold is 3, and a standard eye fixation path should be: 1-2-3-4-5-6-7-8-1. It can start from any node. At the same time, since the counterclockwise rotation path is a cycle, the nodes need to be processed first. The target nodes in the area of the counterclockwise rotation path are constantly increasing. When identifying the eye gaze path of the current user, if the third path obtained by connecting the target labels in the eye gaze path of the current user is: 1-2-3-4-5-6-7-8-1, when it is determined that the target label of the next node of the current node is less than the target label of the current node, the target label of the next node of the current node is incremented by 8 to obtain the fourth path: 1-2-3-4-5-6-7-8-9. After obtaining the fourth path, if the difference between the target label of the next node of the current node in the fourth path and the target label of the current node is less than or equal to 3, it is considered that the eye exercise action is correct; otherwise, it is considered that the eye exercise action is incorrect, and the starting position is re-determined. When the loop returns to the starting position, it is considered that one action is up to standard, and the count is incremented by 1, and so on.
[0249] Further, the determination of whether the target label of the next node of the current node in the third path is less than the target label of the current node further includes:
[0250] When the target label of the next node of the current node in the third path is greater than or equal to the target label of the current node, it is determined that the eye exercise is correct, and the recognition continues in the order of the nodes in the third path.
[0251] Further, the calculation of the difference between the target label of the next node of the current node and the target label of the current node includes:
[0252] When the calculated difference is greater than the preset second threshold, it is determined that the eye exercise action of the current node is incorrect, and the starting position is re-determined.
[0253] In this embodiment, by leveraging the high performance and strong robustness of the target lightweight model, the visual direction value of the current user's eyes is accurately recognized. Then, the eye gaze path of the current user is determined using the visual direction values of the sequential target video frames. Based on the eye gaze path of the current user, logical judgments are made to further identify and count the eye exercise actions, avoiding misjudgment and missed counting, and improving the recognition efficiency and accuracy of the eye exercise actions.
[0254] In summary, for the eye exercise action recognition device based on deep learning described in this embodiment, by constructing a pre-trained lightweight model, training the pre-trained lightweight model based on the second face image set to obtain a target lightweight model, and leveraging the high performance and strong robustness of the target lightweight model, the visual direction value of the current user's eyes can be accurately recognized. Based on the eye exercise video to be recognized, multiple target video frames are determined, and each target video frame is sequentially input into the target lightweight model. The eye exercise actions of the current user's eyes are recognized according to the recognized visual direction values of the current user's eyes in each target video frame, avoiding misjudgment and missed counting, and improving the recognition efficiency and accuracy of eye exercise actions.
[0255] Embodiment 3
[0256] Refer to Figure 10 As shown, it is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. In a preferred embodiment of the present invention, the electronic device 3 includes a memory 31, at least one processor 32, at least one communication bus 33, and a transceiver 34.
[0257] Those skilled in the art should understand that Figure 10 The structure of the electronic device shown does not constitute a limitation on the embodiments of the present invention. It can be a bus structure or a star structure. The electronic device 3 may further include more or fewer other hardware or software than shown, or different component arrangements.
[0258] In some embodiments, the electronic device 3 is an electronic device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits, programmable gate arrays, digital signal processors, and embedded devices, etc. The electronic device 3 may further include a client device, and the client device includes, but is not limited to, any electronic product that can perform human-computer interaction with the client through means such as a keyboard, mouse, remote control, touchpad, or voice control device. For example, a personal computer, a tablet computer, a smart phone, a digital camera, etc.
[0259] It should be noted that the electronic device 3 is only an example, and other existing or future possible electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are hereby incorporated by reference.
[0260] In some embodiments, the memory 31 is used to store program codes and various data, such as the deep learning-based eye exercise action recognition device 20 installed in the electronic device 3, and can achieve high-speed and automatic access to programs or data during the operation of the electronic device 3. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data.
[0261] In some embodiments, the at least one processor 32 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The at least one processor 32 is the control core (Control Unit) of the electronic device 3, connects various components of the entire electronic device 3 through various interfaces and circuits, and executes various functions of the electronic device 3 and processes data by running or executing programs or modules stored in the memory 31 and calling data stored in the memory 31.
[0262] In some embodiments, the at least one communication bus 33 is configured to enable connection communication between the memory 31 and the at least one processor 32, etc.
[0263] Although not shown, the electronic device 3 may further include a power source (such as a battery) for powering each component. Optionally, the power source may be logically connected to the at least one processor 32 through a power management device, so as to manage functions such as charging, discharging, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 3 may further include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0264] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0265] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) or a processor to execute a part of the methods described in the various embodiments of the present invention.
[0266] In a further embodiment, in combination with Figure 9 , the at least one processor 32 may execute the operating device of the electronic device 3 and various installed application programs (such as the above-mentioned eye exercise action recognition device 20 based on deep learning), program codes, etc. For example, the above-mentioned various modules.
[0267] Program codes are stored in the memory 31, and the at least one processor 32 may call the program codes stored in the memory 31 to execute related functions. For example, Figure 9 the various modules described in
[0268] are program codes stored in the memory 31 and are executed by the at least one processor 32, so as to implement the functions of the various modules to achieve the purpose of eye exercise action recognition based on deep learning.
[0269] In one embodiment of the present invention, the memory 31 stores a plurality of computer-readable instructions, and the at least one processor 32 executes the plurality of computer-readable instructions to implement the function of recognizing eye exercise movements based on deep learning.
[0270] Specifically, for the specific implementation method of the at least one processor 32 for the above instructions, reference may be made to Figures 1 to 8 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0271] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0272] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0273] In addition, in each embodiment of the present invention, the functional modules can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.
[0274] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights. In addition, it is obvious that the word "comprising" does not exclude other units or, the singular does not exclude the plural. The multiple units or devices described in the present invention can also be implemented by one unit or device through software or hardware. First, second, etc. are used to indicate names and do not represent any specific order.
[0275] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for identifying eye exercise actions based on deep learning, characterized in that, the method includes: Obtain a first face image set, and perform visual direction value annotation on the first face image set to obtain a second face image set; Construct a pre-trained lightweight model, and train the pre-trained lightweight model based on the second face image set to obtain a target lightweight model; the construction of the pre-trained lightweight model includes: obtaining an initial lightweight model, and replacing the last classification layer in the initial lightweight model with a regression layer with two nodes to obtain a pre-trained lightweight model; In response to the received eye exercise recognition request, obtain the eye exercise video to be recognized; the eye exercise types of the eye exercise video to be recognized include a first eye operation type, a second eye operation type, and a third eye operation type; the first eye operation type includes: left-right rotation, up-down rotation, upper-left to lower-right rotation, and upper-right to lower-left rotation; the second eye operation type includes: clockwise rotation; the third eye operation type includes: counterclockwise rotation; Based on the eye exercise video to be recognized, determine a plurality of target video frames, and sequentially input each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame output by the regression layer, where the visual direction value includes a left-right direction value and an up-down direction value; Based on the multiple visual direction values of the current user's eyes in the multiple target video frames, identify the eye exercise actions of the current user's eyes to obtain the gaze path of the current user's eyes; Based on the eye exercise type, identify the gaze path of the current user's eyes and record the number of times the eye exercise actions are up to standard.
2. The method for identifying eye exercise actions based on deep learning according to claim 1, characterized in that, the sequentially inputting each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame includes: According to the time sequence of the multiple target video frames, sequentially input each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame.
3. The method for identifying eye exercise actions based on deep learning according to claim 1, characterized in that, the identifying the eye exercise actions of the current user's eyes based on the multiple visual direction values of the current user's eyes in the multiple target video frames to obtain the gaze path of the current user's eyes includes: Obtain a preset visual direction two-dimensional space map, where the visual direction two-dimensional space map contains multiple regions and the label of each region; Sort the multiple target video frames in chronological order to obtain a target queue; Starting from the head of the target queue, map the visual direction value of the current user's eyes in each target video frame to the visual direction two-dimensional space map one by one, and obtain the target label of the target region where the current user's eyes are located in the visual direction two-dimensional space map for each target video frame; Connect multiple target labels in the order of obtaining the target labels of the current user's eyes in each target video frame to generate the gaze path of the current user's eyes.
4. The deep learning-based eye exercise action recognition method according to claim 1, characterized in that recognizing the gaze path of the current user's eyes and recording the number of times the eye exercise action is completed includes: When the eye exercise type of the eye exercise video to be recognized is the first eye operation type, obtain the valid area, invalid area, and ignored area corresponding to the first eye operation type; Start recognizing from the first node of the gaze path. When it is recognized that the node of the gaze path first falls into the valid area, determine the node where it first falls as the starting position; When it is recognized that the next node of the starting position falls into the valid area and the target label of the next node of the starting position is inconsistent with the target label of the starting position, determine the next node of the starting position as the ending position; or when it is recognized that the next node of the starting position falls into the ignored area, ignore the next node of the starting position, re-determine the next node of the initial node in the order of the nodes of the gaze path and perform recognition; or when it is recognized that the next node of the starting position falls into the invalid area, re-determine the starting position in the order of the nodes of the gaze path and perform recognition; When it is recognized that the next node of the ending position falls into the valid area and is consistent with the target label of the starting position, determine it as a completed eye exercise action and count once. By analogy, record the number of times the eye exercise action is completed.
5. The deep learning-based eye exercise action recognition method according to claim 1, characterized in that recognizing the gaze path of the current user's eyes and recording the number of times the eye exercise action is completed includes: When the eye exercise type of the eye exercise video to be recognized is the second eye operation type, obtain the first target value of the valid area corresponding to the second eye operation type; Obtain the target label of each node and connect them in the order of the obtained target labels to obtain the first path; Determine the first node of the first path as the current node and sequentially recognize each node of the gaze path; Judge whether the target label of the next node of the current node in the first path is greater than the target label of the current node; When the target label of the next node of the current node in the first path is greater than the target label of the current node, calculate the difference between the target label of the next node of the current node and the first target value to obtain the new target label of the next node of the current node; Update the target labels in the first path based on the next node of the current node to obtain the second path; Determine the first node of the second path as the current node and start recognizing sequentially from the current node of the second path; Calculate the difference between the target label of the next node of the current node and the target label of the current node; When the calculated difference is greater than or equal to the preset first threshold, determine the current node as the starting position; Redetermine the next node of the current node in the order of the nodes of the second path as the current node, and repeatedly execute calculating the difference between the target label of the next node of the current node and the target label of the current node. When the calculated difference is greater than or equal to a preset first threshold, determine the current node as the target node at the starting position. Until the target label of the target node is the same as the target label at the starting position, determine it as a qualified eye exercise action and count once. And so on, record the qualified times of the eye exercise action.
6. The method for identifying eye exercise actions based on deep learning according to claim 1, characterized in that identifying the gaze path of the current user's eyes and recording the qualified times of the eye exercise action includes: When the eye exercise type of the eye exercise video to be identified is the third eye operation type, obtain the second target value of the effective area corresponding to the third eye operation type; Obtain the target label of each node, and connect them in the order of the obtained target labels to obtain the third path; Determine the first node of the third path as the current node, and sequentially identify each node of the gaze path; Judge whether the target label of the next node of the current node in the third path is less than the target label of the current node; When the target label of the next node of the current node in the third path is less than the target label of the current node, calculate the sum of the target label of the next node of the current node and the target value to obtain the new target label of the next node of the current node; Update the target label in the third path based on the new target label of the next node of the current node to obtain the fourth path; Determine the first node of the fourth path as the current node, and start identifying sequentially from the current node of the fourth path; Calculate the difference between the target label of the next node of the current node and the target label of the current node; When the calculated difference is less than or equal to a preset second threshold, determine the current node as the starting position; Redetermine the next node of the current node in the order of the nodes of the fourth path as the current node, and repeatedly execute calculating the difference between the target label of the next node of the current node and the target label of the current node. When the calculated difference is less than or equal to a preset second threshold, determine the current node as the target node at the starting position. Until the target label of the target node is the same as the target label at the starting position, determine it as a qualified eye exercise action and count once. And so on, record the qualified times of the eye exercise action.
7. An apparatus for identifying eye exercise actions based on deep learning, characterized in that the apparatus includes: A first acquisition module, configured to acquire a first face image set, and perform visual direction value annotation on the first face image set to obtain a second face image set; A construction module for constructing a pre-trained lightweight model, training the pre-trained lightweight model based on the second face image set to obtain a target lightweight model; constructing the pre-trained lightweight model includes: obtaining an initial lightweight model, and replacing the last classification layer in the initial lightweight model with a regression layer of two nodes to obtain a pre-trained lightweight model; A second acquisition module for acquiring an eye exercise video to be recognized in response to an eye exercise recognition request received; the eye exercise types of the eye exercise video to be recognized include a first eye operation type, a second eye operation type, and a third eye operation type; the first eye operation type includes: left-right rotation, up-down rotation, upper-left to lower-right rotation, upper-right to lower-left rotation; the second eye operation type includes: clockwise rotation; the third eye operation type includes: counterclockwise rotation; A determination module for determining a plurality of target video frames based on the eye exercise video to be recognized, and sequentially inputting each target video frame into the target lightweight model to obtain the visual direction value of the current user's eyes in each target video frame output by the regression layer, where the visual direction value includes a left-right direction value and an up-down direction value; An identification module for identifying the eye exercise action of the current user's eyes based on the plurality of visual direction values of the current user's eyes in the plurality of target video frames to obtain the gaze path of the current user's eyes; An identification and recording module for identifying the gaze path of the current user's eyes based on the eye exercise type and recording the number of times the eye exercise action meets the standard.
8. An electronic device, characterized in that, the electronic device includes a processor and a memory, and the processor is used to implement the deep learning-based eye exercise action recognition method according to any one of claims 1 to 6 when executing a computer program stored in the memory.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, the computer program is used to implement the deep learning-based eye exercise action recognition method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Sight line tracking model training method, and sight line tracking method and device
CN110058694A
Method and system for detecting attention position of eyeballs in real time
CN111209811A
Information processing method and system based on eyeball tracking and payment processing method
CN111667265A
Data processing method and device, operation instruction recognition method and device, equipment and medium
CN112270210A