Swallowing rehabilitation training method, device and equipment and storage medium
By introducing gamification design and facial image detection technology into swallowing rehabilitation training, the operation process is simplified, the effectiveness of swallowing rehabilitation training and patient participation are improved, and the problems of inconvenient operation and high pressure on medical staff in existing technologies are solved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-10
AI Technical Summary
Existing swallowing rehabilitation training methods are inconvenient to operate, resulting in heavy workload for medical staff, poor rehabilitation effects, and patients cannot directly observe their training status in real time.
By responding to swallowing rehabilitation training games selected by users, the system obtains the corresponding training process and qualification assessment standards. It uses facial image detection to determine the qualification of training actions, displays the training process, and provides guidance through games, simplifying the operation and reducing reliance on medical staff.
It improves the effectiveness of swallowing rehabilitation training and patient participation, reduces the workload of medical staff, and allows patients to train in any location.
Smart Images

Figure CN121641331A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of rehabilitation training technology, and in particular relates to a swallowing rehabilitation training method, device, equipment and storage medium. Background Technology
[0002] Dysphagia can easily lead to aspiration, aspiration pneumonia, malnutrition, and psychological and social difficulties, resulting in decreased immune response and increased mortality. Therefore, targeted swallowing training is crucial for patients with dysphagia. For example, strengthening the facial, tongue, and jaw muscles through rehabilitation training can significantly improve swallowing function and enhance patients' quality of life.
[0003] Currently, swallowing training is often conducted during hospitalization or in rehabilitation hospitals both domestically and internationally, with a one-on-one training model. Most rehabilitation training devices on the market that target facial muscles, tongue muscles, or jaw muscles are dental fixation devices, lip or tongue training frames, tongue support surfaces, sensors, etc.
[0004] Current rehabilitation training often requires patients to place their head or tongue on a training frame or wear a dental brace, which is difficult to operate, uncomfortable, and cumbersome. Patients cannot directly observe their own movement status in real time, resulting in heavy workload for medical staff and poor rehabilitation effects. Summary of the Invention
[0005] This application provides a swallowing rehabilitation training method, device, equipment, and storage medium, which can solve the technical problems of heavy workload for medical staff and poor rehabilitation effect caused by the inconvenience of operation of existing swallowing rehabilitation training methods.
[0006] In a first aspect, embodiments of this application provide a swallowing rehabilitation training method, including:
[0007] In response to the user's selected start operation of the swallowing rehabilitation training game, the swallowing rehabilitation training process corresponding to the swallowing rehabilitation training game is obtained, wherein the swallowing rehabilitation training process includes at least one training action and a qualification assessment standard corresponding to at least one training action.
[0008] Obtain the current training action contained in the user's current face image;
[0009] If the current training action is deemed qualified according to the qualification assessment criteria corresponding to the current training action, the number of times the current training action is qualified is determined.
[0010] The swallowing rehabilitation training process is displayed according to the number of times the current training action is passed.
[0011] In one possible implementation of the first aspect, determining the current training action as qualified according to the qualification assessment criteria corresponding to the current training action includes:
[0012] Based on the current training action, determine the qualification assessment criteria corresponding to the current training action;
[0013] Facial key point detection is performed on the current face image to obtain multiple facial key points;
[0014] Based on the multiple facial key points, determine whether the current training action meets the qualification assessment criteria corresponding to the current training action;
[0015] When the current training action meets the qualification assessment criteria corresponding to the current training action, the current training action is determined to be qualified.
[0016] In one possible implementation of the first aspect, displaying the swallowing rehabilitation training process according to the number of qualified attempts of the current training action includes:
[0017] Based on the number of successful attempts at the current training action, determine whether the user has completed the current training action;
[0018] When the user has not completed the current training action, obtain the next frame face image corresponding to the current face image;
[0019] When the user completes the current training action, a voice prompt is issued to prompt the user to choose whether to continue the swallowing rehabilitation training game;
[0020] When the user selects to continue the swallowing rehabilitation training game, the system acquires the user's next frame of facial image in response to the user's selection.
[0021] In one possible implementation of the first aspect, when the swallowing rehabilitation training game selected by the user is a cheek rehabilitation training game, the step of determining whether the current training action meets the qualification assessment criteria corresponding to the current training action based on the plurality of facial key points includes:
[0022] Based on the multiple facial key points, determine whether the user's mouth is currently closed;
[0023] If it is determined that the user's mouth is currently closed, the user's current mouth expression coefficient is calculated based on the multiple facial key points;
[0024] Based on the user's current mouth expression coefficient, determine whether the current training action meets the corresponding qualification evaluation standard;
[0025] The facial expression coefficients include at least one of the following: left corner of mouth backward displacement coefficient, right corner of mouth backward displacement coefficient, lip contraction and opening shape coefficient, lip leftward displacement coefficient, lip rightward displacement coefficient, left lower lip upward compression coefficient, right lower lip upward compression coefficient, lip contraction and compression coefficient, and lower lip movement towards the oral cavity coefficient.
[0026] In one possible implementation of the first aspect, when the swallowing rehabilitation training game selected by the user is a tongue rehabilitation training game, the step of determining whether the current training action meets the qualification assessment criteria corresponding to the current training action based on the plurality of facial key points includes:
[0027] Based on the multiple facial key points, determine whether the user's mouth is currently open;
[0028] If it is determined that the user's mouth is currently open, the current face image is converted to a different format to obtain the current BGR face image and the current HSV face image corresponding to the current face image.
[0029] Based on the current BGR face image and the current HSV face image corresponding to the current face image, determine the H channel value and B channel value of each of the facial key points;
[0030] From the plurality of facial key points, select a plurality of target key points corresponding to the current training action;
[0031] Based on the H-channel and B-channel values of multiple target key points corresponding to the current training action, it is determined whether the current training action meets the qualification evaluation criteria corresponding to the current training action.
[0032] In one possible implementation of the first aspect, when the current training action is an upward or downward tongue movement, determining whether the current training action meets the qualification evaluation criteria corresponding to the current training action based on the H-channel and B-channel values of multiple target key points corresponding to the current training action includes:
[0033] Based on the H-channel and B-channel values of multiple target key points corresponding to the current training action, it is determined whether the user's current tongue state is an extended state;
[0034] When it is determined that the user's current tongue state is an extended state, the gradient of the multiple target key points corresponding to the current training action is calculated based on the pixel values of the multiple target key points corresponding to the current training action.
[0035] Based on the gradients of the multiple target key points corresponding to the current training action, determine whether the current training action meets the qualification evaluation criteria corresponding to the current training action.
[0036] In one possible implementation of the first aspect, when the swallowing rehabilitation training game selected by the user is a neck rehabilitation training game, the step of determining whether the current training action meets the qualification assessment criteria corresponding to the current training action based on the plurality of facial key points includes:
[0037] Based on the multiple facial key points, a 3D transformation matrix corresponding to the current face image is calculated;
[0038] When the current training action is a head-down action, the z-axis rotation component in the 3D transformation matrix corresponding to the current face image is used to determine whether the current training action meets the qualification evaluation criteria corresponding to the current training action.
[0039] Secondly, embodiments of this application provide a swallowing rehabilitation training device, comprising:
[0040] The user operation response module is used to respond to the user's selected swallowing rehabilitation training game start operation, and to obtain the swallowing rehabilitation training process corresponding to the swallowing rehabilitation training game, wherein the swallowing rehabilitation training process includes at least one training action and a qualification assessment standard corresponding to at least one training action.
[0041] A face image acquisition module is used to acquire the current training action contained in the user's current face image;
[0042] The training action qualification judgment module is used to determine the number of times the current training action is qualified when the current training action is determined to be qualified according to the qualification evaluation standard corresponding to the current training action.
[0043] The training action completion judgment module is used to display the swallowing rehabilitation training process according to the number of qualified attempts of the current training action.
[0044] Thirdly, embodiments of this application provide a swallowing rehabilitation training device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any of the above-mentioned embodiments.
[0045] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any of the preceding claims.
[0046] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute any of the methods described above.
[0047] The beneficial effects of this application embodiment compared with the prior art are as follows: By responding to the user's selected swallowing rehabilitation training game, the system obtains the swallowing rehabilitation training process corresponding to the game. This process includes at least one training action and a corresponding qualification assessment standard. The system acquires the current training action from the user's current facial image. If the current training action is deemed qualified according to the qualification assessment standard, the system determines the number of times the current training action has been qualified. Through a game-like approach, users can select the corresponding swallowing rehabilitation training game based on their needs. This is simple and easy to operate, avoiding conventional sensor positioning or training frame support. Rehabilitation training can be completed using a computer with a built-in camera. Furthermore, the swallowing rehabilitation training process is displayed according to the number of qualified actions, making the user's training actions more precise and improving the effectiveness of the swallowing rehabilitation training. In addition, rehabilitation training based on the swallowing rehabilitation training process does not require full-time guidance from medical personnel, and the location of the rehabilitation training is not limited. This solves the technical problems of existing swallowing rehabilitation training methods, which are inconvenient to operate, leading to heavy workloads for medical personnel and poor rehabilitation effects. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a schematic flowchart of a swallowing rehabilitation training method provided in one embodiment of this application;
[0050] Figure 2 This is a schematic diagram of an interface for swallowing rehabilitation training provided in one embodiment of this application;
[0051] Figure 3 This is a schematic diagram of the result of facial key point detection provided in another embodiment of this application;
[0052] Figure 4This is a flowchart illustrating a cheek swallowing rehabilitation training method according to another embodiment of this application;
[0053] Figure 5 This is a flowchart illustrating a tongue swallowing rehabilitation training method according to another embodiment of this application;
[0054] Figure 6 This is a flowchart illustrating a neck swallowing rehabilitation training method according to another embodiment of this application;
[0055] Figure 7 This is a schematic diagram of the structure of a swallowing rehabilitation training device provided in one embodiment of this application. Detailed Implementation
[0056] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0057] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0058] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0059] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0060] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0061] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0062] Dysphagia is a clinical manifestation of impaired structure or function of the jaw, lips, tongue, soft palate, pharynx, esophagus, or other organs, making it impossible to safely and effectively move food from the mouth to the stomach. Decreased muscle mass, reduced saliva production, and oropharyngeal muscle incoordination and delayed nerve reflexes caused by neurological disorders can all lead to dysphagia. Dysphagia can easily lead to aspiration, aspiration pneumonia, malnutrition, and psychological and social difficulties, resulting in a weakened immune response and increased mortality. Therefore, targeted swallowing training is crucial for patients with dysphagia.
[0063] Timely and effective swallowing training for patients with dysphagia helps improve the coordination of swallowing-related muscles, stimulates the central nervous system, expands the sensory range of the cerebral cortex, and promotes the recovery and reconstruction of the swallowing reflex arc. This can significantly improve patients' swallowing ability, reduce their burden, shorten hospital stays, lower the incidence of complications and mortality, and improve their quality of life. After food enters the mouth, it is manipulated and chewed through the contraction of the lips, cheek muscles, and orbicularis oris muscles. Therefore, stretching, strengthening, or other exercises that improve the basic motor characteristics of muscles are very common in swallowing training. For example, improving the strength and coordination of oral and facial muscles helps patients improve chewing efficiency; tongue movement exercises can improve swallowing pressure and tongue muscle strength, effectively enhancing tongue muscle strength and coordination; and tongue elevation and lateral tongue movements can improve oral transport function. Therefore, strengthening rehabilitation training of the facial, tongue, and mandibular muscles can significantly improve patients' swallowing function and quality of life.
[0064] Currently, swallowing training, both domestically and internationally, is often conducted during patient hospitalization or in rehabilitation hospitals, using a one-on-one training model. This cumbersome process leads to heavy workloads for medical staff, hindering remote guidance and intelligent intervention for swallowing rehabilitation. Furthermore, patients may experience emotional changes due to the slow progress of rehabilitation, gradually shifting from initial positive cooperation to negativity and difficulty in maintaining the regimen, often failing to meet the standards set by therapists or nurses.
[0065] Applying gamification to swallowing rehabilitation training involves selecting appropriate gamification intentions, creating a world suitable for patients with swallowing disorders, and providing scenarios that suit different patient experiences. This allows patients to engage in various intuitive and natural real-time sensory interactions, such as visual, auditory, and tactile senses, making the swallowing rehabilitation process enjoyable and participatory, which is of paramount importance.
[0066] This application provides a swallowing rehabilitation training method, as shown in the figure. Figure 1 This is a flowchart illustrating a swallowing rehabilitation training method according to an embodiment of this application, including:
[0067] Step S11: In response to the user's selected swallowing rehabilitation training game start operation, obtain the swallowing rehabilitation training process corresponding to the swallowing rehabilitation training game, wherein the swallowing rehabilitation training process includes at least one training action and a qualification assessment standard corresponding to at least one training action.
[0068] Step S12: Obtain the current training action contained in the user's current face image;
[0069] Step S13: If the current training action is deemed qualified according to the qualification assessment standard corresponding to the current training action, determine the number of times the current training action is qualified.
[0070] Step S14: Based on the number of qualified attempts for the current training action, display the corresponding swallowing rehabilitation training process.
[0071] It should be noted that the swallowing rehabilitation training method described in this embodiment can be implemented through a swallowing rehabilitation training system. The swallowing rehabilitation training system can be installed on an intelligent electronic device with human-computer interaction capabilities. This system includes a voice function for providing voice prompts to the user; a display function for showing the captured images and the current content of the swallowing rehabilitation training system on a screen; a microphone for capturing the user's voice; and a camera function for capturing facial images. The frequency of capturing facial images can be set according to the needs of the scenario, for example, 30 frames per second.
[0072] Specifically, when a user clicks on the swallowing rehabilitation training system, swallowing rehabilitation training games appear on the screen. Different swallowing rehabilitation training games correspond to rehabilitation training for different body parts. These games include, but are not limited to, cheek rehabilitation training games, tongue rehabilitation training games, and neck rehabilitation training games. For example, "Collect Carrots" is a cheek rehabilitation training game for cheek rehabilitation, "Maze Challenge" is a tongue rehabilitation training game for tongue rehabilitation, and "Birds Flying" is a neck rehabilitation training game for neck rehabilitation.
[0073] When a user selects a swallowing rehabilitation training game to begin, in response to the user's selected start action, the system retrieves the swallowing rehabilitation training process corresponding to that game. This process includes one or more training actions, the sequence of these actions, and the corresponding pass / fail criteria for each action. These criteria are used to assess whether the user's training actions are performed correctly. (Reference) Figure 2 , Figure 2 This is a schematic diagram of a swallowing rehabilitation training interface provided in an embodiment of this application. When the user selects the tongue rehabilitation training game "Maze Challenge", the game page is displayed on the screen. The right side of the game page is a maze route map, and the game character is a panda. Each time the panda character moves to a position, a corresponding training action appears. This training action is used to guide the user to perform swallowing rehabilitation training.
[0074] When the user selects a cheek rehabilitation training game, the swallowing rehabilitation training process corresponding to the cheek rehabilitation training game is obtained. The swallowing rehabilitation training process corresponding to the cheek rehabilitation training game includes multiple cheek puffing movements, such as: front cheek puffing movement, left cheek puffing movement and right cheek puffing movement, training thresholds and order of multiple cheek puffing movements, and the qualification assessment criteria corresponding to each cheek puffing movement. Among them, the training threshold of the cheek puffing movement represents the threshold of the number of qualified cheek puffing movements, which is pre-built.
[0075] When the user selects a tongue rehabilitation training game, the swallowing rehabilitation training process corresponding to the tongue rehabilitation training game is obtained. The swallowing rehabilitation training process corresponding to the tongue rehabilitation training game includes multiple tongue protrusion movements, such as: upward tongue protrusion movement, downward tongue protrusion movement, left tongue protrusion movement, and right tongue protrusion movement. The training thresholds, order, and conditions for the occurrence of multiple tongue protrusion movements, as well as the corresponding qualification assessment criteria for each tongue protrusion movement, are obtained. Among them, the training threshold for tongue protrusion movement represents the threshold for the number of qualified tongue protrusion movements, which is pre-constructed.
[0076] When the user selects a neck rehabilitation training game, the swallowing rehabilitation training process corresponding to the neck rehabilitation training game is obtained. The swallowing rehabilitation training process corresponding to the neck rehabilitation training game includes head-down movement, the training threshold for head-down movement, and the qualification assessment standard for head-down movement. Among them, the training threshold for head-down movement represents the threshold for the number of qualified head-down movements, which can be pre-built or randomly generated during game initialization.
[0077] Based on the swallowing rehabilitation training game selected by the user, the game initializes, acquires the swallowing rehabilitation training process, and uses the camera function to capture the current frame image in real time, which is then synchronously displayed on the game page. Before acquiring the current frame image, after acquiring the swallowing rehabilitation training process, the game interface also displays the standard action corresponding to the first movement in the swallowing rehabilitation training process to guide the user in performing that first movement. The first movement is determined according to the order of the training movements in the swallowing rehabilitation training process.
[0078] Based on the acquired current frame image, the system detects whether a face exists in the current frame image. If a face exists in the current frame image, it is used as the current face image. The current face image represents the image captured when the user performs the current training action according to the standard action corresponding to the first action. The pass / fail evaluation standard corresponding to the current training action contained in the current face image is the pass / fail evaluation standard corresponding to the first action, and each current training action contained in the current face image corresponds to a unique first action. If no face exists in the current frame image, the system acquires the next frame image corresponding to the current frame image as the current frame image at the next moment, and detects whether a face exists in the current frame image at the next moment. This process is repeated until the current face image is acquired.
[0079] refer to Figure 2 The left side of the game screen displays real-time captured facial images, allowing users to see their own movements at any time during gameplay and adjust their actions accordingly based on the real-time video feedback. When users are performing tongue rehabilitation training through the "Maze Challenge" tongue rehabilitation game, the right side of the game screen also features game buttons: camera mirror, restart, pause, and exit. Additionally, voice prompts are provided; when the user is ready, a voice prompt to "stick out your tongue" reminds them to perform the training movement.
[0080] It's understandable that when users actively perform facial movements or their faces change naturally, their facial expressions will change, and the key facial features in the relevant areas of the face image will also change. For example, when a user puffs out their right cheek or sticks out their tongue, they will compress the jaw muscles, and the angle of their head will change. However, when a user actively performs training movements, their face undergoes active movement, and the changes in the key facial features will exhibit specific, regular patterns and be more pronounced. For example, when a user actively puffs out their cheeks from side to side, the corners of their mouth will move significantly; when a user actively sticks out their tongue, the tongue will change the color and texture of the area it covers. When a user's face undergoes natural changes, the changes are usually smaller and more irregular, such as small blinks or changes in expression. Changes in the key facial features can be used to determine whether a user's facial movements are active.
[0081] Each training action has a corresponding pass / fail evaluation standard. To avoid misidentifying natural changes in the user's face as training actions performed by the user, the current facial image is processed according to the pass / fail evaluation standard corresponding to the current training action to determine whether the user's current training action is qualified.
[0082] If the user's current training action is determined to be unqualified, the next frame image corresponding to the current face image is obtained as the current frame image at the next moment, and the detection of whether there is a face in the current frame image continues.
[0083] When a user selects to start the swallowing rehabilitation training game, the number of successful attempts for the current training action is set to 0. Each time the current training action is confirmed to be successful, the number of successful attempts for the current training action is incremented by 1, thus accumulating the number of successful attempts for the current training action and obtaining the number of successful attempts for the current training action at the current moment.
[0084] Judging whether a user has completed the current training action based on the number of qualified attempts can reduce the probability of action misidentification. This requires the user to maintain the current training action for a certain period of time, improving the accuracy of judging whether the user has completed the current training action, which is more conducive to achieving rehabilitation goals.
[0085] If the user has not completed the current training action, determine whether the standard action displayed in the game interface is the standard action corresponding to the first action of the current training action, remind the user to continue to complete the first action, obtain the next frame face image corresponding to the current face image, use it as the current face image, and continue to prompt the user to complete the first action.
[0086] When the user completes the current training action, a voice prompt is issued to ask the user whether to continue the swallowing rehabilitation training game. Through multi-dimensional stimulation including visual and auditory elements, the system prompts the user to complete the swallowing rehabilitation training game. When the user chooses to continue, in response to this choice, the system determines the next training action as the first action of the next training step, displays the standard action corresponding to the first action of the next training step, acquires the next frame of the user's face, and repeats steps S12-S14 in a loop.
[0087] In one optional example, the user uses tongue movements to control the movement of a panda character in the game. When the user extends their tongue and performs the current training action, the system determines whether the current training action is satisfactory and provides feedback. In the game interface, when the user performs the first action (extending the tongue, extending the tongue upwards, downwards, leftwards, or rightwards), and it is deemed satisfactory, the number of satisfactory actions for the current training action next to the current face image in the game interface is incremented by one. When the number of satisfactory actions for the current training action reaches the threshold corresponding to that action, the panda character in the maze is activated by the user's tongue movement and moves up, down, left, or right in accordance with that first action. When the user's first action is extending their tongue, the panda character in the game will move one space in the direction the user's tongue is extended. When the user completes the current training action, they can choose to rest for n seconds, retract their tongue, close their mouth, and prepare for the next training session to continue the current rehabilitation training game. The current rehabilitation training game ends when the panda character reaches the bamboo at the end of the maze.
[0088] It should be noted that the swallowing rehabilitation training game process and route can also be designed according to the user's rehabilitation needs. Furthermore, the swallowing rehabilitation training process is randomly generated each time it is opened, which can increase the user's interest.
[0089] It is understood that the embodiments of this application start operation in response to the user's selection of a swallowing rehabilitation training game, and obtain the swallowing rehabilitation training process corresponding to the swallowing rehabilitation training game. The swallowing rehabilitation training process includes at least one training action and a qualification assessment standard corresponding to the training action. The current training action contained in the user's current facial image is obtained. If the current training action is qualified according to the qualification assessment standard corresponding to the current training action, the number of times the current training action is qualified is determined. Through the game, users can select the corresponding swallowing rehabilitation training game according to the part of rehabilitation training they need. It is simple and easy to operate, avoiding conventional sensor positioning or training rack support. Rehabilitation training can be completed using a computer with a built-in camera. And according to the number of qualified times of the current training action, the swallowing rehabilitation training process is displayed, making the training actions performed by the user more accurate and improving the effectiveness of swallowing rehabilitation training. In addition, rehabilitation training according to the swallowing rehabilitation training process does not require full guidance from medical staff, and the location of rehabilitation training is not limited. Thus, it solves the technical problems of heavy workload for medical staff and poor rehabilitation effect caused by the inconvenience of operation of existing swallowing rehabilitation training methods.
[0090] In one possible implementation, step S13, determining whether the current training action is qualified based on the qualification evaluation criteria corresponding to the current training action, includes:
[0091] Step S131: Determine the qualification assessment criteria corresponding to the current training action based on the current training action;
[0092] Step S132: Perform facial key point detection on the current face image to obtain multiple facial key points;
[0093] Step S133: Based on multiple facial key points, determine whether the current training action meets the qualification assessment standard corresponding to the current training action;
[0094] Step S134: When the current training action meets the qualification assessment criteria corresponding to the current training action, the current training action is determined to be qualified.
[0095] Specifically, the current training movement is deemed qualified based on the qualification assessment criteria corresponding to that movement, including:
[0096] Based on the current training movement, determine the corresponding qualification assessment criteria. For example: when the current training movement is a forward cheek puffing movement, determine the qualification assessment criteria corresponding to a forward cheek puffing movement; when the current training movement is a left cheek puffing movement, determine the qualification assessment criteria corresponding to a left cheek puffing movement; when the current training movement is a right cheek puffing movement, determine the qualification assessment criteria corresponding to a right cheek puffing movement. When the current training movement is a left tongue extension movement, determine the qualification assessment criteria corresponding to a left tongue extension movement; when the current training movement is a right tongue extension movement, determine the qualification assessment criteria corresponding to a right tongue extension movement; when the current training movement is an upward tongue extension movement, determine the qualification assessment criteria corresponding to an upward tongue extension movement; when the current training movement is a downward tongue extension movement, determine the qualification assessment criteria corresponding to a downward tongue extension movement. When the current training movement is a head-down movement, determine the qualification assessment criteria corresponding to a head-down movement.
[0097] A face detection model is used to detect facial landmarks in the current face image, resulting in multiple facial landmarks. These facial landmarks are 3D feature points used to represent facial features in the mouth area. The face detection model is trained on a large number of face images and can accurately identify facial landmarks in face images.
[0098] It should be noted that facial landmark detection can be performed on the current face image using the Mediapipe face detection model, or other face detection models; this embodiment does not specifically limit the method. The Mediapipe face detection model is a face detection model built on the Mediapipe machine learning model application framework developed by Google.
[0099] In an optional example, after detecting a face in the current face image, facial landmarks are extracted from the current face image to achieve a complete face mapping, outputting 468 3D facial landmarks. From these, facial landmarks in the mouth region are selected as facial key points, such as... Figure 3 As shown, Figure 3 This is a schematic diagram of the result of facial key point detection provided in another embodiment of this application, where the obtained facial key points are... Figure 3 The key points shown in the image.
[0100] Specifically, from the multiple facial key points detected above, relevant key points corresponding to the current training action are selected. Based on the relevant key points corresponding to the current training action, it is determined whether the current training action meets the corresponding qualification assessment standard. When the user's current training action meets the corresponding qualification assessment standard, the user's current training action is determined to be qualified; when the current training action does not meet the corresponding qualification assessment standard, the user's current training action is determined to be unqualified.
[0101] In an optional example, when the current training action is the cheek-puffing action, relevant key points corresponding to the cheek-puffing action are selected from multiple detected facial key points. Based on the relevant key points corresponding to the cheek-puffing action, it is determined whether the user's current training action meets the qualification assessment criteria corresponding to the cheek-puffing action. If the current training action meets the qualification assessment criteria corresponding to the cheek-puffing action, the user's current training action is determined to be qualified; if the current training action does not meet the qualification assessment criteria corresponding to the cheek-puffing action, the user's current training action is determined to be unqualified.
[0102] In one possible implementation, step S14 involves displaying the swallowing rehabilitation training process based on the number of qualified attempts for the current training action, including:
[0103] Based on the number of successful attempts at the current training action, determine whether the user has completed the current training action;
[0104] If the user has not completed the current training action, obtain the next frame of the face image corresponding to the current face image;
[0105] When the user completes the current training action, a voice prompt is issued to ask the user whether to continue the swallowing rehabilitation training game.
[0106] When the user selects to continue the swallowing rehabilitation training game, the system acquires the user's next frame of facial image in response to the user's selection.
[0107] Specifically: Based on the number of successful repetitions of the current training movement, the swallowing rehabilitation training process is displayed, including:
[0108] In the swallowing rehabilitation training game, there are thresholds for the number of times each training action can be performed. The number of times a user performs a training action at the current moment is compared with the threshold. If the number of times a user performs a training action at the current moment is greater than or equal to the threshold, the user is considered to have completed the current training action. If the number of times a user performs a training action at the current moment is less than the threshold, the user is considered to have failed to complete the current training action.
[0109] If the user has not completed the current training action, the system retrieves the next frame of the face image corresponding to the current face image and issues a voice prompt to remind the user to continue to complete the current training action.
[0110] When the user completes the current training action, a voice prompt is issued, asking the user to choose whether to continue the swallowing rehabilitation training game. If the user chooses to continue, the system acquires the next frame of the user's face. If the user chooses to end the swallowing rehabilitation training game, the system terminates the current session.
[0111] In an optional example, when the game is a tongue-stretching rehabilitation training game and the current training action is tongue protrusion, the game provides real-time feedback on whether the user has completed the tongue protrusion action. A threshold of 3 is set for the number of successful tongue protrusion attempts. If the user's successful tongue protrusion attempt at the current moment is greater than or equal to 3, the user is considered to have completed the tongue protrusion action, and the successful attempts for all other training actions are set to 0. Next, it checks if the panda character has reached the finish line. If the panda character has not reached the finish line, a voice prompt is given to ask the user whether to continue the swallowing rehabilitation training game. If the panda character has reached the finish line, the current tongue-stretching rehabilitation training game ends. If the successful tongue protrusion attempt at the current moment is less than 3, the user is considered to have failed the current tongue protrusion action. The next frame of the face image corresponding to the current face image is then retrieved, and the standard action corresponding to the current tongue protrusion action continues to be displayed in the game interface.
[0112] In one possible implementation, when the user selects a cheek rehabilitation training game, step S133 involves determining, based on multiple facial key points, whether the current training action meets the corresponding qualification assessment criteria, including:
[0113] Based on multiple facial key points, determine whether the user's mouth is currently closed;
[0114] Given that the user's mouth is currently closed, the user's current mouth expression coefficient is calculated based on multiple facial key points.
[0115] Based on the user's current mouth expression coefficient, determine whether the current training action meets the corresponding qualification assessment standard.
[0116] Among them, the mouth expression coefficient includes at least one of the following: left corner of mouth backward displacement coefficient, right corner of mouth backward displacement coefficient, lip contraction and opening shape coefficient, lip left displacement coefficient, lip right displacement coefficient, left lower lip upward compression coefficient, right lower lip upward compression coefficient, lip contraction and compression coefficient, and lower lip movement towards the oral cavity coefficient.
[0117] Specifically, when the user selects a cheek rehabilitation training game, the system uses multiple facial key points to determine whether the current training movement meets the corresponding qualification assessment criteria, including:
[0118] The cheek rehabilitation training game includes several cheek-puffing movements. Since puffing the cheeks requires closing the mouth, after steps S11, S12, S131, and S132, when determining whether the user's current training movement meets the corresponding qualification assessment criteria for cheek-puffing movements, reference is made to... Figure 4 , Figure 4 This is a flowchart illustrating a cheek swallowing rehabilitation training method according to another embodiment of this application. It requires first determining whether the user's mouth is currently closed. From multiple facial key points, lip key points are selected. When the mouth is open, the vertical distance between the upper and lower lips increases, especially in the middle of the lips. Therefore, the key point in the middle of the upper and lower lips is selected as the lip key point. Based on the positional information between the lip key points, it is determined whether the user's mouth is currently closed.
[0119] In an optional example, such as Figure 1 As shown, among multiple facial key points, two key points 12 and 13 in the middle of the upper lip, and one key point 14 in the middle of the lower lip, are selected as lip key points. The vertical distance d1 between key points 12 and 13, and the vertical distance d2 between key points 13 and 14 are calculated. Whether the mouth is closed or not has little impact on d1. When the mouth is closed, d1 is many times larger than d2. Therefore, one-third of d1 is selected as the comparison value of d1. When d2 is greater than the comparison value of d1, the user's mouth is determined to be open; when d2 is not greater than the comparison value of d1, the user's mouth is determined to be closed.
[0120] Specifically, if the user's mouth is currently open, the current training action is determined not to meet the passing evaluation criteria for the cheek-puffing action, and the next frame of the face image corresponding to the current face image is acquired again. If the user's mouth is currently closed, cheek key points are selected from multiple facial key points, including corner mouth key points and lip key points. Based on the position information of the user's cheek key points, the user's current mouth expression coefficient is calculated. The mouth expression coefficient can be represented by the fusion deformation score of the mouth; the higher the fusion deformation score, the greater the influence on the expression. The mouth expression coefficient can include at least one of the following: left corner mouth displacement coefficient, right corner mouth displacement coefficient, lip contraction and opening shape coefficient, lip left displacement coefficient, lip right displacement coefficient, left lower lip upward compression coefficient, right lower lip upward compression coefficient, lip contraction and compression coefficient, and lower lip movement towards the oral cavity coefficient. Among them, the left corner of the mouth backward movement coefficient is the coefficient for the left corner of the mouth to move backward, the right corner of the mouth backward movement coefficient is the coefficient for the right corner of the mouth to move backward, the lip contraction and opening shape coefficient is the coefficient for the lip contraction into an open shape, the lip left movement coefficient is the coefficient for the lip to move to the left simultaneously, the lip right movement coefficient is the coefficient for the lip to move to the right simultaneously, the left lower lip upward compression coefficient is the coefficient for the left lower lip to compress upward, the right lower lip upward compression coefficient is the coefficient for the right lower lip to compress upward, the lip contraction and compression coefficient is the coefficient for the contraction and compression of the lip, and the lower lip movement into the oral cavity coefficient is the coefficient for the lower lip to move into the oral cavity.
[0121] It should be noted that the user's current mouth expression coefficient can be calculated using the Blendshape prediction model based on cheek key points to obtain the fusion deformation score of the mouth, which is then used as the user's current mouth expression coefficient. Other prediction models can also be used to calculate the user's current mouth expression coefficient; this embodiment does not specifically limit this method. Specifically, the Blendshape prediction model is used to predict 52 fusion deformation scores based on the received cheek key points.
[0122] Specifically, different cheek-puffing movements will cause certain changes in the left and right cheeks and the left and right corners of the mouth, which will affect the judgment of some facial expressions and the fusion deformation score of the mouth will also be different. Therefore, for different cheek-puffing movements, different mouth expression coefficients are selected and multiple combinations are made. The combined mouth expression coefficients are used as the qualified evaluation coefficients in the qualified evaluation criteria for the cheek-puffing movement. The qualified thresholds corresponding to the qualified evaluation coefficients of the cheek-puffing movement are used as the qualified evaluation coefficient thresholds in the qualified evaluation criteria for the cheek-puffing movement.
[0123] According to the passing evaluation coefficient in the passing evaluation criteria for the current training action, select multiple current mouth expression coefficients corresponding to the current training action from the calculated current mouth expression coefficients of the user. By comparing the values of the multiple expression coefficients with the passing evaluation coefficient threshold in the passing evaluation criteria for the current training action, determine whether the user's current training action meets the passing evaluation criteria corresponding to the current training action.
[0124] In an optional example, calculate the coefficient of the left corner of the mouth moving backward (MDL), the coefficient of the right corner of the mouth moving backward (MDR), the coefficient of the lips contracting into an open shape (MF), the coefficient of the lips moving left simultaneously (ML), the coefficient of the left lower lip compressing upward (MPL), the coefficient of the right lower lip compressing upward (MPR), the coefficient of the lips contracting and compressing (MP), the coefficient of the lips moving right simultaneously (MR), and the coefficient of the lower lip moving into the oral cavity (MRL).
[0125] If the current training action is the positive puffing action, when the mouth expression coefficients meet the following conditions, it is determined that the user's current training action meets the passing evaluation criteria corresponding to the positive puffing action: MDL ≤ 0.05, MDR ≤ 0.05, MF > 0.004, ML ≤ 0.05, MP > 0.85, MR ≤ 0.05, 0.1 < MRL < 0.2. Otherwise, it is determined that the user's current training action does not meet the passing evaluation criteria corresponding to the positive puffing action.
[0126] If the current training action is the left puffing action, when the mouth expression coefficients meet the following conditions, it is determined that the user's current training action meets the passing evaluation criteria corresponding to the left puffing action: MDL > 0.05, MDR ≤ 0.05, ML > 0.05, MPL ≤ 0.5, MR ≤ 0.05. Otherwise, it is determined that the user's current training action does not meet the passing evaluation criteria corresponding to the left puffing action.
[0127] If the current training action is the right puffing action, when the mouth expression coefficients meet the following conditions, it is determined that the user's current training action meets the passing evaluation criteria corresponding to the right puffing action: MDL ≤ 0.05, MDR > 0.05, ML ≤ 0.05, MPR ≤ 0.5, MR > 0.05. Otherwise, it is determined that the user's current training action does not meet the passing evaluation criteria corresponding to the right puffing action.
[0128] In a possible implementation manner, when the swallowing rehabilitation training game selected by the user is the tongue rehabilitation training game, step S133, determining whether the current training action meets the passing evaluation criteria corresponding to the current training action according to multiple facial key points, includes:
[0129] Step S1331, determine whether the user's mouth is currently in an open state according to multiple facial key points;
[0130] Step S1332: If it is determined that the user's mouth is currently open, the current face image is converted to a different format to obtain the current BGR face image and the current HSV face image corresponding to the current face image.
[0131] Step S1333: Determine the H channel value and B channel value of each facial key point based on the current BGR face image and the current HSV face image corresponding to the current face image;
[0132] Step S1334: Select multiple target key points corresponding to the current training action from multiple facial key points;
[0133] Step S1335: Based on the H-channel and B-channel values of multiple target key points corresponding to the current training action, determine whether the current training action meets the qualification evaluation criteria corresponding to the current training action.
[0134] Specifically, when the user selects a tongue rehabilitation training game, the system uses multiple facial key points to determine whether the current training action meets the corresponding qualification assessment criteria, including:
[0135] Similar to the cheek rehabilitation training game, the tongue rehabilitation training game includes multiple tongue-protrusion movements. After steps S11, S12, S131, and S132, when determining whether the user's current training movement meets the corresponding qualification assessment criteria for tongue protrusion movements, reference is made. Figure 5 , Figure 5 This is a flowchart illustrating a tongue swallowing rehabilitation training method provided in another embodiment of this application. First, based on multiple facial key points, it is determined whether the user's mouth is currently open. The method for determining whether the user's mouth is currently open is the same as the method for determining whether the user's mouth is currently closed, and will not be repeated here.
[0136] Assuming the user's mouth is currently open, the current face image is format-converted to obtain the current BGR face image and the current HSV face image. The current BGR face image contains the B, G, and R channel values for each facial keypoint. The current HSV face image contains the H, S, and V channel values for each facial keypoint.
[0137] Based on the current BGR face image and the current HSV face image corresponding to the current face image, determine the H channel value and B channel value of each facial key point.
[0138] When the current training action is a right tongue protrusion, the left inner corner of the mouth is unobstructed by the tongue, while the right inner corner is obstructed. The color difference of key points in the area inward from the left and right corners of the mouth can be used to determine whether the user's current training action meets the passing evaluation criteria for the right tongue protrusion action. From multiple facial key points, the key points in the area inward from the left and right corners of the mouth are selected as target key points corresponding to the right tongue protrusion action. Based on the H-channel and B-channel values of these target key points, it is determined whether the user's current training action meets the passing evaluation criteria for the right tongue protrusion action.
[0139] In an optional example, from multiple facial key points, such as Figure 3 As shown, key points 78 and 95 on the left corner of the mouth and key points 308 and 318 on the right corner of the mouth are selected. The ordinate of key point 78 is used as the ordinate of the center point on the left, and the abscissa of key point 95 is used as the abscissa of the center point on the left, thus obtaining the center point on the left. Four facial key points adjacent to the center point on the left, above, below, left, and right are then obtained. These five points are used as the target key points for the left corner of the mouth. Similarly, the target key points for the right corner of the mouth are obtained based on key points 308 and 318. The target key points for the left and right corners of the mouth are used as multiple target key points corresponding to the right tongue-out action. When the H-channel and B-channel values of the multiple target key points corresponding to the right tongue-out action meet either of the following two conditions, the user's current training action is determined to not meet the qualification evaluation criteria for the right tongue-out action:
[0140] There is at least one B-channel value greater than 50 for the first left corner of the mouth target key point, and the H-channel values of all the first left corner of the mouth target key points are not in the range of 9-149;
[0141] All B-channel values of the first right corner of the mouth target key points are less than or equal to 50, or all H-channel values of the first right corner of the mouth target key points are between 8 and 150.
[0142] When the H-channel and B-channel values of multiple target key points corresponding to the right tongue extension action do not meet the above two conditions, the user's current training action is determined to meet the qualified evaluation criteria corresponding to the right tongue extension action.
[0143] Specifically, when the current training action is a left tongue protrusion, the right inner corner of the mouth is unobstructed by the tongue, while the left inner corner is obstructed. The color difference of key points in the area inward from the left and right corners of the mouth can be used to determine whether the user's current training action meets the passing evaluation criteria for the left tongue protrusion. From multiple facial key points, the key points in the area inward from the left and right corners of the mouth are selected as multiple target key points corresponding to the left tongue protrusion. Based on the H-channel and B-channel values of these target key points corresponding to the left and right tongue protrusion actions, it is determined whether the user's current training action meets the passing evaluation criteria for the left tongue protrusion.
[0144] In an optional example, from multiple facial key points, such as Figure 3 As shown, key points 78 and 88 on the left corner of the mouth and key points 308 and 324 on the right corner of the mouth are selected. The ordinate of key point 78 is used as the ordinate of the center point on the left side, and the abscissa of key point 88 is used as the abscissa of the center point on the left side, thus obtaining the center point on the left side. Four facial key points adjacent to the center point on the left side are then obtained. These five points—the center point on the left side and its four adjacent facial key points—are used as the target key points for the left corner of the mouth. Similarly, based on key points 308 and 324 on the right corner of the mouth, the target key points for the right corner of the mouth are obtained. These target key points on the left and right corners of the mouth are used as multiple target key points corresponding to the left tongue-out action. When the H-channel and B-channel values of the multiple target key points corresponding to the left tongue-out action meet either of the following two conditions, the user's current training action is determined to not meet the qualification evaluation criteria for the left tongue-out action:
[0145] All B-channel values of the second left corner of the mouth target key points are less than or equal to 50, or all H-channel values of the second left corner of the mouth target key points are in the range of 8-150.
[0146] All B-channel values of the second right corner of the mouth target key points are greater than 50, and all H-channel values of the second right corner of the mouth target key points are not in the range of 9-149.
[0147] If the H-channel and B-channel values of multiple target key points corresponding to the left tongue extension action do not meet the above two conditions, the user's current training action is determined to meet the qualified evaluation criteria corresponding to the left tongue extension action.
[0148] Specifically, when the current training action is a tongue-protrusion movement, the user's mouth is open, resulting in a large gradient change on the upper side of the lower lip. This area is obscured by the tongue, causing a smaller gradient in that region, while the gradient at the tip of the tongue is relatively larger. By calculating the color and gradient of these relevant areas, it can be determined whether the user's current training action meets the passing evaluation criteria for the tongue-protrusion movement. From multiple facial key points, the key points in the area of the lower lip obscured by the tongue are selected as multiple target key points corresponding to the tongue-protrusion movement. (Reference) Figure 5 Based on the H-channel and B-channel values of multiple target key points corresponding to the tongue-protruding action, the system first determines whether the user's tongue is currently extended. If the tongue is extended, the gradients of the multiple target key points corresponding to the tongue-protruding action are calculated. Based on these gradients, the system determines whether the user's current training action meets the pass / fail criteria for the tongue-protruding action. If the tongue is not extended, the system determines that the user's current training action does not meet the pass / fail criteria for the tongue-protruding action and then obtains the next frame image corresponding to the current face image.
[0149] In an optional example, refer to Figure 3 When the current training action is the tongue-down movement, from multiple facial key points, select key points 13, 14, 15 and 18 in the area above the lower lip and the tip of the tongue that are covered by the tongue. Calculate the center point Pt of the line connecting key point 13 and key point 14. Select the first tongue key point in the area connecting the center point Pt and key point 15 from the facial key points. Use the first tongue key point as multiple target key points corresponding to the tongue-down movement.
[0150] Specifically, when the current training action is the tongue-up movement, the user's mouth is open. The gradient changes significantly on both sides of the upper lip. The area below the upper lip is obscured by the tongue, causing a smaller gradient in that area, while the gradient at the tip of the tongue is relatively larger. By calculating the color and gradient of these areas, it can be determined whether the user's current training action meets the passing evaluation criteria for the tongue-up movement. From multiple facial key points, the key points in the area below the upper lip obscured by the tongue are selected as multiple target key points corresponding to the tongue-up movement. (Reference) Figure 5Based on the H-channel and B-channel values of multiple target key points corresponding to the downward tongue protrusion action, it is first determined whether the user's tongue is currently extended. If the user's tongue is extended, the gradients of multiple target key points corresponding to the upward tongue protrusion action are calculated. Based on the gradients of these target key points, it is determined whether the user's current training action meets the qualification evaluation criteria for the upward tongue protrusion action. If the user's tongue is not extended, it is determined that the user's current training action does not meet the qualification evaluation criteria for the upward tongue protrusion action, and the next frame image corresponding to the current face image is obtained.
[0151] In an optional example, refer to Figure 3 When the current training action is the tongue extension action, from multiple facial key points, select key points 13, 14, 12 and 164 in the area of the lower side of the upper lip and the tip of the tongue that are covered by the tongue. Calculate the center point Pb of the line connecting key point 13 and key point 14. Select the second tongue key point in the area of the line connecting the center point Pb and key point 12 from the facial key points. Use the second tongue key point as multiple target key points corresponding to the tongue extension action.
[0152] In one possible implementation, when the current training action is an upward or downward tongue movement, in step S1335, based on the H-channel and B-channel values of multiple target key points corresponding to the current training action, it is determined whether the current training action meets the corresponding qualification evaluation criteria, including:
[0153] Based on the H-channel and B-channel values of multiple target key points corresponding to the current training action, determine whether the user's current tongue state is extended.
[0154] When the user's current tongue state is determined to be an extended state, the gradients of the multiple target key points corresponding to the current training action are calculated based on the pixel values of the multiple target key points corresponding to the current training action.
[0155] Based on the gradients of multiple target key points corresponding to the current training action, determine whether the current training action meets the corresponding qualification assessment criteria.
[0156] Specifically, when the current training action is an upward or downward tongue movement, the H-channel and B-channel values of multiple target key points corresponding to the current training action are used to determine whether the current training action meets the corresponding qualification assessment criteria, including:
[0157] When the current training action is a tongue-down movement, the system determines whether the user's current tongue state is extended based on the H-channel and B-channel values of multiple target keypoints corresponding to the tongue-down movement. If the user's current tongue state is determined to be extended, the system calculates the gradients of the multiple target keypoints corresponding to the tongue-down movement based on their pixel values. Figure 5 Based on the gradients of multiple target key points corresponding to the tongue-out movement, it is determined whether the user's current training movement meets the qualification evaluation criteria corresponding to the tongue-out movement.
[0158] Continuing with the example above, when the current training action is a downward tongue extension, the first tongue keypoint is used as one of the multiple target keypoints corresponding to the downward tongue extension action. Based on the H-channel and B-channel values of the first tongue keypoint, it is determined whether the user's current tongue state is extended. The user's current tongue state is determined to be non-extended when the H-channel and B-channel values of the first tongue keypoint meet the following conditions:
[0159] All B-channel values of the first tongue key points are less than or equal to 50, or all H-channel values of the first tongue key points are in the range of 8-150.
[0160] If the H-channel and B-channel values of the first tongue keypoint do not meet the above conditions, the user's current tongue state is determined to be an extended state. All first gradient keypoints on the line connecting center point Pt to keypoint 18 are selected. Based on the pixel values of all first gradient keypoints, their gradient values are calculated. The first gradient keypoint with the largest gradient value is selected as the first final target keypoint. If the ordinate of the first final target keypoint is greater than the ordinate of keypoint 15, the user's current training action is determined to meet the qualification evaluation criteria for the extended tongue action. If the ordinate of the first final target keypoint is not greater than the ordinate of keypoint 15, the user's current training action is determined to not meet the qualification evaluation criteria for the extended tongue action.
[0161] Specifically, when the current training action is tongue extension, the system determines whether the user's current tongue state is extended based on the H-channel and B-channel values of multiple target keypoints corresponding to the tongue extension action. If the user's current tongue state is determined to be extended, the system calculates the gradients of these target keypoints based on their pixel values, and then references the provided text. Figure 5 Based on the gradients of multiple target key points corresponding to the tongue-up movement, it is determined whether the user's current training action meets the qualification evaluation criteria corresponding to the tongue-up movement.
[0162] Continuing with the example above, when the current training action is tongue extension, the second tongue keypoint is used as multiple target keypoints corresponding to the tongue extension action. Based on the H-channel and B-channel values of the second tongue keypoint, it is determined whether the user's current tongue state is extended. The user's current tongue state is determined to be non-extended when the H-channel and B-channel values of the second tongue keypoint meet the following conditions:
[0163] All B-channel values of the second tongue key points are less than or equal to 50, or all H-channel values of the second tongue key points are in the range of 8-150.
[0164] If the H-channel and B-channel values of the second tongue keypoint do not meet the above conditions, the user's current tongue state is determined to be an extended state. All second gradient keypoints on the line connecting the center point Pb to keypoint 164 are selected. Based on the pixel values of all second gradient keypoints, the gradient values of all second gradient keypoints are calculated. The second gradient keypoint with the largest gradient value is selected as the second final target keypoint. When the ordinate of the second final target keypoint is less than the ordinate of keypoint 12, the user's current training action is determined to meet the qualification evaluation criteria for the extended tongue action. When the ordinate of the second final target keypoint is not less than the ordinate of keypoint 12, the user's current training action is determined to not meet the qualification evaluation criteria for the extended tongue action.
[0165] In one possible implementation, when the user selects a swallowing rehabilitation training game as a neck rehabilitation training game, step S133 involves determining, based on multiple facial key points, whether the current training action meets the corresponding qualification assessment criteria, including:
[0166] Based on multiple facial key points, the 3D transformation matrix corresponding to the current face image is calculated;
[0167] When the current training action is a head-down action, the z-axis rotation component in the 3D transformation matrix corresponding to the current face image is used to determine whether the current training action meets the corresponding qualification evaluation criteria.
[0168] Specifically, when the user selects a neck rehabilitation training game, after steps S11, S12, S131, and S132, if the current training action is a head-down movement, it is determined whether the user's current training action meets the qualification assessment criteria corresponding to the head-down movement, referring to... Figure 6 , Figure 6 This is a flowchart illustrating a neck swallowing rehabilitation training method according to another embodiment of this application. Based on multiple facial key points, it determines whether the current training movement meets the corresponding qualification assessment criteria, including:
[0169] Based on multiple detected facial key points, a 3D transformation matrix corresponding to the current face image is calculated using the FaceMarkerer matrix model. The FaceMarkerer matrix model is a deep learning model used to transform facial key points from a canonical face model into the detected face using the transformation matrix, allowing the user to apply the detected facial key points. The 3D transformation matrix includes component information along various coordinate axes, such as rotation, translation, and scaling components.
[0170] When the current training action is a head-down action, the system determines whether the user's current training action meets the qualification evaluation criteria for a head-down action based on the z-axis rotation component in the 3D transformation matrix corresponding to the current face image.
[0171] In an optional example, based on multiple detected facial key points, a 4x4 homogeneous 3D transformation matrix corresponding to the current face image is calculated using the FaceMarkerer matrix model. M20, M21, and M22 in the 4x4 homogeneous transformation matrix represent the rotation components along the z-axis in 3D coordinates. Head-down or head-up actions affect the value of M21. The value of M21 is selected to determine whether the user's current training action meets the pass / fail criteria for head-down actions. When the value of M21 is greater than or equal to 0.22, the user's current training action is deemed to meet the pass / fail criteria for head-down actions; when the value of M21 is less than 0.22, the user's current training action is deemed not to meet the pass / fail criteria for head-down actions.
[0172] It should be noted that when a user selects to start the "Bird Flying" game, a threshold number of qualified head-down actions is randomly generated, and the length of the corresponding pipe is generated according to this threshold number of qualified actions.
[0173] In an optional example, refer to Figure 6 When the current training action meets the pass / fail evaluation criteria for the head-down action, the user's current training action is deemed passable. The number of passable head-down actions is calculated, and it is determined whether the number of passable head-down actions meets the pass / fail threshold. If the number of passable head-down actions meets the pass / fail threshold, the user is deemed to have completed the head-down action. If the number of passable head-down actions does not meet the pass / fail threshold, the user is deemed not to have completed the head-down action, and the next frame of the face image is acquired. When the user is deemed to have completed the head-down action, the user is prompted whether to continue the "Bird Flying" game. If the user selects no, the game ends. If the user selects to continue, the next frame of the face image is acquired.
[0174] Understandably, in this embodiment, training the model with a large number of images is unnecessary; a single frame is sufficient to determine different training actions. By selecting target facial key points related to the current training action, such as the corners of the mouth and the edges of the lips, the relative changes between these target facial key points (such as distance, color, and gradient) are calculated. Combined with auxiliary data such as expression coefficients and transformation matrices, the calculation results are compared with predetermined thresholds. When a specific threshold or pattern is met, it can be determined whether the user's current training action is qualified. Furthermore, through gamification and competition, the compliance and enjoyment of rehabilitation training are increased, improving patients' acceptance of video games in swallowing training and enhancing the improvement in patients' oral muscle strength and endurance, as well as swallowing function. This embodiment also incorporates voice prompts to alleviate the workload of medical staff in the current "one-on-one" rehabilitation model. Medical staff can also use voice prompts to determine whether the patient's swallowing movements are qualified, making it intuitive and easy to use. This allows both patients and medical staff to quickly get started, effectively solving the technical problems of heavy workload for medical staff and poor rehabilitation effects caused by the inconvenience of operation in existing swallowing rehabilitation training methods.
[0175] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0176] Corresponding to the method in the above embodiments, Figure 7 A schematic diagram of a swallowing rehabilitation training device provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.
[0177] Reference Figure 7 The device includes:
[0178] User operation response module 71 is used to respond to the user's selected swallowing rehabilitation training game start operation, and obtain the swallowing rehabilitation training process corresponding to the swallowing rehabilitation training game. The swallowing rehabilitation training process includes at least one training action and a qualification assessment standard corresponding to at least one training action.
[0179] The face image acquisition module 72 is used to acquire the current training action contained in the user's current face image;
[0180] The training action qualification judgment module 73 is used to determine the number of times the current training action is qualified when the current training action is determined to be qualified according to the qualification evaluation standard corresponding to the current training action.
[0181] The training action completion judgment module 74 is used to display the swallowing rehabilitation training process according to the number of qualified attempts for the current training action.
[0182] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0183] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0184] This application also provides a swallowing rehabilitation training device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in any of the above method embodiments.
[0185] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0186] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.
[0187] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0188] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0189] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0190] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0191] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0192] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method of swallowing rehabilitation training, characterized by, The method comprises the following steps: In response to a user-selected swallowing rehabilitation training game start operation, a swallowing rehabilitation training flow corresponding to the swallowing rehabilitation training game is acquired, wherein the swallowing rehabilitation training flow comprises at least one training action and a qualified evaluation standard corresponding to the at least one training action; A current training action contained in a current face image of the user is acquired; In a case where the current training action is determined to be qualified according to the qualified evaluation standard corresponding to the current training action, a qualified number of the current training action is determined; According to the qualified number of the current training action, the swallowing rehabilitation training flow is correspondingly displayed.
2. The swallowing rehabilitation training method according to claim 1, wherein The determination of the current training action to be qualified according to the qualified evaluation standard corresponding to the current training action comprises: According to the current training action, the qualified evaluation standard corresponding to the current training action is determined; Face key point detection is performed on the current face image to obtain a plurality of face key points; According to the plurality of face key points, it is judged whether the current training action conforms to the qualified evaluation standard corresponding to the current training action; When the current training action conforms to the qualified evaluation standard corresponding to the current training action, it is determined that the current training action is qualified.
3. The swallowing rehabilitation training method according to claim 1, wherein The display of the swallowing rehabilitation training flow according to the qualified number of the current training action comprises: According to the qualified number of the current training action, it is judged whether the user has completed the current training action; When the user has not completed the current training action, a next frame of face image corresponding to the current face image is acquired; When the user has completed the current training action, a voice prompt is issued to prompt the user to select whether to continue the swallowing rehabilitation training game; When the user selects a swallowing rehabilitation training game continuation operation, in response to the user-selected swallowing rehabilitation training game continuation operation, a next frame of face image of the user is acquired.
4. The swallowing rehabilitation training method according to claim 2, wherein When the user-selected swallowing rehabilitation training game is a buccal rehabilitation training game, the judgment of whether the current training action conforms to the qualified evaluation standard corresponding to the current training action according to the plurality of face key points comprises: According to the plurality of face key points, it is determined whether the user's mouth is currently in a closed state; In a case where it is determined that the user's mouth is currently in a closed state, according to the plurality of face key points, a current mouth expression coefficient of the user is calculated; According to the current mouth expression coefficient of the user, it is judged whether the current training action conforms to the qualified evaluation standard corresponding to the current training action; The mouth expression coefficient comprises at least one of the following: left corner of the mouth backward movement coefficient, right corner of the mouth backward movement coefficient, double-lip contraction and opening shape coefficient, double-lip left movement coefficient, double-lip right movement coefficient, left lower lip upward compression coefficient, right lower lip upward compression coefficient, double-lip contraction compression coefficient, and lower lip moving to the oral cavity coefficient.
5. The swallowing rehabilitation training method according to claim 2, wherein When the swallowing rehabilitation training game selected by the user is a tongue rehabilitation training game, the step of judging whether the current training action conforms to the qualified evaluation standard corresponding to the current training action according to the plurality of facial key points comprises: determining whether the mouth of the user is currently in an open state according to the plurality of facial key points; in a case where it is determined that the mouth of the user is currently in an open state, performing format conversion on the current face image to obtain a current BGR face image and a current HSV face image corresponding to the current face image; determining the H channel value and the B channel value of each facial key point according to the current BGR face image and the current HSV face image corresponding to the current face image; selecting a plurality of target key points corresponding to the current training action from the plurality of facial key points; judging whether the current training action conforms to the qualified evaluation standard corresponding to the current training action according to the H channel value and the B channel value of the plurality of target key points corresponding to the current training action.
6. The swallowing rehabilitation training method according to claim 5, wherein When the current training action is an up tongue action or a down tongue action, the step of judging whether the current training action conforms to the qualified evaluation standard corresponding to the current training action according to the H channel value and the B channel value of the plurality of target key points corresponding to the current training action comprises: judging whether the tongue state of the user is currently in an extended state according to the H channel value and the B channel value of the plurality of target key points corresponding to the current training action; when it is determined that the tongue state of the user is currently in an extended state, calculating the gradient of the plurality of target key points corresponding to the current training action according to the pixel value of the plurality of target key points corresponding to the current training action, respectively; judging whether the current training action conforms to the qualified evaluation standard corresponding to the current training action according to the gradient of the plurality of target key points corresponding to the current training action.
7. The swallowing rehabilitation training method according to claim 2, wherein When the swallowing rehabilitation training game selected by the user is a neck rehabilitation training game, the step of judging whether the current training action conforms to the qualified evaluation standard corresponding to the current training action according to the plurality of facial key points comprises: calculating a 3D transformation matrix corresponding to the current face image according to the plurality of facial key points; when the current training action is a head-down action, judging whether the current training action conforms to the qualified evaluation standard corresponding to the current training action according to the z-axis rotation component in the 3D transformation matrix corresponding to the current face image.
8. A swallowing rehabilitation training device, characterized by, comprise: a user operation response module, configured to, in response to a swallowing rehabilitation training game selected by a user starting to operate, acquire a swallowing rehabilitation training process corresponding to the swallowing rehabilitation training game, wherein the swallowing rehabilitation training process comprises at least one training action and a qualified evaluation standard corresponding to at least one training action; a face image acquisition module, configured to acquire a current training action contained in a current face image of the user; The training action qualification judgment module is configured to determine a qualified number of the current training action when it is determined that the current training action is qualified according to the qualified evaluation standard corresponding to the current training action. The training action completion judgment module is configured to correspondingly display the swallowing rehabilitation training process according to the qualified number of the current training action.
9. A swallowing rehabilitation training device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method of any one of claims 1 to 7.