Swallowing training game execution system and method based on computer vision

Through a computer vision-based swallowing training game execution system, the Unity engine and deep learning technology are used to achieve swallowing training without external devices, improving the convenience and fun of swallowing training.

CN120268053APending Publication Date: 2025-07-08HUNAN PROVINCIAL TUMOR HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510248612.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing swallowing training game system relies on external devices, resulting in insufficient convenience and fun.

Method used

The swallowing training game execution system based on computer vision is adopted, and the Unity engine and deep learning technology are used to collect facial motion information through mobile phone cameras, identify user actions and drive game character movements, realizing swallowing training without the need for surface electromyography sensors and facial motion recognition devices.

Benefits of technology

It improves the convenience and fun of swallowing training, simplifies operations, and only needs a smartphone to realize swallowing training game, solving the problem of sensor dependence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120268053A_ABST
    Figure CN120268053A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computer vision, and relates to a swallowing training game execution system and method based on computer vision, and the system comprises an interaction module which is used for recognizing a user action; the system module is used for managing player login authentication, level loading and selection, friend interaction, player chatting, player data ranking information and player integral acquisition and use; the core module is used for managing a role state machine of a player, behavior interaction and dress changing, a camera picture, a game interface and communication between a client and a server, and processing event triggering in the game, sound in the game and the camera picture in the game; the server side is used for a user to carry out login authentication, process game logic and store data; and the Unity engine is used for game client development of the system. The swallowing training game system is simple in structure and convenient to operate, the problem that a swallowing training game system depends on equipment such as a sensor is solved, and the convenience of a swallowing training game and the interestingness of swallowing training are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a swallowing training game execution system and method based on computer vision. Background Art

[0002] Common swallowing training includes sensory stimulation, active swallowing training and passive swallowing training. Among them, common active swallowing training includes buccal movement, tongue movement, mandibular movement, etc. When an individual conducts swallowing training, the game client collects facial movement information through the mobile phone camera and uploads it to the face recognition module developed by computer vision technology for dynamic recognition of facial feature points, and then returns an action identifier to the interaction module. According to the matching of the game interaction configuration, corresponding control of the game end is performed to drive the character to move, jump, etc., and the game scene objects make corresponding interaction feedback.

[0003] The prior art uses a computer screen and a face recognition device for swallowing training. An individual will face a computer and a face recognition device, which is constructed by a facial landmark detector in the artificial intelligence model library. It will generate the three-dimensional positions of 468 facial feature points on the facial image, including information on facial regions such as the cheeks, mouth, and chin. By detecting the changes in the three-dimensional positions of the facial feature points, the changes in the facial muscles and expressions of the participant are detected, and the computer obtains biofeedback through real-time game control calculation. The prior art requires not only the equipment of the game end but also relies on surface electromyography sensors or facial action recognition devices. The placement position of the sensors is greatly restricted and can only be attached to the body surface position. For example, for the movement of the tongue, which is a key part of swallowing training, the sensors cannot achieve induction, and the face recognition device also limits the accessibility of the swallowing training game, resulting in a reduction in the convenience and practicality of swallowing training. The patent with the publication number CN110037695B provides a swallowing trainer, including: a piezoresistive pressure sensor placed in the patient's throat during swallowing detection, used to detect the movement strength of the patient's laryngeal muscles when the patient swallows food and generate an electrical signal representing the movement strength; a laryngeal muscle trainer placed in the patient's throat during the patient's swallowing training; a controller respectively connected to the output end of the piezoresistive pressure sensor, the control end of the laryngeal muscle trainer, and the upper computer, used to control the laryngeal muscle trainer to train the patient's laryngeal muscles after receiving a swallowing training instruction; when performing swallowing detection on the patient, uploading the electrical signal to the upper computer for display for professional treatment doctors to view. In this patent, swallowing training is carried out by externally placing sensors at the throat. The placement position of the sensors is restricted, and the convenience and practicality of swallowing training are relatively low, having the same drawbacks as the prior art.

[0004] Therefore, how to solve the dependence of swallowing training games on devices such as sensors to improve the convenience of swallowing training games and at the same time have training interest is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the object of the present invention is to provide a computer vision-based swallowing training game execution system to solve the problems of dependence on external devices and lack of interest in swallowing training games in the prior art; in addition, the present invention also provides a computer vision-based swallowing training game execution method.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a computer vision-based swallowing training game execution system, including:

[0008] An interaction module for recognizing user actions to drive the movement of game characters;

[0009] A system module for managing player login authentication, loading and selection of levels of different difficulties, adding, deleting, viewing and interacting with friends, private chatting between different players and group chatting of multiple players, player data ranking information, obtaining and using of player points;

[0010] A core module for managing the character state machine, behavior interaction and dressing change of players, collecting camera screen, displaying and interacting with the game interface, communicating between the client and the server, handling collisions and event triggers between objects in the game, loading, playing and stopping sounds in the game, and recognizing the camera screen in the game;

[0011] A server for user login authentication, handling game logic, storing player account data, character dressing data, level completion data, friend data, chat data, player point data and leaderboard data;

[0012] The Unity engine for developing the game client of the system.

[0013] Further, the interaction module includes a head and neck relaxation exercise interaction module, a buccal movement interaction module, a tongue movement interaction module, a mandibular movement interaction module, a swallowing technique interaction module and a CTAR training interaction module.

[0014] Further, the system module includes a login module, a level module, a friend module, a chat module, a leaderboard module and a point module.

[0015] Further, the core module includes a character module, a camera module, a UI module, a network module, a physics module, an audio module and a face recognition module.

[0016] Further, the server communicates with the client using Socket, the protocol uses Protobuf, and the database uses MySQL.

[0017] In a second aspect, the present invention also provides a method for executing a swallowing training game based on computer vision, including the following steps:

[0018] S10. Develop a game client using the Unity3D engine, including an interaction module, a core module, and a system module;

[0019] S20. Develop a facial action recognition model using Python based on deep learning and convolutional neural networks in computer vision technology;

[0020] S30. Develop the login server, logic server, and database;

[0021] S40. Transmit the frame images captured by the mobile phone camera to the face recognition module for dynamic recognition of facial feature points. The face recognition module returns an action identifier as an interactive input source, and controls the corresponding in-game Avatar according to the matching of the game interaction configuration, driving the game scene characters and objects to make corresponding interactive feedback. After all operations are completed, a single game settlement is performed and communication and storage are carried out with the server.

[0022] Compared with the prior art, the swallowing training game execution system and method based on computer vision provided by the present invention has at least the following beneficial effects:

[0023] In the prior art, in addition to the equipment on the game side, it also needs to rely on surface electromyography sensors or facial action recognition devices. The placement positions of the sensors are greatly limited and can only be attached to the body surface positions. For example, for the key parts of swallowing training such as the movement of the tongue, the sensors cannot achieve induction, and the facial recognition devices also limit the accessibility of the swallowing training game, resulting in a reduction in the convenience and practicality of the swallowing training game. The structure of the present invention is simple and the operation is convenient. By using computer vision technology to develop a facial action recognition model, it does not require surface electromyography sensors, computers, and facial action recognition devices. Only a smart phone can be used to implement the swallowing training game, solving the dependence of the swallowing training game system on devices such as sensors, and having a high degree of simplicity and intelligence, greatly improving the convenience and interest of gamified swallowing training. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the solution of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 The architecture diagram of a computer vision-based swallowing training game execution system provided by an embodiment of the present invention;

[0026] Figure 2 The flowchart of a computer vision-based swallowing training game execution method provided by an embodiment of the present invention. Detailed implementation manners

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs; the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. For example, terms such as "length", "width", "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or position based on the orientation or position shown in the drawings, which are only for convenience of description and should not be construed as a limitation to the technical solution of the present application.

[0028] The terms "comprising" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion; the terms "first", "second", etc. in the specification and claims of the present invention or the above-mentioned drawings are used to distinguish different objects and not for describing a specific order. In the specification and claims of the present invention and the above-mentioned drawings, when an element is referred to as "fixed to" or "mounted on" or "disposed on" or "connected to" another element, it may be directly or indirectly located on the other element. For example, when an element is referred to as "connected to" another element, it may be directly or indirectly connected to the other element.

[0029] In addition, the mention of "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0030] The present invention provides a computer vision-based swallowing training game execution system, which is applied to the process of active swallowing training (such as buccal movement, tongue movement, mandibular movement, etc.). The computer vision-based swallowing training game execution system includes:

[0031] An interaction module for recognizing user actions to drive the movement of game characters; a system module for managing player login authentication, loading and selection of levels with different difficulties, adding, deleting, viewing and interacting with friends, private chatting between different players, group chatting among multiple players, player data ranking information, obtaining and using player points; a core module for managing the player's character state machine, behavior interaction and dressing up, capturing camera images, displaying and interacting with the game interface, communicating between the client and the server, handling collisions and event triggers between objects in the game, loading, playing and stopping sounds in the game, and recognizing camera images in the game; a server for authenticating user logins, processing game logic, storing player account data, character dressing data, level completion data, friend data, chat data, player point data, and leaderboard data; the Unity engine for developing the game client of the system.

[0032] The structure of the present invention is simple and the operation is convenient, which solves the dependence of the swallowing training game system on devices such as sensors, and greatly improves the convenience and interest of swallowing training.

[0033] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0034] The present invention provides a computer vision-based swallowing training game execution system, which is applied to the process of active swallowing training (such as buccal movement, tongue movement, mandibular movement, etc.). Figure 1 As shown, in this embodiment, the computer vision-based swallowing training game execution system includes: a game client, a facial action recognition model, and a server.

[0035] Specifically, in this embodiment, the game client is developed using the Unity3D engine and includes an interaction module, a core module, and a system module. The interaction module includes a head and neck relaxation exercise module, a buccal movement module, a tongue movement module, a mandibular movement module, and a swallowing technique module, which are used to recognize user actions to drive the movement of game characters. The system module includes: a login module, which is responsible for managing relevant logics such as player login authentication; a level module, which is responsible for managing the loading and selection of levels with different difficulties; a friend module, which is responsible for managing friend addition, deletion, viewing, and interaction; a chat module, which is responsible for managing private chats between different players and group chat logics for multiple players; a leaderboard module, which is responsible for managing player data ranking information; and a points module, which is responsible for managing logics such as player points acquisition and usage. The core module includes: a character module, which manages the player's character state machine, behavior interaction, and dressing change capabilities; a camera module, which supports camera screen capture to provide source data for human recognition; a UI module, which is responsible for game interface display and interaction; a network module, which is responsible for communication between the client and the server; a physics module, which is responsible for handling collisions, event triggers, etc. between all objects in the game; an audio system, which is responsible for managing logics such as sound loading, playback, and stopping in the game; and a face recognition module, which is responsible for recognizing the camera screen in the game to drive character movement. The server is used for user login authentication, processing game logics, storing player account data, character dressing data, level completion data, friend data, chat data, player points data, and leaderboard data; the Unity engine is used for the development of the game client of the system.

[0036] Furthermore, the facial action recognition model is developed using Python and is based on two key technologies in computer vision technology, namely deep learning and convolutional neural networks. The specific steps are as follows:

[0037] Data collection and annotation: Collect a large amount of data with different human face postures and expressions, and annotate this data for use in training and testing the action model;

[0038] Data preprocessing: Preprocess the collected human face image data, including noise removal, image alignment, face detection, illumination and color correction, etc., to ensure the quality and consistency of the input data;

[0039] Feature extraction: Use appropriate feature extraction methods, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs) in deep learning, or traditional feature extraction algorithms, to extract features from the preprocessed human face images for action recognition;

[0040] Model training: Design and train a model for recognizing facial actions. Deep learning models can be used, or traditional machine learning algorithms can be combined to train the model to improve performance and accuracy;

[0041] Model verification and evaluation: Verify and evaluate the trained facial recognition action model to ensure the accuracy and robustness of the model;

[0042] Interface design and development: Design and develop the calling interface for the client to enable the direct calling and use of facial recognition on the client.

[0043] Furthermore, in this embodiment, the server includes a login server, a logic server, and a database. The client and the server use Socket communication, and the protocol uses Protobuf, while the database uses MySQL. The login server is used for users to perform login authentication and obtain game qualifications. The client inputs the account password or the mobile phone number verification code and then calls the login RPC. Then, the server performs authentication. If the authentication is passed, it returns to the client that the login is successful; otherwise, the login fails, and the corresponding prompt message is popped up. The logic server is mainly used to process game logic, etc. After the client completes a level, it communicates with the server through RPC. The server stores information such as the level completion information, duration information, score acquisition, and leaderboard in the database for the client to query. Players can also view the normal task data display by pulling the server data in the task panel. Other general logics also need to be developed and designed, such as adding and deleting friends of players, chatting communication between players, obtaining the leaderboard of players, configuring the costumes of characters, obtaining and using scores, and configuring and distributing game levels. The database mainly stores player account data, character costume data, level completion data, friend data, chat data, player score data, and leaderboard data, etc.

[0044] In this embodiment, the head and neck relaxation exercise module corresponds to a game named "Celebrating Spring with Lion Dances". The user corresponds to the game character of an acrobat, and the action is the "rice character exercise" for the head and neck. The specific operation method is that the acrobat controls the lion to jump over the stools by moving the head up, down, left, and right. The buccal movement module corresponds to a game named "Spring Sowing and Autumn Harvest". The user corresponds to the game character of a farmer, and the actions are closed-lip movement, lip-stretching movement, lip-puckering movement, lip-compressing movement, bilateral cheek retraction movement, and closed-lip cheek puffing movement. The specific operation method is that the farmer reaches the vegetable garden by puffing out the cheeks. When the bilateral cheeks retract, carrots are successfully planted. Closing the lips provides sunlight, stretching the lips for weeding, puckering the lips for pest control, and compressing the lips for fertilization. After completing the above series of actions and puffing out the cheeks again, the carrots will automatically fall into the basket. The tongue movement module corresponds to a game named "Searching for the Treasure Chest". The user corresponds to the game character of a treasure hunter, and the actions are tongue upward movement, tongue downward movement, tongue leftward movement, tongue rightward movement, tongue circular movement, and tongue clicking movement. The specific operation method is that the treasure hunter moves the tongue up, down, left, and right to reach the location of the treasure chest. After arriving, the tongue circular movement is completed to open the treasure chest, and the tongue clicking movement locks the treasure chest. The mandibular movement module corresponds to a game named "Regardless of Wind and Rain". The user corresponds to the game character of a food delivery person, and the actions are mandibular forward and backward movement, mandibular up and down movement, mandibular left and right movement, and mandibular grinding movement. The specific operation method is that the food delivery person moves the mandible forward, backward, left, and right to avoid obstacles on the road and deliver the food to the designated location. The swallowing technique module corresponds to a game named "Defending the City Wall". The user corresponds to the game character of a soldier, and the actions are supraglottic swallowing method, supra-supraglottic swallowing method, forceful swallowing method, and Mendelsohn swallowing method. The specific operation method is that the soldier is in battle. After completing the corresponding swallowing actions according to the prompts, bombs can be successfully dropped. Some bombs need to complete an involuntary cough before they can be detonated. The CTAR training module corresponds to a game named "Flying in the Blue Sky". The user corresponds to the game character of an athlete, and the action is to sit upright and place a rubber ball (the size of a fist) between the mandible and the manubrium sterni. The mandible presses the rubber ball forcefully against the manubrium sterni. The specific operation method is that the paraglider athlete controls the height of the paraglider by squeezing the rubber ball at the manubrium sterni with the chin. Pressing the rubber ball for 2 s will cause the paraglider to rise 2 m. The paraglider can stay in the air during the break in the middle.

[0045] Furthermore, in this embodiment, the technical link of the swallowing training game mainly transmits the frame images collected by the mobile phone camera to the face recognition module for dynamic recognition of facial feature points. The module returns the action identifier as the interactive input source. According to the matching of the game interaction configuration, the corresponding control of the in-game Avatar is performed, driving the character to move, jump, etc., and the game scene objects make corresponding interactive feedback. After all operations are completed, a single game settlement is performed and communication and storage are carried out with the server. The technical link of the gamified swallowing training system.

[0046] Second aspect, an embodiment of the present invention further provides a method for executing a swallowing training game based on computer vision, which is applied during active swallowing training (such as buccal movement, tongue movement, mandibular movement, etc.). For example, Figure 2 As shown, in this embodiment, the method for executing a swallowing training game based on computer vision includes the following steps:

[0047] S10. Develop a game client using the Unity3D engine, including an interaction module, a core module, and a system module; S20. Develop a facial action recognition model using Python based on deep learning and convolutional neural networks in computer vision technology; S30. Develop a login server, a logic server, and a database; S40. Transmit the frame images captured by the mobile phone camera to the face recognition module for dynamic recognition of facial feature points. The face recognition module returns an action identifier as an interactive input source, and controls the corresponding in-game Avatar according to the matching of the game interaction configuration, driving the game scene characters and objects to make corresponding interactive feedback. After all operations are completed, a single game settlement is performed and communication and storage are carried out with the server.

[0048] Compared with the prior art, the above-described system and method for executing a swallowing training game based on computer vision require, in addition to the equipment on the game side, a surface electromyogram sensor or a facial action recognition device. The placement position of the sensor is greatly restricted and can only be attached to the body surface position. For example, for the key parts of swallowing training such as tongue movement, the sensor cannot achieve induction, and the facial recognition device also limits the accessibility of the swallowing training game, resulting in a reduction in the convenience and practicality of swallowing training. The structure of the present invention is simple and the operation is convenient. By using computer vision technology to develop a facial action recognition model, it does not require a surface electromyogram sensor, a computer, and a facial action recognition device. Only a smart phone can be used to implement the swallowing training game, solving the dependence of the swallowing training game system on devices such as sensors, and having a high degree of simplicity and intelligence, greatly improving the convenience and interest of swallowing training.

[0049] Obviously, the above-described embodiments are only the preferred embodiments of the present invention, not all of the embodiments. The preferred embodiments of the present invention are shown in the drawings, but do not limit the patent scope of the present invention. The present invention can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present invention more thorough and comprehensive. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present invention in other related technical fields shall be within the scope of the patent protection of the present invention by the same token.

Claims

1. A swallowing training game execution system based on computer vision, characterized in that Including: An interaction module, which is used to identify user actions to drive the movement of game characters; A system module, which is used to manage player login authentication, loading and selection of levels with different difficulties, adding, deleting, viewing and interacting with friends, private chatting between different players and group chatting among multiple players, player data ranking information, obtaining and using of player points; A core module, which is used to manage the character state machine of players, behavior interaction and dressing change, camera screen capture, display and interaction of the game interface, communication between the client and the server, handling of collisions and event triggers between objects in the game, loading, playing and stopping of sounds in the game, and recognition of camera screens in the game; A server, which is used for users to perform login authentication, handle game logic, store player account data, character dressing data, level completion data, friend data, chat data, player point data and ranking list data; The Unity engine, which is used for the game client development of the system.

2. The execution system of a swallowing training game based on computer vision according to claim 1, characterized in that, The interaction module includes a head and neck relaxation exercise interaction module, a buccal movement interaction module, a tongue movement interaction module, a mandibular movement interaction module, a swallowing technique interaction module and a CTAR training interaction module.

3. The execution system of a swallowing training game based on computer vision according to claim 2, wherein, The system module includes a login module, a level module, a friend module, a chat module, a ranking list module and a point module.

4. A computer vision-based swallowing training game execution system according to claim 1, characterized in that The core module includes a character module, a camera module, a UI module, a network module, a physics module, an audio module and a face recognition module.

5. A computer vision-based swallowing training game execution system according to claim 1, wherein, The server and the client communicate using Socket, the protocol uses Protobuf, and the database uses MySQL.

6. A method applied to the system according to any one of claims 1 to 5, characterized in that, Including the following steps: S10. Develop a game client using the Unity3D engine, including an interaction module, a core module and a system module; S20. Develop a facial action recognition model using Python based on deep learning and convolutional neural networks in computer vision technology; S30. Develop the login server, logic server and database; S40. Transmit the captured mobile phone camera frame images to the face recognition module for dynamic recognition of facial feature points. The face recognition module returns an action identifier as the interaction input source. According to the matching of the game interaction configuration, the corresponding control of the in-game Avatar is performed, driving the game scene characters and objects to make corresponding interaction feedback. After all operations are completed, a single game session settlement is performed and communication and storage are carried out with the server.

Citation Information

Patent Citations

  • Swallowing training device

    CN110037695B