Team vision training system and method with augmented reality, speech and motion recognition

CN117427319BActive Publication Date: 2026-09-22NEUINX CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210819031.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-12
Publication Date
2026-09-22
Estimated Expiration
2042-07-12

AI Technical Summary

Technical Problem

[0003]以篮球赛事为例,球员需要注意的目标不仅仅是球,同时还需掌握球场上其余九位球员的一举一动,而大多数球员在持球且有防守者的状态下,往往只专注于单侧的近距离队友或防守者,忽略了远边真正有空档的队友或是远边埋伏的防守者,导致原先规划好的战术无法顺利执行甚至造成失误

Benefits of technology

[0043]由上述实施方式可知,本发明具有下列优点:其一,运用延展实境头盔结合语音互动与动作辨识技术,可有效协助运动员进行视野训练,让球员在瞬息万变的球场上,更容易掌握住队友的动向,进而帮助队伍得分获胜,以避免现有技术需仰赖实体球场上的多人反复练习而导致训练上所需花费的人力成本较高的问题。其二,让使用者穿戴延展实境头盔在模拟情境中进行第一人称的战术执行,并可搭配简易的动作捕捉系统(惯性感测器或影像感测器)记录使用者动作。当使用者观看模拟内容完成视野训练任务时,动作捕捉系统会即时辨识使用者动作,判断其是否可同步进行稳定的指定运球动作,进而训练使用者的运球稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117427319B_ABST
    Figure CN117427319B_ABST
Patent Text Reader

Abstract

The present application provides a team vision training system with augmented reality, voice and motion recognition, which is used to train the vision and motion of a user. A head-mounted display device plays a virtual task scenario image and senses a voice signal of the user to generate voice information. A motion capture device captures the motion of the user to generate motion information. An operation server generates the virtual task scenario image according to a scenario setting parameter set, generates a voice recognition result by recognizing the voice information according to a voice recognition program, judges the scenario setting parameter set and the voice recognition result to generate a vision training result, generates a motion recognition result by recognizing the motion information according to a motion recognition program, and judges the scenario setting parameter set and the motion recognition result to generate a motion training result. In this way, the user can be effectively assisted to perform vision and motion training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a team vision training system and method, and more particularly to a team vision training system and method with extended reality, voice and motion recognition capabilities. Background Technology

[0002] Athlete training encompasses various aspects, including skills, reaction time, tactics, and cognitive psychology. Team sports (such as basketball or football) place particular emphasis on tactics and teamwork.

[0003] Taking basketball as an example, players need to pay attention not only to the ball, but also to the actions of the other nine players on the court. However, most players, when in possession of the ball and under the presence of a defender, tend to focus only on their close teammates or defenders on one side, ignoring teammates with open spaces on the far side or defenders lurking there. This can lead to the failure to execute the originally planned tactics or even result in mistakes.

[0004] Therefore, it can be seen that there is currently a lack of systems and methods on the market that can assist athletes in visual training, and relevant researchers are seeking solutions. Summary of the Invention

[0005] Therefore, the purpose of this invention is to provide a team vision training system and method with extended reality, voice and motion recognition. It uses an extended reality helmet combined with voice interaction and motion recognition technology to effectively assist athletes in vision training, making it easier for players to grasp the movements of their teammates on the ever-changing court, thereby helping the team score and win. This avoids the problem of high manpower costs in training caused by the reliance on repeated practice by multiple people on a physical court in existing technologies.

[0006] According to one embodiment of the present invention, a team vision training system with extended reality, voice, and motion recognition is provided for training a user's vision and movements. The team vision training system with extended reality, voice, and motion recognition includes a head-mounted display device, a motion capture device, and a computing server. The head-mounted display device is disposed on the user and includes a task scenario playback module, a voice sensing module, and a gesture sensing module. The task scenario playback module plays virtual task scenario images. The voice sensing module senses the user's voice signals and generates voice information. The gesture sensing module senses the user's gestures and generates gesture sensing results. The motion capture device captures the user's movements and generates motion information. The computing server is signal-connected to the head-mounted display device and the motion capture device, and stores scenario setting parameter sets and receives motion and voice information. The computing server includes a task scenario generation module, a voice recognition module, and a motion recognition module. The task scenario generation module generates virtual task scenario images and task parameter sets according to the scenario setting parameter sets, and transmits the virtual task scenario images to the head-mounted display device for the user to view and generate voice signals and movements. The speech recognition module receives speech information and generates speech recognition results based on the speech recognition program. It also determines the visual field training results by considering the task parameters set from the task context generation module and the speech recognition results. The motion recognition module receives motion information and generates motion recognition results based on the motion recognition program. It also determines the motion training results by considering the context setting parameter set and the motion recognition results. The visual field training results and motion training results are used to determine whether the user has met the training requirements. The context setting parameter set includes task execution parameters and motion execution parameters. The task execution parameters include numerical summation items, which are selected based on gesture sensing results. The motion execution parameters include one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, with one of these four selected based on gesture sensing results. The motion capture device is an inertial sensor, which is installed on the user and senses the user's movements to generate motion information, which is then transmitted to the motion recognition module on the computing server. The motion recognition module is an inertial motion recognition module. It identifies motion information and generates motion recognition results. It then determines whether the motion recognition results are the same as or similar to one of the following motion parameters in the scenario setting parameter group: one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, thus generating a motion training result. The virtual task scenario image contains multiple virtual objects and multiple numbers.When a digital summation item is selected based on the gesture sensing results, the numbers are displayed around these virtual objects for the user to view and generate a voice signal. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the digital summation value of the digital summation item in the context setting parameter group, and generates a vision training result, where the digital summation value is equal to the sum of these numbers.

[0007] Therefore, the team vision training system with extended reality, voice and motion recognition of the present invention uses an extended reality helmet combined with voice interaction and motion recognition technology to assist athletes in vision training, making it easier for players to grasp the movements of their teammates on the ever-changing field.

[0008] Other embodiments of the aforementioned implementation are as follows: The scenario setting parameter group further includes player tactical parameters and defensive player generation parameters. The player tactical parameters include enabling and disabling tactical items, with one of these options selected based on gesture sensing results. The defensive player generation parameters include enabling and disabling defensive items, with one of these options selected based on gesture sensing results. The task execution parameters further include a color change option, which is selected based on gesture sensing results.

[0009] Other embodiments of the aforementioned implementation are as follows: The aforementioned virtual task scenario image further includes a first color and a second color, the first color and the second color being different. When a color change item is selected based on the gesture sensing result, numbers are displayed around these virtual objects respectively, and one of these virtual objects changes from the first color to the second color, so that a voice signal is generated after the user views it. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the color change number of the color change item in the scenario setting parameter group and generates a visual training result, wherein the color change number is equal to one of the numbers displayed around this virtual object.

[0010] Other embodiments of the aforementioned implementation are as follows: The aforementioned scenario setting parameter group further includes task difficulty adjustment parameters, and the computing server further includes a task difficulty adjustment module, which adjusts the selection of player tactical parameters (enabling and disabling tactical items), defensive player parameters (enabling and disabling defensive items), numerical summation and color change items of execution task parameters, and the selection of execution motion parameters (one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling).

[0011] According to one embodiment of the present invention, a team vision training method with extended reality, voice, and motion recognition is provided to train a user's vision and motion. The team vision training method with extended reality, voice, and motion recognition includes a virtual task scenario playback step, a voice recognition step, and a motion recognition step. The virtual task scenario playback step includes setting a head-mounted display device on the user, driving a task scenario generation module of a computing server to generate a virtual task scenario image and a task parameter set according to a scenario setting parameter set, transmitting the virtual task scenario image to the head-mounted display device, and then driving a task scenario playback module of the head-mounted display device to play the virtual task scenario image for the user to view and generate voice signals and motions. The voice recognition step includes driving a voice sensing module of the head-mounted display device to sense the user's voice signal and generate voice information, then driving a voice recognition module of the computing server to receive the voice information and recognize the voice information according to a voice recognition program to generate a voice recognition result, and judging the task parameter set of the task scenario generation module and the voice recognition result to generate a vision training result. Furthermore, the motion recognition step includes driving a motion capture device to capture the user's movements to generate motion information, then driving a motion recognition module to receive the motion information and identify the motion information according to the motion recognition program to generate motion recognition results, and judging the context setting parameter group and motion recognition results to generate exercise training results. Visual field training results and exercise training results are used to determine whether the user has met the training requirements. The head-mounted display device also includes a gesture sensing module, which is used to sense the user's gestures and generate gesture sensing results. The context setting parameter group includes task execution parameters and exercise execution parameters. The task execution parameters include numerical summation items, which are selected based on the gesture sensing results. The exercise execution parameters include one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, with one of these four items selected based on the gesture sensing results. In the motion recognition step, the motion capture device is an inertial sensor. The inertial sensor is installed on the user and senses the user's movements to generate motion information, which is then transmitted to the motion recognition module of the computing server. The motion recognition module is an inertial motion recognition module. This module identifies the motion information, generates motion recognition results, and determines whether the motion recognition results are the same as or similar to one of the following: one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, based on the scenario setting parameter group. This results in a motion training result. The virtual task scenario image includes multiple virtual objects and multiple numbers.When a digital summation item is selected based on the gesture sensing results, the numbers are displayed around these virtual objects for the user to view and generate a voice signal. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the digital summation value of the digital summation item in the context setting parameter group, and generates a vision training result, where the digital summation value is equal to the sum of these numbers.

[0012] Therefore, the team vision training method with extended reality, voice and motion recognition of the present invention uses an extended reality helmet combined with voice interaction and motion recognition technology to assist athletes in vision training, making it easier for players to grasp the movements of their teammates on the ever-changing field.

[0013] Other embodiments of the aforementioned implementation are as follows: The scenario setting parameter group further includes player tactical parameters and defensive player generation parameters. The player tactical parameters include enabling and disabling tactical items, with one of these options selected based on gesture sensing results. The defensive player generation parameters include enabling and disabling defensive items, with one of these options selected based on gesture sensing results. The task execution parameters further include a color change option, which is selected based on gesture sensing results.

[0014] Other embodiments of the aforementioned implementation are as follows: The aforementioned virtual task scenario image further includes a first color and a second color, the first color and the second color being different. When a color change item is selected based on the gesture sensing result, numbers are displayed around these virtual objects respectively, and one of these virtual objects changes from the first color to the second color, so that a voice signal is generated after the user views it. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the color change number of the color change item in the scenario setting parameter group and generates a visual training result, wherein the color change number is equal to one of the numbers displayed around this virtual object.

[0015] Other embodiments of the aforementioned implementation are as follows: The aforementioned scenario setting parameter group further includes task difficulty adjustment parameters, and the virtual task scenario playback step further includes the task difficulty adjustment module of the driving computing server adjusting the player's tactical parameters according to the task difficulty adjustment parameters, such as enabling and disabling tactical items, enabling and disabling defensive items, and selecting the numerical summation and color change items of the execution task parameters, as well as the selection of the execution motion parameters such as one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling.

[0016] According to one embodiment of the present invention, a team vision training system with extended reality, voice, and motion recognition is provided for training a user's vision and movements. The team vision training system with extended reality, voice, and motion recognition includes a head-mounted display device, a motion capture device, and a computing server. The head-mounted display device is disposed on the user and includes a task scenario playback module, a voice sensing module, and a gesture sensing module. The task scenario playback module plays virtual task scenario images. The voice sensing module senses the user's voice signals and generates voice information. The gesture sensing module senses the user's gestures and generates gesture sensing results. The motion capture device captures the user's movements and generates motion information. The computing server is signal-connected to the head-mounted display device and the motion capture device, and stores scenario setting parameter sets and receives motion and voice information. The computing server includes a task scenario generation module, a voice recognition module, and a motion recognition module. The task scenario generation module generates virtual task scenario images and task parameter sets according to the scenario setting parameter sets, and transmits the virtual task scenario images to the head-mounted display device for the user to view and generate voice signals and movements. The speech recognition module receives speech information and generates speech recognition results based on the speech recognition program. It also determines the visual field training results by considering the task parameters set from the task context and the speech recognition results. The motion recognition module receives motion information and generates motion recognition results based on the motion recognition program. It further determines the motion training results by considering the context setting parameters set and the motion recognition results. The visual field training results and motion training results are used to determine whether the user has met the training requirements. The context setting parameters set includes task execution parameters and motion execution parameters. The task execution parameters include numerical summation items, which are selected based on gesture sensing results. The motion execution parameters include one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, with one of these four selected based on gesture sensing results. The motion capture device is an image sensor, which includes a camera corresponding to the user. The image sensor captures the user's movements through the camera to generate motion information, and transmits the motion information to the motion recognition module of the computing server. The motion recognition module is an image motion recognition module. The image motion recognition module identifies the motion information to generate motion recognition results, and determines whether the motion recognition results are the same as or similar to one of the execution motion parameters in the scenario setting parameter group, namely, one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, and generates a motion training result. The virtual task scenario image includes multiple virtual objects and multiple numbers.When a digital summation item is selected based on the gesture sensing results, the numbers are displayed around these virtual objects for the user to view and generate a voice signal. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the digital summation value of the digital summation item in the context setting parameter group, and generates a vision training result, where the digital summation value is equal to the sum of these numbers.

[0017] According to one embodiment of the present invention, a team vision training system with extended reality, voice, and motion recognition is provided for training a user's vision and movements. The team vision training system with extended reality, voice, and motion recognition includes a head-mounted display, a motion capture device, and a computing server. The head-mounted display is disposed on the user and includes a task scenario playback module, a voice sensing module, and a gesture sensing module. The task scenario playback module plays virtual task scenario images. The voice sensing module senses the user's voice signal and generates voice information. The gesture sensing module senses the user's gestures and generates gesture sensing results. The motion capture device captures the user's movements and generates motion information, and the motion capture device includes an inertial sensor and an image sensor. The inertial sensor is disposed on the user and senses the user's movements to generate inertial motion information, which is then transmitted to the motion recognition module of the computing server. The image sensor includes a camera corresponding to the user; the image sensor captures the user's movements through the camera to generate image motion information, which is then transmitted to the motion recognition module of the computing server. The computing server connects the head-mounted display and the motion capture device. It stores scenario setting parameters and receives motion and voice information. The computing server includes a task scenario generation module, a voice recognition module, and a motion recognition module. The task scenario generation module generates a virtual task scenario image and a set of task parameters based on the scenario setting parameters and transmits the virtual task scenario image to the head-mounted display for the user to view, generating voice signals and movements. The voice recognition module receives voice information and identifies it using a voice recognition program to generate a voice recognition result. It also determines the visual field training result by analyzing the task parameter set from the task scenario generation module and the voice recognition result. The motion recognition module receives motion information and identifies it using a motion recognition program to generate a motion recognition result. It also determines the motion training result by analyzing the scenario setting parameters and the motion recognition result. Motion information includes inertial motion information and image motion information. The visual field training result and motion training result are used to determine whether the user has met the training requirements. The scenario setting parameter group includes task execution parameters and motion execution parameters. The task execution parameters include numerical summation items, which are selected based on gesture sensing results. The motion execution parameters include one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, with one of these four selected based on gesture sensing results. The virtual task scenario image includes multiple virtual objects and multiple numbers.When a digital summation item is selected based on the gesture sensing results, the numbers are displayed around these virtual objects for the user to view and generate a voice signal. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the digital summation value of the digital summation item in the context setting parameter group, and generates a vision training result, where the digital summation value is equal to the sum of these numbers.

[0018] Other embodiments of the aforementioned implementation are as follows: The aforementioned motion recognition module includes an inertial motion recognition module and an image motion recognition module. The inertial motion recognition module identifies inertial motion information to generate an inertial motion recognition result, and determines whether the inertial motion recognition result is the same as or similar to one of the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling items in the scenario setting parameter group, thereby generating a first motion training result. The image motion recognition module identifies motion information to generate a motion recognition result, and determines whether the motion recognition result is the same as or similar to one of the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling items in the scenario setting parameter group, thereby generating a second motion training result. The motion training result includes the first motion training result and the second motion training result.

[0019] According to one embodiment of the present invention, a team vision training method with extended reality, voice, and motion recognition is provided to train a user's vision and motion. The team vision training method with extended reality, voice, and motion recognition includes a virtual task scenario playback step, a voice recognition step, and a motion recognition step. The virtual task scenario playback step includes setting a head-mounted display device on the user, driving a task scenario generation module of a computing server to generate a virtual task scenario image and a task parameter set according to a scenario setting parameter set, transmitting the virtual task scenario image to the head-mounted display device, and then driving a task scenario playback module of the head-mounted display device to play the virtual task scenario image for the user to view and generate voice signals and motions. The voice recognition step includes driving a voice sensing module of the head-mounted display device to sense the user's voice signal and generate voice information, then driving a voice recognition module of the computing server to receive the voice information and recognize the voice information according to a voice recognition program to generate a voice recognition result, and judging the task parameter set of the task scenario generation module and the voice recognition result to generate a vision training result. Furthermore, the motion recognition step includes driving a motion capture device to capture the user's movements to generate motion information, then driving a motion recognition module to receive the motion information and identify the motion information according to the motion recognition program to generate motion recognition results, and judging the context setting parameter group and motion recognition results to generate exercise training results. Visual field training results and exercise training results are used to determine whether the user has met the training requirements. The head-mounted display device also includes a gesture sensing module, which is used to sense the user's gestures and generate gesture sensing results. The context setting parameter group includes task execution parameters and exercise execution parameters. The task execution parameters include numerical summation items, which are selected based on the gesture sensing results. The exercise execution parameters include one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, with one of these four items selected based on the gesture sensing results. In the motion recognition step, the motion capture device is an image sensor, which includes a camera corresponding to the user. The image sensor captures the user's movements through the camera to generate motion information and transmits the motion information to the motion recognition module of the computing server. The motion recognition module is an image motion recognition module, which identifies the motion information to generate motion recognition results and determines whether the motion recognition results are the same as or similar to one of the execution motion parameters in the scenario setting parameter group, namely, one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, thus generating a motion training result. The virtual task scenario image includes multiple virtual objects and multiple numbers.When a digital summation item is selected based on the gesture sensing results, the numbers are displayed around these virtual objects for the user to view and generate a voice signal. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the digital summation value of the digital summation item in the context setting parameter group, and generates a vision training result, where the digital summation value is equal to the sum of these numbers.

[0020] According to one embodiment of the present invention, a team vision training method with extended reality, voice, and motion recognition is provided to train a user's vision and motion. The team vision training method with extended reality, voice, and motion recognition includes a virtual task scenario playback step, a voice recognition step, and a motion recognition step. The virtual task scenario playback step includes setting a head-mounted display device on the user, driving a task scenario generation module of a computing server to generate a virtual task scenario image and a task parameter set according to a scenario setting parameter set, transmitting the virtual task scenario image to the head-mounted display device, and then driving a task scenario playback module of the head-mounted display device to play the virtual task scenario image for the user to view and generate voice signals and motions. The voice recognition step includes driving a voice sensing module of the head-mounted display device to sense the user's voice signal and generate voice information, then driving a voice recognition module of the computing server to receive the voice information and recognize the voice information according to a voice recognition program to generate a voice recognition result, and judging the task parameter set of the task scenario generation module and the voice recognition result to generate a vision training result. Furthermore, the motion recognition step includes driving a motion capture device to capture the user's movements to generate motion information, then driving a motion recognition module to receive the motion information and identify the motion information according to the motion recognition program to generate motion recognition results, and judging the context setting parameter group and motion recognition results to generate exercise training results. Visual field training results and exercise training results are used to determine whether the user has met the training requirements. The head-mounted display device also includes a gesture sensing module, which is used to sense the user's gestures and generate gesture sensing results. The context setting parameter group includes task execution parameters and exercise execution parameters. The task execution parameters include numerical summation items, which are selected based on the gesture sensing results. The exercise execution parameters include one-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling, with one of these four items selected based on the gesture sensing results. In the motion recognition step, the motion capture device includes an inertial sensor and an image sensor. The inertial sensor is positioned on the user and senses the user's movements to generate inertial motion information, which is then transmitted to the motion recognition module of the computing server. The image sensor includes a camera corresponding to the user; it captures the user's movements through the camera to generate image motion information, which is also transmitted to the motion recognition module of the computing server. The virtual task scenario image includes multiple virtual objects and multiple numbers. When a number summation item is selected based on the gesture sensing result, these numbers are displayed around these virtual objects for the user to view and generate a voice signal. The voice recognition module determines whether the voice recognition result of the corresponding voice signal is the same as the sum of the numbers in the scenario setting parameter group, thus generating a visual training result, where the sum of the numbers equals the sum of all the numbers. Motion information includes inertial motion information and image motion information.

[0021] Other embodiments of the aforementioned implementation are as follows: In the aforementioned motion recognition step, the motion recognition module includes an inertial motion recognition module and an image motion recognition module. The inertial motion recognition module identifies inertial motion information to generate an inertial motion recognition result, and determines whether the inertial motion recognition result is the same as or similar to one of the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling items in the scenario setting parameter group, thereby generating a first motion training result. The image motion recognition module identifies motion information to generate a motion recognition result, and determines whether the motion recognition result is the same as or similar to one of the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling items in the scenario setting parameter group, thereby generating a second motion training result. The motion training result includes the first motion training result and the second motion training result. Attached Figure Description

[0022] Figure 1 This is a block diagram illustrating a team vision training system with extended reality, voice and motion recognition according to a first embodiment of the present invention. Figure 2 This is a schematic diagram illustrating a team vision training system with extended reality, voice and motion recognition according to a second embodiment of the present invention; Figure 3 It is a drawing Figure 2 A block diagram of a team vision training system with extended reality, voice and motion recognition; Figure 4 It is a drawing Figure 2 A schematic diagram of the scenario setting parameter group stored in the computing server; Figure 5 It is a drawing Figure 2 A schematic diagram of an embodiment of a virtual task context image of a head-mounted display device; Figure 6 It is a drawing Figure 2 A schematic diagram of another embodiment of virtual task context imagery for a head-mounted display device; Figure 7 It is a drawing Figure 2 A schematic diagram of the image captured by the image sensor and the movement trajectory of the sphere in the image identified by the computing server. Figure 8 This is a flowchart illustrating a team vision training method with extended reality, voice and motion recognition according to a third embodiment of the present invention. Figure 9 This is a flowchart illustrating a team vision training method with extended reality, voice and motion recognition according to a fourth embodiment of the present invention. Figure 10This is a block diagram illustrating a team vision training system with extended reality, voice, and motion recognition according to a fifth embodiment of the present invention; and Figure 11 This is a schematic diagram illustrating a team vision training system with extended reality, voice and motion recognition according to a sixth embodiment of the present invention. The reference numerals in the attached figures are explained as follows: 100, 100a, 100b, 100c: Team vision training system with extended reality, voice and motion recognition. 110: User 120: sphere 122: Movement trajectory 200: Head-mounted display device 202, 204: Virtual mission scenario images 2022, 2024, 2026, 2028, 2042, 2044, 2046, 2048: Virtual Objects 210: Task Context Playback Module 220: Voice sensing module 230: Gesture sensing module 300, 300a: Motion capture device 310: Inertial Sensor 320: Image Sensor 322: Images 400, 400a, 400b, 400c: Computing servers 402: Scenario Setting Parameter Group 4021: Player Tactical Parameters 4021a: Activate Tactical Project 4021b: Close tactical projects 4022: Parameters generated by the defending player 4022a: Activate defensive features 4022b: Disable defensive items 4023: Execution Task Parameters 4023a: Number Summation Project 4023b: Color Change Project 4024: Execution of motion parameters 4024a: One-handed dribbling event 4024b: Crossover dribbling event 4024c: Crossover dribbling event 4024d: Behind-the-back dribbling event 4025: Task difficulty adjustment parameter 410: Task Context Generation Module 420: Speech Recognition Module 430, 430a: Action recognition module 432: Inertial Motion Recognition Module 434: Image Motion Recognition Module 440: Task Difficulty Adjustment Module 500, 600: Team Vision Training Methods with Extended Reality, Voice and Motion Recognition S02: Virtual Task Context Playback Steps S04: Speech Recognition Steps S06: Action Recognition Steps S12, S14, S16, S18: Steps Detailed Implementation

[0023] Several embodiments of the present invention will now be described with reference to the accompanying drawings. For clarity, many practical details will be set forth in the following description. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential. Furthermore, for the sake of simplicity in the drawings, some conventionally used structures and elements will be illustrated in a simple schematic manner; and repeated elements may be denoted by the same reference numerals.

[0024] Furthermore, in this document, when a component (or unit or module, etc.) is "connected" to another component, it can mean that the component is directly connected to the other component, or that the component is indirectly connected to the other component, meaning that there is another component between the component and the other component. Only when it is explicitly stated that a component is "directly connected" to another component does it indicate that there is no other component between the component and the other component. The terms "first," "second," and "third" are only used to describe different components and do not limit the components themselves; therefore, "first component" can also be referred to as "second component." Moreover, the combinations of components / units / circuits in this document are not combinations generally known, conventional, or existing in this field. Whether the component / unit / circuit itself is existing cannot be used to determine whether its combination relationship is easily accomplished by someone of ordinary skill in the art.

[0025] Please see Figure 1 , Figure 1This is a block diagram illustrating a team vision training system 100 with extended reality, voice, and motion recognition according to a first embodiment of the present invention. The team vision training system 100 with extended reality, voice, and motion recognition is used to train a user's vision and movements, and includes a head-mounted display device 200, a motion capture device 300, and a computing server 400. The head-mounted display device 200 is disposed on the user and includes a task scenario playback module 210 and a voice sensing module 220, wherein the task scenario playback module 210 is used to play virtual task scenario images. The voice sensing module 220 senses the user's voice signal and generates voice information. The motion capture device 300 captures the user's movements and generates motion information. Furthermore, the computing server 400 is signal-connected to the head-mounted display device 200 and the motion capture device 300. The computing server 400 stores scenario setting parameter sets and receives motion and voice information. The computing server 400 includes a task scenario generation module 410, a voice recognition module 420, and a motion recognition module 430. The task scenario generation module 410 generates a virtual task scenario image and a task parameter set based on the scenario setting parameter set, and transmits the virtual task scenario image to the head-mounted display device 200 for the user to view and generate voice signals and actions. The voice recognition module 420 receives voice information and recognizes the voice information according to the voice recognition program to generate a voice recognition result. The voice recognition module 420 also judges the task parameter set of the task scenario generation module 410 and the voice recognition result to generate a visual field training result. The motion recognition module 430 receives motion information and recognizes the motion information according to the motion recognition program to generate a motion recognition result. The motion recognition module 430 also judges the scenario setting parameter set and the motion recognition result to generate a motion training result. The visual field training result and the motion training result are used to determine whether the user has met the training requirements. Therefore, the team vision training system 100 with extended reality, voice, and motion recognition of the present invention utilizes an extended reality helmet combined with voice interaction and motion recognition technology to effectively assist athletes in vision training, allowing players to more easily grasp the movements of their teammates on the ever-changing court, thereby helping the team score and win. Furthermore, the present invention allows for individual training, avoiding the problem of high labor costs associated with existing technologies that rely on repeated practice by multiple people on a physical court. The following detailed embodiments illustrate the details of the above-mentioned devices.

[0026] Please refer to the following: Figure 1 , Figure 2 , Figure 3 and Figure 4 ,in Figure 2 This is a schematic diagram illustrating a team vision training system 100a with extended reality, voice and motion recognition according to a second embodiment of the present invention; Figure 3 It is a drawing Figure 2 A block diagram of the team vision training system 100a with extended reality, voice and motion recognition; and Figure 4 It is a drawing Figure 2 A schematic diagram of the scenario setting parameter group 402 stored in the computing server 400a. As shown in the figure, the team vision training system 100a with extended reality, voice and motion recognition is used to train the vision and motion of the user 110, and includes a head-mounted display device 200, a motion capture device 300a and a computing server 400a.

[0027] A head-mounted display device 200 is disposed on the user 110 and includes a task scenario playback module 210, a voice sensing module 220, and a gesture sensing module 230. The task scenario playback module 210 is used to play virtual task scenario images. The voice sensing module 220 senses the user 110's voice signal to generate voice information. The gesture sensing module 230 is used to sense the user 110's gestures to generate gesture sensing results. In one embodiment, the head-mounted display device 200 may be a mixed reality (MR) headset or a virtual reality (VR) headset, and may be worn on the user 110's head and transmit relevant information (including virtual task scenario images transmitted to the task scenario playback module 210 by the computing server 400a) wirelessly (e.g., via wireless network or Bluetooth) or via wired means. The task scenario playback module 210 can be a screen, the voice sensing module 220 can be a microphone, and the gesture sensing module 230 can be a camera. After the user 110 puts on the head-mounted display device 200, the eyes can view MR or VR images (i.e., virtual task scenario images) corresponding to the screen, and the mouth can pick up sound for subsequent processing corresponding to the microphone, but the present invention is not limited thereto.

[0028] Motion capture device 300a is used to capture the movements of user 110 and generate motion information. Specifically, motion capture device 300a includes an inertial sensor 310 and an image sensor 320. The inertial sensor 310 is mounted on user 110 and senses user 110's movements to generate inertial motion information, which is then transmitted to the motion recognition module 430a of computing server 400a. For example, when user 110 is dribbling a ball and the inertial sensor 310 is worn on user 110's hand, the inertial sensor 310 captures the dribbling motion of user 110's hand, and the generated motion information is equivalent to the ball's movement information; in other words, during dribbling, when the ball contacts the hand, the hand's movement trajectory is equal to the ball's movement trajectory. Furthermore, the image sensor 320 includes a camera corresponding to user 110. The image sensor 320 captures user 110's movements through the camera to generate image motion information, which is then transmitted to the motion recognition module 430a of computing server 400a. Motion information includes inertial motion information and image motion information. The image sensor 320 can be a camera or a mobile phone. It is also worth mentioning that if the team sport is basketball, the inertial sensor 310 is worn on the user 110's hand; if the team sport is soccer, the inertial sensor 310 is worn on the user 110's foot, depending on the training needs.

[0029] The computing server 400a is connected to the head-mounted display device 200 and the motion capture device 300a. The computing server 400a stores the scenario setting parameter group 402 and receives motion and voice information. The scenario setting parameter group 402 includes player tactical parameters 4021, defensive player generation parameters 4022, task execution parameters 4023, motion execution parameters 4024, and task difficulty adjustment parameters 4025. The player tactical parameters 4021 include enabling tactical items 4021a and disabling tactical items 4021b. Enabling tactical items 4021a and disabling tactical items 4021b is selected based on gesture sensing results. Enabling tactical items 4021a means the virtual player will move within the virtual task scenario image; disabling tactical items 4021b means the virtual player will remain stationary. Furthermore, the defensive player generation parameter 4022 includes enabling defensive items 4022a and disabling defensive items 4022b. Enabling defensive items 4022a and disabling defensive items 4022b are selected based on gesture sensing results. The defending team is the opposing team. Enabling defensive items 4022a means that virtual defensive players will appear in the virtual mission scenario image; that is, the virtual mission scenario image will simultaneously display multiple virtual teammates and multiple virtual defensive players. For example, in basketball, when enabling defensive items 4022a is selected, the virtual mission scenario image will display 4 virtual teammates and 5 virtual defensive players. Disabling defensive items 4022b means that the virtual mission scenario image will only display virtual teammates and no virtual defensive players.

[0030] The execution task parameter 4023 includes a number summation item 4023a and a color change item 4023b, one of which is selected based on the gesture sensing result. The number summation item 4023a represents the display of numbers around (e.g., above the heads) of virtual objects (such as virtual teammates) in the virtual task scenario image, which generates a voice signal after being viewed by the user 110. The color change item 4023b represents the display of numbers around (e.g., above the heads) of virtual objects (such as virtual teammates), and the color of the clothing of one of the virtual objects changes from a first color to a second color, which also generates a voice signal after being viewed by the user 110. Additionally, the execution motion parameter 4024 includes a one-handed dribbling item 4024a, a crossover dribbling item 4024b, a between-the-legs dribbling item 4024c, and a behind-the-back dribbling item 4024d. One of the following dribbling actions—one-handed dribbling, one-handed crossover dribbling, one-legged crossover dribbling, and one-behind-the-back dribbling—is selected based on the gesture sensing results. One-handed dribbling 4024a indicates that user 110 should perform a one-handed dribbling action; one-handed crossover dribbling 4024b indicates that user 110 should perform a one-handed crossover dribbling action; one-legged crossover dribbling 4024c indicates that user 110 should perform a crossover dribbling action; and one-behind-the-back dribbling 4024d indicates that user 110 should perform a behind-the-back dribbling action, for the purpose of judging subsequent dribbling posture and dribbling stability.

[0031] The task difficulty adjustment parameter 4025 represents the adjustment parameters to control the difficulty of the task. Adjustable parameters include the aforementioned player tactical parameters 4021, defensive player generation parameters 4022, task execution parameters 4023, movement action parameters 4024, and the virtual player's movement speed or voice interaction time limit, but this invention is not limited thereto. As described above, the scenario setting parameter group 402 can be displayed in the virtual task scenario image. Combined with virtual reality and action selection, it allows the user 110 to select the desired scenario parameters in the virtual task scenario image. In one embodiment, the virtual task scenario image changes the selection box color and checked content according to the position of the user 110's virtual hand, thereby completing the selection of scenario parameters. Furthermore, the task difficulty can be set by the coach. For example, the coach uses a specific device (such as an MR / VR headset, mobile device, or tablet) to set the task difficulty. The specific device and the computing server 400a can transmit the relevant parameters corresponding to the task difficulty wirelessly or via wired connection.

[0032] The computing server 400a includes a task context generation module 410, a speech recognition module 420, a motion recognition module 430a, and a task difficulty adjustment module 440. The task context generation module 410 generates a virtual task context image and a task parameter set according to the context setting parameter set 402, and transmits the virtual task context image to the head-mounted display device 200 for the user 110 to view and generate speech signals and actions. The speech recognition module 420 receives speech information and recognizes the speech information according to a speech recognition program to generate a speech recognition result. The speech recognition module 420 also judges the task parameter set and speech recognition result from the task context generation module 410 to generate a visual field training result. In one embodiment, the speech recognition program may be Microsoft's speech recognition software (Azure Cognitive Service), but this invention is not limited thereto.

[0033] The motion recognition module 430a receives motion information and identifies the motion information according to the motion recognition program to generate motion recognition results. The motion recognition module 430a also determines the exercise training results by comparing the scenario setting parameter group 402 with the motion recognition results. The visual training results and the exercise training results are used to determine whether the user 110 has met the training requirements. The motion recognition program is implemented using computer vision, signal processing, and artificial intelligence technologies. Specifically, the motion recognition module 430a includes an inertial motion recognition module 432 and an image motion recognition module 434. The inertial motion recognition module 432 identifies inertial motion information to generate inertial motion recognition results and determines whether the inertial motion recognition results are the same as or similar to one of the following (i.e., the item selected by the user 110): one-handed dribbling 4024a, crossover dribbling 4024b, between-the-legs dribbling 4024c, and behind-the-back dribbling 4024d, in the scenario setting parameter group 402, and generates a first exercise training result. Furthermore, the image motion recognition module 434 identifies motion information to generate motion recognition results, and determines whether the motion recognition results are the same as or similar to the single-handed dribbling item 4024a, crossover dribbling item 4024b, between-the-legs dribbling item 4024c, and behind-the-back dribbling item 4024d in the execution motion parameters 4024 of the scenario setting parameter group 402, thereby generating a second motion training result. The motion training result includes the first motion training result and the second motion training result. This invention can effectively improve the accuracy of recognition through dual recognition of inertial motion and image motion.

[0034] The task difficulty adjustment module 440 adjusts the selection of player tactical parameters 4021 (activating tactical items 4021a and deactivating tactical items 4021b), defensive player generation parameters 4022 (activating defensive items 4022a and deactivating defensive items 4022b), execution task parameters 4023 (numerical summation items 4023a and color change items 4023b), and execution movement parameters 4024 (one-handed dribbling item 4024a, crossover dribbling item 4024b, between-the-legs dribbling item 4024c, and behind-the-back dribbling item 4024d) based on the task difficulty adjustment parameter 4025, so as to execute tasks of different difficulty levels. For example, high-difficulty tasks can be matched by enabling tactical item 4021a, enabling defensive item 4022a, numerical summation item 4023a and / or behind-the-back dribbling item 4024d; low-difficulty tasks can be matched by disabling tactical item 4021b, disabling defensive item 4022b, color change item 4023b and / or one-handed dribbling item 4024a.

[0035] The computing server 400a includes a memory and a high-performance image processing processor. The memory can store scenario setting parameter sets 402, multiple virtual motion scenes, voice recognition programs, and motion recognition programs. The high-performance image processing processor is used to process MR or VR images (i.e., virtual task scenario images) in real time, such as a central processing unit (CPU) or a graphics processing unit (GPU). The computing server 400a can be a computer, mobile device, or other high-speed electronic computing device, but the present invention is not limited thereto. Therefore, the team vision training system 100a of the present invention, which combines extended reality, voice, and motion recognition, effectively assists athletes in vision training by using an extended reality helmet combined with voice interaction and motion recognition technology. This allows players to more easily grasp the movements of their teammates on the ever-changing court, thereby helping the team score and win. Furthermore, the present invention allows for individual training, avoiding the problem of high labor costs associated with existing technologies that rely on repeated practice by multiple people on a physical court.

[0036] Please refer to the following: Figure 2 , Figure 3 , Figure 4 and Figure 5 ,in Figure 5 It is a drawing Figure 2This is a schematic diagram of an embodiment of a virtual task context image 202 of a head-mounted display device 200. As shown, the virtual task context image 202 includes multiple virtual objects 2022, 2024, 2026, 2028 and multiple numbers. When a number summation item 4023a is selected based on a gesture sensing result, these numbers are displayed around the virtual objects 2022, 2024, 2026, and 2028 respectively, so that the user 110 can view them and generate a voice signal. The voice recognition module 420 determines whether the voice recognition result of the corresponding voice signal is the same as the sum of the numbers in the task parameter group of the task context generation module 410 and generates a visual training result, wherein the sum of the numbers is equal to the sum of these numbers. For example, taking basketball as an example, the virtual objects 2022, 2024, 2026, and 2028 are virtual teammates, and the numbers above their heads are 5, 7, 7, and 8 respectively, and the sum of these numbers is equal to 27. When the speech recognition module 420 determines that the speech recognition result is the same as the sum of the numbers, the visual training result is "the sum of the numbers read by the user is the correct answer", and it is determined that the user 110 has met the training requirements (this belongs to the cognitive training of visual training); when the speech recognition module 420 determines that the speech recognition result is different from the sum of the numbers, the visual training result is "the sum of the numbers read by the user is not the correct answer", and it is determined that the user 110 has not met the training requirements.

[0037] Please refer to the following: Figure 2 , Figure 3 , Figure 4 and Figure 6 ,in Figure 6 It is a drawing Figure 2This is a schematic diagram of another embodiment of the virtual task context image 204 of the head-mounted display device 200. As shown, the virtual task context image 204 includes a plurality of virtual objects 2042, 2044, 2046, 2048, a plurality of numbers, a first color, and a second color. The first color and the second color are different. When a color change item 4023b is selected based on a gesture sensing result, the numbers are displayed around the virtual objects 2042, 2044, 2046, 2048, and one of the virtual objects 2042, 2044, 2046, 2048 changes from the first color to the second color, so that the user 110 can view it and generate a voice signal. The speech recognition module 420 determines whether the speech recognition result of the corresponding speech signal is the same as a color-changing number in the task parameter group of the task context generation module 410, and generates a vision training result, wherein the color-changing number is equal to one of the numbers displayed around these virtual objects 2042, 2044, 2046, and 2048. For example, taking a basketball as an example, virtual objects 2042, 2044, 2046, and 2048 are virtual teammates, and the numbers above their heads are 5, 7, 5, and 2, respectively. When the speech recognition module 420 determines that the speech recognition result is the same as the color-changing number (the teammate whose clothes change color is a virtual object 2042, whose color-changing number is 5), the vision training result is "the color-changing number recited by the user is the correct answer", and it is determined that the user 110 has met the training requirements (this belongs to the reaction training of vision training); when the speech recognition module 420 determines that the speech recognition result is different from the color-changing number, the vision training result is "the color-changing number recited by the user is not the correct answer", and it is determined that the user 110 has not met the training requirements.

[0038] Please refer to the following: Figure 2 , Figure 3 , Figure 4 and Figure 7 ,in Figure 7 It is a drawing Figure 2The diagram illustrates the image 322 captured by the image sensor 320 and the movement trajectory 122 of the ball 120 identified by the computing server 400a in the image 322. As shown, the image 322 captured by the image sensor 320 is transmitted to the computing server 400a for identification. The image motion recognition module 434 identifies the motion information of the user 110 in the image 322 and generates a motion recognition result. It then determines whether the motion recognition result is the same as or similar to one of the following: single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling in the execution motion parameters 4024 of the scenario setting parameter group 402, and generates a second motion training result to analyze the dribbling posture of the user 110. In addition, the image motion recognition module 434 identifies the ball 120 in the image 322 and obtains the movement trajectory 122 of the ball 120 to analyze the stability of the user 110's dribbling. The stability is positively correlated with the waveform frequency of the movement trajectory 122. For example, taking basketball as an example, user 110 dribbles with one hand. When the image motion recognition module 434 determines that the motion recognition result is the same as the one-handed dribbling item 4024a, and the waveform frequency of the movement trajectory 122 is within a preset range, the second exercise training result is "the user's dribbling posture is correct" and "the dribbling stability is high", and it is determined that user 110 has met the training requirements; when the image motion recognition module 434 determines that the motion recognition result is different from the one-handed dribbling item 4024a, and the waveform frequency of the movement trajectory 122 exceeds a preset range, the second exercise training result is "the user's dribbling posture is incorrect" and "the dribbling stability is low", and it is determined that user 110 has not met the training requirements.

[0039] Please refer to the following: Figure 2 and Figure 8 ,in Figure 8 This is a flowchart illustrating a team vision training method 500 with extended reality, voice, and motion recognition according to a third embodiment of the present invention. As shown in the figure, the team vision training method 500 with extended reality, voice, and motion recognition is applied to... Figure 1A team vision training system 100 with extended reality, voice, and motion recognition is used to train the vision and motion of user 110, and includes a virtual task scenario playback step S02, a voice recognition step S04, and a motion recognition step S06. The virtual task scenario playback step S02 includes setting the head-mounted display device 200 on user 110, driving the task scenario generation module 410 of the computing server 400 to generate a virtual task scenario image and a task parameter set according to a scenario setting parameter set, and transmitting the virtual task scenario image to the head-mounted display device 200. Then, the task scenario playback module 210 of the head-mounted display device 200 is driven to play the virtual task scenario image so that user 110 can generate voice signals and actions after watching it. Furthermore, the voice recognition step S04 includes driving the voice sensing module 220 of the head-mounted display device 200 to sense the voice signal of the user 110 and generate voice information. Then, it drives the voice recognition module 420 of the computing server 400 to receive the voice information and recognize the voice information according to the voice recognition program to generate a voice recognition result. Finally, it determines the task parameter set of the task context generation module 410 and the voice recognition result to generate a visual field training result. In addition, the motion recognition step S06 includes driving the motion capture device 300 to capture the user 110's movements and generate motion information. Then, it drives the motion recognition module 430 to receive the motion information and recognize the motion information according to the motion recognition program to generate a motion recognition result. Finally, it determines the context setting parameter set and the motion recognition result to generate a motion training result. The visual field training result and the motion training result are used to determine whether the user 110 has met the training requirements. Therefore, the team vision training method 500 with extended reality, voice, and motion recognition of the present invention utilizes an extended reality helmet combined with voice interaction and motion recognition technology to effectively assist athletes in vision training, allowing players to more easily grasp the movements of their teammates on the ever-changing court, thereby helping the team score and win. Furthermore, the present invention allows for individual training, avoiding the problem of high labor costs associated with existing technologies that rely on repeated practice by multiple people on a physical court.

[0040] Please refer to the following: Figure 2 , Figure 3 and Figure 9 ,in Figure 9This is a flowchart illustrating a team vision training method 600 with extended reality, voice, and motion recognition according to a fourth embodiment of the present invention. As shown in the figure, the team vision training method 600 with extended reality, voice, and motion recognition includes steps S12, S14, S16, and S18. Step S12 is "recording voice information," which means that the voice sensing module 220 of the head-mounted display device 200 senses the voice signal of the user 110 to generate and record voice information. Step S14 is "recording inertial motion information," which means that the inertial sensor 310 of the motion capture device 300a senses the motion of the user 110 to generate and record inertial motion information. Step S16 is "recording image motion information," which means that the image sensor 320 of the motion capture device 300a captures the motion of the user 110 through a camera to generate and record image motion information. Step S18 is "transmission server identification," which means that the voice sensing module 220, inertial sensor 310, and image sensor 320 transmit voice information, inertial motion information, and image motion information to the voice recognition module 420, motion recognition module 430a, and image motion recognition module 434 of the computing server 400a for identification, thereby generating visual field training results and motion training results. In this way, the team visual field training method 600 of the present invention, with extended reality, voice, and motion recognition, utilizes the interaction of extended reality helmets, voice interaction, and motion recognition technologies to effectively assist athletes in visual field training and motion training. This allows players to more easily grasp the movements of their teammates on the ever-changing court, thereby helping the team score and win, thus avoiding the problem of high labor costs associated with existing technologies that rely on repeated practice by multiple people on a physical court.

[0041] Please refer to the following: Figure 1 , Figure 3 , Figure 4 and Figure 10 ,in Figure 10 This is a block diagram illustrating a team vision training system 100b with extended reality, voice, and motion recognition according to a fifth embodiment of the present invention. As shown, the team vision training system 100b with extended reality, voice, and motion recognition is used to train the vision and movements of user 110, and includes a head-mounted display device 200, a motion capture device, and a computing server 400b. The head-mounted display device 200 and... Figure 1The head-mounted display device 200 is the same. The motion capture device is an inertial sensor 310, which is installed on the user 110 and senses the user 110's movements to generate motion information, which is then transmitted to the motion recognition module of the computing server 400b. The computing server 400b stores the scenario setting parameter group 402 and includes a task scenario generation module 410, a voice recognition module 420, a motion recognition module, and a task difficulty adjustment module 440. The scenario setting parameter group 402, the task scenario generation module 410, the voice recognition module 420, and the task difficulty adjustment module 440 are the same as those in Figures 3 and 4, respectively, and will not be described again. Specifically, the motion recognition module of the computing server 400b is an inertial motion recognition module. This inertial motion recognition module identifies motion information and generates motion recognition results. It then determines whether the motion recognition results are the same as or similar to the single-handed dribbling item 4024a, crossover dribbling item 4024b, between-the-legs dribbling item 4024c, and behind-the-back dribbling item 4024d in the execution motion parameters 4024 of the scenario setting parameter group 402, thereby generating exercise training results. Thus, the team vision training system 100b of the present invention, which features extended reality, voice, and motion recognition, can achieve team vision training and exercise training in single-player mode using only the inertial sensor 310 and the inertial motion recognition module of the computing server 400b, and is simple and convenient to set up.

[0042] Please refer to the following: Figure 1 , Figure 3 , Figure 4 and Figure 11 ,in Figure 11 This is a schematic diagram illustrating a team vision training system 100c with extended reality, voice, and motion recognition according to a sixth embodiment of the present invention. As shown, the team vision training system 100c with extended reality, voice, and motion recognition is used to train the vision and movements of user 110, and includes a head-mounted display device 200, a motion capture device, and a computing server 400c. The head-mounted display device 200 and Figure 1The head-mounted display device 200 is the same. The motion capture device is an image sensor 320, which includes a camera corresponding to the user 110. The image sensor 320 captures the user 110's movements through the camera to generate motion information and transmits the motion information to the motion recognition module of the computing server 400c. The computing server 400c stores the scenario setting parameter group 402 and includes a task scenario generation module 410, a voice recognition module 420, a motion recognition module, and a task difficulty adjustment module 440. The scenario setting parameter group 402, the task scenario generation module 410, the voice recognition module 420, and the task difficulty adjustment module 440 are the same as those in Figures 3 and 4, respectively, and will not be described again. Specifically, the motion recognition module of the computing server 400c is an image motion recognition module. The image motion recognition module identifies motion information and generates motion recognition results, and determines whether the motion recognition results are the same as or similar to the single-handed dribbling item 4024a, crossover dribbling item 4024b, between-the-legs dribbling item 4024c, and behind-the-back dribbling item 4024d of the execution motion parameters 4024 in the scenario setting parameter group 402, and generates exercise training results. In this way, the team vision training system 100c with extended reality, voice, and motion recognition of the present invention can realize team vision training and exercise training in single-person mode using only the image sensor 320 and the image motion recognition module of the computing server 400c, and is simple and convenient to set up.

[0043] As can be seen from the above embodiments, the present invention has the following advantages: First, by using an augmented reality (AR) helmet combined with voice interaction and motion recognition technology, it can effectively assist athletes in visual training, making it easier for players to grasp the movements of their teammates on the ever-changing court, thereby helping the team score and win. This avoids the problem of high manpower costs in training caused by existing technologies that rely on repeated practice by multiple people on a physical court. Second, it allows users to perform first-person tactical execution in simulated scenarios while wearing an AR helmet, and can be paired with a simple motion capture system (inertial sensor or image sensor) to record user movements. When the user completes visual training tasks while watching simulated content, the motion capture system will instantly recognize the user's movements and determine whether they can synchronously perform a stable designated dribbling action, thereby training the user's dribbling stability.

[0044] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Any person skilled in the art can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A team vision training system with extended reality, voice, and motion recognition capabilities, used to train a user's vision and motion, characterized in that, This team vision training system, featuring extended reality, voice, and motion recognition, includes: A head-mounted display device, installed on the user and comprising: A task scenario playback module plays a virtual task scenario video; A voice sensing module senses a user's voice signal and generates voice information; and A gesture sensing module is used to sense a gesture of the user and generate a gesture sensing result; A motion capture device captures a user's movement and generates motion information; and A computing server, signal-connected to the head-mounted display and the motion capture device, stores a set of contextual parameters and receives the motion information and the voice information. The computing server includes: A task scenario generation module generates a virtual task scenario image and a task parameter set based on the scenario setting parameter set, and transmits the virtual task scenario image to the head-mounted display device so that the user can generate the voice signal and the action after viewing it. A speech recognition module receives the speech information and recognizes the speech information according to a speech recognition program to generate a speech recognition result. The speech recognition module also determines a visual field training result by comparing the task parameter set of the task context generation module with the speech recognition result. A motion recognition module receives the motion information and recognizes the motion information according to a motion recognition program to generate a motion recognition result. The motion recognition module also judges the situation setting parameter group and the motion recognition result to generate a sports training result. The visual training results and the motor training results are used to determine whether the user has met the training requirements. The scenario setting parameter group includes a task execution parameter and a motion execution parameter. The task execution parameter includes a numerical summation item, which is selected based on the gesture sensing result. The motion execution parameter includes a one-handed dribbling item, a crossover dribbling item, a crossover dribbling item, and a behind-the-back dribbling item, and one of the one-handed dribbling item, the crossover dribbling item, the crossover dribbling item, and the behind-the-back dribbling item is selected based on the gesture sensing result. The motion capture device is an inertial sensor. The inertial sensor is installed on the user and senses the user's movements to generate motion information, which is then transmitted to the motion recognition module of the computing server. The motion recognition module is an inertial motion recognition module. The inertial motion recognition module recognizes the motion information and generates the motion recognition result. It then determines whether the motion recognition result is the same as or similar to the single-hand dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling parameters of the execution motion parameters in the scenario setting parameter group, and generates the motion training result accordingly. The virtual task scenario image includes multiple virtual objects and multiple numbers; When the sum of numbers is selected based on the gesture sensing result, the sum of numbers is displayed around the virtual objects for the user to view and generate the voice signal. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as the sum of numbers of the sum of numbers in the context setting parameter group and generates the vision training result, wherein the sum of numbers is equal to the sum of the sum of the multiple numbers.

2. The team vision training system with extended reality, voice, and motion recognition as described in claim 1, characterized in that, This scenario setting parameter group further includes: A player's tactical parameters include an active tactical option and an active tactical option, one of which is selected based on the gesture sensing result; and A defensive player generates parameters including an on defensive option and an off defensive option, one of which is selected based on the gesture sensing result; The execution task parameters further include a color change item, which is selected based on the gesture sensing result.

3. The team vision training system with extended reality, voice, and motion recognition as described in claim 1, characterized in that, The virtual task scenario image further includes a first color and a second color, wherein the first color and the second color are different; and When a color change item is selected based on the gesture sensing result, the plurality of numbers are displayed around the plurality of virtual objects respectively. One of the plurality of virtual objects changes from the first color to the second color. After the user views it, the voice signal is generated. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as a color change number of the color change item in the context setting parameter group and generates the vision training result, wherein the color change number is equal to one of the plurality of numbers displayed around the virtual object.

4. The team vision training system with extended reality, voice, and motion recognition as described in claim 1, characterized in that, The scenario setting parameter group further includes a task difficulty adjustment parameter, and the computing server further includes: A task difficulty adjustment module adjusts the selection of an on and off tactical items for a player's tactical parameters, an on and off defensive items for a defending player's generated parameters, a numerical summation item and a color change item for the task execution parameters, and the selection of the one-handed dribbling item, the crossover dribbling item, the between-the-legs dribbling item, and the behind-the-back dribbling item for the execution motion parameters, based on the task difficulty adjustment parameters.

5. A team vision training method with extended reality, voice and motion recognition, for training a user's vision and motion, characterized in that, This team vision training method, which incorporates extended reality, voice, and motion recognition, includes the following steps: A virtual task scenario playback step includes setting a head-mounted display device on the user, driving a task scenario generation module of a computing server to generate a virtual task scenario image and a task parameter group according to a scenario setting parameter group, transmitting the virtual task scenario image to the head-mounted display device, and then driving a task scenario playback module of the head-mounted display device to play the virtual task scenario image so that the user can generate a voice signal and an action after watching it. A voice recognition step includes driving a voice sensing module of the head-mounted display device to sense the user's voice signal and generate voice information, then driving a voice recognition module of the computing server to receive the voice information and recognize the voice information according to a voice recognition program to generate a voice recognition result, and judging the task parameter group of the task context generation module and the voice recognition result to generate a visual field training result. as well as A motion recognition step includes driving a motion capture device to capture the user's motion to generate motion information, then driving a motion recognition module to receive the motion information and recognize the motion information according to a motion recognition program to generate a motion recognition result, and judging the scenario setting parameter group and the motion recognition result to generate a sports training result. The visual training results and the motor training results are used to determine whether the user has met the training requirements. The head-mounted display device further includes a gesture sensing module, which is used to sense a gesture of the user and generate a gesture sensing result; The scenario setting parameter group includes a task execution parameter and a motion execution parameter. The task execution parameter includes a numerical summation item, which is selected based on the gesture sensing result. The motion execution parameter includes a one-handed dribbling item, a crossover dribbling item, a crossover dribbling item, and a behind-the-back dribbling item. One of the one-handed dribbling item, the crossover dribbling item, the crossover dribbling item, and the behind-the-back dribbling item is selected based on the gesture sensing result. In the motion recognition step, the motion capture device is an inertial sensor. The inertial sensor is installed on the user and senses the user's movements to generate motion information, which is then transmitted to the motion recognition module of the computing server. The motion recognition module is an inertial motion recognition module. The inertial motion recognition module recognizes the motion information to generate the motion recognition result, and determines whether the motion recognition result is the same as or similar to the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling parameters of the scenario setting parameter group, thereby generating the exercise training result. The virtual task scenario image includes multiple virtual objects and multiple numbers; When the sum of numbers is selected based on the gesture sensing result, the sum of numbers is displayed around the virtual objects for the user to view and generate the voice signal. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as the sum of numbers of the sum of numbers in the context setting parameter group and generates the vision training result, wherein the sum of numbers is equal to the sum of the sum of the multiple numbers.

6. The team vision training method with extended reality, voice and motion recognition as described in claim 5, characterized in that, This scenario setting parameter group further includes: A player's tactical parameters include an active tactical option and an active tactical option, one of which is selected based on the gesture sensing result; and A defensive player generates parameters including an on defensive option and an off defensive option, one of which is selected based on the gesture sensing result; The execution task parameters further include a color change item, which is selected based on the gesture sensing result.

7. The team vision training method with extended reality, voice and motion recognition as described in claim 5, characterized in that, The virtual task scenario image further includes a first color and a second color, wherein the first color and the second color are different; and When a color change item is selected based on the gesture sensing result, the plurality of numbers are displayed around the plurality of virtual objects respectively. One of the plurality of virtual objects changes from the first color to the second color. After the user views it, the voice signal is generated. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as a color change number of the color change item in the context setting parameter group and generates the vision training result, wherein the color change number is equal to one of the plurality of numbers displayed around the virtual object.

8. The team vision training method with extended reality, voice and motion recognition as described in claim 5, characterized in that, The scenario setting parameter group also includes a task difficulty adjustment parameter, and the virtual task scenario playback steps further include: A task difficulty adjustment module driving the computing server adjusts, based on the task difficulty adjustment parameters, the selection of an enabled and disabled tactical item for a player's tactical parameters, an enabled and disabled defensive item for a defending player's generated parameters, the numerical summation item and color change item for the executed task parameters, and the selection of the single-handed dribbling item, the crossover dribbling item, the between-the-legs dribbling item, and the behind-the-back dribbling item for the executed movement parameters.

9. A team vision training system with extended reality, voice, and motion recognition capabilities, used to train a user's vision and motion, characterized in that... This team vision training system, featuring extended reality, voice, and motion recognition, includes: A head-mounted display device, installed on the user and comprising: A task scenario playback module plays a virtual task scenario video; A voice sensing module senses a user's voice signal and generates voice information; and A gesture sensing module is used to sense a gesture of the user and generate a gesture sensing result; A motion capture device captures a user's movement and generates motion information; and A computing server, signal-connected to the head-mounted display and the motion capture device, stores a set of contextual parameters and receives the motion information and the voice information. The computing server includes: A task scenario generation module generates a virtual task scenario image and a task parameter set based on the scenario setting parameter set, and transmits the virtual task scenario image to the head-mounted display device so that the user can generate the voice signal and the action after viewing it. A speech recognition module receives the speech information and recognizes the speech information according to a speech recognition program to generate a speech recognition result. The speech recognition module also determines a visual field training result by comparing the task parameter set of the task context generation module with the speech recognition result. A motion recognition module receives the motion information and recognizes the motion information according to a motion recognition program to generate a motion recognition result. The motion recognition module also judges the situation setting parameter group and the motion recognition result to generate a sports training result. The visual training results and the motor training results are used to determine whether the user has met the training requirements. The scenario setting parameter group includes a task execution parameter and a motion execution parameter. The task execution parameter includes a numerical summation item, which is selected based on the gesture sensing result. The motion execution parameter includes a one-handed dribbling item, a crossover dribbling item, a crossover dribbling item, and a behind-the-back dribbling item, and one of the one-handed dribbling item, the crossover dribbling item, the crossover dribbling item, and the behind-the-back dribbling item is selected based on the gesture sensing result. The motion capture device is an image sensor, which includes a camera facing the user. The image sensor captures the user's movements through the camera to generate motion information and transmits the motion information to the motion recognition module of the computing server. The motion recognition module is an image motion recognition module. The image motion recognition module recognizes the motion information and generates the motion recognition result. It then determines whether the motion recognition result is the same as or similar to the single-hand dribbling, crossover dribbling, crossover dribbling, and behind-the-back dribbling parameters of the execution motion parameters in the scenario setting parameter group, and generates the exercise training result. The virtual task scenario image includes multiple virtual objects and multiple numbers; When the sum of numbers is selected based on the gesture sensing result, the sum of numbers is displayed around the virtual objects for the user to view and generate the voice signal. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as the sum of numbers of the sum of numbers in the context setting parameter group and generates the vision training result, wherein the sum of numbers is equal to the sum of the sum of the multiple numbers.

10. A team vision training system with extended reality, voice, and motion recognition capabilities, used to train a user's vision and motion, characterized in that... This team vision training system, featuring extended reality, voice, and motion recognition, includes: A head-mounted display device, installed on the user and comprising: A task scenario playback module plays a virtual task scenario video; A voice sensing module senses a user's voice signal and generates voice information; and A gesture sensing module is used to sense a gesture of the user and generate a gesture sensing result; A motion capture device that captures a user's movement to generate motion information, and the motion capture device includes: An inertial sensor is installed on the user and senses the user's movements to generate inertial motion information, which is then transmitted to a motion recognition module of a computing server; and An image sensor includes a camera facing the user. The image sensor captures the user's movements through the camera to generate image motion information and transmits the image motion information to the motion recognition module of the computing server; and The computing server is signal-connected to the head-mounted display and the motion capture device. The computing server stores a set of contextual parameters and receives the motion information and the voice information. The computing server includes: A task scenario generation module generates a virtual task scenario image and a task parameter set based on the scenario setting parameter set, and transmits the virtual task scenario image to the head-mounted display device so that the user can generate the voice signal and the action after viewing it. A speech recognition module receives the speech information and recognizes the speech information according to a speech recognition program to generate a speech recognition result. The speech recognition module also determines a visual field training result by comparing the task parameter set of the task context generation module with the speech recognition result. The action recognition module receives the action information and recognizes the action information according to an action recognition program to generate an action recognition result. The action recognition module also judges the situation setting parameter group and the action recognition result to generate a sports training result. The motion information includes the inertial motion information and the image motion information; The visual training results and the motor training results are used to determine whether the user has met the training requirements. The scenario setting parameter group includes a task execution parameter and a motion execution parameter. The task execution parameter includes a numerical summation item, which is selected based on the gesture sensing result. The motion execution parameter includes a one-handed dribbling item, a crossover dribbling item, a crossover dribbling item, and a behind-the-back dribbling item, and one of the one-handed dribbling item, the crossover dribbling item, the crossover dribbling item, and the behind-the-back dribbling item is selected based on the gesture sensing result. The virtual task scenario image includes multiple virtual objects and multiple numbers; When the sum of numbers is selected based on the gesture sensing result, the sum of numbers is displayed around the virtual objects for the user to view and generate the voice signal. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as the sum of numbers of the sum of numbers in the context setting parameter group and generates the vision training result, wherein the sum of numbers is equal to the sum of the sum of the multiple numbers.

11. The team vision training system with extended reality, voice, and motion recognition as described in claim 10, characterized in that, This action recognition module includes: An inertial motion recognition module identifies the inertial motion information to generate an inertial motion recognition result, and determines whether the inertial motion recognition result is the same as or similar to the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling parameters of the execution motion parameters in the scenario setting parameter group, thereby generating a first exercise training result; and An image motion recognition module identifies the motion information and generates the motion recognition result, and determines whether the motion recognition result is the same as or similar to the single-hand dribbling, crossover dribbling, crossover dribbling and behind-the-back dribbling parameters of the execution motion parameters in the scenario setting parameter group, and generates a second motion training result. The exercise training results include the first exercise training results and the second exercise training results.

12. A team vision training method with extended reality, voice, and motion recognition, used to train a user's vision and motion, characterized in that, This team vision training method, which incorporates extended reality, voice, and motion recognition, includes the following steps: A virtual task scenario playback step includes setting a head-mounted display device on the user, driving a task scenario generation module of a computing server to generate a virtual task scenario image and a task parameter group according to a scenario setting parameter group, transmitting the virtual task scenario image to the head-mounted display device, and then driving a task scenario playback module of the head-mounted display device to play the virtual task scenario image so that the user can generate a voice signal and an action after watching it. A voice recognition step includes driving a voice sensing module of the head-mounted display device to sense the user's voice signal and generate voice information, then driving a voice recognition module of the computing server to receive the voice information and recognize the voice information according to a voice recognition program to generate a voice recognition result, and judging the task parameter group of the task context generation module and the voice recognition result to generate a visual field training result. as well as A motion recognition step includes driving a motion capture device to capture the user's motion to generate motion information, then driving a motion recognition module to receive the motion information and recognize the motion information according to a motion recognition program to generate a motion recognition result, and judging the scenario setting parameter group and the motion recognition result to generate a sports training result. The visual training results and the motor training results are used to determine whether the user has met the training requirements. The head-mounted display device further includes a gesture sensing module, which is used to sense a gesture of the user and generate a gesture sensing result; The scenario setting parameter group includes a task execution parameter and a motion execution parameter. The task execution parameter includes a numerical summation item, which is selected based on the gesture sensing result. The motion execution parameter includes a one-handed dribbling item, a crossover dribbling item, a crossover dribbling item, and a behind-the-back dribbling item. One of the one-handed dribbling item, the crossover dribbling item, the crossover dribbling item, and the behind-the-back dribbling item is selected based on the gesture sensing result. In the motion recognition step, the motion capture device is an image sensor, which includes a camera facing the user. The image sensor captures the user's movements through the camera to generate motion information and transmits the motion information to the motion recognition module of the computing server. The motion recognition module is an image motion recognition module, which identifies the motion information to generate the motion recognition result and determines whether the motion recognition result is the same as or similar to the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling parameters of the scenario setting parameter group, thereby generating the exercise training result. The virtual task scenario image includes multiple virtual objects and multiple numbers; When the sum of numbers is selected based on the gesture sensing result, the sum of numbers is displayed around the virtual objects for the user to view and generate the voice signal. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as the sum of numbers of the sum of numbers in the context setting parameter group and generates the vision training result, wherein the sum of numbers is equal to the sum of the sum of the multiple numbers.

13. A team vision training method with extended reality, voice, and motion recognition capabilities, for training a user's vision and motion, characterized in that... This team vision training method, which incorporates extended reality, voice, and motion recognition, includes the following steps: A virtual task scenario playback step includes setting a head-mounted display device on the user, driving a task scenario generation module of a computing server to generate a virtual task scenario image and a task parameter group according to a scenario setting parameter group, transmitting the virtual task scenario image to the head-mounted display device, and then driving a task scenario playback module of the head-mounted display device to play the virtual task scenario image so that the user can generate a voice signal and an action after watching it. A voice recognition step includes driving a voice sensing module of the head-mounted display device to sense the user's voice signal and generate voice information, then driving a voice recognition module of the computing server to receive the voice information and recognize the voice information according to a voice recognition program to generate a voice recognition result, and judging the task parameter group of the task context generation module and the voice recognition result to generate a visual field training result. as well as A motion recognition step includes driving a motion capture device to capture the user's motion to generate motion information, then driving a motion recognition module to receive the motion information and recognize the motion information according to a motion recognition program to generate a motion recognition result, and judging the scenario setting parameter group and the motion recognition result to generate a sports training result. The visual training results and the motor training results are used to determine whether the user has met the training requirements. The head-mounted display device further includes a gesture sensing module, which is used to sense a gesture of the user and generate a gesture sensing result; The scenario setting parameter group includes a task execution parameter and a motion execution parameter. The task execution parameter includes a numerical summation item, which is selected based on the gesture sensing result. The motion execution parameter includes a one-handed dribbling item, a crossover dribbling item, a crossover dribbling item, and a behind-the-back dribbling item. One of the one-handed dribbling item, the crossover dribbling item, the crossover dribbling item, and the behind-the-back dribbling item is selected based on the gesture sensing result. In the motion recognition step, the motion capture device includes an inertial sensor and an image sensor. The inertial sensor is located on the user and senses the user's movements to generate inertial motion information, which is then transmitted to the motion recognition module of the computing server. The image sensor includes a camera facing the user. The image sensor captures the user's movements through the camera to generate image motion information, which is then transmitted to the motion recognition module of the computing server. The virtual task scenario image includes multiple virtual objects and multiple numbers; When the sum of numbers is selected based on the gesture sensing result, the sum of numbers is displayed around the virtual objects for the user to view and generate the voice signal. The voice recognition module determines whether the voice recognition result corresponding to the voice signal is the same as the sum of numbers of the sum of numbers in the context setting parameter group and generates the vision training result, wherein the sum of numbers is equal to the sum of the sum of numbers. The motion information includes the inertial motion information and the image motion information.

14. The team vision training method with extended reality, voice and motion recognition as described in claim 13, characterized in that, In this action recognition step, the action recognition module includes: An inertial motion recognition module identifies the inertial motion information to generate an inertial motion recognition result, and determines whether the inertial motion recognition result is the same as or similar to the single-handed dribbling, crossover dribbling, between-the-legs dribbling, and behind-the-back dribbling parameters of the execution motion parameters in the scenario setting parameter group, thereby generating a first exercise training result; and An image motion recognition module identifies the motion information and generates the motion recognition result, and determines whether the motion recognition result is the same as or similar to the single-hand dribbling, crossover dribbling, crossover dribbling and behind-the-back dribbling parameters of the execution motion parameters in the scenario setting parameter group, and generates a second motion training result. The exercise training results include the first exercise training results and the second exercise training results.

Citation Information

Patent Citations

  • Football training method

    CN109091837A

  • Systems and methods for tracking dribbling and passing performance in sporting environments

    US20180099201A1