Multi-mode interactive voice gesture collaborative dealing system and method
The multimodal interactive voice and gesture collaborative card dealing system solves the shortcomings of existing card dealing devices in voice and gesture collaborative control, realizes efficient and natural card dealing operation, and improves user experience and operation smoothness.
Patent Information
- Application Number
- CN202510977462.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing card dealing devices are inadequate in terms of voice and gesture-based collaborative control, failing to meet users' needs for a natural interactive experience. Furthermore, the lack of clearly defined collaborative triggering conditions limits the flexibility and convenience of user scenarios.
By establishing a conflict resolution mechanism between voice and gesture, integrating multi-sensor data from microphones and cameras, and defining clear collaborative triggering conditions, such as the combination of "hand hovering + voice command", a multimodal interactive voice and gesture collaborative card dealing system is realized.
It improves the collaborative control efficiency and accuracy of the card dealing device, enhances interactive flexibility and scene adaptability, optimizes user experience, resolves command conflict issues, and ensures smooth operation.
Smart Images

Figure CN120872148A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of human-computer interaction and intelligent control technology, specifically a multimodal interactive voice and gesture collaborative card dealing system and method. Background Technology
[0002] With the intelligent development of tabletop gaming devices, card dealing devices are gradually evolving from simple mechanical operation to multimodal interaction. However, existing card dealing devices still have shortcomings in voice and gesture-based collaborative control, making it difficult to meet users' needs for a natural interactive experience.
[0003] A high-efficiency card dealing machine and method, disclosed in CN113813588B, achieves efficient card dealing through a combination of a horizontal rotation mechanism, a drive device, and distance and counting sensors. It can continuously load and deal cards at a preset frequency while adjusting the dealing direction by rotation. Although this technical solution demonstrates excellent performance in mechanical structure optimization and efficiency improvement, its human-computer interaction is still limited to traditional button or manual triggering, failing to incorporate voice commands or gesture recognition as triggering conditions. Furthermore, the device lacks clearly defined collaborative triggering conditions (such as a combination of "palm hover + voice command"), limiting the flexibility and convenience of user scenarios. Simultaneously, the solution does not adequately consider how to handle potential conflicts between voice and gesture commands, which may affect the smoothness of operation in practical applications.
[0004] The aforementioned problems indicate that existing card dealing devices still have significant shortcomings in multimodal interaction support, voice and gesture collaborative control mechanism design, and the application of multi-sensor data fusion technology. Therefore, this invention provides a multimodal interactive voice and gesture collaborative card dealing system and method. It aims to achieve a more natural, flexible, and efficient card dealing operation by establishing a voice and gesture conflict resolution mechanism (such as priority rules), fusing multi-sensor data from microphones and cameras, and defining clear collaborative triggering conditions (such as a combination of "hand hover + voice command"), thereby meeting the intelligent interaction needs of modern tabletop game devices. Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal interactive voice and gesture collaborative card dealing system and method, which solves the shortcomings of existing card dealing devices in tabletop game devices in terms of voice and gesture collaborative control, conflict resolution mechanism design, and application of multi-sensor data fusion technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a multimodal interactive voice and gesture-based collaborative card dealing method, the method comprising:
[0007] Determine the parameter range of the card dealing device in the target interaction scenario, set different trigger modes within the parameter range, and the target card dealing device responds to user commands according to the different trigger modes set.
[0008] Signal acquisition is performed on user commands that have undergone different triggering modes within the target card dealing device to obtain input feature information of user commands under different triggering modes, and then an interactive behavior distribution map under different triggering modes is generated based on all the input feature information.
[0009] The operation status data of the card dealing device under different triggering modes is collected. Based on all the operation status data, the response time volatility of the card dealing device under different triggering modes is determined. Priority weight fitting is performed on all response time volatility to obtain the interaction efficiency coefficient of the card dealing device.
[0010] Based on the interaction efficiency coefficient and the distribution map of all interaction behaviors, the matching degree of user instructions under different triggering modes is determined, and the optimal trigger combination when the user instruction completes the response is determined by all matching degrees.
[0011] Preferably, the signal acquisition of user commands that have undergone different trigger modes within the target card dealing device, and the acquisition of input feature information of user commands under different trigger modes specifically include:
[0012] For each user command in the triggering mode, acquire multimodal data consisting of audio signals transmitted by the microphone array and visual signals transmitted by the camera array in the card dealing environment;
[0013] The input feature information of the user command under each trigger mode is parsed from the multimodal data.
[0014] Preferably, generating an interaction behavior distribution map under different triggering modes based on all input feature information specifically includes:
[0015] For each user command under each trigger mode, a behavior modeling tool is used to analyze the execution trajectory of the user command and obtain a dynamic mapping of the interaction behavior.
[0016] Based on the user's input feature information and the dynamic mapping graph, an interaction behavior distribution map is generated for each trigger mode.
[0017] Preferably, determining the response time volatility of the card dealing device under different triggering modes based on all operational status data specifically includes:
[0018] For card dealing devices under different triggering modes, the execution time of each instruction node is determined based on the operating status data of the card dealing device;
[0019] Determine the standard time interval for processing instructions by the card dealing device;
[0020] Calculate the time deviation value of each instruction node based on all execution times and the standard time interval;
[0021] The response time volatility of the card dealing device is determined by all time deviation values, and then the response time volatility of the card dealing device under different triggering modes is obtained.
[0022] Preferably, determining the matching degree of user commands under different triggering modes based on the interaction efficiency coefficient and the distribution map of all interaction behaviors specifically includes:
[0023] Obtain an ideal distribution map of behavior when user commands are responded to smoothly;
[0024] For user commands under different triggering modes, the distribution difference value between the actual distribution and the ideal distribution of the user commands is calculated based on the interaction behavior distribution map and the ideal distribution map.
[0025] The matching degree of user instructions is calculated based on the interaction efficiency coefficient and the distribution difference value, thereby obtaining the matching degree of user instructions under different triggering modes.
[0026] Preferably, the user instructions are basic interactive instructions obtained by integrating data using a multimodal data fusion protocol based on voice signals, gestures, and environmental parameters.
[0027] Preferably, the environmental parameters are auxiliary parameters corresponding to the hand hovering position or the voice intensity.
[0028] Preferably, calculating the distribution difference between the actual distribution and the ideal distribution of user commands based on the interaction behavior distribution map and the ideal distribution map specifically includes:
[0029] For each user command in the triggering mode, extract the set of actual behavior trajectory points in the interaction behavior distribution map and the set of target behavior trajectory points in the ideal distribution map;
[0030] The trajectory deviation values between the actual behavior trajectory point set and the corresponding positions in the target behavior trajectory point set are statistically analyzed, and the distribution difference value is determined by the weighted average of all trajectory deviation values.
[0031] Preferably, the standard time interval for determining the card-dealing device's processing instructions specifically includes:
[0032] Acquire the time records of when the card dealing device completes similar instructions during historical interactions;
[0033] Calculate the average of all time records and use the average as the standard time interval for the card dealing device to process instructions.
[0034] Preferably, the present invention also includes a multimodal interactive voice and gesture collaborative card dealing system, the system including a user command response optimization device, the user command response optimization device comprising:
[0035] The mode setting module is used to determine the parameter range of the card dealing device in the target interaction scenario, set different trigger modes within the parameter range, and instruct the target card dealing device to respond to user commands according to the different trigger modes set.
[0036] The signal analysis module is used to control the acquisition of signals from user commands that have undergone different triggering modes within the target card dealing device, obtain the input feature information of user commands under different triggering modes, and then generate an interactive behavior distribution map under different triggering modes based on all the input feature information.
[0037] The signal analysis module is also used to collect the operating status data of the card dealing device under different triggering modes, determine the response time fluctuation rate of the card dealing device under different triggering modes based on all the operating status data, and perform priority weight fitting on all response time fluctuation rates to obtain the interaction efficiency coefficient of the card dealing device.
[0038] The strategy optimization module is used to determine the matching degree of user instructions under different triggering modes based on the interaction efficiency coefficient and the distribution map of all interaction behaviors, and to determine the optimal trigger combination when the user instructions complete the response based on all matching degrees.
[0039] The beneficial effects of the multimodal interactive voice and gesture collaborative card dealing system and method are mainly reflected in the following aspects:
[0040] 1. Improve the efficiency and accuracy of collaborative control: By integrating multimodal data fusion (integrating voice signals, gestures and environmental parameters such as hand hovering position and voice intensity), combined with response time fluctuation analysis, interaction efficiency coefficient calculation and matching degree optimization, efficient collaboration between voice and gesture commands is achieved, reducing the error of single-modal interaction and improving the response speed and accuracy of the card dealing device to user commands.
[0041] 2. Enhanced interactive flexibility and scenario adaptability: Breaking through the limitations of existing card dealing devices that rely on traditional buttons or manual triggering, it introduces multiple modes such as voice triggering, gesture triggering, and "voice + gesture" collaborative triggering, and defines clear collaborative triggering conditions (such as "palm hovering + voice command"), which greatly expands the flexibility and convenience of user scenarios.
[0042] 3. Optimize user experience: Generate an interaction behavior distribution map through behavior modeling tools, calculate the difference value by combining it with the ideal distribution map, and then determine the optimal trigger combination based on the interaction efficiency coefficient. This can adapt to user habits and dynamically adjust parameters, making the interaction more in line with the natural operation logic and significantly improving the user's natural interaction experience.
[0043] 4. Resolving command conflict issues: By fitting the response time fluctuation rate with priority weights, a conflict resolution mechanism for voice and gesture commands was established, avoiding confusion when multimodal commands are triggered simultaneously and ensuring the smoothness of card dealing operations. Attached Figure Description
[0044] Figure 1 The schematic diagram of the module structure of the multimodal interactive voice and gesture collaborative card dealing system provided in the embodiment of the present invention shows the composition of each functional module in the system and their connection relationship.
[0045] Figure 2 This is a flowchart illustrating the user instruction input feature information collection process under different triggering modes provided in embodiments of the present invention.
[0046] Figure 3 The logical block diagram of the interaction behavior distribution map generation method provided in the embodiments of the present invention.
[0047] Figure 4 A flowchart illustrating the response time volatility calculation method provided in this embodiment of the invention.
[0048] Figure 5 A flowchart illustrating the matching degree calculation and optimal trigger combination determination method provided in this embodiment of the invention. Detailed Implementation
[0049] This invention provides a multimodal interactive voice and gesture collaborative card dealing system and method. The specific embodiments of this invention will be described in detail below with reference to the accompanying drawings. Figure 1 The schematic diagram of the module structure of the multimodal interactive voice and gesture collaborative card dealing system provided in the embodiment of the present invention shows the composition of each functional module in the system and their connection relationship; Figures 2 to 5 The flowcharts, corresponding to different steps, are used to explain in detail the functional implementation process of each module.
[0050] like Figure 1As shown, this system includes a mode setting module, a signal analysis module, and a strategy optimization module. The mode setting module and the signal analysis module are connected via a data transmission interface, and the signal analysis module is further connected to the strategy optimization module. Together, these three components constitute a complete user command response optimization device. In addition, the system includes a microphone array, a camera array, a behavior modeling tool, a runtime data acquisition unit, and a priority weight fitting unit. These components are connected to the signal analysis module to collect, process, and analyze multimodal data. Physically, the microphone array and camera array are mounted on the top of the card dealing device to ensure coverage of the user's voice and gestures. The behavior modeling tool is deployed inside the signal analysis module to generate dynamic mapping maps and interactive behavior distribution maps. The runtime data acquisition unit is directly connected to the control unit of the card dealing device to acquire real-time runtime data, while the priority weight fitting unit is embedded in the strategy optimization module to calculate response time volatility and perform priority weight fitting.
[0051] In actual operation, the mode setting module first determines the parameter range of the card dealing device under the target interaction scenario and sets different trigger modes within this range. For example, in a tabletop game scenario, the target interaction scenario might include the user issuing a voice command "Please deal three cards" or pointing in a certain direction with a gesture to indicate the card dealing position. The mode setting module presets three trigger modes according to the scenario requirements: voice trigger mode, gesture trigger mode, and voice-gesture combined trigger mode. Each trigger mode has a corresponding parameter range; for example, the sound intensity threshold for the voice trigger mode needs to reach above 60 decibels, and the hand hovering position for the gesture trigger mode needs to be between 20 and 50 centimeters in front of the card dealing device. The parameter range setting is based on experimental data and user habit statistics to ensure the rationality and feasibility of the trigger modes. The mode setting module sends the above parameter range and trigger mode information to the signal parsing module for subsequent response to user commands.
[0052] The signal analysis module is responsible for acquiring signals from user commands that have undergone different trigger modes within the target card dealing device and generating input feature information. For example... Figure 2As shown, the signal analysis module first acquires audio and visual signals through a microphone array and a camera array, respectively. The microphone array consists of four high-sensitivity microphones, installed at the four corners of the top of the card dealing device, to capture the user's voice commands. The camera array includes two wide-angle cameras, located on either side of the top of the card dealing device, to capture the user's hand gestures. After receiving the audio signal from the microphone array, the signal analysis module uses a speech recognition algorithm to extract the speech content and sound intensity parameters. Simultaneously, it performs image processing on the visual signal received from the camera array to identify the gesture trajectory and the hand's hovering position. Based on this, the signal analysis module integrates the audio and visual signals according to a multimodal data fusion protocol to generate the input feature information of the user's command for each trigger mode. For example, in the voice-gesture collaborative trigger mode, the input feature information might include the voice content "deal five cards," a sound intensity of 65 decibels, a gesture pointing azimuth angle of 30 degrees, and a hand hovering position 35 centimeters away from the card dealing device.
[0053] The signal parsing module also utilizes behavioral modeling tools to generate interactive behavior distribution maps under different triggering modes. For example... Figure 3 As shown, the behavior modeling tool first analyzes the execution trajectory of user commands to generate a dynamic mapping map. The dynamic mapping map records the entire process of a user command from issuance to completion, including the start and end times of the voice command, the starting and ending coordinates of the gesture, and the change curve of the hand's hovering position. Subsequently, the behavior modeling tool combines input feature information and the dynamic mapping map to generate an interaction behavior distribution map. The interaction behavior distribution map is presented in the form of a two-dimensional histogram, with the horizontal axis representing the time axis and the vertical axis representing the frequency of user behavior under different triggering modes. For example, under voice triggering mode, the interaction behavior distribution map shows that most user commands are completed within 1 to 3 seconds after the voice command is issued, while under gesture triggering mode, the completion time of user commands is mainly distributed within the range of 2 to 4 seconds. The interaction behavior distribution map can intuitively reflect the execution patterns of user commands under different triggering modes, providing data support for subsequent optimization.
[0054] The signal analysis module further collects operational status data of the card dealing device under different trigger modes and calculates the response time volatility based on this data. For example... Figure 4As shown, the operation status data acquisition unit monitors the operation status of the card dealing device in real time, including motor speed, displacement of the card dealing mechanism, and execution time of command nodes. The signal analysis module determines the execution time of each command node based on the operation status data. For example, the execution time of the voice command analysis node is 0.5 seconds, the gesture recognition node is 0.8 seconds, and the card dealing action execution node is 1.2 seconds. The signal analysis module also acquires the time records of the card dealing device completing similar commands in historical interactions and calculates the average of all time records as the standard time interval. For example, for the command "deal three cards," the average historical time record is 3 seconds, which is then used as the standard time interval. The signal analysis module calculates the time deviation value of each command node based on all execution times and the standard time interval. For example, the time deviation value of the voice command analysis node is 0.5 seconds minus one-third of the standard time interval, i.e., 1 second, resulting in -0.5 seconds. By summarizing all time deviation values, the response time volatility is obtained, thereby determining the response time volatility of the card dealing device under different triggering modes. For example, in voice-triggered mode, the response time fluctuation rate is positive 0.2 seconds, while in gesture-triggered mode, the response time fluctuation rate is negative 0.3 seconds.
[0055] The strategy optimization module determines the matching degree of user commands under different triggering modes based on the interaction efficiency coefficient and the distribution map of all interaction behaviors, and determines the optimal trigger combination when the user command completes the response based on the matching degree. For example... Figure 5As shown, the strategy optimization module first obtains the ideal distribution map of user commands when responses are smooth. The ideal distribution map is a standard behavior distribution derived from a large amount of user interaction data; its horizontal axis also represents the time axis, and the vertical axis represents the frequency of user behavior under different triggering modes. The strategy optimization module uses a priority weight fitting unit to fit all response time fluctuations to obtain the interaction efficiency coefficient of the card-dealing device. For example, the interaction efficiency coefficient for the voice triggering mode is 0.85, the interaction efficiency coefficient for the gesture triggering mode is 0.75, and the interaction efficiency coefficient for the voice-gesture collaborative triggering mode is 0.9. The strategy optimization module further calculates the distribution difference between the actual distribution and the ideal distribution of user commands. Taking the voice triggering mode as an example, the strategy optimization module extracts the set of actual behavior trajectory points in the interaction behavior distribution map and the set of target behavior trajectory points in the ideal distribution map, and calculates the trajectory deviation value at corresponding positions. For example, the behavior frequency at a certain time point in the actual behavior trajectory point set is 0.6, while the behavior frequency at the same time point in the target behavior trajectory point set is 0.8, and the trajectory deviation value is 0.2. The distribution difference value is obtained by taking the weighted average of all trajectory deviation values; for example, the distribution difference value for the voice-triggered mode is 0.15. The strategy optimization module calculates the matching degree of the user command based on the interaction efficiency coefficient and the distribution difference value; for example, the matching degree of the voice-triggered mode is 0.85 minus 0.15 equals 0.7. Finally, the strategy optimization module determines the optimal trigger combination when the user command completes the response by comparing the matching degrees of all trigger modes. For example, in a specific application scenario, the voice-gesture collaborative trigger mode has the highest matching degree and is therefore selected as the optimal trigger combination.
[0056] In practical applications, this system can be deployed on desktop gaming devices to enhance the user experience. For example, in a multiplayer poker game, players can request to be dealt cards by using the voice command "Please deal me two cards" or by gesturing towards their position. The system automatically selects the optimal trigger combination based on the current scenario parameters. For instance, when a player simultaneously issues a voice command and makes a gesture, the system will prioritize the voice-gesture coordinated trigger mode, thereby achieving fast and accurate card dealing. Furthermore, the system can dynamically adjust the parameter range and trigger mode according to user habits, further improving interaction efficiency and user satisfaction.
[0057] To enable those skilled in the art to fully understand and implement this invention, the specific implementation principles of this invention are further supplemented below with a specific application scenario.
[0058] In practical applications, card dealing devices are deployed in tabletop gaming devices to enhance the player's interactive experience. For example, in a multiplayer poker game, the system uses microphone and camera arrays to capture players' voice commands and gestures in real time, and automatically selects the optimal trigger combination based on the current scene parameters. When a player simultaneously issues the voice command "Please deal me two cards" and gestures towards their position, the system prioritizes the voice-gesture coordinated trigger mode, thereby quickly and accurately completing the card dealing operation. The following will explain the operating principle of this process in detail with reference to the module labels and flowcharts in the accompanying drawings.
[0059] First, the mode setting module determines the parameter range applicable to the current game environment based on the requirements of the target interaction scenario and sets different trigger modes. For example, in a poker game, the voice trigger mode requires a sound intensity of 60 decibels or higher, while the gesture trigger mode requires the palm to be hovered between 20 and 50 centimeters in front of the card dealing device. These parameter ranges are derived based on experimental data and statistical results of user habits, ensuring the rationality and feasibility of the trigger conditions. The mode setting module then sends the parameter range and trigger mode information to the signal parsing module, providing a basis for subsequent user command responses.
[0060] Next, the signal analysis module acquires audio and visual signals through a microphone array and a camera array, respectively. For example... Figure 2 As shown, the microphone array consists of four high-sensitivity microphones mounted at the four corners of the top of the card dealing device to capture voice commands issued by the player. The camera array includes two wide-angle cameras located on either side of the top of the card dealing device to capture the player's hand gestures. The signal analysis module performs speech recognition processing on the received audio signal, extracting the voice content "Please deal me two cards" and a sound intensity of 65 decibels. Simultaneously, it performs image processing on the visual signal, recognizing the hand gesture's azimuth angle as 30 degrees and the hand's hovering position as 35 centimeters away from the card dealing device. Based on this, the signal analysis module integrates the audio and visual signals according to a multimodal data fusion protocol to generate input feature information. This feature information includes not only the voice content and sound intensity parameters but also data such as the hand gesture trajectory and hand hovering position, providing support for subsequent interactive behavior analysis.
[0061] Subsequently, the signal parsing module uses behavioral modeling tools to generate a distribution map of interactive behaviors under different triggering modes. For example... Figure 3As shown, the behavior modeling tool first analyzes the execution trajectory of user commands to generate a dynamic mapping map. The dynamic mapping map records the entire process of a user command from issuance to completion, including the start and end times of the voice command, the starting and ending coordinates of the gesture, and the change curve of the hand's hovering position. Combining input feature information with the dynamic mapping map, the behavior modeling tool generates an interaction behavior distribution map. This distribution map is presented in the form of a two-dimensional histogram, with the horizontal axis representing the time axis and the vertical axis representing the frequency of user behavior under different triggering modes. For example, in the voice triggering mode, user commands are mainly completed within 1 to 3 seconds after the voice command is issued; while in the gesture triggering mode, the completion time of user commands is distributed within the range of 2 to 4 seconds. These distribution maps can intuitively reflect the execution patterns of user commands under different triggering modes, providing data support for optimizing trigger combinations.
[0062] Meanwhile, the signal analysis module monitors the operating status data of the card dealing device in real time through the operating status data acquisition unit, including motor speed, displacement of the card dealing mechanism, and execution time of the command nodes. For example... Figure 4 As shown, the signal analysis module determines the execution time of each instruction node based on the operating status data. For example, the execution time of the voice instruction analysis node is 0.5 seconds, the gesture recognition node is 0.8 seconds, and the card dealing execution node is 1.2 seconds. Furthermore, the signal analysis module also obtains the time records of the card dealing device completing similar instructions during historical interactions and calculates the average of all time records as the standard time interval. For example, for the instruction "deal three cards," the average historical time record is 3 seconds, which is then used as the standard time interval. The signal analysis module calculates the time deviation value of each instruction node based on all execution times and the standard time interval. For example, the time deviation value of the voice instruction analysis node is 0.5 seconds minus one-third of the standard time interval, i.e., 1 second, resulting in -0.5 seconds. By summarizing all time deviation values, the signal analysis module determines the response time fluctuation rate of the card dealing device under different triggering modes. For example, in the voice triggering mode, the response time fluctuation rate is positive 0.2 seconds, while in the gesture triggering mode, the response time fluctuation rate is negative 0.3 seconds.
[0063] Finally, the strategy optimization module determines the matching degree of user commands under different triggering modes based on the interaction efficiency coefficient and the distribution map of all interaction behaviors, and determines the optimal trigger combination when the user command completes the response based on the matching degree. For example... Figure 5As shown, the strategy optimization module first obtains the ideal distribution map of user commands when responses are smooth. The ideal distribution map is a standard behavior distribution derived from a large amount of user interaction data; its horizontal axis also represents the time axis, and the vertical axis represents the frequency of user behavior under different triggering modes. The strategy optimization module uses a priority weight fitting unit to fit all response time fluctuations to obtain the interaction efficiency coefficient of the card-dealing device. For example, the interaction efficiency coefficient for the voice triggering mode is 0.85, the interaction efficiency coefficient for the gesture triggering mode is 0.75, and the interaction efficiency coefficient for the voice-gesture collaborative triggering mode is 0.9. The strategy optimization module further calculates the distribution difference between the actual distribution and the ideal distribution of user commands. Taking the voice triggering mode as an example, the strategy optimization module extracts the set of actual behavior trajectory points in the interaction behavior distribution map and the set of target behavior trajectory points in the ideal distribution map, and calculates the trajectory deviation value at corresponding positions. For example, the behavior frequency at a certain time point in the actual behavior trajectory point set is 0.6, while the behavior frequency at the same time point in the target behavior trajectory point set is 0.8, and the trajectory deviation value is 0.2. The distribution difference value is obtained by taking the weighted average of all trajectory deviation values; for example, the distribution difference value for the voice trigger mode is 0.15. The strategy optimization module calculates the matching degree of the user command based on the interaction efficiency coefficient and the distribution difference value; for example, the matching degree of the voice trigger mode is 0.85 minus 0.15 equals 0.7. Finally, the strategy optimization module determines the optimal trigger combination when the user command completes the response by comparing the matching degrees of all trigger modes. In the poker game scenario mentioned above, the voice-gesture collaborative trigger mode has the highest matching degree and is therefore selected as the optimal trigger combination.
[0064] Through the steps described above, the system achieves efficient card dealing operations through coordinated voice and gesture control. When a player simultaneously issues a voice command and makes a gesture, the system can quickly analyze multimodal data, generate input feature information, and determine the optimal trigger combination by combining the interaction behavior distribution map and response time volatility. This design not only improves the response speed and accuracy of the card dealing device but also significantly enhances the user experience. Furthermore, the system can dynamically adjust the parameter range and trigger mode based on user habits, further improving interaction efficiency and user satisfaction.
[0065] All content not described in detail in this specification is prior art known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited; conventional equipment can be used. Electrical control components not mentioned in this technical solution are not shown in the figures because they are prior art, and will not be described further here.
[0066] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0067] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal interactive voice and gesture-based collaborative card dealing method, characterized in that, Includes the following steps: Determine the parameter range of the card dealing device in the target interaction scenario, set different trigger modes within the parameter range, and the target card dealing device responds to user commands according to the different trigger modes set. Signal acquisition is performed on user commands that have undergone different triggering modes within the target card dealing device to obtain input feature information of user commands under different triggering modes, and then an interactive behavior distribution map under different triggering modes is generated based on all the input feature information. The operation status data of the card dealing device under different triggering modes is collected. Based on all the operation status data, the response time volatility of the card dealing device under different triggering modes is determined. Priority weight fitting is performed on all response time volatility to obtain the interaction efficiency coefficient of the card dealing device. Based on the interaction efficiency coefficient and the distribution map of all interaction behaviors, the matching degree of user instructions under different triggering modes is determined, and the optimal trigger combination when the user instruction completes the response is determined by all matching degrees.
2. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 1, characterized in that, Signal acquisition is performed on user commands that have undergone different trigger modes within the target card dealing device. The specific input characteristic information of user commands under different trigger modes includes: For each user command in the triggering mode, acquire multimodal data consisting of audio signals transmitted by the microphone array and visual signals transmitted by the camera array in the card dealing environment; The input feature information of the user command under each trigger mode is parsed from the multimodal data.
3. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 1, characterized in that, The generation of interaction behavior distribution maps under different triggering modes based on all input feature information specifically includes: For each user command under each trigger mode, a behavior modeling tool is used to analyze the execution trajectory of the user command and obtain a dynamic mapping of the interaction behavior. Based on the user's input feature information and the dynamic mapping graph, an interaction behavior distribution map is generated for each trigger mode.
4. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 1, characterized in that, Based on all operational status data, the response time volatility of the card dealing device under different triggering modes is determined, specifically including: For card dealing devices under different triggering modes, the execution time of each instruction node is determined based on the operating status data of the card dealing device; Determine the standard time interval for processing instructions by the card dealing device; Calculate the time deviation value of each instruction node based on all execution times and the standard time interval; The response time volatility of the card dealing device is determined by all time deviation values, and then the response time volatility of the card dealing device under different triggering modes is obtained.
5. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 1, characterized in that, Determining the matching degree of user commands under different triggering modes based on the interaction efficiency coefficient and the distribution map of all interaction behaviors specifically includes: Obtain an ideal distribution map of behavior when user commands are responded to smoothly; For user commands under different triggering modes, the distribution difference value between the actual distribution and the ideal distribution of the user commands is calculated based on the interaction behavior distribution map and the ideal distribution map. The matching degree of user instructions is calculated based on the interaction efficiency coefficient and the distribution difference value, thereby obtaining the matching degree of user instructions under different triggering modes.
6. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 1, characterized in that, The user commands are basic interactive commands obtained by integrating data based on voice signals, gestures, and environmental parameters using a multimodal data fusion protocol.
7. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 6, characterized in that, The environmental parameters are auxiliary parameters corresponding to the hand hovering position or the voice intensity.
8. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 5, characterized in that, The calculation of the distribution difference between the actual and ideal distributions of user commands based on the interaction behavior distribution map and the ideal distribution map specifically includes: For each user command in the triggering mode, extract the set of actual behavior trajectory points in the interaction behavior distribution map and the set of target behavior trajectory points in the ideal distribution map; The trajectory deviation values between the actual behavior trajectory point set and the corresponding positions in the target behavior trajectory point set are statistically analyzed, and the distribution difference value is determined by the weighted average of all trajectory deviation values.
9. The multimodal interactive voice and gesture collaborative card dealing method as described in claim 4, characterized in that, The standard time interval for determining the card-dealing device's processing instructions specifically includes: Acquire the time records of when the card dealing device completes similar instructions during historical interactions; Calculate the average of all time records and use the average as the standard time interval for the card dealing device to process instructions.
10. A multimodal interactive voice and gesture collaborative card dealing system, characterized in that, It includes a user command response optimization device, the user command response optimization device comprising: The mode setting module is used to determine the parameter range of the card dealing device in the target interaction scenario, set different trigger modes within the parameter range, and instruct the target card dealing device to respond to user commands according to the different trigger modes set. The signal analysis module is used to control the acquisition of signals from user commands that have undergone different triggering modes within the target card dealing device, obtain the input feature information of user commands under different triggering modes, and then generate an interactive behavior distribution map under different triggering modes based on all the input feature information. The signal analysis module is also used to collect the operating status data of the card dealing device under different triggering modes, determine the response time fluctuation rate of the card dealing device under different triggering modes based on all the operating status data, and perform priority weight fitting on all response time fluctuation rates to obtain the interaction efficiency coefficient of the card dealing device. The strategy optimization module is used to determine the matching degree of user instructions under different triggering modes based on the interaction efficiency coefficient and the distribution map of all interaction behaviors, and to determine the optimal trigger combination when the user instructions complete the response based on all matching degrees.
Citation Information
Patent Citations
A high-efficiency card dealing machine and card dealing method
CN113813588B