Multi-mode cooperative processing method, control device and terminal equipment

By analyzing the vibration characteristics of audio signals to generate tactile feedback signals, the problem of the lack of integration between audio spatial information and tactile feedback is solved, realizing the fusion of hearing and touch, and improving user experience and spatial positioning accuracy.

CN122002209APending Publication Date: 2026-05-08GOERTEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GOERTEK INC
Filing Date
2025-12-30
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, audio spatial information and haptic feedback are not effectively combined, resulting in poor haptic feedback, especially at different distances, which affects the user experience.

Method used

By extracting the target vibration signal from the audio signal and analyzing its spectral characteristics to determine the spatial information of the sound source, and based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, a signal to drive the vibration unit is generated to achieve dual-sensory fusion of hearing and touch.

Benefits of technology

It improves the perception of sound sources at different distances, reduces the disconnect between hearing and touch, optimizes the user experience, and enhances the immersiveness of interaction and the accuracy of spatial positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002209A_ABST
    Figure CN122002209A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode cooperative processing method, a control device and terminal equipment, and relates to the technical field of audio processing, and the multi-mode cooperative processing method comprises the following steps: obtaining an audio signal; extracting a target vibration signal in the audio signal, performing spatial information analysis on the target vibration signal, and determining a spectrum feature of the target vibration signal; determining target distance information in the spatial information of the sound source based on the spectrum characteristics of the target vibration signal; determining a tactile feedback parameter corresponding to the spatial information of the sound source; and generating a driving signal for driving the vibration unit to work so as to generate tactile feedback based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameter. The sound source space position is converted into perceptible directional tactile experience, and the technical problem that the tactile feedback effect on different distance positions is poor due to the fact that audio space information and tactile feedback are not combined is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a multimodal collaborative processing method, control device, and terminal equipment. Background Technology

[0002] 3D Spatial Audio uses technology to simulate the effect of sound propagation in real three-dimensional space, allowing listeners to perceive the location, distance, and trajectory of sound, creating an immersive auditory experience; its core is to reproduce the human ear's sound localization mechanism through the Head Related Transfer Function (HRTF).

[0003] In related technologies, mobile phones, game consoles, smart glasses, and VR (Virtual Reality) devices use motors to achieve haptic feedback. However, this haptic feedback generally only provides single vibration feedback and is not combined with spatial audio information, resulting in a disconnect between the user's auditory and tactile senses. The actual haptic feedback effect is not good, especially in that it cannot distinguish sounds at different distances, affecting the user experience.

[0004] Therefore, a solution is needed that can combine audio spatial information to provide haptic feedback, thereby optimizing the user experience. Summary of the Invention

[0005] The main objective of this application is to provide a multimodal collaborative processing method, control device, and terminal equipment, which aims to solve the technical problem of poor tactile feedback at different distances and positions due to the failure to combine audio spatial information with tactile feedback.

[0006] On the one hand, a multimodal collaborative processing method is provided, including the following steps: Acquire audio signals; Extract the target vibration signal from the audio signal, and perform spatial information analysis on the target vibration signal to determine the spectral characteristics of the target vibration signal; Based on the spectral characteristics of the target vibration signal, the target distance information in the spatial information of the sound source is determined; Determine the tactile feedback parameters corresponding to the spatial information of the sound source; Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, a driving signal is generated to drive the vibration unit to produce tactile feedback. In one embodiment, determining the target distance information in the spatial information of the sound source based on the spectral characteristics of the target vibration signal includes: Based on the spectral characteristics of the target vibration signal, the energy ratio of high-frequency and low-frequency signals in the target vibration signal is calculated, and the target distance information in the spatial information of the sound source is determined. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the working intensity of the motor is determined according to the target distance information of the sound source. Generate a drive signal to drive the vibration unit to operate at the corresponding working intensity to produce tactile feedback.

[0007] In one embodiment, determining the target distance information in the spatial information of the sound source based on the spectral characteristics of the target vibration signal includes: Based on the spectral characteristics of the target vibration signal, the trend of the target vibration signal over time is calculated, and the target distance information in the spatial information of the sound source is determined. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the working intensity of the motor is determined according to the target distance information of the sound source. Generate a drive signal to drive the vibration unit to operate at the corresponding working intensity to produce tactile feedback.

[0008] In one embodiment, the vibration unit includes multiple motors, and the tactile feedback parameters include the motor position, motor operating intensity, and motor operating frequency corresponding to each motor. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: The motor position is obtained, and the target motor is determined based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters. The drive signal is generated to control the target motor to produce tactile feedback based on the motor's operating intensity and frequency.

[0009] In one embodiment, the spatial information analysis of the target vibration signal further includes: The target vibration signal is processed by framing to obtain time-frequency information and determine the target frequency of the sound source. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the operating frequency of the motor is determined according to the target frequency of the sound source. Generate a drive signal to drive the target motor at the corresponding operating frequency to produce tactile feedback.

[0010] In one embodiment, prior to performing the step of generating a drive signal for driving the vibration unit to produce tactile feedback, the multimodal collaborative processing method includes: Obtain configuration information and update the target tactile feedback parameters based on the configuration information. The obtained configuration information includes any one or more of the following combinations: Retrieve configuration information triggered based on the user settings interface; Obtain configuration information updated based on machine learning; Obtain configuration information determined based on scenario mode instructions.

[0011] In one embodiment, prior to performing the step of generating a drive signal for driving the vibration unit to produce tactile feedback, the multimodal collaborative processing method includes: Extract the envelope features from the audio signal to determine the amplitude characteristics of the sound source; The target tactile feedback parameters are updated based on the amplitude characteristics of the sound source.

[0012] In one embodiment, the multimodal collaborative processing method further includes the following steps: Extract the target acoustic signal from the audio signal; Based on the target acoustic signal, an acoustic drive signal for controlling the operation of a loudspeaker unit is generated and output, wherein the loudspeaker unit includes at least one loudspeaker component.

[0013] On the other hand, a control device is also provided, the device comprising: a memory, a processor, and a multimodal cooperative processing program stored in the memory and executable on the processor, the multimodal cooperative processing program being configured to implement the steps of the multimodal cooperative processing method as described above.

[0014] On the other hand, a terminal device is also provided, including: The vibration unit includes at least one motor; A loudspeaker unit, including at least one loudspeaker assembly; The processor is used to control the operation of the vibration unit and the speaker unit, and also to implement the steps of the multimodal collaborative processing method described above.

[0015] One or more technical solutions proposed in this application have at least the following technical effects: By extracting the target vibration signal from the audio signal and performing spatial information analysis on the target vibration signal, the spectral characteristics of the target vibration signal are determined. Based on the spectral characteristics of the target vibration signal, the target distance information in the spatial information of the sound source can be determined. Based on the analyzed spatial information of the sound source, the corresponding tactile feedback parameters are determined. Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the determined tactile feedback parameters are converted into driving signals that can be recognized by the vibration unit. By driving the vibration unit to work, the spatial information of the sound source is transformed into a perceptible tactile experience, which improves the perceptual experience of sound sources at different distances. By coupling the spatial information of the sound source, the dual perception fusion of hearing and touch is realized, solving the technical problem of poor tactile feedback effect caused by not combining audio spatial information with tactile feedback, especially the poor tactile feedback effect at different distance positions. This further reduces the possible disconnect between hearing and touch and optimizes the user experience. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating an embodiment of the multimodal collaborative processing method of this application is provided; Figure 2 A schematic diagram of audio signal processing is provided as an embodiment of the multimodal collaborative processing method of this application; Figure 3 A schematic flowchart of target vibration signal processing is provided for an embodiment of the multimodal collaborative processing method of this application; Figure 4 This is a detailed flowchart of step S200 in one embodiment of this application; Figure 5 This is a detailed flowchart of step S500 of one embodiment of this application; Figure 6 A schematic flowchart illustrating the generation of driving signals is provided for an embodiment of the multimodal collaborative processing method of this application; Figure 7 This is a detailed flowchart of step S500 in another embodiment of this application; Figure 8This is a detailed flowchart of step S500 in another embodiment of this application; Figure 9 A partial flowchart is provided for another embodiment of the multimodal collaborative processing method of this application; Figure 10 This is a partial flowchart illustrating another embodiment of the multimodal collaborative processing method of this application.

[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0022] 3D spatial audio is a technology that simulates the propagation of sound in real three-dimensional space, allowing listeners to perceive the location, distance, and trajectory of sound, creating an immersive auditory experience. Its core is replicating the human ear's sound localization mechanism through the Head-Related Transfer Function (HRTF). Related technologies utilize HRTF technology to personalize the simulation of a three-dimensional sound field, optimizing sound localization accuracy by measuring the user's auricular structure to ensure the most precise 3D spatial audio effect for all users. It also supports dynamic tracking, primarily by adjusting the sound field in real-time using gyroscopes and AI algorithms based on head or device movement. For example, if the sound field initially lies directly in front of the user, when the user turns to the left, the sound image will shift to the left synchronously, making the spatial sound effect more realistic. However, in these technologies, this type of application is limited to the auditory dimension.

[0023] The embodiments of this application are mainly used in the consumer electronics field, such as headphones, VR devices, etc., which support the setting of dynamic head tracking spatial audio. In devices such as mobile phones, game consoles, smart glasses, and VR devices, vibration functions are implemented through motors, providing services such as haptic feedback and incoming call reminders. The working principle of a linear resonant actuator (LRA) is to drive an internal mass block to vibrate linearly through electromagnetic drive, thereby generating haptic feedback (such as mobile phone vibration). The signal driving the motor vibration is an AC signal, and the frequency is usually in the resonant frequency range (such as 40Hz to 300Hz). However, haptic feedback in related technologies (such as mobile phone linear motors, VR controllers, etc.) is mostly a single vibration mode, which has not yet formed a deep coupling with the spatial information of the audio signal, and has not combined the spatial information of the audio to determine the corresponding haptic feedback information. In order to solve the technical problem of poor haptic feedback effect caused by the failure to combine audio spatial information with haptic feedback, the embodiments of this application propose a multimodal collaborative processing method, control device, and terminal device. In this application, for ease of description, the main focus is on control devices such as controllers and control modules with control functions.

[0024] like Figure 1 , Figure 2 , Figure 3 As shown, the multimodal collaborative processing method includes the following steps: Step S100: Acquire audio signal.

[0025] Audio signals are used to represent the mechanical vibration of a sound source and the sound waves formed by the sound source propagating in the air. The mechanical vibration of the sound source (such as the vibration of a loudspeaker diaphragm, the impact of an explosion, etc.) will cause air vibration, forming air pressure waves (acoustic signals) containing characteristics such as harmonics and timbre. Audio signals can characterize the mechanical vibration of the sound source. The mechanical vibration of the sound source itself can be directly perceived by touch. Therefore, vibration signals reflecting the vibration characteristics can be extracted from audio signals to be further combined with tactile feedback.

[0026] The audio signal may be acquired by at least one or more of the following methods: acquiring sound waves in the air by a microphone; acquiring vibrations of the sound source, carrier, etc. by using multiple sensors such as an accelerometer; or acquiring it from an audio database containing pre-stored audio signals. No limitation is imposed here.

[0027] Step S200: Extract the target vibration signal from the audio signal, and perform spatial information analysis on the target vibration signal to determine the spectral characteristics of the target vibration signal.

[0028] The extraction of target vibration signals in audio signals can be achieved, but is not limited to, directly capturing the electrical signals corresponding to the physical parameters of the sound source, such as displacement and acceleration, thereby extracting the target vibration signals in the audio signals; or by extracting target vibration features through filtering and separation algorithms, thereby extracting the target vibration signals in the audio signals.

[0029] Spatial information analysis of the target vibration signal is used to determine the spatial location information (such as orientation and distance) of the sound source in three-dimensional space. For example, the spatial location of the sound source can be analyzed by inferring the three-dimensional coordinates based on real-time changes in the sound source (such as position movement and intensity attenuation); the spatial location of the sound source can be analyzed by processing multi-channel signals through an audio engine; the spatial location of the sound source can be analyzed by directly determining the location of the sound source based on the device operating data caused by audio propagation; no further limitations are specified here.

[0030] Step S300: Based on the spectral characteristics of the target vibration signal, determine the target distance information in the spatial information of the sound source.

[0031] Because air absorbs sound to varying degrees at different frequencies, determining the spectral characteristics of the target vibration signal allows us to ascertain the target's distance within the spatial information of the sound source.

[0032] Step S400: Determine the tactile feedback parameters corresponding to the spatial information of the sound source.

[0033] Understandably, in the embodiments of this application, the corresponding tactile feedback parameters (such as vibration triggering direction, vibration intensity, vibration frequency, vibration duration, etc.) can be determined based on the parsed spatial information of the sound source (such as target distance information, and possibly also orientation). For example, the mapping relationship between the spatial information of different sound sources and tactile feedback parameters can be established in advance through experimental calibration to determine the tactile feedback parameters corresponding to the spatial information of the sound source; the tactile feedback parameters can be dynamically adjusted according to the real-time acquired spatial information of the sound source (such as orientation changes, distance changes, etc.); this is not limited here.

[0034] Step S500: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, generate a driving signal to drive the vibration unit to produce tactile feedback.

[0035] Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the determined tactile feedback parameters are converted into driving signals that the vibration unit can recognize. By driving the vibration unit, the spatial information of the sound source is transformed into a perceptible tactile experience, thereby improving the perception of sound sources at different distances. By coupling the spatial information of the sound source, dual-sensory fusion of hearing and touch is achieved, solving the technical problem of poor tactile feedback effect caused by not combining audio spatial information with tactile feedback, especially the poor tactile feedback effect at different distances. This further reduces the possibility of a disconnect between hearing and touch, optimizing the user experience. For example, in game scenarios, users can directly perceive explosions, sound source movement, etc. through touch, thereby significantly improving the interactive immersion and spatial positioning accuracy, and optimizing the user experience.

[0036] In the embodiments of this application, the multimodal collaborative processing method can be applied to terminal devices such as mobile phones, game consoles, smart glasses, and VR devices. In addition to outputting the generated drive signal used to drive the vibration unit to work and generate haptic feedback to the vibration unit configured on the corresponding device to control the vibration unit configured on the device to work, it can also output the generated drive signal to the vibration unit on other devices and control the vibration unit configured on other devices to work; this is not limited here.

[0037] For example, air absorbs sound, and the degree of absorption varies for different frequencies. Air absorption attenuates mid-to-high frequency signals more than low-frequency signals; high-frequency components in mid-to-high frequency signals are more easily absorbed by air, and the greater the distance to the sound source, the more significant the high-frequency loss. Based on this characteristic, the distance to the sound source can be deduced by analyzing the energy ratio of mid-to-high frequency signals to low-frequency signals in the target vibration signal.

[0038] like Figure 3 As shown, in one embodiment, determining the target distance information in the spatial information of the sound source based on the spectral characteristics of the target vibration signal in step S300 includes: Based on the spectral characteristics of the target vibration signal, the energy ratio of high-frequency and low-frequency signals in the target vibration signal is calculated, and the target distance information in the spatial information of the sound source is determined.

[0039] For example, the target vibration signal can be processed using methods such as Fast Fourier Transform (FFT analysis) to present the amplitude changes of the target vibration signal in the time domain. Then, based on the spectrum obtained from the FFT, the spectral characteristics of the target vibration signal, such as the intensity distribution at different frequencies, can be determined. After determining the spectral characteristics of the target vibration signal, the target distance information of the sound source can be further determined.

[0040] For example, a frequency threshold can be set first, classifying target vibration signals with frequencies not lower than the set threshold as mid-to-high frequency signals and those lower as low-frequency signals. Alternatively, a frequency range can be set, classifying target vibration signals with frequencies not lower than the upper limit of the set range as mid-to-high frequency signals and those lower than the lower limit as low-frequency signals (the specific method of determination is not limited here). Taking a set frequency threshold of 5kHz as an example, target vibration signals with frequencies not lower than 5kHz can be identified as mid-to-high frequency signals, and those lower than 5kHz as low-frequency signals.

[0041] In the embodiments of this application, the target distance information of the sound source can be determined based on the energy ratio of mid-to-high frequency signals to low-frequency signals in the target vibration signal: when the sound source is very close, the energy of the mid-to-high frequency signals is greater than that of the low-frequency signals; when the sound source is very far away, the energy of the mid-to-high frequency signals is less than that of the low-frequency signals; as the sound source gradually approaches, the proportion of energy of the mid-to-high frequency signals increases accordingly; as the sound source gradually moves away, the proportion of energy of the mid-to-high frequency signals decreases accordingly. This method can improve the positioning accuracy of the sound source's distance and is suitable for audio signal processing scenarios with complex sound fields.

[0042] like Figure 3 As shown, in another embodiment, determining the target distance information in the spatial information of the sound source based on the spectral characteristics of the target vibration signal in step S300 includes: Based on the spectral characteristics of the target vibration signal, the trend of the target vibration signal over time is calculated, and the target distance information in the spatial information of the sound source is determined.

[0043] For example, the spectral characteristics of the target vibration signal, such as the intensity distribution at different frequencies, can be determined based on the spectrum obtained by the Fast Fourier Transform. Furthermore, the trend of the amplitude across the entire frequency band over time can be analyzed: when the amplitude of the target vibration signal generally increases across the entire frequency band, it can be determined that the sound source is gradually approaching; conversely, when the amplitude generally decreases across the entire frequency band, it can be determined that the sound source is gradually moving away.

[0044] In the embodiments of this application, the target distance information of the sound source can be determined based on the change trend of the amplitude of the target vibration signal across the entire frequency band over time. This method effectively determines whether the sound source is approaching or moving away and is applicable to audio signal processing scenarios with complex sound fields.

[0045] Furthermore, when the sound source is stereo, spatial information analysis of the target vibration signal can further determine the sound source's location information within the spatial information of the sound source. For example, when the sound source is biased to the left, the sound energy received on the left (e.g., the left ear) is stronger, while the sound energy received on the right (e.g., the right ear) is relatively weaker; conversely, when the sound source is biased to the right, the sound energy received on the right (e.g., the right ear) is stronger, while the sound energy received on the left (e.g., the left ear) is relatively weaker; when the sound source is located directly in front or behind, the sound energy received on the left (e.g., the left ear) and right (e.g., the right ear) is basically the same.

[0046] like Figure 3 , Figure 4 As shown, in one embodiment, step S200, which involves extracting the target vibration signal from the audio signal and performing spatial information parsing on the target vibration signal, includes: Step S211: Extract the target vibration signal from the audio signal, perform spatial information analysis on the target vibration signal, and extract the left and right audio signals from the stereo audio signal when the target vibration signal is a stereo audio signal. Step S212: Determine the sound source location information from the energy difference between the left and right audio signals.

[0047] Understandably, spatial information analysis of the target vibration signal involves extracting the left and right audio signals (i.e., LR signals) from the stereo audio signal when the target vibration signal is a stereo audio signal. Then, the energy difference between the left and right audio signals is calculated to determine whether the sound source is emanating from the left, the right, or moving from the left to the right, etc., to determine the sound source's directional information.

[0048] In the embodiments of this application, a target time period in the target vibration signal can be selected, and audio frames can be split according to a set time interval, a set frequency band, etc.

[0049] In one embodiment, step S212, determining the sound source location information from the energy difference between the left and right audio signals, includes: Based on the energy difference between the left and right audio signals at multiple time points within the target time period, the location information of the sound source at multiple time points within the target time period is determined.

[0050] For example, a target time period can be selected from the target vibration signal, and audio frames can be split according to a set time interval. The energy of the left and right audio signals at different time points can be calculated, and the sound source location information can be determined based on the energy difference between the two audio signals. For example, when the sound source is biased to the left, the sound energy received by the left ear is stronger, and that of the right ear is relatively weaker; the opposite is true when the sound source is biased to the right; when the sound source is located directly in front or behind, the sound energy received by the left and right ears is basically the same. Through these characteristics, the sound source location information can be obtained, such as whether the sound source is emitted from the left, from the right, or has moved from the left to the right.

[0051] In the embodiments of this application, not only can the location information of the sound source at the corresponding time point be directly determined, but the movement path of the sound source within the target time period can also be determined based on the energy difference between the left and right audio signals at multiple consecutive time points within the target time period. This reduces data processing volume, improves real-time performance, and adapts to the processing needs of sound sources with dynamically changing locations.

[0052] Step S400: Determine the tactile feedback parameters corresponding to the spatial information of the sound source, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the target motor at the corresponding time point is determined according to the sound source location information at multiple time points in the target time period.

[0053] Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the corresponding target motor can be selected from the known motor positions of each motor according to the determined sound source location information. For example, if the sound source is determined to be biased to the left, the motor located on the left is set as the target motor; if the sound source is determined to be biased to the right, the motor located on the right is set as the target motor; when determining the movement path of the sound source within the target time period, the order in which the target motors switch operation at multiple consecutive time points within the target time period is determined accordingly; no restrictions are imposed here.

[0054] In another embodiment, step S212, determining the sound source location information from the energy difference between the left and right audio signals, includes: Based on the energy difference between the left and right audio signals in multiple frequency bands within the target time period, the location information of the sound source in multiple frequency bands within the target time period is determined.

[0055] The target time period in the target vibration signal is selected, and the left and right audio signals are processed by frequency band division using methods such as Fast Fourier Transform (FFT analysis). The energy difference between the two signals in each frequency band is then calculated to analyze and determine the sound source location information.

[0056] For example, in the high-frequency band, the greater the energy difference between the signals on both sides, the more the sound source is biased towards the side with the stronger signal. For instance, when the sound source is biased to the left, the signal energy received by the left ear is stronger; the opposite is true when the sound source is biased to the right. When the sound source is directly in front or behind, the signal energy of the left and right ears is basically the same. Through these characteristics, we can obtain information about the sound source's location, such as whether it is emitted from the left, from the right, or has moved from the left to the right. This can improve the accuracy of sound source localization and is applicable to audio signal processing scenarios with complex sound fields.

[0057] Step S400: Determine the tactile feedback parameters corresponding to the spatial information of the sound source, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the target motor of the corresponding frequency band is determined according to the sound source directional information of multiple frequency bands in the target time period.

[0058] Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the corresponding target motor can be selected from the known motor positions of each motor according to the determined sound source location information. For example, if the sound source is determined to be biased to the left, the motor located on the left is set as the target motor; if the sound source is determined to be biased to the right, the motor located on the right is set as the target motor; when determining the movement path of the sound source within the target time period, the order in which the target motors switch operation at multiple consecutive time points within the target time period is determined accordingly; no restrictions are imposed here.

[0059] like Figure 5 As shown, in one embodiment, step S500, generating a driving signal for driving the vibration unit to produce tactile feedback based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, includes: Step S511: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, determine the working intensity of the motor according to the target distance information of the sound source; Step S512: Generate a drive signal to drive the vibration unit to work at the corresponding working intensity to generate tactile feedback.

[0060] Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the operating intensity of the vibration unit's motor can be adjusted according to the determined target distance information of the sound source. For example, when it is determined that the sound source is gradually moving away, the motor operating intensity is reduced; when it is determined that the sound source is gradually moving closer, the motor operating intensity is increased; when it is determined that the relative distance between the sound source and the current position remains unchanged, the motor operating intensity is maintained.

[0061] like Figure 3 , Figure 6 As shown, in one embodiment, the spatial information analysis of the target vibration signal in step S200 further includes: The target vibration signal is processed by framing to obtain time-frequency information and determine the target frequency of the sound source.

[0062] The target vibration signal can be segmented into frames using methods such as Short Time Fourier Transform (STFT) time-frequency analysis to extract time-frequency information and determine the target frequency of the sound source (e.g., determine the specific frequency range of the sound source). wait).

[0063] like Figure 7 As shown, step S500, generating a driving signal to drive the vibration unit to produce tactile feedback based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, includes: Step S521: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, determine the motor operating frequency according to the target frequency of the sound source.

[0064] Based on the mapping relationship between sound source spatial information and tactile feedback parameters, for example, a pre-established mapping relationship between sound source frequency and motor operating frequency can be invoked, according to the frequency range of the sound source (e.g., Where A represents audio, i.e., the audio information of the sound source), correspondingly determining the operating frequency range of the motor (e.g. , where H represents haptic vibration, that is, the vibration of a motor or other vibrating unit.

[0065] For example, the spatial information of the sound source and the mapping relationship of the haptic feedback parameters can be presented according to the following preset spatial information and haptic feedback parameter mapping rules. Through frequency remapping, the target frequency of each frequency point of the sound source can be mapped one by one to the corresponding operating frequency of the motor:

[0066]

[0067] Furthermore, an inverse transform synthesis is performed to generate a signal that represents the motor's operating frequency over time. For example, a short-time Fourier transform (ISTFT) can be used to generate this signal. This signal corresponds to the target vibration acceleration signal Acc(t) of the motor and can be used to ultimately determine the motor's actual operating frequency.

[0068] Step S522: Generate a drive signal to drive the target motor to operate at the corresponding operating frequency to generate tactile feedback.

[0069] For example, based on the motor operating frequency in the target haptic feedback parameters, the target vibration acceleration signal of the motor is determined. Then, based on the motor motion mode (such as the motor's dynamic parameter model), the motor acceleration and voltage conversion relationship (acc to voltage) corresponding to the target vibration acceleration signal are calculated. This is further used to generate a drive signal containing parameters such as the drive voltage V(t), which is then transmitted through a power amplifier to drive the motor and other vibration units. For the same motor or a device equipped with that motor, the motor's dynamic parameter model remains consistent across different game or sound scenarios. This dynamic parameter model is derived from the motor's vibration formula and is essentially a mathematical model describing the vibration characteristics of the motor during operation, primarily used to analyze, predict, and control the motor's vibration behavior. The specific dynamic parameter model can be determined based on the actual motor model selected, and is not limited here.

[0070] In one embodiment, the driving vibration unit includes multiple motors. These multiple motors can be distributed in multiple different locations, and their directions of motion can be the same or different. For example, the multiple motors can be arranged in a mixed layout of various types to achieve a variety of tactile effects. The motors can be any type of linear motor or rotor motor; for example, a Z-axis linear motor for simulating a "tapping" or "pulsating" sensation (such as the effect of being hit by a bullet), which can be positioned perpendicular to the user's skin; or an X / Y-axis linear motor for simulating a "scratching" or "flowing" sensation (such as wind blowing across the neck), which can be positioned parallel to the user's skin; no limitation is imposed herein.

[0071] In the embodiments of this application, the tactile feedback parameters include the motor position, motor operating intensity, and motor operating frequency corresponding to each motor. For example, the spatial information of the sound source includes the sound source orientation information, the sound source distance information, the sound source frequency, etc., and can determine the motor position corresponding to the sound source orientation information, the motor operating intensity corresponding to the sound source distance information, the motor operating frequency corresponding to the sound source frequency, etc., without limitation.

[0072] like Figure 3 , Figure 8 As shown, step S500, generating a driving signal to drive the vibration unit to produce tactile feedback based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, includes: Step S531: Obtain the motor position. Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, determine the target motor according to the obtained motor position.

[0073] Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the corresponding target motor can be selected from the known motor positions of each motor according to the determined sound source location information. For example, if the sound source is determined to be biased to the left, the motor located on the left is set as the target motor; if the sound source is determined to be biased to the right, the motor located on the right is set as the target motor; when determining the movement path of the sound source within a target time period, the order in which the target motors switch operation at multiple consecutive time points within the target time period is determined accordingly. By accurately matching the motor positions, the generated and output drive signals are ensured to act only on the target motor corresponding to the sound source location information, rather than all motors, reducing potential drive failure problems.

[0074] Step S532: Generate a drive signal to control the target motor to operate based on the motor's operating intensity and frequency, thereby generating tactile feedback.

[0075] Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the working intensity of the motor of the vibration unit is adjusted according to the determined target distance information of the sound source. Furthermore, based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, a pre-established mapping relationship between the sound source frequency and the motor's working frequency is invoked. Based on the target frequency, such as the frequency range of the sound source, the corresponding motor's working frequency range is determined. In this way, a signal can be sent through a power amplifier to drive the target motor of the corresponding channel (controlling the output of the corresponding channel's vibration unit). This ensures, to a certain extent, that the drive signal only acts on the target motor corresponding to the sound source's location information, and controls the target motor's working intensity and frequency. When simultaneously driving motors at different positions and in different directions of motion, different tactile feedback effects can be achieved. Thus, the generated tactile feedback corresponds to the sound source distance and intensity, breaking through the single-dimensional experience of traditional audio and adapting to multiple scenarios such as headphones, VR, AR, game controllers, and in-vehicle entertainment, comprehensively enhancing the user's immersive experience in different application scenarios.

[0076] In one embodiment, before performing the step S500 of generating a drive signal for driving the vibration unit to operate and generate tactile feedback, the multimodal collaborative processing method includes: acquiring configuration information and updating the target tactile feedback parameters according to the configuration information.

[0077] In the embodiments of this application, obtaining configuration information includes any one or more of the following combinations: obtaining configuration information triggered by the user settings interface; obtaining configuration information updated based on machine learning; obtaining configuration information determined based on scene mode instructions.

[0078] Acquiring configuration information triggered by the user settings interface: This configuration information can be generated by the user through device control panels or other settings interfaces (such as mobile apps or VR settings interfaces), or settings interfaces with input modules such as touch modules, input buttons, and voice modules, during the calibration of haptic sensitivity preferences. For example, the overall vibration intensity of the vibration unit, or at least the vibration intensity of some motors, can be graded. Due to differences in user perception, even for the same vibration intensity, different users will have different experiences (for example, the same vibration intensity may be felt too strong by some users, while it may be just right by others). By analyzing the motor operating parameters, such as the overall operating intensity of the vibration unit or at least some motors, and based on the configuration information triggered by the user settings interface, a drive signal can be generated to drive the vibration unit to produce corresponding haptic feedback. This enables personalized haptic configuration to meet the haptic experience needs of different users.

[0079] Obtain configuration information updated based on machine learning: This configuration information can be the output information when updating the configuration through AI machine learning and other methods. For example, in different game scenarios, haptic feedback can be automatically adapted by adjusting vibration feedback thresholds through machine learning, and a corresponding drive signal can be generated (such as dynamically adjusting the vibration feedback threshold, such as the motor working intensity, according to the game scenario) to meet the real-time experience requirements of different application scenarios such as games.

[0080] The system obtains configuration information based on scene mode commands: Multiple scene modes can be pre-set. These preset scene modes can be configured according to different application environments or user behaviors (e.g., movie mode, music mode, game mode, or cinema mode, concert mode, game mode, etc.); they can also be pre-set according to different application behaviors (e.g., game mode, office mode, etc.); or they can be pre-set according to application time or time period (e.g., day mode, night mode, etc.). These preset scene modes can be pre-set by terminals such as mobile phones, game consoles, smart glasses, VR devices, or other related devices and programs before leaving the factory, or they can be selected or modified by the user based on these devices, or updated by the user through network download or cloud download. The triggering conditions for scene modes include at least one or more of the following: receiving a corresponding scene mode command, sensors detecting that the conditions for switching to the corresponding scene mode are met, and the processor recognizing that user behavior meets the switching conditions of the corresponding scene mode. When the triggering conditions are met, the system switches to the corresponding scene mode and determines relevant configuration information according to the triggered scene mode command. In the embodiments of this application, different vibration effects can be configured for different modes such as movie mode, music mode, and game mode to achieve scene adaptation. For example, when the algorithm processor confirms that it is currently in game mode, it can enhance directional vibration (such as the gunshot location indicator in some games: when a gunshot comes from the left, the user only feels vibration on the left side, etc.); when the algorithm processor confirms that it is currently in cinema mode, it can weaken the vibration intensity and focus on restoring the continuous tactile sensation of ambient sound (such as the gradually increasing vibration of thunder, or the continuous vibration from left to right achieved by multiple motors of the vibration unit when water flows from left to right).

[0081] In related technologies, while the vibration effect is realistic and effective, it can also affect the interactive experience. To further optimize the interactive experience, in the embodiments of this application, before executing step S500, which generates a drive signal to drive the vibration unit to produce tactile feedback, the multimodal collaborative processing method includes: like Figure 9 As shown, step S610 involves extracting the envelope features from the audio signal to determine the amplitude features of the sound source. For example, the audio signal can be rectified and filtered to extract the envelope features in the audio signal. The envelope features are used to represent the waveform profile of the audio signal amplitude changing over time, the duration of different amplitude changes, etc., thereby determining the amplitude characteristics of the sound source.

[0082] Step S620: Update the target tactile feedback parameters based on the amplitude characteristics of the sound source.

[0083] The target tactile feedback parameters are updated based on the amplitude characteristics of the sound source. This update is used to determine different vibration effects based on different envelope characteristics. The updated tactile feedback parameters are then used to generate a driving signal for the vibration unit to produce tactile feedback, based on the determined tactile feedback parameters corresponding to the spatial information of the sound source and the tactile feedback parameters updated based on the amplitude characteristics of the sound source.

[0084] In the embodiments of this application, the generated tactile driving signal not only accurately reflects the spatial location of the sound source, but also realistically simulates the dynamic details of its physical interaction (such as instantaneous impact, continuous action, etc.). This effectively solves the problem of limited realism and poor immersion caused by monotonous tactile feedback and lack of dynamic changes in related technologies. Thus, based on accurate spatial positioning, it is possible to achieve realism in the dynamic details and continuity of vibration effects, thereby greatly enhancing the realism and immersion of the interactive experience.

[0085] like Figure 2 , Figure 10 As shown, in one embodiment, the multimodal cooperative processing method further includes the following steps: Step S710: Extract the target acoustic signal from the audio signal; Step S720: Based on the target acoustic signal, generate and output an acoustic drive signal for controlling the operation of the loudspeaker unit, wherein the loudspeaker unit includes at least one loudspeaker component.

[0086] Audio signals represent the mechanical vibrations of a sound source and the sound waves formed by the sound source propagating in the air. The mechanical vibrations of the sound source (such as the vibration of a speaker diaphragm or the impact of an explosion) induce air vibrations, forming air pressure waves (acoustic signals) containing characteristics such as harmonics and timbre. The target acoustic signal is extracted from the audio signal. Based on the target acoustic signal, an acoustic drive signal is generated and output to control the operation of the speaker unit. This signal can be amplified to control the speaker unit or other acoustic drive units, thereby playing the full-frequency range or the desired original music signal.

[0087] It should be noted that the multimodal collaborative processing method of this application can be applied to terminal devices such as mobile phones, game consoles, smart glasses, and VR devices. Speaker components and vibration units can be located on the same terminal device or on different terminal devices. When located on different devices, the acoustic components such as speaker components can use external speaker devices, headphones, etc., and can be interconnected with the vibration units via wired or wireless connections (such as Bluetooth or Wi-Fi). This allows the device's sound to be emitted through external speakers or other sound devices instead of its own speaker components. In the embodiments of this application, a drive signal can be generated to drive the vibration unit to transmit haptic feedback. The generated drive signal is output to the vibration units located on the same or different terminal devices to control the operation of the vibration units on different terminal devices. Besides providing haptic feedback for head-mounted devices, this can also extend haptic feedback to other parts of the body (such as synchronized vibration of the controller with the direction of headphone vibration). Furthermore, an acoustic drive signal can be generated to control the operation of the speaker units, and the generated acoustic drive signal is output to the speaker components located on the same or different terminal devices to control the operation of the speaker components on the corresponding different devices. This is not limited to any specific method.

[0088] Embodiments of this application also provide a control device, the device including: a memory, a processor, and a multimodal cooperative processing program stored in the memory and executable on the processor, the multimodal cooperative processing program being configured to implement the steps of the multimodal cooperative processing method described above.

[0089] The control device in this application embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Specific configurations can be made according to actual conditions and should not impose any limitations on the functionality and scope of application of this application embodiment.

[0090] The processor can execute various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). RAM also stores various programs and data required for the operation of the control device. The processor, ROM, and RAM are interconnected via a bus. Input / output interfaces are also connected to the bus. Typically, the following systems can be connected to the input / output interface: input devices including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. Communication devices allow the control device to communicate wirelessly or wiredly with other devices to exchange data.

[0091] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a multimodal coprocessor product comprising a multimodal coprocessor carried on a computer-readable medium, the multimodal coprocessor containing program code for performing the methods shown in the flowcharts. In such embodiments, the multimodal coprocessor can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a read-only memory. When the multimodal coprocessor is executed by a processor, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0092] The control device provided in this application employs the multimodal collaborative processing method described in the above embodiments, which can solve the technical problem of poor tactile feedback caused by the failure to combine audio spatial information with tactile feedback. Compared with the prior art, the beneficial effects of the control device provided in this application are the same as those of the multimodal collaborative processing method provided in the above embodiments, and other technical features of the control device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0093] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0094] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0095] Embodiments of this application also provide a terminal device, which includes a vibration unit, a speaker unit, and a processor. The vibration unit includes at least one motor; the speaker unit includes at least one speaker assembly.

[0096] When the vibration unit includes multiple motors, a multi-axis vibration unit array can be used. Taking VR devices as an example, a built-in multi-axis motor array can include four micro vibration units: a first motor corresponding to the left ear, a second motor corresponding to the left side of the head, a third motor corresponding to the right side of the head, and a fourth motor corresponding to the right ear, etc. The motors can be either linear motors or rotor motors. For example, a Z-axis linear motor can be used to simulate a "tapping" or "pulsating" sensation (such as the effect of a bullet hitting), and it can be set perpendicular to the user's skin; an X / Y-axis linear motor can be used to simulate a "scratching" or "flowing" sensation (such as wind blowing across the neck), and it can be set parallel to the user's skin; this is not limited here. The corresponding motor is driven to vibrate according to the direction of the sound, enhancing spatial awareness in games / VR. For example, when simulating an explosion sound on the left, the first motor corresponding to the left ear can be controlled to vibrate at a high frequency, and the vibration intensity can be increased. Simultaneously, the second motor corresponding to the left side of the head can also be controlled to vibrate, with its vibration intensity and starting time set lower than the first motor; this is not limited here.

[0097] A loudspeaker unit comprises multiple loudspeaker components, which can be arranged in an array, depending on the actual setup and without limitation.

[0098] The processor is used to control the operation of the vibration unit and the speaker unit, and also to implement the steps of the multimodal collaborative processing method described above.

[0099] The terminal device provided in this application includes a processor that controls the operation of a vibration unit and a speaker unit, and also implements the steps of the multimodal collaborative processing method described in the above embodiments. This solves the technical problem of poor tactile feedback caused by the failure to combine audio spatial information with tactile feedback. Compared with the prior art, the beneficial effects of the terminal device provided in this application are the same as those of the terminal device control method provided in the above embodiments. Specific implementation methods of the terminal device's control device can be found in the descriptions of the aforementioned embodiments, and other technical features in the terminal device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0100] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A multimodal collaborative processing method, characterized in that, Includes the following steps: Acquire audio signals; Extract the target vibration signal from the audio signal, and perform spatial information analysis on the target vibration signal to determine the spectral characteristics of the target vibration signal; Based on the spectral characteristics of the target vibration signal, the target distance information in the spatial information of the sound source is determined; Determine the tactile feedback parameters corresponding to the spatial information of the sound source; Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, a driving signal is generated to drive the vibration unit to produce tactile feedback.

2. The multimodal collaborative processing method as described in claim 1, characterized in that, The determination of target distance information in the spatial information of the sound source based on the spectral characteristics of the target vibration signal includes: Based on the spectral characteristics of the target vibration signal, the energy ratio of high-frequency and low-frequency signals in the target vibration signal is calculated, and the target distance information in the spatial information of the sound source is determined. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the working intensity of the motor is determined according to the target distance information of the sound source. Generate a drive signal to drive the vibration unit to operate at the corresponding working intensity to produce tactile feedback.

3. The multimodal collaborative processing method as described in claim 1, characterized in that, The determination of target distance information in the spatial information of the sound source based on the spectral characteristics of the target vibration signal includes: Based on the spectral characteristics of the target vibration signal, the trend of the target vibration signal over time is calculated, and the target distance information in the spatial information of the sound source is determined. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the working intensity of the motor is determined according to the target distance information of the sound source. Generate a drive signal to drive the vibration unit to operate at the corresponding working intensity to produce tactile feedback.

4. The multimodal collaborative processing method as described in claim 1, characterized in that, The vibration unit includes multiple motors, and the tactile feedback parameters include the motor position, motor operating intensity, and motor operating frequency corresponding to each motor. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: The motor position is obtained, and the target motor is determined based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters. The drive signal is generated to control the target motor to produce tactile feedback based on the motor's operating intensity and frequency.

5. The multimodal collaborative processing method as described in claim 1, characterized in that, The spatial information analysis of the target vibration signal also includes: The target vibration signal is processed by framing to obtain time-frequency information and determine the target frequency of the sound source. The mapping relationship between the spatial information of the sound source and the tactile feedback parameters is used to generate a driving signal for driving the vibration unit to produce tactile feedback, including: Based on the mapping relationship between the spatial information of the sound source and the tactile feedback parameters, the operating frequency of the motor is determined according to the target frequency of the sound source. Generate a drive signal to drive the target motor at the corresponding operating frequency to produce tactile feedback.

6. The multimodal collaborative processing method as described in any one of claims 1 to 5, characterized in that, Prior to performing the step of generating the drive signal for driving the vibration unit to produce tactile feedback, the multimodal collaborative processing method includes: Obtain configuration information and update the target tactile feedback parameters based on the configuration information. The obtained configuration information includes any one or more of the following combinations: Retrieve configuration information triggered based on the user settings interface; Obtain configuration information updated based on machine learning; Obtain configuration information determined based on scenario mode instructions.

7. The multimodal collaborative processing method according to any one of claims 1 to 5, characterized in that, Prior to performing the step of generating the drive signal for driving the vibration unit to produce tactile feedback, the multimodal collaborative processing method includes: Extract the envelope features from the audio signal to determine the amplitude characteristics of the sound source; The target tactile feedback parameters are updated based on the amplitude characteristics of the sound source.

8. The multimodal collaborative processing method as described in any one of claims 1 to 5, characterized in that, The multimodal collaborative processing method further includes the following steps: Extract the target acoustic signal from the audio signal; Based on the target acoustic signal, an acoustic drive signal for controlling the operation of a loudspeaker unit is generated and output, wherein the loudspeaker unit includes at least one loudspeaker component.

9. A control device, characterized in that, The apparatus includes: a memory, a processor, and a multimodal collaborative processing program stored in the memory and executable on the processor, the multimodal collaborative processing program being configured to implement the steps of the multimodal collaborative processing method as described in any one of claims 1 to 8.

10. A terminal device, characterized in that, include: The vibration unit includes at least one motor; A loudspeaker unit, including at least one loudspeaker assembly; The processor is used to control the operation of the vibration unit and the speaker unit, and is also used to implement the steps of the multimodal collaborative processing method as described in any one of claims 1 to 8.