Toy sound perception and voice recognition fusion evaluation method
By analyzing the pressure and deformation distribution during toy interaction, and dynamically adjusting sensor signals and speech recognition algorithms, the interference of plush material changes on speech recognition was resolved, enabling AI toys to respond accurately under different levels of interaction.
Patent Information
- Application Number
- CN202511087828.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-12-09
Smart Images

Figure CN121096366A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, and in particular to a toy sound perception and speech recognition fusion evaluation method. BACKGROUND
[0002] Artificial intelligence-driven plush pets, as a new interactive toy, are gradually becoming an important part of family entertainment and emotional companionship, thanks to their realistic movements and voice responses. This technology combines mechanical movement, sound processing, and intelligent algorithms to provide an immersive experience for users. However, its core lies in accurately responding to user interaction behavior, which puts high demands on the system's perception ability and dynamic adaptability. Studying the correlation between sound perception and interactive response not only improves user experience but also has key significance for promoting the application of intelligent interaction technology in consumer products. Existing methods rely solely on single sensor input when processing user interaction, without considering the interference of toy material and structural changes on signal transmission, and ignoring the influence of the complexity of the physical structure of plush pets on the perception effect, resulting in insufficient response accuracy of the system in different use scenarios, especially when the user interaction force changes. For example, existing toys often have decreased voice recognition stability when facing different touch pressures due to material deformation, affecting the accurate recognition of user commands. The core challenge lies in the interference of the dynamic changes of plush pet materials on the sound transmission path. Compression of the plush fibers enhances the damping effect of sound propagation, resulting in a decrease in the signal strength received by the sensor, which in turn affects the stability of voice recognition. This is because the compression of fibers changes the density distribution and pore structure of the sound wave propagation medium, causing more scattering and energy loss of sound waves during propagation, thereby reducing the transmission efficiency of acoustic signals. Further, the change in this damping effect causes uneven distribution of the internal filler density, causing the drift of the sound resonance frequency, making it difficult for the system to accurately distinguish between user commands and environmental noise. For example, when users gently touch or press the plush pet, the degree of fiber compression varies, and the sound propagation path changes, causing the toy to misjudge the command or react slowly. These two factors are interrelated: fiber compression directly changes the sound transmission characteristics, and changes in transmission characteristics further affect the stability of the resonance frequency, thereby forming a technical problem for dynamic adjustment of voice recognition. SUMMARY
[0003] The present application provides a toy sound perception and speech recognition fusion evaluation method, mainly including: Obtain the local pressure peak distribution and deformation distribution when the user touches, hugs, and presses, analyze the influence of the difference in the degree of compression of the plush fibers caused by the change in the user interaction force on the sound propagation path, and obtain the spatial distribution data of the compression of the plush fibers; According to the spatial distribution data of the compression of the plush fibers, calculate the compression directivity index; The density change rate and local density peak point of the internal filling of the AI toy are analyzed, a filling density uneven distribution pattern caused by the change of damping effect is matched, and a quantitative value of the filling density change is obtained; The influence degree of the sound resonance frequency drift phenomenon caused by the uneven distribution of the filling density on the frequency spectrum of the speech signal is analyzed, and the frequency drift dynamic range is determined; According to the frequency drift dynamic range, the sensor signal intensity is dynamically adjusted, the enhanced speech signal data is generated, and the discrimination result of the user voice command and the environmental noise is determined; The discrimination result of the user voice command and the environmental noise is detected in quality, the sensor array signal fusion is optimized, and the optimized AI toy command recognition result is obtained; According to the optimized AI toy command recognition result, the action of the AI toy plush pet is adjusted, and the dynamic response output conforming to the user interaction force is obtained.
[0004] Further, the local pressure peak distribution and deformation distribution when the user touches, hugs, and presses are obtained, the influence of the difference in the compression degree of the plush fiber caused by the change in the user interaction force on the sound propagation path is analyzed, and the spatial distribution data of the compression of the plush fiber is obtained, including: The pressure values of each sensing point in the user interaction process are obtained, the local pressure gradient is calculated according to the pressure difference between adjacent sensing points, the deformation amount of each position is calculated in combination with the elastic modulus parameter of the plush material, and a three-dimensional deformation distribution map is constructed; according to the three-dimensional deformation distribution map, the ratio of the deformation amount of each position to the original plush thickness is calculated to obtain the compression rate value, the spatial dispersion degree of the compression rate is calculated by standard deviation, and a compression uniformity coefficient distribution map is generated; according to the compression uniformity coefficient distribution map, the ratio of the fiber density of each region to the original density is calculated, the sound wave propagation speed change rate is determined in combination with the change of the fiber arrangement direction, and the compression degree distribution field is generated by spatial interpolation according to the sound wave propagation speed change rate and the sensing point position coordinates to determine the compression degree level of each spatial position, and the spatial distribution data of the compression of the plush fiber is obtained.
[0005] Further, the compression directionality index is calculated according to the spatial distribution data of the compression of the plush fiber, including: According to the spatial distribution data of the compression of the plush fiber, the density gradient value is obtained by dividing the fiber density difference between adjacent positions by the distance, the components of the compression in each direction of the three-dimensional space are calculated, the maximum component direction is determined as the main compression direction, and the compression directionality index is calculated as the cosine value of the angle between the compression direction of each position and the main direction.
[0006] Further, the analysis AI toy internal filling density change rate and local density peak point, match out the filling density uneven distribution mode caused by the change of damping effect, get the quantization value of the filling density change, including: Calculate the filling volume change ratio, get the density change value, determine the density change rate by the ratio of adjacent area density difference and interval, mark the position of local density peak point whose density change rate exceeds the threshold value; According to the density change rate and the local density peak point, construct a three-dimensional distribution feature vector, generate a digital coding sequence; Through the digital coding sequence in the preset database, calculate the similarity with the standard coding, extract the density distribution parameter set; According to the density distribution parameter set, calculate the filling density value of each spatial position, and the difference between the standard density value, get the quantization value of the filling density change.
[0007] Further, the analysis of the influence degree of sound resonance frequency drift phenomenon caused by the uneven distribution of filling density on the frequency spectrum of speech signal, determine the frequency drift dynamic range, including: According to the quantization value of the filling density change, calculate the density gradient vector, determine the sound wave deflection degree, calculate the local resonance frequency value; According to the difference between the local resonance frequency value and the initial resonance frequency, get the resonance frequency drift value, identify the frequency component change amplitude through spectrum decomposition; According to the frequency component change amplitude, calculate the frequency spectrum distortion degree, generate the frequency spectrum influence degree distribution diagram; According to the frequency spectrum influence degree distribution diagram, count the drift value range, determine the frequency drift dynamic range.
[0008] Further, according to the frequency drift dynamic range, dynamically adjust the sensor signal intensity, generate the enhanced speech signal data, judge the discrimination result of user voice command and environmental noise, including: According to the frequency drift dynamic range, calculate the density recovery rate, determine the hardness change rule, calculate the sensor gain adjustment coefficient by the product of recovery rate and hardness increment, generate the enhanced speech signal data; According to the enhanced speech signal data, extract the speech feature vector, combine the filling density change rate to correct the low frequency component, adjust the speech recognition decision threshold value; According to the decision threshold value, calculate the matching degree of audio feature vector and standard instruction feature vector, distinguish user voice command and environmental noise.
[0009] Further, the quality detection is carried out on the discrimination result of user voice command and environmental noise, the sensor array signal fusion is optimized, and the optimized AI toy instruction recognition result is obtained, including: According to the discrimination result of the user voice instruction and the environmental noise, the recognition category proportion in the time window is calculated, the recognition stability is evaluated, and the recognition quality score is generated; according to the recognition quality score and the pressure distribution, the sensor signal noise covariance is estimated, the weight initial value is calculated, and the weight coefficient is adjusted combined with the filler density change rate; according to the weight coefficient, the sensor audio signal is weighted and summed to generate a fusion signal, and gain compensation is applied to generate a spectrum-compensated audio signal; according to the spectrum-compensated audio signal, the standard instruction feature is matched to obtain the optimized AI toy instruction recognition result.
[0010] Further, the AI toy instruction recognition result is adjusted according to the optimized AI toy instruction recognition result to obtain a dynamic response output that conforms to the user interaction force, comprising: According to the optimized AI toy instruction recognition result, the instruction type code is extracted, the interaction direction is determined combined with the compression directionality, the preset mapping relationship is queried, and the action parameter set is obtained; according to the action parameter set and the material hardness change value, the head rotation angle correction coefficient and the four limb swing amplitude ratio are calculated, and the action control parameter is determined; according to the action control parameter, the rotation angle of the servo motor is controlled to generate a sound output signal; the mechanical action signal is synchronized with the sound output signal to obtain the dynamic response output.
[0011] The technical scheme provided by the embodiment of the application can include the following beneficial effects: The application discloses a toy sound perception and voice recognition fusion evaluation method, which collects user interaction data through an AI toy built-in pressure sensor array, and fuses an intelligent voice interaction optimization method of the compression characteristics of plush fibers and the sound propagation damping effect. In view of the problems of uneven compression of plush fibers, change of filler density and enhancement of sound propagation damping caused by user interaction such as petting and hugging, the application analyzes the pressure peak value distribution and the deformation distribution, calculates the fiber compression rate and the density gradient, quantifies the sound wave refraction and scattering intensity, determines the material damping coefficient and the resonance frequency drift value, and then adaptively adjusts the sensor signal intensity and the voice recognition algorithm parameter to optimize the instruction and noise discrimination accuracy. Meanwhile, combined with the compression directionality index and the material hardness change, the action response of the AI toy is dynamically adjusted to ensure accurate voice recognition and response under different interaction forces. The application significantly improves the voice instruction recognition robustness and response accuracy of the AI toy in a complex interaction scene, overcomes the interference of the damping effect of the plush material on sound propagation, and provides a natural and smooth interaction experience for the user. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 The flowchart of the toy sound perception and voice recognition fusion evaluation method of the application. DETAILED DESCRIPTION
[0013] For a further understanding of the present application, reference will be made to the following detailed description of the application taken in conjunction with the accompanying drawings. The following detailed description of the application is provided as an example to explain the application and should not be considered as limiting the application. It should also be noted that the drawings are only intended to illustrate the present application and not to limit the present application.
[0014] As Figure 1 The toy sound perception and voice recognition fusion evaluation method of the embodiment can specifically include the following steps. In step S101, the local pressure peak value distribution and deformation distribution when the user pats, hugs, and presses are obtained, the influence of the difference in the compression degree of the plush fiber caused by the change in the user interaction force on the sound propagation path is analyzed, and the spatial distribution data of the compression of the plush fiber is obtained.
[0015] The pressure values of each sensing point in the user interaction process are obtained through the pressure sensor array, the local pressure gradient is calculated according to the pressure difference of adjacent sensing points, the deformation amount of each position is calculated through the ratio of the pressure value to the elastic modulus parameter of the plush material, and a complete three-dimensional deformation distribution map is constructed. Based on the three-dimensional deformation distribution map data, the compression rate value of each position is obtained by dividing the deformation amount by the original plush thickness, the dispersion degree of the compression rate in space is evaluated by using the standard deviation calculation method, if the standard deviation is less than a preset threshold, it is determined that the compression is uniform, otherwise the compression non-uniform region is identified, and a compression uniformity coefficient distribution map is generated. According to the compression uniformity coefficient distribution map, for different compression degree regions, the fiber density change amount of each region is calculated through the ratio relationship of the original density and the fiber density of the plush fiber, and the change degree of the fiber arrangement direction is combined to evaluate the propagation speed change rate of the sound wave under different compression states. The propagation speed change rate data is combined with the position coordinates of each sensing point, and a continuous compression degree distribution field is generated by using the space interpolation method, and the compression degree level of each spatial position is determined according to the comprehensive weight of the pressure peak value, the deformation amount and the fiber density change, and the spatial distribution data of the compression of the plush fiber is obtained.
[0016] In one possible implementation, the arrangement of the pressure sensor array adopts a matrix arrangement, and the spacing between the sensing points is kept within the range of 5-10 millimeters, so as to ensure that the pressure changes caused by different contact areas such as the user's fingers and palms can be captured. When the user pats the toy, each sensing point collects pressure values in real time, and the pressure gradient is calculated through the pressure difference of adjacent sensing points. This gradient calculation method can reflect the change trend of the pressure in space, and provide basic data support for subsequent deformation analysis.
[0017] Specifically, the calculation process of the deformation variable involves the analysis of the elastic properties of the plush material. Different materials of plush fibers have different elastic moduli, and the elastic modulus of cotton plush is smaller, while the elastic modulus of synthetic fiber plush is relatively larger. By pre-determining the elastic modulus parameters and combining the measured pressure values, the deformation variable at each position can be accurately calculated. The advantage of this calculation method is that it can convert discrete pressure point data into continuous deformation distribution field, forming a complete three-dimensional deformation distribution map.
[0018] It should be noted that the spatial distribution evaluation of the compression ratio uses standard deviation as the uniformity judgment index. Standard deviation can quantitatively reflect the degree of dispersion of data. When the difference in compression ratio at each position is small, the standard deviation value is also small. By setting a reasonable threshold, the uniform compression area and the non-uniform compression area can be effectively distinguished. This distinction is of great significance for subsequent acoustic characteristic analysis, because different compression degrees will directly affect the propagation characteristics of sound waves.
[0019] In one embodiment, the evaluation process of fiber density change is based on the volume change relationship before and after compression. When the plush fibers are compressed, the gap between the fibers decreases, the number of fibers in a unit volume increases, resulting in an increase in local density. By establishing the corresponding relationship between the compression degree and the density change, the density distribution state of each region can be accurately evaluated. This density change directly affects the propagation speed of sound waves, and the greater the density of the region, the speed of sound wave propagation will usually change accordingly.
[0020] Preferably, the application of spatial interpolation method enables the discrete sensor point data to be converted into continuous distribution field. Commonly used interpolation methods include Kriging interpolation, inverse distance weighted interpolation, etc. These methods can calculate the value of unknown positions according to the data of known points. By considering multiple factors such as pressure peak value, deformation variable and fiber density change, and assigning different weight coefficients, the generated spatial distribution data can fully reflect the compression state distribution characteristics of the plush fibers under user interaction.
[0021] Step S102, according to the spatial distribution data of the compression of the plush fibers, the compression directionality index is calculated.
[0022] According to the spatial distribution data of the compression of the plush fibers, the density gradient value is obtained by dividing the difference in fiber density between adjacent positions by the distance, the components of compression in each direction in three-dimensional space are calculated by using vector decomposition method, the maximum component direction is determined as the main direction of compression, and the cosine value of the angle between the compression direction of each position and the main direction is calculated as the compression directionality index.
[0023] In one possible implementation, the calculation of the density gradient value is based on the spatial variation characteristics of the density of the plush fibers in the compressed state. When the user presses different positions of the toy, the fiber density will present a non-uniform distribution, and the density difference between adjacent positions reflects the gradual variation characteristics of the compression degree. By calculating the density variation within a unit distance, the degree of this gradual variation can be quantified. The greater the density gradient value, the more intense the compression variation in this region, and the more significant the impact on the propagation of sound waves.
[0024] It should be noted that the vector decomposition method plays a key role in determining the main direction of compression. The compression of plush fibers is not simply a vertical deformation, but presents complex distribution characteristics in three-dimensional space. By decomposing the compression vector into components in the x, y, and z directions, the dominant direction of compression can be accurately identified. This determination of the main direction provides a reference benchmark for subsequent calculation of the compression directionality index, enabling uniform quantitative characterization of the compression characteristics of each position.
[0025] Specifically, the difference in propagation speed of sound waves in different density media is the fundamental reason for the refraction phenomenon. When sound waves propagate from a low-density region to a high-density region, the propagation speed changes, causing the propagation direction to deflect. The size of this deflection angle depends on the degree of density difference between the two regions. Based on the density gradient value and the compression directionality index, according to the difference in propagation speed of sound waves in different density media, the sound wave refraction angle is calculated by the ratio of the sine of the incident angle to the sine of the refraction angle, and the sound wave propagation deflection angle data of each position is obtained. By establishing the mathematical relationship between the incident angle and the refraction angle, the actual propagation path of the sound wave in the plush layer can be predicted, which is of great significance for optimizing the sound output effect of the toy.
[0026] In one embodiment, the logarithmic relationship between sound wave amplitude attenuation and propagation distance reflects the regularity of energy loss. In the compressed state of the plush fibers, the contact area between fibers increases, and the friction effect enhances, leading to intensified energy loss of sound waves during propagation. This attenuation is not a linear relationship, but presents a logarithmic attenuation characteristic, i.e., the attenuation is faster at the initial stage of propagation, and the attenuation rate gradually decreases as the distance increases. By analyzing this attenuation law, the damping effect of different compression states on sound propagation can be accurately evaluated. Using the sound wave propagation deflection angle data, combined with the energy loss characteristics generated by the friction between fibers in the compressed state of the plush material, the attenuation coefficient reflecting the degree of propagation damping enhancement is calculated through the logarithmic relationship between sound wave amplitude attenuation and propagation distance. The attenuation coefficient of sound wave propagation can be calculated by the following formula:
[0027] , α represents the attenuation coefficient of sound wave propagation, A0 represents the initial sound wave amplitude, A dThe amplitude of the sound wave after the propagation distance d is represented, d represents the propagation distance of the sound wave, and the attenuation coefficient reflecting the degree of propagation damping enhancement is calculated by the logarithmic relationship between the sound wave amplitude attenuation and the propagation distance. According to the attenuation coefficient and the randomness of the fiber arrangement, the scattering probability is calculated by the ratio of the number of collisions of the sound wave with the fiber interface to the total propagation path. The scattering probability formula can be:
[0028] , P s represents the scattering probability, N c represents the number of collisions of the sound wave with the fiber interface, L total represents the total propagation path length, R f represents the randomness coefficient of the fiber arrangement, and the formula calculates the scattering probability by combining the randomness of the fiber with the ratio of the number of collisions to the total propagation path. By combining the ratio of the wavelength of the sound wave of different frequencies to the fiber spacing, the material damping coefficient is determined, and the formula is:
[0029] , P s represents the scattering probability of the sound wave, r represents the fiber radius, and d represents the fiber spacing. The formula describes the relationship between the scattering probability of the sound wave in the fiber material and the wavelength and fiber geometric parameters. When the ratio of the wavelength to the fiber spacing is small, the scattering effect is stronger.
[0030] Step S103, analyze the density change rate and local density peak point of the internal filling of the AI toy, match the uneven distribution mode of the filling density caused by the change of damping effect, and obtain the quantitative value of the filling density change.
[0031] The density change value of the filler at different positions relative to the uncompressed state is calculated by the material damping coefficient numerical range corresponding to the volume change ratio of the filler after being compressed. The density change rate is obtained by dividing the density difference of adjacent regions by the region spacing. The positions where the density change rate exceeds the preset threshold are determined as the local density peak points. Based on the density change rate and the spatial coordinates of the local density peak points, a three-dimensional distribution feature vector including the number of peak points, the average distance between peak points, and the main direction angle of the density gradient is constructed. A digital coding sequence of the uneven distribution mode of the filler density is generated by combining the components of the feature vector according to a predetermined rule. The digital coding sequence is used to search in a preset density distribution database of the plush pet filler. The database stores standard density distribution modes and corresponding parameters in different compressed states. The similarity is calculated by calculating the Euclidean distance between the matching coding sequence and the standard coding in the database. If the similarity is less than the matching threshold, the corresponding density distribution parameter set is extracted. According to the density correction coefficient and the spatial distribution function in the density distribution parameter set, and in combination with the measured damping effect data, the current density value of the filler at each spatial position is obtained by interpolation calculation. The quantified value of the density change of the filler is obtained by subtracting the standard density value of the filler when the toy is shipped from the current density value.
[0032] Specifically, when the sound wave propagates in the filler, a higher damping coefficient indicates greater energy loss, which usually corresponds to a high-density state of the filler after being compressed. A mapping relationship between the damping coefficient interval and the volume compression ratio is established. For example, when the damping coefficient is in the range of 0.3-0.5, the volume compression ratio is 20%-30%. This quantitative relationship provides a reliable basis for subsequent density calculation.
[0033] In one possible implementation, the calculation of the density change rate is realized by a difference method for accurate quantization. The density difference of adjacent regions reflects the uneven degree of the filler in space, and the region spacing determines the severity of the change. When the density change rate exceeds the set threshold, it indicates that there is a significant density mutation at this position, and these mutation positions are the local density peak points. The identification of these peak points plays a key role in understanding the overall distribution characteristics of the filler.
[0034] It should be noted that the construction process of the three-dimensional distribution feature vector involves the extraction of multiple key parameters. The number of peak points reflects the complexity of the density distribution, and the more the number, the more uneven the distribution; the average distance between peak points embodies the spatial scale characteristics of the density change; and the main direction angle of the density gradient reveals the dominant direction of the compression effect. These three parameters describe the distribution characteristics of the filler from different dimensions, and by combining them into a feature vector, the digital representation of the complex distribution mode is realized.
[0035] Specifically, the generation of the digital code sequence follows a specific coding rule. For example, the number of peak points is mapped to two digits, the distance between peaks is mapped to one digit, and the main direction angle interval is mapped to two digits, forming a five-digit code sequence. This coding method not only preserves the original feature information, but also facilitates fast retrieval and matching in the database.
[0036] In one embodiment, the construction of the plush pet filler density distribution database is based on a large amount of measured data. The database stores standard distribution patterns of fillers of different materials and different initial densities in various compressed states, each pattern containing a complete set of density distribution parameters, including a density correction coefficient matrix and a spatial distribution function expression. By calculating the similarity between the to-be-matched code and the standard code using the Euclidean distance, the closest standard pattern can be found, and the corresponding parameters can be obtained for subsequent calculations.
[0037] It should be noted that interpolation calculation plays an important role in obtaining continuous density distribution. Since the measured data is usually discrete point data, the density value at any spatial position can be calculated by interpolation method. Combined with the density correction coefficient and spatial distribution function extracted from the database, the standard pattern can be adjusted to a density distribution that conforms to the actual measurement results. By comparing with the factory standard density value, the specific change amount of the filler density at each position is obtained.
[0038] Step S104, analyze the influence degree of the sound resonance frequency shift phenomenon caused by the uneven distribution of the filler density on the frequency spectrum of the voice signal, and determine the frequency shift dynamic range.
[0039] According to the quantized value of the filler density change, the density gradient vector of each spatial position is calculated, the deflection degree of the sound wave is determined by the angle between the density gradient direction and the sound wave propagation direction, and the local resonance frequency value at different positions is calculated combined with the linear relationship coefficient between the compression modulus and the density of the filler material. Based on the difference between the local resonance frequency value and the initial resonance frequency of the toy design, the resonance frequency shift value is obtained, the frequency spectrum of the voice signal is decomposed by using fast Fourier transform, and the change amplitude of the frequency component corresponding to the resonance frequency shift value in the frequency spectrum is identified. According to the proportion of the change amplitude of the frequency component to the original amplitude, the frequency spectrum distortion degree is calculated, if the distortion degree exceeds the preset threshold, the frequency band range and its distortion degree value that are most seriously affected are recorded, and a frequency spectrum influence degree distribution map is formed. Using the distortion degree value in the frequency spectrum influence degree distribution map and the corresponding frequency shift value, the frequency shift dynamic range is determined by the difference between the maximum value and the minimum value of the shift value of all measurement points, and the reverse compensation gain curve is generated according to the range, realizing the frequency compensation of the interference of the sound propagation characteristics caused by the change of the damping.
[0040] In one possible implementation, the calculation of the density gradient vector involves the directional derivative of the density field in three-dimensional space. When the filler is subjected to uneven compression, the density exhibits a gradient distribution in space. By calculating the density difference between adjacent measurement points and dividing by the spatial distance, the rate of change of density in each direction can be obtained. The vector composed of these rates of change is the density gradient vector, whose direction points to the direction of the fastest density increase, and whose magnitude reflects the degree of density change. When a sound wave encounters a density gradient during propagation, refraction occurs, and the degree of refraction is closely related to the angle between the gradient direction and the sound wave propagation direction.
[0041] It should be noted that the linear relationship coefficient between the compression modulus and the density reflects the mechanical properties of the material. For a plush filler, as the density increases, the stiffness of the material also increases accordingly. This relationship can be approximated as linear within a certain range. By experimentally determining the compression modulus values at different densities, the slope coefficient of this linear relationship can be determined. The local resonant frequency is directly related to the compression modulus of the material. The greater the modulus, the higher the resonant frequency. This relationship provides a theoretical basis for subsequent frequency drift calculations.
[0042] Specifically, the fast Fourier transform plays a core role in speech signal spectrum analysis. This transform can convert time-domain speech signals into frequency-domain representations, revealing the amplitude and phase information of each frequency component in the signal. When the density distribution of the filler inside the toy is uneven, the differences in resonant characteristics at different positions will cause certain frequency components to be enhanced or attenuated. By comparing and analyzing the differences in the frequency spectrum before and after the transformation, the frequency components affected by the resonant frequency drift and their change amplitudes can be accurately identified.
[0043] In one embodiment, the quantitative evaluation of the degree of spectral distortion uses the relative change rate method. By calculating the ratio of the current amplitude of the affected frequency component to the original amplitude, the distortion coefficient of each frequency point can be obtained. When the distortion coefficients of multiple frequency points exceed a predetermined threshold, it indicates that the frequency band is significantly affected. Recording the start and end frequencies of these affected frequency bands, as well as the corresponding distortion degree values, forms a complete frequency spectrum impact degree distribution map, providing a basis for subsequent compensation.
[0044] It should be noted that the design of the inverse compensation gain curve is based on the statistical properties of the frequency drift. By analyzing the distribution of frequency drift values of all measurement points, the dynamic range of the drift, i.e., the difference between the maximum and minimum drift values, is determined. According to this range, the corresponding gain curve is designed to produce an opposite gain effect to the distortion in the affected frequency band.
[0045] For example, if a frequency band is attenuated by 3 decibels due to resonant drift, the compensation curve should provide a gain of 3 decibels in that frequency band.
[0046] Step S105, for the dynamic range of frequency drift, dynamically adjust the sensor signal intensity, generate enhanced voice signal data, and judge the distinction result of user voice instruction and environmental noise.
[0047] For the dynamic range of frequency drift, the ratio value of the density recovery of the filler in unit time after pressure release is obtained as the density recovery rate, the hardness change rule is determined according to the corresponding relationship between the material compression time and the hardness increment, the product of the recovery rate and the hardness increment is calculated as the sensor gain adjustment coefficient, and the adjustment coefficient is multiplied by the original sensor signal to generate enhanced voice signal data. Based on the enhanced voice signal data, the Mel frequency cepstrum coefficient is extracted to obtain the voice feature vector, the low-frequency components in the feature vector are weighted and corrected according to the density change rate of the filler, and the voice recognition decision threshold value is adjusted according to the corresponding relationship between the material damping coefficient and the frequency. Using the adjusted decision threshold value, the cosine similarity between the input audio feature vector and the stored standard voice instruction feature vector is calculated as the matching degree, if the matching degree is higher than the decision threshold value, it is recognized as the user voice instruction, otherwise it is classified as environmental noise, and the preliminary recognition result is obtained. According to the recognition consistency of the continuous frames in the preliminary recognition result and the change trend of the short-time energy of the audio signal, and combining the different attenuation ratios of the high-frequency band and the low-frequency band due to the damping effect, the distinction result of the user voice instruction and the environmental noise is judged.
[0048] In one possible implementation, the measurement of the density recovery rate is based on the elastic recovery characteristics of the filler. When the user stops pressing the toy, the compressed filler will gradually recover to the original state. By measuring the density value of the filler at different time points, the density change amount per unit time can be calculated.
[0049] For example, the cotton filler may recover 30% of the original density in the first second after the pressure is released, and 20% in the second second, which reflects the viscoelastic characteristics of the material. At the same time, the hardness of the material will change with repeated compression, and the filler of a new toy is relatively fluffy, and will become compact after long-term use, and the hardness increment and the use time present a logarithmic relationship.
[0050] It should be noted that the calculation of the sensor gain adjustment coefficient considers both the recovery rate and the hardness change. When the recovery rate of the filler is fast, the sound wave propagation environment changes dramatically, and higher gain is needed to compensate for signal attenuation; while the increase in hardness will cause the sound wave propagation speed to increase, and the gain needs to be reduced. By multiplying these two factors, the adjustment coefficient obtained can dynamically balance the signal intensity in different states, ensuring that the voice signal maintains a stable volume level under various use conditions.
[0051] Specifically, the mel-frequency cepstral coefficient is a standard feature extraction method in the field of speech recognition. This method simulates the auditory characteristics of the human ear, maps the frequency spectrum of the speech signal according to the mel scale, and then obtains the feature vector through the cepstrum transform. In the case of uneven filler density, the low-frequency component is more significantly affected because the wavelength of low-frequency sound waves is longer and more susceptible to density changes. By weighting and correcting the low-frequency component, the spectral distortion caused by density changes can be compensated for, improving the accuracy of feature extraction.
[0052] In one embodiment, the cosine similarity calculation provides an effective pattern matching method. The standard voice command feature vector is pre-recorded and extracted under ideal conditions and stored in the memory of the toy. When a user input is received, the system extracts its feature vector and then calculates the cosine value of the included angle between the two vectors. The closer the cosine value is to 1, the more similar the two vectors are, i.e., the more the input audio matches the standard command. The decision threshold value is dynamically adjusted according to the damping coefficient, and the threshold value is appropriately lowered when the damping is greater to compensate for the effects of signal quality degradation.
[0053] It should be noted that the continuous frame recognition consistency check and short-time energy analysis provide double protection for the final judgment. Speech commands usually have stable timing characteristics, and the recognition results of adjacent frames should remain consistent; environmental noise is often random, and the recognition results will change frequently. The trend of short-time energy change can also reflect the essential characteristics of the audio, and the energy of the speech signal presents a regular fluctuation, while the energy of the noise changes more chaotically. Combined with the differences in the effects of the damping effect on different frequency bands and the characteristics of the more severe attenuation of high frequency bands, the user's voice command and environmental noise can be accurately distinguished, ensuring that the AI toy can correctly respond to the user's interactive needs.
[0054] Step S106, quality detection is performed on the discrimination results of the user's voice command and the environmental noise, the sensor array signal fusion is optimized, and the optimized AI toy command recognition result is obtained.
[0055] The quality of the discrimination result of the user voice instruction and the environmental noise is detected, a confidence is obtained by calculating the proportion of the category with the highest frequency of occurrence in the recognition result in a continuous time window, and a recognition stability is calculated by calculating the variance of the confidence values of multiple windows. If the variance exceeds a preset threshold, it is marked as a low-quality recognition segment, and a recognition quality score of each time segment is obtained. According to the correspondence between the recognition quality score and the local pressure peak value distribution position, a Kalman filtering algorithm is used to estimate the noise covariance of the sensor signal, the weight initial value of each sensor channel is calculated based on the reciprocal of the covariance, the weight is normalized and adjusted combined with the filler density change rate, and the dynamic weight coefficient of the sensor array is determined. Based on the dynamic weight coefficient, the fusion signal is obtained by weighted sum of the audio signals of each sensor, the attenuation amount of each frequency band is determined according to the material damping coefficient calculated in the foregoing, and the material damping coefficient and the attenuation amount of each frequency band have a preset corresponding relationship. The gain compensation opposite to the attenuation amount is applied to the fusion signal to generate the audio signal after spectral compensation. The audio signal after spectral compensation is used to obtain a new recognition result by matching with the standard voice instruction feature, if the new recognition result is consistent with the original recognition result, the original judgment is maintained, if the new recognition result is not consistent with the original recognition result, the result with higher matching degree is selected, and an optimized AI toy instruction recognition result is obtained.
[0056] Specifically, the recognition quality detection evaluates the reliability of the recognition result through statistical methods. In a continuous time window, the system will recognize the same audio multiple times, and each window may contain 5-10 recognition results. By calculating the proportion of the frequency of occurrence of a certain category in the total number of recognition times, the confidence value of the window is obtained.
[0057] For example, in 10 recognitions, 8 are recognized as "play music" instructions, and the confidence is 0.8. When the confidence values of multiple windows fluctuate greatly, it indicates that the recognition result is unstable, and may be affected by environmental interference or signal quality problems.
[0058] In one possible implementation, the Kalman filtering algorithm plays an important role in sensor signal processing. This algorithm recursively estimates the statistical properties of system states and measurement noise by establishing a state space model. For a sensor array, the noise levels of each sensor are different, and the degree of influence of the pressure distribution is also different. By estimating the noise covariance of each sensor through Kalman filtering, the signal quality of each sensor can be quantified. The smaller the covariance, the more reliable the signal of the sensor, and a higher weight should be given. This weight allocation method based on signal quality can adaptively optimize the multi-sensor fusion effect.
[0059] It should be noted that the weight normalization adjustment process takes into account the impact of the change in the density of the filling. When the rate of change of the filling density in a certain area is large, the acoustic characteristics of this area are unstable, and the weight of the corresponding sensor should be appropriately reduced. Normalization processing ensures that the sum of all weights is 1, maintaining the energy balance of the fused signal. Through this dynamic adjustment, the system can optimize the signal acquisition strategy according to the real-time state of the toy.
[0060] In one embodiment, the implementation of spectral compensation is based on the pre-calculated material damping coefficient. Different frequency bands are affected by damping to different degrees, and generally high-frequency signals decay more severely. In the compensation process, the system designs a corresponding gain curve according to the attenuation amount of each frequency band.
[0061] For example, if the 1000Hz frequency band attenuates by 5 decibels, a gain of 5 decibels is applied in that frequency band; if the 2000Hz frequency band attenuates by 8 decibels, a gain of 8 decibels is applied. This frequency-selective compensation method can accurately restore the spectral characteristics of the original speech and improve the accuracy of subsequent recognition.
[0062] It should be noted that the final recognition result optimization uses a comparison verification mechanism. The compensated audio signal is re-extracted and pattern-matched to obtain a new recognition result. The system compares the matching degree values of the two recognitions before and after compensation, and selects the result with higher matching degree as the final output. This double recognition mechanism not only improves the accuracy of recognition, but also verifies the effectiveness of the compensation process. When the two recognition results are consistent, it means that the original recognition is already reliable enough; when the results are different, the recognition after compensation is usually more accurate because it eliminates the interference of the damping effect.
[0063] Step S107, adjusting the action of the AI plush pet according to the optimized AI toy instruction recognition result to obtain a dynamic response output consistent with the user's interactive force.
[0064] According to the optimized AI toy instruction recognition result, the instruction type code and confidence value are extracted, the main direction of user interaction is determined in combination with the compression directionality index, the preset mapping relationship between the instruction code and the corresponding action type, amplitude and speed parameters is queried to obtain a basic action parameter set containing an action type identifier and an initial amplitude value. Based on the initial amplitude value in the basic action parameter set and the material hardness change value, a correction coefficient of the head rotation angle is calculated by the ratio of the hardness change value to the preset reference value, a scaling ratio of the limb swing amplitude is calculated using an inverse proportional relationship, and if the hardness increases by more than a preset threshold, the action amplitude is multiplied by a safety factor to determine the adjusted action control parameter. The adjusted action control parameter is used to control the rotation angle of the servo motor through the duty cycle of the pulse width modulation signal to realize head rotation and limb swing, and the volume gain and tone offset are linearly mapped according to the interaction force value to generate a sound output signal matched with the pressure perception. The mechanical action control signal is sent synchronously with the sound output signal after a delay compensation time, and the compensation time is determined according to the signal transmission delay calculated based on the aforementioned damping coefficient, to obtain a dynamic response output that meets the user interaction force.
[0065] In a possible implementation, the mapping relationship between the instruction code and the action parameter is constructed based on a pre-designed interaction logic. Each voice instruction corresponds to a unique code, for example, the "look left" instruction code is 001, and the corresponding action parameters include a 45-degree left turn of the head and a rotation speed of 30 degrees per second. Such a mapping table is stored in the flash memory of the toy and contains the correspondence between all identifiable instructions and corresponding actions. The introduction of the compression directionality index enables the system to perceive the spatial characteristics of user interaction, and when the user touches the toy from the left side, the action instruction to the left is preferentially responded, improving the naturalness of interaction.
[0066] It should be noted that the influence of material hardness change on action execution is reflected in multiple aspects. The filling of a new toy is relatively soft, and the servo motor can easily drive the mechanical structure to complete the preset action. As the use time increases, the filling gradually compacts, the material hardness increases, and the action amplitude will decrease under the same motor torque. By establishing a mathematical relationship between the hardness change value and the action correction coefficient, the control parameter can be dynamically adjusted.
[0067] For example, when the hardness increases by 20%, the correction coefficient of the head rotation is 0.85, that is, the original 45-degree rotation is adjusted to about 38 degrees, ensuring the smoothness of action execution.
[0068] Specifically, the principle of pulse width modulation signal controlling the servo motor is to adjust the average voltage of the motor by changing the duty cycle of the square wave signal. The greater the duty cycle, the higher the average power obtained by the motor, and the greater the rotation angle. In the AI toy, the controller generates a corresponding PWM signal according to the adjusted action parameter.
[0069] For example, to achieve a 38-degree head rotation, the system calculates the corresponding duty cycle to be approximately 42% and continuously outputs a square wave signal with that duty cycle until the motor reaches the target position.
[0070] In one embodiment, the mapping between interaction force and sound response embodies the concept of emotional design. A gentle touch corresponds to a mild volume and low pitch, conveying the toy's docile nature; a forceful hug triggers a larger volume and lively pitch change, representing the toy's excited state. The volume gain is calculated using a linear mapping, with pressure values ranging from 0-100 Pascals corresponding to a volume adjustment range of 0-20 dB. Pitch shift is achieved by changing the audio sampling rate; for every 10 Pascal increase in pressure, the pitch rises by approximately 50 Hz.
[0071] It should be noted that the delay compensation mechanism in the synchronous signal transmission solves the timing matching problem between mechanical and acoustic responses. Because mechanical components have inertia, the execution of actions requires a certain amount of time, while sound can be played instantly. Based on the damping coefficient, the response delay of the mechanical action can be predicted.
[0072]
[0073] , t comp ζ represents the compensation time, ζ represents the material damping coefficient, L represents the signal transmission distance, and ω represents the signal transmission distance. n This represents the system's natural frequency. This formula is used to calculate the compensation time required to delay the mechanical motion control signal based on the material's damping coefficient.
[0074] For example, in high-damped mode, the head rotation delay is approximately 200 milliseconds. Sending the sound output signal with the same delay ensures that the user hears the corresponding sound as soon as they see the toy turn, creating a consistent and coordinated interactive experience. This precise timing control makes the AI toy's responses more vivid and natural, enhancing the user's emotional connection.
[0075] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for evaluating the fusion of toy sound perception and speech recognition, characterized in that, The method includes: The local pressure peak distribution and deformation distribution when the user touches, hugs, or presses are obtained. The influence of the difference in the degree of plush fiber compression caused by the change in the user's interaction force on the sound propagation path is analyzed, and the spatial distribution data of plush fiber compression is obtained. The compression directionality index is calculated based on the spatial distribution data of plush fiber compression. By analyzing the density change rate and local density peaks of the internal stuffing in AI toys, the non-uniform distribution pattern of stuffing density caused by the change in damping effect is matched, and the quantitative value of the density change of stuffing is obtained. The influence of the sound resonance frequency drift caused by the uneven distribution of filler density on the speech signal spectrum was analyzed, and the frequency drift range was determined. The sensor signal strength is dynamically adjusted to adjust the frequency drift range, and the enhanced voice signal data is generated to determine the distinction between user voice commands and environmental noise. The quality of the user's voice command and environmental noise discrimination results is tested, and the sensor array signal fusion is optimized to obtain optimized AI toy command recognition results; Based on the optimized AI toy command recognition results, the movements of the AI toy plush pet are adjusted to obtain a dynamic response output that matches the intensity of user interaction.
2. The toy sound perception and speech recognition fusion evaluation method according to claim 1, characterized in that, The process involves acquiring the local pressure peak distribution and deformation distribution during user touch, hugging, and pressing, analyzing the impact of differences in plush fiber compression caused by changes in user interaction force on the sound propagation path, and obtaining spatial distribution data of plush fiber compression, including: The system acquires pressure values at each sensor point during user interaction, calculates local pressure gradients based on pressure differences between adjacent sensor points, and calculates deformation at each location using the elastic modulus parameter of the plush material, constructing a three-dimensional deformation distribution map. Based on this map, the ratio of deformation at each location to the original plush thickness is calculated to obtain the compression ratio. The spatial dispersion of the compression ratio is calculated using the standard deviation, generating a compression uniformity coefficient distribution map. Based on this map, the ratio of fiber density in each region to the original density is calculated, and the rate of change of sound wave propagation velocity is determined by considering changes in fiber orientation. Based on the rate of change of sound wave propagation velocity and the sensor point coordinates, a compression degree distribution field is generated through spatial interpolation to determine the compression degree level at each spatial location, obtaining the spatial distribution data of the plush fiber compression.
3. The toy sound perception and speech recognition fusion evaluation method according to claim 1, characterized in that, The calculation of the compression directionality index based on the spatial distribution data of the compressed plush fibers includes: Based on the spatial distribution data of plush fiber compression, the density gradient value is obtained by dividing the fiber density difference between adjacent positions by the distance. The components of compression in each direction in three-dimensional space are calculated, the direction of the maximum component is determined as the main compression direction, and the cosine value of the angle between the compression direction and the main direction at each position is calculated as the compression directionality index.
4. The toy sound perception and speech recognition fusion evaluation method according to claim 1, characterized in that, The analysis of the density change rate and local density peaks of the internal filling material in the AI toy was used to match the non-uniform distribution pattern of the filling material density caused by the damping effect, and to obtain a quantitative value of the filling material density change, including: The density change value is obtained by calculating the proportion of the change in filler volume. The density change rate is determined by the ratio of the density difference between adjacent regions to the distance between them. Locations where the density change rate exceeds a threshold are marked as local density peak points. Based on the density change rate and the local density peak points, a three-dimensional distribution feature vector is constructed to generate a digital coding sequence. The digital coding sequence is searched in a preset database, and the similarity with the standard code is calculated to extract the density distribution parameter set. Based on the density distribution parameter set, the filler density value at each spatial location is calculated, and the difference is made with the standard density value to obtain the quantified value of the filler density change.
5. The toy sound perception and speech recognition fusion evaluation method according to claim 1, characterized in that, The analysis examines the impact of the sound resonance frequency drift caused by the uneven distribution of filler density on the speech signal spectrum, determining the frequency drift range, including: Based on the quantized value of the density change of the filler, the density gradient vector is calculated to determine the degree of sound wave deflection and the local resonant frequency value is calculated. Based on the difference between the local resonant frequency value and the initial resonant frequency, the resonant frequency drift value is obtained, and the frequency component change amplitude is identified through spectral decomposition. Based on the frequency component change amplitude, the degree of spectral distortion is calculated, and a spectral influence distribution map is generated. Based on the spectral influence distribution map, the drift value range is statistically analyzed to determine the frequency drift dynamic range.
6. The toy sound perception and speech recognition fusion evaluation method according to claim 1, characterized in that, The process of dynamically adjusting the sensor signal strength to address the frequency drift range, generating enhanced voice signal data, and determining the distinction between user voice commands and environmental noise includes: Based on the frequency drift dynamic range, the density recovery rate is calculated to determine the hardness change pattern. The sensor gain adjustment coefficient is calculated by multiplying the recovery rate by the hardness increment to generate the enhanced speech signal data. Based on the enhanced speech signal data, speech feature vectors are extracted, and the low-frequency components are corrected by combining the filling density change rate to adjust the speech recognition decision threshold. Based on the decision threshold, the matching degree between the audio feature vector and the standard command feature vector is calculated to distinguish between user voice commands and environmental noise.
7. The toy sound perception and speech recognition fusion evaluation method according to claim 1, characterized in that, The process of performing quality checks on the distinction between user voice commands and environmental noise, optimizing sensor array signal fusion, and obtaining optimized AI toy command recognition results includes: Based on the distinction between the user's voice command and environmental noise, the proportion of recognized categories within a time window is calculated, recognition stability is evaluated, and a recognition quality score is generated. Based on the recognition quality score and the pressure distribution, the sensor signal noise covariance is estimated, initial weight values are calculated, and the weight coefficients are adjusted in conjunction with the stuffing density change rate. Based on the weight coefficients, the sensor audio signals are weighted and summed to generate a fused signal, gain compensation is applied, and a spectrum-compensated audio signal is generated. Based on the spectrum-compensated audio signal, it is matched with standard command features to obtain the optimized AI toy command recognition result.
8. The toy sound perception and speech recognition fusion evaluation method according to claim 1, characterized in that, The step of adjusting the movements of the AI toy plush pet based on the optimized AI toy command recognition results to obtain a dynamic response output that matches the user's interaction intensity includes: Based on the optimized AI toy instruction recognition results, the instruction type code is extracted, and the interaction direction is determined by combining the compression directionality. A preset mapping relationship is queried to obtain the action parameter set. Based on the action parameter set and the material hardness change value, the head rotation angle correction coefficient and the limb swing amplitude ratio are calculated to determine the action control parameters. Based on the action control parameters, the rotation angle of the servo motor is controlled to generate a sound output signal. The mechanical action signal is synchronized with the sound output signal to obtain the dynamic response output.