Voice song requesting interaction data processing method and system based on multi-mode fusion
Through multimodal fusion technology, ambient sound signals and semantic words are dynamically captured, which solves the problem of matching degree calculation of the voice tuning system in complex noise environments, and realizes accurate track push in low signal-to-noise ratio scenarios.
Patent Information
- Application Number
- CN202510765348.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing voice-oriented interactive system has reduced the accuracy of speech signal quality and semantic recognition in complex noise environments, and the matching degree calculation method lacks dynamic adjustment, resulting in mispushing or omissions, and the real matching set of multiple candidate tracks cannot be accurately identified.
Through multimodal fusion technology, the ambient sound signals are dynamically captured, combined with environmental factors and semantic words, dynamically adjust the matching threshold and weights, build timing association and environment perception mechanisms, and optimize the matching degree calculation and push strategy.
In complex noise environments, users' sound field characteristics are accurately captured and matching thresholds are dynamically adjusted, which solves the matching degree distortion problem caused by noise interference, and realizes intelligently optimized track push.
Smart Images

Figure CN120279870A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a voice song ordering interaction data processing method and system based on multimodal fusion. Background Art
[0002] With the rapid development of intelligent voice interaction technology, voice-on-demand song interaction systems have been widely used in scenarios such as home entertainment and learning, smart home, and in-vehicle entertainment. However, in actual deployment of such systems, there are adaptability problems at the data level: First, the anti-interference ability against the acoustic environment is weak. When the user is in a scene with complex background noise (such as a noisy street, a construction site, or a party environment), the quality of the voice signal and the accuracy of semantic recognition significantly decrease, resulting in deviations in the extraction of the core features for track matching. Measured data shows that when the environmental noise exceeds 75 dB, the Top1 push accuracy after existing data processing will drop sharply from 89% in a quiet environment to 52%, and the wrongly pushed tracks often have a semantic gap with the user's intention. Second, existing matching degree calculation methods mostly rely on fixed thresholds or static weight assignments, and it is difficult to dynamically adjust the screening strategy according to the real-time environment. For example, in a high-noise environment, the difference in discrimination between the true matching target and the secondary candidates will be non-uniformly suppressed or amplified, leading to false pushes or omission of valid targets. For example, the correlation between the environmental noise intensity and the user's semantic clarity is not considered, further resulting in a decrease in push accuracy in a low signal-to-noise ratio scenario. Finally, existing solutions lack dynamic quantification means for the recognition of the matching degree selection set - since the fluctuations in the matching degree sequence caused by noise may mask the true matching break, when the matching values of multiple candidate tracks are close, the existing fixed difference threshold division method (for example, if the difference between the top two is > 0.3, it is determined as a break) will cause the system to frequently misjudge in a high-fluctuation environment, and in some special voice input conditions (such as the weak correlation between the user's vague expression and the lyrics text), the conventional matching difference calculation method cannot accurately identify the range of the candidate set to be pushed. Summary of the Invention
[0003] Aiming at the defects in the prior art, the present invention provides a voice song ordering interaction data processing method and system based on multimodal fusion.
[0004] A voice song ordering interaction data processing method based on multimodal fusion, comprising: receiving a song ordering voice signal and obtaining the processing time point when the song ordering voice signal is received, obtaining a semantic word according to the song ordering voice signal, obtaining a processing time period with the end time point being the processing time point, obtaining an ambient sound signal recorded within the processing time period, and obtaining an influence factor according to the ambient sound signal; obtaining the matching degree corresponding to each song according to the influence factor, the semantic word, and the word segmentation of multiple songs in the song library, obtaining multiple matching degrees exceeding the matching threshold, arranging the multiple matching degrees in descending order to form a pre-matching degree sequence; obtaining the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtaining a difference threshold according to the influence factor, and obtaining a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; obtaining an interaction push sequence according to the songs corresponding to the target matching degree sequence, and pushing the songs according to the interaction push sequence.
[0005] Optionally, obtaining a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence includes: obtaining adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; dividing the pre-matching degree sequence according to one or more division positions to form multiple sub-matching degree sequences; obtaining the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.
[0006] Optionally, obtaining an influence factor according to the ambient sound signal includes: presetting multiple noise intensity ranges, and predefined that multiple noise intensity ranges respectively correspond to different influence factors; obtaining the ambient intensity according to the ambient sound signal, and obtaining the corresponding influence factor according to the noise intensity range in which the ambient intensity falls.
[0007] Optionally, obtaining a difference threshold according to the influence factor includes: obtaining an environmental sensitivity coefficient, and obtaining the maximum value of the matching differences between adjacent matching degrees in the pre-matching degree sequence; obtaining a difference threshold according to the influence factor, the environmental sensitivity coefficient, and the maximum value of the matching differences between adjacent matching degrees in the pre-matching degree sequence.
[0008] Optionally, obtaining a difference threshold according to the influence factor is expressed as: ; where is the difference threshold, is the matching difference between the j-th group of adjacent matching degrees in the pre-matching degree sequence, is the number of different adjacent matching degrees in the pre-matching degree sequence, is the influence factor, is the environmental sensitivity coefficient.
[0009] Optionally, obtaining the matching degree corresponding to each track according to the impact factor, semantic word, and word segmentation of multiple tracks in the music library includes: obtaining the maximum fuzzy increment coefficient; obtaining the matching degree corresponding to each track according to the maximum fuzzy increment coefficient, impact factor, semantic word, and word segmentation of multiple tracks in the music library.
[0010] Optionally, the matching degree corresponding to each track obtained according to the impact factor, semantic word, and word segmentation of multiple tracks in the music library is expressed as: ; where is the matching degree corresponding to the i-th track, is the maximum fuzzy increment coefficient, is the impact factor, is the semantic word, is the word segmentation corresponding to the i-th track, is the number of words in the semantic word, is the number of words in the word segmentation corresponding to the i-th track.
[0011] There is also provided a voice song-request interaction data processing system based on multimodal fusion. The system includes: a data interaction module, configured to receive a song-request voice signal, obtain the processing time point when the song-request voice signal is received, obtain the semantic word according to the song-request voice signal, obtain the processing time period with the end time point being the processing time point, obtain the ambient sound signal recorded within the processing time period, and obtain the impact factor according to the ambient sound signal; a data processing module, configured to obtain the matching degree corresponding to each track according to the impact factor, semantic word, and word segmentation of multiple tracks in the music library, obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order to form a pre-matching degree sequence; a data screening module, configured to obtain the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtain the difference threshold according to the impact factor, and obtain the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; a data pushing module, configured to obtain an interaction pushing sequence according to the tracks corresponding to the target matching degree sequence, and push the tracks according to the interaction pushing sequence.
[0012] Optionally, the data screening module is further configured to: obtain adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; divide the pre-matching degree sequence according to one or more division positions to form multiple sub-matching degree sequences; obtain the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.
[0013] Optionally, the data interaction module is further configured to: preset multiple noise intensity ranges, and pre-define that different noise intensity ranges correspond to different impact factors; obtain the ambient intensity according to the ambient sound signal, and obtain the corresponding impact factor according to the noise intensity range into which the ambient intensity falls.
[0014] The beneficial effects of the present invention are embodied in: In the entire voice song-ordering interactive data processing method based on multimodal fusion, first of all, the dynamic perception technology of ambient sound based on time series association can accurately capture the real sound field characteristics when the user speaks, and overcome the defect of the traditional fixed threshold's delayed response to transient interference through decibel grading mapping and burst noise weight adjustment, thereby ensuring the spatiotemporal consistency of environmental factor calculation; further, the dynamic coupling of ambient noise and semantic understanding can expand the semantic association range through fuzzy incremental coefficients in low signal-to-noise ratio scenarios, which can not only strengthen the weight of core keywords (such as singer names, song titles), but also be compatible with the fault-tolerant matching of synonyms and fuzzy expressions, effectively compensating for the feature deviation caused by speech recognition errors; further, the dynamic calculation of the difference threshold that links the environmental sensitivity coefficient with the noise spectrum characteristics is introduced, and the matching fault judgment boundary is adaptively adjusted to solve the problem of matching degree sequence distortion caused by noise suppression; further, the push strategy can combine user portraits with real-time environmental status to achieve intelligent optimization of push results. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.
[0016] Figure 1 A schematic diagram of the steps of a method for processing interactive data of song requesting by voice based on multimodal fusion in one embodiment of the present invention; Figure 2 It is a schematic diagram of a part of the steps of S4 in the voice song request interactive data processing method based on multimodal fusion of the present invention; Figure 3 It is a schematic diagram of a part of the steps of S1 in the voice song request interactive data processing method based on multimodal fusion of the present invention; Figure 4 It is a schematic diagram of a part of the steps of S3 in the voice song request interactive data processing method based on multimodal fusion of the present invention; Figure 5 It is a schematic diagram of a part of the steps of S2 in the voice song-ordering interactive data processing method based on multimodal fusion of the present invention. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0018] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0019] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.
[0020] As Figure 1 shown, a voice song-request interaction data processing method based on multimodal fusion is provided, including: S1. Receive a song-request voice signal, obtain the processing time point when the song-request voice signal is received, obtain the semantic word according to the song-request voice signal, obtain the processing time period with the end time point being the processing time point, obtain the ambient sound signal recorded within the processing time period, and obtain the influence factor according to the ambient sound signal; S2. Obtain the matching degree corresponding to each song according to the influence factor, the semantic word, and the word segmentation of multiple songs in the song library, obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order and form a pre-matching degree sequence; S3. Obtain the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtain the difference threshold according to the influence factor, and obtain the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; S4. Obtain an interactive push sequence according to the songs corresponding to the target matching degree sequence, and push the songs according to the interactive push sequence.
[0021] In this embodiment, it should be noted that in S1, by dynamically capturing the real-time acoustic environment characteristics during voice interaction, a multi-modal data association is constructed to provide an environmental perception basis for subsequent semantic matching and threshold decision-making. First, when receiving a user voice signal, the processing time point of the instruction is accurately recorded, and the ambient sound data of the historical time period with this time point as the end is intercepted based on the sequential logic. This selection mechanism for the processing time period fully considers the real-time requirements of voice interaction. For example, when the user says "Play ABC of XYZ", the ambient audio segment within 1-2 seconds before the end of the instruction is retrieved (this duration can be optimized and adjusted according to the actual scenario), rather than statically setting a fixed time window, so as to accurately reflect the real sound field state at the moment when the user speaks. Secondly, for the ambient sound signal within this time period, a decibel intensity grading mapping mechanism is adopted to discretize the continuous ambient noise intensity into multiple preset quantization intervals (such as low noise, medium noise, high noise and other scenario categories), and each interval corresponds to a predefined environmental impact factor. For example, in a home scenario, the ambient sound may trigger a medium-level impact factor due to the sound intensity of music, TV, conversation or kitchen noise; while the high-decibel noise generated in a vehicle scenario can be mapped to a high-order impact factor. This grading mechanism effectively avoids the problem of insensitivity of traditional fixed thresholds to environmental mutations.
[0022] Furthermore, by combining sequential association with environmental quantization, it is possible to dynamically perceive the acoustic context in which the voice instruction is located. For example, when the user orders a song in a coffee shop, the compound noises such as the clinking of cups and plates and the conversations of people that come and go during the processing time period are captured, and the ambient sound intensity level is determined through spectrum analysis and energy integration. If a sudden high-decibel noise (such as the start of a coffee machine) appears during this time period, the weight of the impact factor is automatically adjusted according to the proportion of the duration of this noise, rather than simply taking the average value. This processing method can not only reflect the instantaneous interference intensity, but also avoid the excessive impact of occasional peak noises on the overall environmental assessment. At the same time, the acquisition of the ambient sound signal adopts multi-channel fusion technology, combined with the spatial filtering characteristics of the device microphone array, to effectively separate the user voice and the ambient noise, ensuring the accuracy of subsequent impact factor calculations.
[0023] In S2, a dynamic weight mechanism is established to quantify the impact of environmental noise on semantic understanding as a regulatory parameter for match calculation, achieving elastic matching in noisy scenarios. First, the basic matching value is calculated through the text correlation between semantic words and the segmented words in the music library. For example, when the user says "Play ABC by XYZ", semantic words such as "Play", "XYZ", and "ABC" will be extracted and compared with the segmented words in text fields such as the title and singer of the songs in the music library to calculate the basic vocabulary overlap. In this process, a dynamic amplification factor based on environmental factors is introduced: in a low-noise scenario (such as a home environment), an exact matching strategy is adopted, requiring a high overlap between semantic words and the segmented words of the track; when a high-noise impact factor is detected (such as a vehicle scenario), the matching conditions are appropriately relaxed through a fuzzy increment coefficient, allowing candidates with partial semantic associations but not exact matches to enter the preselection sequence. For example, when the user says "Play DEF", it may be recognized as "Play DEG", and at this time, the weights of near synonyms such as "DE" in the track lyrics will be automatically increased according to the noise level, so that "DEF" can still obtain a high matching degree through extended semantic associations.
[0024] Furthermore, a collaborative mechanism for environmental perception and semantic understanding is constructed to eliminate the interference of matching degree fluctuations caused by noise through the threshold decision of environmental perception. The calculation dimension of the matching degree is dynamically adjusted according to the real-time noise intensity: in a quiet environment, full-text matching is emphasized, such as considering features such as the integrity of the song; while in a high-noise scenario, enhanced matching of core keywords is focused on. For example, when the environmental factor reaches the threshold, the weights of overlapping keyword fields such as the singer name and song title will be amplified. This dynamic strategy not only retains the accuracy of basic text matching but also compensates for the semantic loss caused by noise through the environmental perception mechanism. For example, when the user vaguely expresses "That HIJKL" in a party environment, the matching degree weights of song titles containing "HI" or "KL" will be amplified during the matching calculation through the impact factor, so that the matching degrees of songs containing "HI" or "KL" are at the same level as those of songs containing "HIKL". Finally, multiple matching degrees are sorted in descending order to form a pre-matching degree sequence, providing data support for subsequent threshold decisions.
[0025] In S3, a dynamic coupling mechanism between environmental noise and matching fault recognition is established to eliminate the distortion of the matching degree sequence caused by noise interference through an adaptive difference threshold. First, analyze the distribution of the matching degree differences between adjacent candidate tracks in the pre-matching degree sequence. For example, when a user requests "MNOPQ" at a construction site, environmental noise may compress the matching degree difference between the true matching track and similar candidates (such as "MN", "MNOP") from 0.4 in a quiet environment to 0.15. The traditional fixed threshold of 0.3 will be misjudged as no fault matching in this scenario, resulting in incorrect pushing. This solution dynamically adjusts the difference threshold through environmental factors: when noise above 75dB is detected, the threshold is automatically reduced from 0.3 to 0.12 based on the noise intensity coefficient, enabling the compressed true fault (0.15 difference) to trigger an effective division and accurately lock the target track.
[0026] Furthermore, the collaborative optimization of noise characteristics and threshold sensitivity. Specifically, a dual adjustment mechanism is constructed by combining environmental sensitivity and the fluctuation characteristics of the matching degree: in a continuous low-frequency noise scenario (such as a vehicle environment), a progressive threshold attenuation strategy can be adopted, making the adjustment amplitude of the difference threshold positively correlated with the noise duration; for sudden impulse noise (such as the sound of tableware dropping), an instantaneous response mode can be activated to dynamically correct the threshold through the spectral energy distribution of the impulse noise.
[0027] In S4, a dynamic candidate set pushing mechanism is constructed to optimize the final pushing result through matching fault recognition and context association. First, arrange the tracks in the target matching degree sequence in descending order of the matching degree. For example, when a user requests "ST" in a high-noise environment, the target sequence may contain a matching distribution of "ST" 0.82, "AT" 0.79, "SA" 0.78, "DE" 0.41. Traditionally, the Top1 would be directly pushed, but in a scenario where the matching degree is compressed by noise, it will be detected that the differences among the top three are all less than the dynamic threshold (such as 0.05), and it is determined that there is no significant fault. Then, the top three are packaged to form an interactive pushing sequence, and through the secondary confirmation mode of "Do you want to listen to 'ST', 'AT' or 'SA'", both incorrect pushing is avoided and the candidate range of the user's true intention is retained.
[0028] Furthermore, the pushing method is dynamically selected according to the structural characteristics of the target sequence: when a significant fault is detected (such as the first-place matching degree fault leading the second place by more than 0.3), the optimal track is directly played; if the sequence shows a multi-peak distribution (such as "ABCD" 0.75, "ABC" 0.73, "BCD" 0.71, "CDE" 0.68), then all are pushed, or semantic enhancement is performed in combination with the user's historical request records, and the fault subset with a higher degree of fit to the user's preferences is preferentially pushed. For example, when it is detected that the user often requests XYZ songs, even if the matching degree of "AB" is slightly lower, it will still be promoted to the front of the push due to its strong relevance to the user profile.
[0029] To sum up, in the entire voice song-ordering interactive data processing method based on multimodal fusion, first of all, the dynamic perception technology of ambient sound based on time series association can accurately capture the real sound field characteristics when the user speaks, and overcomes the defect of the traditional fixed threshold's delayed response to transient interference through decibel grading mapping and burst noise weight adjustment, ensuring the spatiotemporal consistency of environmental factor calculation; further, the dynamic coupling of ambient noise and semantic understanding expands the semantic association range through the fuzzy incremental coefficient in low signal-to-noise ratio scenarios, which can not only strengthen the weight of core keywords (such as singer name, song title), but also be compatible with the fault-tolerant matching of synonyms and fuzzy expressions, effectively compensating for the feature deviation caused by speech recognition errors; further, the dynamic calculation of the difference threshold linked to the environmental sensitivity coefficient and the noise spectrum feature is introduced, and the matching fault judgment boundary is adaptively adjusted to solve the problem of matching degree sequence distortion caused by noise suppression; further, the push strategy can combine user portraits with real-time environmental status to achieve intelligent optimization of push results.
[0030] like Figure 2 As shown, in one embodiment, obtaining the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence in S4 includes: S41, obtaining adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold and using them as the division position; S42, dividing the pre-matching degree sequence according to one or more division positions to form a plurality of sub-matching degree sequences; S43: Obtain a sub-matching degree sequence containing the maximum matching degree and use it as a target matching degree sequence.
[0031] In this embodiment, it should be noted that in S41, significant fault boundaries in the matching sequence are dynamically identified. The matching differences of all adjacent candidate tracks in the pre-matching sequence are traversed, and when it is detected that a group of adjacent matching differences exceeds the dynamically calculated difference threshold, the position is marked as a division point. For example, when a user orders "MNOPQ" at a construction site, the environmental noise causes the matching sequence to present a distribution of "MNOP" 0.78, "NOP" 0.62, and "OPQ" 0.60. After the threshold is reduced from 0.3 to 0.15 based on the noise intensity, it is detected that the first group of differences of 0.16 exceeds the threshold, that is, a division point is established between the first and second candidates, accurately isolating the real target from other interference items.
[0032] In S42, the candidate set topology is reconstructed through division points to form multiple subsequences with internal consistency. Taking the previous example, at the first division point, the sequence is split into two subsequences: the first subsequence contains only the highest matching degree of "MNOPQ", and the second subsequence contains sub-optimal candidates such as "NOP" and "OPQ". This division mechanism effectively distinguishes the candidate groups before and after the fault. For example, when the user vaguely expresses "play fast songs of XYZ", the matching degree sequence may form a distribution of "XYZ Rap Collection" 0.75, "XYZ Slow Songs" 0.71, "XYZ Live Version" 0.70, "XYZ Collection" 0.45. A difference fault is detected between the third and fourth candidates, and the sequence is divided into a high matching degree subset and a low matching degree subset.
[0033] In S43, accurate push is achieved through the intelligent selection of the subsequence with the highest matching degree. The subsequence containing the highest matching degree is preferentially selected as the final push set. If this subsequence contains multiple candidates (such as when there is no fault), the complete sequence is retained for subsequent processing. For example, when the user requests "BC" in a vehicle scenario, due to wind noise interference, the matching degree sequence may present a dense distribution of "ABCD" 0.68, "ABC" 0.67, "BCD" 0.66. At this time, it is determined that there is no significant fault, and the complete sequence is passed to S4 for multi-candidate interactive push. This mechanism can quickly lock the target when the fault is clear, and retain semantically related candidates in fuzzy scenarios. For example, when the user misspoke and said "play ABCXYZ", related tracks such as "ABC" and "XYZ Collection" are retained through the complete sequence to improve the fault tolerance ability.
[0034] As Figure 3 shown, in one embodiment, the influencing factors obtained according to the ambient sound signal in S1 include: S11, preset multiple noise intensity ranges, and predefined that multiple noise intensity ranges respectively correspond to different influencing factors; S12, obtain the ambient intensity according to the ambient sound signal, and obtain the corresponding influencing factor according to the noise intensity range in which the ambient intensity falls.
[0035] In this embodiment, it should be noted that in S11, a hierarchical quantization system for the acoustic environment intensity is established, and the standardized mapping of environmental characteristics is realized by presetting multiple noise intensity intervals. According to the acoustic characteristics of different application scenarios (such as home, vehicle, public places), the continuous decibel values are divided into discrete intervals with clear semantics. For example, 20 - 50 dB is divided into a low noise level corresponding to the home scenario, and the corresponding influence factor is 0.3; 50 - 80 dB is divided into a medium noise level corresponding to the coffee shop environment, and the corresponding influence factor is 0.6; above 80 dB is divided into a high noise level corresponding to the construction site, and the corresponding influence factor is 0.9. It should be emphasized that according to the application in the subsequent expressions, the influence factor can only be allocated within the range of 0 to 1. At the same time, 0 represents completely no sound, and 1 represents the extreme noise environment. Each interval is associated with a predefined environmental influence factor, forming a gradient response mechanism from quiet to extreme noise. For example, in the smart home scenario, the 70 dB noise when the vacuum cleaner is working may be classified as a medium noise level, triggering a medium influence factor, while the knocking sound on the window during a heavy rainstorm may be dynamically adjusted to a higher noise level due to large fluctuations in the decibel value.
[0036] In S12, the time - frequency domain joint analysis of the environmental sound signal is carried out through the existing technology to calculate the average sound pressure decibel within a specific time period. In some special cases, for example, when the user orders a song by the street, the continuous traffic noise will be recognized as the reference noise level, and the sudden car horn sound will trigger the transient detection module: if the honking duration exceeds the preset ratio (such as accounting for 30% of the analysis period), the influence factor level will be automatically increased; if it is a short - term interference (such as only accounting for 5%), the reference noise level assessment will be maintained. Capture the main characteristics of the environmental noise to avoid evaluation deviation caused by instantaneous spike interference.
[0037] As Figure 4 shown, in one embodiment, obtaining the difference threshold according to the influence factor in S3 includes: S31. Obtain the environmental sensitivity coefficient and obtain the maximum value of the matching difference between adjacent matching degrees in the pre - matching degree sequence; S32. Obtain the difference threshold according to the influence factor, the environmental sensitivity coefficient, and the maximum value of the matching difference between adjacent matching degrees in the pre - matching degree sequence.
[0038] In this embodiment, it should be noted that in S31, an association mechanism between environmental sensitivity and matching degree fluctuation is established, and the response characteristics to acoustic environment changes are reflected through the environmental sensitivity coefficient independent of the noise intensity. The environmental sensitivity coefficient is preset according to the device type and application scenario. For example, vehicles may be set with a higher sensitivity due to the need to adapt to rapidly changing road noise, while home audio systems use a lower sensitivity to maintain stability.
[0039] Further, by traversing the matching degree differences between all adjacent candidate tracks in the pre-matching degree sequence, the maximum fluctuation amplitude is captured. For example, in a KTV scenario, when a user requests "FRTY", background vocal interference may cause the matching degree sequence to present a distribution of "FRTY" 0.75, "RTY" 0.72, and "FRT" 0.70. At this time, the maximum adjacent difference of 0.05 is identified as the basic parameter for subsequent threshold calculation.
[0040] In S32, by fusing the environmental factor, the sensitivity coefficient, and the maximum fluctuation value, a dynamic threshold generation model is constructed. This model makes the threshold decrease as the noise intensity increases, and is simultaneously adjusted by the sensitivity coefficient: high-sensitivity devices trigger a larger threshold reduction amplitude under the same noise, so as to accurately identify the compressed matching fault in the noise interference. This dual adjustment mechanism not only maintains the environmental adaptability of the threshold adjustment but also takes into account the characteristic differences of different devices.
[0041] In one implementation, the difference threshold obtained according to the influencing factor in S3 is expressed as: ; where is the difference threshold, is the matching difference between the j-th group of adjacent matching degrees in the pre-matching degree sequence, is the number of different adjacent matching degrees in the pre-matching degree sequence, is the influencing factor, is the environmental sensitivity coefficient.
[0042] In this implementation, it should be noted that in the entire expression, is the extraction of the dynamic reference value, capturing the maximum difference amplitude of the current matching degree sequence, and using this as the reference for fault judgment; in a quiet environment, semantic matching is clear, and the matching degree difference between adjacent tracks is large (such as 0.4); while in a high-noise scenario, noise interference will cause the matching degree difference to be compressed (such as 0.15); by taking the maximum value, it is ensured that the threshold reference always reflects the actual matching fluctuation level in the current environment.
[0043] Further, is the noise attenuation term, which dynamically reduces the threshold according to the noise intensity; where controls the noise intensity, the greater the noise, is closer to 1, resulting in being smaller, being larger; regulates the sensitivity. High-sensitivity devices (such as in-vehicle, ) under the same noise, is smaller, is also smaller, the difference threshold reduction amplitude is larger, and the fault is more stringent; low-sensitivity devices (such as household, ) Under the same noise, is larger, is also larger, the reduction of the difference threshold is smaller, and the fault is looser.
[0044] Furthermore, when is closer to 0 (no noise), is closer to 0, the difference threshold approaches 0, and a very small difference in matching degree is required to trigger a fault, which means that only the track corresponding to the highest matching degree needs to be pushed. When is closer to 1 (extreme noise), and at the same time , the difference threshold , and the difference threshold is compressed to a small extent.
[0045] For example, Figure 5 As shown, in one embodiment, obtaining the matching degree corresponding to each track according to the influence factor, semantic word, and word segmentation of multiple tracks in the music library in S2 includes: S21. Obtain the maximum fuzzy increment coefficient; S22. Obtain the matching degree corresponding to each track according to the maximum fuzzy increment coefficient, influence factor, semantic word, and word segmentation of multiple tracks in the music library.
[0046] In this embodiment, it should be noted that in S21, the boundary constraint of semantic expansion is set, and the fault tolerance upper limit of the matching degree calculation in the noise environment is controlled by the maximum fuzzy increment coefficient. This coefficient is preset according to the device type and application scenario. For example, in-vehicle devices need to cope with sudden road noise, so a higher maximum fuzzy increment coefficient (such as 1.3) is set to allow a wider semantic association; while home devices use a lower maximum fuzzy increment coefficient (such as 1.1) to prevent misjudgment caused by excessive fuzzy matching. This preset mechanism ensures that both speech recognition errors can be compensated and the semantic span can be prevented from getting out of control in different scenarios. For example, in a concert scenario, the user's true intention "play XYZGH" may be recognized as "play KLO of XYZ". By dynamically adjusting the weight range of near-sounding words such as "GH" and "KLO" in the lyrics through the maximum fuzzy increment coefficient value, the correct track can still enter the candidate set.
[0047] In S22, through the synergistic effect of the influence factor and the maximum fuzzy increment coefficient, the dynamic coupling of the noise intensity and semantic expansion is realized. In a low-noise scenario, the matching degree calculation focuses on exact text coincidence, and only when the semantic word completely matches the track segmentation, a high weight is obtained; as the noise increases, the fuzzy increment is gradually activated, allowing partial matches to improve the score through semantic relevance. For example, in a vehicle-mounted scenario, when the user says "Play ABC" and it is recognized as "Play ABD" due to noise interference, at this time, a high influence factor value triggers the fuzzy increment, so that the tracks in the music library that contain both the lyrics of "ABC" and the homophonic title of "ABD" can all have an increased matching degree. Then, through the maximum fuzzy increment coefficient to constrain a large increase in the matching degree, multiple tracks are successfully recommended, ensuring that the tracks that truly meet the needs will not be eliminated discontinuously.
[0048] In one implementation, in S2, the matching degree corresponding to each track is obtained according to the influence factor, the semantic word, and the segmentation of multiple tracks in the music library, which is expressed as: ; where is the matching degree corresponding to the i-th track, is the maximum fuzzy increment coefficient, is the influence factor, is the semantic word, is the segmentation corresponding to the i-th track, is the number of words in the semantic word, is the number of words in the segmentation corresponding to the i-th track.
[0049] In this implementation, it should be noted that in the whole expression, is the basic matching degree calculation term, which measures the text coincidence ratio between the user's semantic word and the track segmentation, and avoids the false high matching degree caused by a large number of segmentations in a long-text track. Example: In a noisy environment, the voice signal of the song request input by the user is converted into a voice word of "Play BCD of XYZ" (A = {Play, um, XYZ, B, C, D}, |A| = 6), and the title of track 1 is "BCE" - XYZ ( , ), then , and the basic matching degree is ; the title of track 2 is "OABT" - QWE ( , ), then , and the basic matching degree is
[0050] Furthermore, is the fuzzy increment compensation term. Among them, realizes that the greater the noise, the greater the compensation value (such as when ), 1), so that the greater the noise, the more extra points are allowed for the tracks with incomplete semantic matches; for example: when the user says "Play DEF" (recognized as "Play DEG"), the matching degree of the track "DEF" is compensated. Among them, the basic matching degree is an objective quantitative index of semantic association, which is used to measure the exact coincidence degree between the user's semantic words and the track word segmentation. The whole The min function is incorporated to retain the weight of the original matching information during the noise compensation process, and to avoid the matching degree being too high or too low due to complete dependence on the compensation term; for example, assume that when the user says "ABC" it is recognized as "ABE", and the basic matching degree of the actual desired track "ABC" is 0.67. If the compensation term (such as 0.8) is directly used and the basic matching degree is ignored, the matching degree of "ABC" will be lower, and may be much lower than that of other partially matched tracks. However, through the formula design, the final matching degree is: , in this way, the matching degree of this track can be significantly improved. Among them, and the maximum fuzzy increment coefficient The setting of is used to reflect the rigid ceiling of the matching degree increase, limit the maximum allowable range of noise compensation, and prevent the semantic expansion from getting out of control due to excessive environmental noise; just as in the above example, if there is no restriction, then the directly calculated result is , resulting in the matching degree of a track with a low matching degree being immediately over-increased, leading to out-of-control matching.
[0051] In summary, semantic fault tolerance in a noisy environment is achieved. When noise causes speech recognition errors (such as "ABC" being misrecognized as "ABE"), this embodiment can still raise the matching degree of the track with the true intention. Further, it also achieves the prevention of over-fuzzy matching; over-compensation may cause irrelevant tracks to be recommended due to a single-word coincidence (such as "ABEF" being mis-recommended because it contains "ABE"). Further, it also achieves dynamic weight and device adaptation. For example, in vehicle-mounted devices in device differentiation, set , and for home devices, set etc.
[0052] A voice song selection interaction data processing system based on multimodal fusion is also provided. The system includes: A data interaction module, which is used to receive the song selection voice signal, obtain the processing time point when the song selection voice signal is received, obtain the semantic words according to the song selection voice signal, obtain the processing time period with the end time point being the processing time point, obtain the environmental sound signal recorded during the processing time period, and obtain the influencing factor according to the environmental sound signal; A data processing module, configured to obtain the matching degree corresponding to each track according to the impact factor, semantic word, and word segmentation of multiple tracks in the music library, and obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order, and form a pre-matching degree sequence; A data screening module, configured to obtain the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtain a difference threshold according to the impact factor, and obtain a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; A data pushing module, configured to obtain an interactive push sequence according to the tracks corresponding to the target matching degree sequence, and push the tracks according to the interactive push sequence.
[0053] In one embodiment, the data screening module is further configured to: obtain adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; divide the pre-matching degree sequence according to one or more division positions to form multiple sub-matching degree sequences; obtain the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.
[0054] In one embodiment, the data interaction module is further configured to: preset multiple noise intensity ranges, and pre-define that different noise intensity ranges correspond to different impact factors; obtain the ambient intensity according to the ambient sound signal, and obtain the corresponding impact factor according to the noise intensity range in which the ambient intensity falls.
[0055] In this embodiment, it should be noted that regarding the above-mentioned voice song request interaction data processing system based on multimodal fusion, the specific manner of performing operations has been described in detail in the embodiments of the method for voice song request interaction data processing based on multimodal fusion, and will not be elaborated here.
[0056] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0057] In addition, it should be noted that in the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination methods.
[0058] It should be further noted that, in the above embodiments, the description of the consecutive use of one or more of "A, B, C, D, E, F, G, H, I, J, K, L, M, N, O, P, Q, R, S, T, U, V, W, X, Y, Z" is used to refer to one or more characters in the current song track name and one or more characters in the singer name; at the same time, the same letters between two embodiments may have no association.
[0059] In addition, any combination can be made among various different embodiments of the present disclosure, as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.
[0060] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. A method for processing voice song-request interaction data based on multimodal fusion, characterized in that, Including: Receiving a song-request voice signal, obtaining the processing time point when the song-request voice signal is received, obtaining the semantic word according to the song-request voice signal, obtaining the processing time period with the end time point being the processing time point, obtaining the ambient sound signal recorded within the processing time period, and obtaining the influence factor according to the ambient sound signal; Obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library, obtaining multiple matching degrees exceeding the matching threshold, arranging the multiple matching degrees in descending order successively to form a pre-matching degree sequence; Obtaining the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtaining the difference threshold according to the influence factor, and obtaining the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; Obtaining an interactive push sequence according to the tracks corresponding to the target matching degree sequence, and pushing the tracks according to the interactive push sequence.
2. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, wherein The obtaining the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence includes: Obtaining adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; Dividing the pre-matching degree sequence according to one or more division positions to form multiple sub-matching degree sequences; Obtaining the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.
3. The method for processing voice karaoke interaction data based on multimodal fusion according to claim 1, wherein The obtaining the influence factor according to the ambient sound signal includes: Presetting multiple noise intensity ranges, and predefined that multiple noise intensity ranges respectively correspond to different influence factors; Obtaining the ambient intensity according to the ambient sound signal, and obtaining the corresponding influence factor according to the noise intensity range where the ambient intensity falls.
4. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, wherein The obtaining the difference threshold according to the influence factor includes: Obtaining the environmental sensitivity coefficient, and obtaining the maximum value of the matching differences of adjacent matching degrees in the pre-matching degree sequence; Obtaining the difference threshold according to the influence factor, the environmental sensitivity coefficient, and the maximum value of the matching differences of adjacent matching degrees in the pre-matching degree sequence.
5. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, wherein The obtaining the difference threshold according to the influence factor is expressed as: ; wherein, is the difference threshold, is the matching difference between the j-th group of adjacent matching degrees in the pre-matching degree sequence, is the number of different adjacent matching degrees in the pre-matching degree sequence, is the influence factor, is the environmental sensitivity coefficient.
6. The method for processing voice-based song ordering interaction data based on multimodal fusion according to claim 1, wherein The obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library includes: Obtaining the maximum fuzzy increment coefficient; Obtaining the matching degree corresponding to each track according to the maximum fuzzy increment coefficient, the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library.
7. The method for processing voice song ordering interaction data based on multimodal fusion according to claim 1, characterized in that The obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library is expressed as: ; wherein, is the matching degree corresponding to the i-th track, is the maximum fuzzy increment coefficient, is the influence factor, is the semantic word, is the word segmentation corresponding to the i-th track, is the number of words in the semantic word, is the number of words in the word segmentation corresponding to the i-th track.
8. A voice song ordering interaction data processing system based on multimodal fusion, characterized in that, The system includes: A data interaction module, configured to receive a song-request voice signal, obtain the processing time point when the song-request voice signal is received, obtain the semantic word according to the song-request voice signal, obtain the processing time period with the end time point being the processing time point, obtain the ambient sound signal recorded within the processing time period, and obtain the influence factor according to the ambient sound signal; A data processing module, configured to obtain the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library, obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order successively to form a pre-matching degree sequence; A data screening module, configured to obtain a matching difference between adjacent matching degrees in a pre-matching degree sequence, obtain a difference threshold according to an influence factor, and obtain a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; A data push module, configured to obtain an interactive push sequence according to the track corresponding to the target matching degree sequence, and push the track according to the interactive push sequence.
9. The voice-based song ordering interactive data processing system based on multimodal fusion according to claim 8, wherein The data screening module is further configured to: Obtain adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as division positions; Divide the pre-matching degree sequence according to one or more division positions to form a plurality of sub-matching degree sequences; Obtain the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.
10. The voice-based song ordering interactive data processing system based on multimodal fusion according to claim 9, wherein The data interaction module is further configured to: Preset a plurality of noise intensity ranges, and predefined that different influence factors correspond to different noise intensity ranges respectively; Obtain the environmental intensity according to the ambient sound signal, and obtain the corresponding influence factor according to the noise intensity range in which the environmental intensity falls.
Citation Information
Patent Citations
Semantic processing method and device for semantic comprehension model and storage medium
CN110807333A
Voice recognition method and system for adaptive endpoint detection, and intelligent equipment
CN111816217A
Voice interaction method and device and storage medium
CN116670760A
Voice control method of intelligent household electrical appliance and intelligent household electrical appliance
CN119993146A
Fuzzy instruction analysis method and system based on dual-channel noise reduction and dynamic semantic map
CN120048269A