Voice-on-demand interaction data processing method and system based on multimodal fusion

Through multimodal fusion technology, dynamically capture the ambient sound signals and semantic understanding, the problem of matching degree calculation of the voice lyrics system in complex noise environments is solved, and accurate track push in low signal-to-noise ratio scenarios is realized, which improves the system's adaptability and accuracy.

CN120279870BActive Publication Date: 2025-08-01CHENGDU XIAOCHANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510765348.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-01
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing voice-based interactive system has reduced the quality of speech signal and semantic recognition accuracy in complex noise environments. The matching degree calculation method relies on fixed thresholds and is difficult to dynamically adjust, resulting in mispushing or omissions, and lacks dynamic quantization methods to identify real matching faults.

Method used

Through multimodal fusion technology, ambient sound signals are dynamically captured, combined with timing correlation and decibel hierarchical mapping, the noise weight and difference threshold are dynamically adjusted, and the coupling of environmental factors and semantic understanding is realized, matching fault judgment is adaptively adjusted, and push strategy is optimized.

Benefits of technology

In complex noise environments, users' voice characteristics are accurately captured, voice recognition errors are dynamically compensated, match sequence consistency, and intelligently optimized track pushing is realized, which improves push accuracy and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279870B_ABST
    Figure CN120279870B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for processing voice song selection interaction data based on multi-modal fusion, which relates to the technical field of data processing. The method includes: receiving a voice song selection signal and obtaining a processing time point, obtaining a semantic word, obtaining a processing time period, obtaining an environmental sound signal, and obtaining an influencing factor; obtaining the matching degree corresponding to each track according to the influencing factor, the semantic word, and the word segmentation of multiple tracks in the song library, obtaining multiple matching degrees exceeding the matching threshold, and forming a pre-matching degree sequence; obtaining the matching difference between adjacent matching degrees, obtaining a difference threshold according to the influencing factor, and obtaining a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; obtaining an interactive push sequence according to the tracks corresponding to the target matching degree sequence, and pushing the tracks according to the interactive push sequence. The present invention has the advantages of multi-modal data correlation influence, elastic semantic fault tolerance matching, and adaptive optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for processing voice song requesting interactive data based on multimodal fusion. Background Art

[0002] With the rapid development of intelligent voice interaction technology, voice-on-demand song interaction systems have been widely used in scenarios such as home entertainment and learning, smart homes, and in-car entertainment. However, in actual deployment, such systems have adaptability issues at the data level: First, the ability to resist interference from acoustic environments is weak. When users are in scenes with complex background noise (such as noisy streets, construction sites, or gatherings), the quality of voice signals and the accuracy of semantic recognition drop significantly, resulting in deviations in the extraction of core features for track matching. Actual measured data shows that when the ambient noise exceeds 75dB, the accuracy of the existing Top1 push after data processing will drop sharply from 89% in a quiet environment to 52%, and the incorrectly pushed tracks often have a semantic gap with the user's intentions; secondly, the existing matching degree calculation methods mostly rely on fixed thresholds or static weight allocations, which makes it difficult to dynamically adjust the screening strategy according to the real-time environment. For example, in a high-noise environment, the real matching target and the secondary candidate The difference in discrimination between them will be unevenly suppressed or amplified, causing erroneous push or omission of valid targets. For example, the correlation between environmental noise intensity and user semantic clarity is not taken into account, which further leads to a decrease in push accuracy in low signal-to-noise ratio scenarios. Finally, the existing solutions lack dynamic quantitative means for identifying matching selections - since the fluctuations in the matching sequence caused by noise may mask the real matching faults, when the matching values of multiple candidate tracks are close, the existing fixed difference threshold division method (for example, the difference between the top two is greater than 0.3, which is considered a fault) will cause the system to frequently misjudge in a highly volatile environment. Under certain special voice input conditions (such as the weak correlation between the user's vague expression and the lyrics text), the conventional matching difference calculation method cannot accurately identify the range of candidate sets to be pushed. Summary of the Invention

[0003] In view of the defects in the prior art, the present invention provides a method and system for processing voice song requesting interactive data based on multimodal fusion.

[0004] A voice song ordering interaction data processing method based on multimodal fusion, comprising: receiving a song ordering voice signal, obtaining the processing time point when the song ordering voice signal is received, obtaining a semantic word according to the song ordering voice signal, obtaining a processing time period with the end time point being the processing time point, obtaining an ambient sound signal recorded within the processing time period, and obtaining an influence factor according to the ambient sound signal; obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library, obtaining multiple matching degrees exceeding the matching threshold, arranging the multiple matching degrees in descending order to form a pre-matching degree sequence; obtaining the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtaining a difference threshold according to the influence factor, and obtaining a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; obtaining an interactive push sequence according to the tracks corresponding to the target matching degree sequence, and pushing the tracks according to the interactive push sequence.

[0005] Optionally, obtaining a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence includes: obtaining adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; dividing the pre-matching degree sequence according to one or more division positions to form multiple sub-matching degree sequences; obtaining the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.

[0006] Optionally, obtaining an influence factor according to the ambient sound signal includes: presetting multiple noise intensity ranges, and predefined that multiple noise intensity ranges respectively correspond to different influence factors; obtaining the ambient intensity according to the ambient sound signal, and obtaining the corresponding influence factor according to the noise intensity range into which the ambient intensity falls.

[0007] Optionally, obtaining a difference threshold according to the influence factor includes: obtaining an environmental sensitivity coefficient, and obtaining the maximum value of the matching differences between adjacent matching degrees in the pre-matching degree sequence; obtaining a difference threshold according to the influence factor, the environmental sensitivity coefficient, and the maximum value of the matching differences between adjacent matching degrees in the pre-matching degree sequence.

[0008] Optionally, obtaining a difference threshold according to the influence factor is expressed as: ; where is the difference threshold, is the matching difference between the jth group of adjacent matching degrees in the pre-matching degree sequence, is the number of different adjacent matching degrees in the pre-matching degree sequence, is the influence factor, is the environmental sensitivity coefficient.

[0009] Optionally, obtaining the matching degree corresponding to each track according to the impact factor, semantic word, and word segmentation of multiple tracks in the music library includes: obtaining the maximum fuzzy increment coefficient; obtaining the matching degree corresponding to each track according to the maximum fuzzy increment coefficient, impact factor, semantic word, and word segmentation of multiple tracks in the music library.

[0010] Optionally, the matching degree corresponding to each track obtained according to the impact factor, semantic word, and word segmentation of multiple tracks in the music library is expressed as: ; where is the matching degree corresponding to the i-th track, is the maximum fuzzy increment coefficient, is the impact factor, is the semantic word, is the word segmentation corresponding to the i-th track, is the number of words in the semantic word, is the number of words in the word segmentation corresponding to the i-th track.

[0011] There is also provided a voice song selection interaction data processing system based on multimodal fusion. The system includes: a data interaction module, configured to receive a song selection voice signal, obtain the processing time point when the song selection voice signal is received, obtain the semantic word according to the song selection voice signal, obtain the processing time period with the end time point being the processing time point, obtain the ambient sound signal recorded within the processing time period, and obtain the impact factor according to the ambient sound signal; a data processing module, configured to obtain the matching degree corresponding to each track according to the impact factor, semantic word, and word segmentation of multiple tracks in the music library, obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order and form a pre-matching degree sequence; a data screening module, configured to obtain the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtain the difference threshold according to the impact factor, and obtain the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; a data push module, configured to obtain an interactive push sequence according to the tracks corresponding to the target matching degree sequence, and push the tracks according to the interactive push sequence.

[0012] Optionally, the data screening module is further configured to: obtain adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; divide the pre-matching degree sequence according to one or more division positions and form multiple sub-matching degree sequences; obtain the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.

[0013] Optionally, the data interaction module is further configured to: preset multiple noise intensity ranges, and predefined that different noise intensity ranges correspond to different impact factors; obtain the ambient intensity according to the ambient sound signal, and obtain the corresponding impact factor according to the noise intensity range in which the ambient intensity falls.

[0014] The beneficial effects of the present invention are embodied in:

[0015] In the entire voice-based song ordering interactive data processing method based on multi-modal fusion, first, based on the time-sequence correlation-based ambient sound dynamic perception technology, the real sound field characteristics when the user makes a sound are accurately captured. Through decibel-level mapping and sudden noise weight adjustment, the defect of the traditional fixed threshold's late response to transient interference is overcome, ensuring the spatio-temporal consistency of environmental factor calculation; further, the dynamic coupling of environmental noise and semantic understanding expands the semantic association range through a fuzzy increment coefficient in a low signal-to-noise ratio scenario, which can not only strengthen the weights of core keywords (such as singer names and song titles), but also be compatible with the fault-tolerant matching of synonyms and fuzzy expressions, effectively compensating for the feature deviation caused by speech recognition errors; further, the introduction of the dynamic calculation of the difference threshold linked to the environmental sensitivity coefficient and the noise spectrum characteristics solves the problem of the distortion of the matching degree sequence caused by noise suppression by adaptively adjusting the matching fault determination boundary; further, the push strategy can combine the user profile and the real-time environmental state to achieve the intelligent optimization of the push results. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0017] Figure 1 It is a schematic diagram of the steps of the voice-based song ordering interactive data processing method based on multi-modal fusion of the present invention in one embodiment;

[0018] Figure 2 It is a schematic diagram of a part of the steps of S4 in the voice-based song ordering interactive data processing method based on multi-modal fusion of the present invention;

[0019] Figure 3 It is a schematic diagram of a part of the steps of S1 in the voice-based song ordering interactive data processing method based on multi-modal fusion of the present invention;

[0020] Figure 4 It is a schematic diagram of a part of the steps of S3 in the voice-based song ordering interactive data processing method based on multi-modal fusion of the present invention;

[0021] Figure 5 It is a schematic diagram of a part of the steps of S2 in the voice-based song ordering interactive data processing method based on multi-modal fusion of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein generally can be arranged and designed in a variety of different configurations.

[0023] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0024] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0025] As Figure 1 shown, a voice song ordering interaction data processing method based on multimodal fusion is provided, including:

[0026] S1. Receive a song ordering voice signal, obtain the processing time point when the song ordering voice signal is received, obtain a semantic word according to the song ordering voice signal, obtain a processing time period with the end time point being the processing time point, obtain an ambient sound signal recorded within the processing time period, and obtain an influence factor according to the ambient sound signal;

[0027] S2. Obtain the matching degree corresponding to each song according to the influence factor, the semantic word, and the word segmentation of multiple songs in the song library, obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order and form a pre-matching degree sequence;

[0028] S3. Obtain the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtain a difference threshold according to the influence factor, and obtain a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence;

[0029] S4. Obtain an interactive push sequence according to the songs corresponding to the target matching degree sequence, and push the songs according to the interactive push sequence.

[0030] In this embodiment, it should be noted that in S1, by dynamically capturing the real-time acoustic environment characteristics during voice interaction, a multi-modal data association is constructed to provide an environmental perception basis for subsequent semantic matching and threshold decision-making. First, when receiving a user voice signal, the processing time point of the instruction is accurately recorded, and the environmental sound data of the historical time period with this time point as the end is intercepted based on temporal logic. This processing time period selection mechanism fully considers the real-time requirements of voice interaction. For example, when the user says "Play ABC of XYZ", the environmental audio segment within 1-2 seconds before the end of the instruction will be retrieved (this duration can be optimized and adjusted according to the actual scenario), rather than statically setting a fixed time window, so as to accurately reflect the real sound field state at the moment when the user makes a sound. Secondly, for the environmental sound signal within this time period, a decibel intensity grading mapping mechanism is adopted to discretize the continuous environmental noise intensity into multiple preset quantization intervals (such as low noise, medium noise, high noise and other scenario categories), and each interval corresponds to a predefined environmental impact factor. For example, in a home scenario, the environmental sound may trigger a medium-level impact factor due to the sound intensity of music, TV, conversation or kitchen noise; while the high-decibel noise generated in a vehicle scenario can be mapped to a high-order impact factor. This grading mechanism effectively avoids the problem of insensitivity of traditional fixed thresholds to environmental mutations.

[0031] Furthermore, by combining temporal association with environmental quantization, it is possible to dynamically perceive the acoustic context in which the voice instruction is located. For example, when the user orders a song in a coffee shop, the compound noises such as the clinking of cups and dishes and the conversations of people rising and falling within the processing time period will be captured, and the environmental sound intensity level will be determined through spectral analysis and energy integration. If a sudden high-decibel noise (such as the start of a coffee machine) appears during this time period, the weight of the impact factor will be automatically adjusted according to the duration ratio of the noise, rather than simply taking the average value. This processing method can not only reflect the instantaneous interference intensity, but also avoid the excessive influence of occasional spike noises on the overall environmental assessment. At the same time, the acquisition of environmental sound signals adopts multi-channel fusion technology, combined with the spatial filtering characteristics of the device microphone array, to effectively separate the user's voice and environmental noise, ensuring the accuracy of subsequent impact factor calculations.

[0032] In S2, a dynamic weight mechanism is established to quantify the impact of environmental noise on semantic understanding as a regulatory parameter for match degree calculation, achieving flexible matching in noisy scenarios. First, the basic match value is calculated through the text correlation between semantic words and the segmented words in the music library. For example, when the user says "Play ABC by XYZ", semantic words such as "Play", "XYZ", and "ABC" will be extracted and compared with the segmented words in the text fields such as the title and singer of the songs in the music library to calculate the basic vocabulary coincidence degree. In this process, a dynamic amplification coefficient based on environmental factors is introduced: in a low-noise scenario (such as a home environment), an exact matching strategy is adopted, requiring a high coincidence between semantic words and the segmented words of the track; when a high-noise impact factor is detected (such as a vehicle-mounted scenario), the matching conditions are appropriately relaxed through a fuzzy increment coefficient, allowing candidates with partial semantic associations but not exact matches to enter the pre-selection sequence. For example, when the user says "Play DEF", it may be recognized as "Play DEG". At this time, the weights of synonyms such as "DE" in the track lyrics will be automatically increased according to the noise level, so that "DEF" can still obtain a high match degree through extended semantic associations.

[0033] Furthermore, a collaborative mechanism for environmental perception and semantic understanding is constructed to eliminate the interference of match degree fluctuations caused by noise through threshold decision-making of environmental perception. The calculation dimension of the match degree is dynamically adjusted according to the real-time noise intensity: in a quiet environment, comprehensive text matching is emphasized, such as considering features such as the integrity of the song; while in a high-noise scenario, enhanced matching of core keywords is focused. For example, when the environmental factor reaches the threshold, the weights of overlapping keyword fields such as the singer name and song title will be amplified. This dynamic strategy not only retains the accuracy of basic text matching but also compensates for semantic losses caused by noise through the environmental perception mechanism. For example, when the user vaguely expresses "That song HIJKL" in a party environment, the match degree weights of the song titles containing "HI" or "KL" will be amplified during the match calculation through the impact factor, so that the match degrees of the song tracks containing "HI" or "KL" and the song track containing "HIKL" are at the same match degree level. Finally, multiple match degrees are sorted in descending order to form a pre-match degree sequence, providing data support for subsequent threshold decision-making.

[0034] In S3, a dynamic coupling mechanism is established for environmental noise and matching gap identification. Distortion of the matching sequence caused by noise interference is eliminated through an adaptive difference threshold. First, the distribution of matching gaps between adjacent candidate tracks in the pre-matching sequence is analyzed. For example, when a user orders "MNOPQ" at a construction site, environmental noise may compress the matching gap between the actual matching track and similar candidates (such as "MN" and "MNOP") from 0.4 in a quiet environment to 0.15. A traditional fixed threshold of 0.3 would misjudge a match without gaps in this scenario, resulting in push errors. This solution dynamically adjusts the difference threshold based on environmental factors: when noise above 75dB is detected, the threshold is automatically reduced from 0.3 to 0.12 based on the noise intensity coefficient, allowing the compressed actual gap (0.15 difference) to trigger effective segmentation and accurately lock onto the target track.

[0035] Furthermore, noise characteristics and threshold sensitivity are collaboratively optimized. Specifically, a dual adjustment mechanism is constructed by combining environmental sensitivity and matching degree fluctuation characteristics: in scenarios with persistent low-frequency noise (such as in-vehicle environments), a progressive threshold attenuation strategy can be adopted, ensuring that the adjustment amplitude of the difference threshold is positively correlated with the noise duration; while for sudden impulse noise (such as the sound of falling cutlery), a transient response mode can be activated to dynamically adjust the threshold based on the spectral energy distribution of the impulse noise.

[0036] In S4, a dynamic candidate set push mechanism is constructed, optimizing the final push results through matching gap identification and contextual association. First, the tracks in the target match sequence are sorted in descending order of match score. For example, if a user requests "ST" in a high-noise environment, the target sequence might contain a match distribution of "ST" 0.82, "AT" 0.79, "SA" 0.78, and "DE" 0.41. Traditionally, the top 1 track would be pushed directly. However, in scenarios where noise causes matching gaps, the top three tracks are detected to have a difference less than a dynamic threshold (e.g., 0.05), indicating no significant gap. These top three tracks are then packaged into an interactive push sequence. By using a secondary confirmation mechanism to determine whether you want to listen to "ST," "AT," or "SA," this prevents mispush attempts while preserving the candidate range that reflects the user's true intent.

[0037] Furthermore, the push method is dynamically selected based on the structural characteristics of the target sequence: when a significant gap is detected (e.g., the top track leads the second by more than 0.3), the optimal track is played directly; if the sequence exhibits a multimodal distribution (e.g., "ABCD" 0.75, "ABC" 0.73, "BCD" 0.71, "CDE" 0.68), all tracks are pushed, or semantic reinforcement is performed based on the user's historical on-demand records, prioritizing a subset of tracks that are more closely aligned with the user's preferences. For example, if it is detected that a user frequently requests song XYZ, even if "AB" has a slightly lower match, it will still be promoted to the top of the push list due to its strong correlation with the user's profile.

[0038] To sum up, in the entire voice song-ordering interactive data processing method based on multimodal fusion, first of all, the dynamic perception technology of ambient sound based on time series association accurately captures the real sound field characteristics when the user speaks, and overcomes the defect of the traditional fixed threshold's delayed response to transient interference through decibel grading mapping and burst noise weight adjustment, ensuring the spatiotemporal consistency of environmental factor calculation; further, the dynamic coupling of ambient noise and semantic understanding expands the semantic association range through the fuzzy incremental coefficient in low signal-to-noise ratio scenarios, which can not only strengthen the weight of core keywords (such as singer name, song title), but also be compatible with the fault-tolerant matching of synonyms and fuzzy expressions, effectively compensating for the feature deviation caused by speech recognition errors; further, the dynamic calculation of the difference threshold linked to the environmental sensitivity coefficient and the noise spectrum characteristics is introduced, and the matching fault judgment boundary is adaptively adjusted to solve the problem of matching degree sequence distortion caused by noise suppression; further, the push strategy can combine user portraits with real-time environmental status to achieve intelligent optimization of push results.

[0039] like Figure 2 As shown, in one embodiment, obtaining the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence in S4 includes:

[0040] S41, obtaining adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than a difference threshold and using them as division positions;

[0041] S42, dividing the pre-matching degree sequence according to one or more division positions to form a plurality of sub-matching degree sequences;

[0042] S43: Obtain a sub-matching degree sequence containing the maximum matching degree and use it as a target matching degree sequence.

[0043] In this embodiment, it should be noted that in S41, significant fault boundaries in the matching sequence are dynamically identified. The matching differences of all adjacent candidate tracks in the pre-matching sequence are traversed, and when it is detected that a group of adjacent matching differences exceeds the dynamically calculated difference threshold, the position is marked as a dividing point. For example, when a user orders "MNOPQ" at a construction site, the environmental noise causes the matching sequence to present a distribution of "MNOP" 0.78, "NOP" 0.62, and "OPQ" 0.60. After reducing the threshold from 0.3 to 0.15 based on the noise intensity, it is detected that the first group of differences of 0.16 exceeds the threshold, that is, a dividing point is established between the first and second candidates, accurately isolating the real target from other interference items.

[0044] In S42, the topological structure of the candidate set is reconstructed through the division points to form multiple subsequences with internal consistency. Taking the previous example, at the first division point, the sequence is split into two subsequences: the first subsequence only contains 《MNOPQ》 with the highest matching degree, and the second subsequence contains sub-optimal candidates such as 《NOP》 and 《OPQ》. This division mechanism effectively distinguishes the candidate groups before and after the fault. For example, when the user vaguely expresses "play fast songs of XYZ", the matching degree sequence may form a distribution of 《XYZ Rap Compilation》 0.75, 《XYZ Slow Songs》 0.71, 《XYZ Live Version》 0.70, 《XYZ Collection》 0.45. A difference fault is detected between the third and fourth candidates, and the sequence is divided into a high-matching degree subset and a low-matching degree subset.

[0045] In S43, accurate push is achieved through the intelligent selection of the subsequence with the maximum matching degree. The subsequence containing the highest matching degree is preferentially selected as the final push set. If this subsequence contains multiple candidates (such as when no fault appears), the complete sequence is retained for subsequent processing. For example, when the user requests 《BC》 in a vehicle scenario, due to wind noise interference, the matching degree sequence may show a dense distribution of 《ABCD》 0.68, 《ABC》 0.67, 《BCD》 0.66. At this time, it is determined that there is no significant fault, and the complete sequence is transmitted to S4 for multi-candidate interactive push. This mechanism can quickly lock the target when the fault is clear and retain semantically related candidates in fuzzy scenarios. For example, when the user mispeaks as "play ABCXYZ", related tracks such as 《ABC》 and 《XYZ Collection》 are retained through the complete sequence, improving the fault tolerance ability.

[0046] As Figure 3 shown, in one embodiment, the influencing factors obtained according to the ambient sound signal in S1 include:

[0047] S11. Preset multiple noise intensity ranges and pre-define that different noise intensity ranges correspond to different influencing factors;

[0048] S12. Obtain the ambient intensity according to the ambient sound signal, and obtain the corresponding influencing factor according to the noise intensity range in which the ambient intensity falls.

[0049] In this embodiment, it should be noted that in S11, a hierarchical quantization system for the acoustic environment intensity is established, and the standard mapping of environmental characteristics is realized by presetting multiple noise intensity intervals. According to the acoustic characteristics of different application scenarios (such as home, vehicle, public places), the continuous decibel values are divided into discrete intervals with clear semantics. For example, 20 - 50 dB is divided into a low noise level corresponding to the home scenario, and the corresponding impact factor is 0.3; 50 - 80 dB is divided into a medium noise level corresponding to the coffee shop environment, and the corresponding impact factor is 0.6; above 80 dB is divided into a high noise level corresponding to the construction site, and the corresponding impact factor is 0.9. It should be emphasized that according to the application in the subsequent expressions, the impact factor can only be assigned within the range of 0 to 1. At the same time, 0 represents complete silence, and 1 represents the extreme noise environment. Each interval is associated with a predefined environmental impact factor, forming a gradient response mechanism from quiet to extreme noise. For example, in the smart home scenario, the 70 dB noise during the operation of the vacuum cleaner may be classified as a medium noise level, triggering a medium impact factor, while the knocking sound on the window during a rainstorm may be dynamically adjusted to a higher noise level due to large fluctuations in the decibel value.

[0050] In S12, the time-frequency domain joint analysis of the environmental sound signal is performed through the existing technology to calculate the average sound pressure decibel within a specific time period. In some special cases, for example, when the user orders a song on the street side, the continuous traffic noise will be recognized as the reference noise level, and the sudden car horn sound will trigger the transient detection module: if the duration of the horn sound exceeds the preset ratio (such as accounting for 30% of the analysis period), the impact factor level will be automatically increased; if it is a short-term interference (such as only accounting for 5%), the reference noise level assessment will be maintained. The main characteristics of the environmental noise are captured to avoid evaluation deviations caused by instantaneous spike interference.

[0051] As Figure 4 shown, in one embodiment, obtaining the difference threshold according to the impact factor in S3 includes:

[0052] S31. Obtain the environmental sensitivity coefficient and obtain the maximum value of the matching difference between adjacent matching degrees in the pre-matching degree sequence;

[0053] S32. Obtain the difference threshold according to the impact factor, the environmental sensitivity coefficient, and the maximum value of the matching difference between adjacent matching degrees in the pre-matching degree sequence.

[0054] In this embodiment, it should be noted that in S31, an association mechanism between environmental sensitivity and matching degree fluctuation is established, and the response characteristics to the change of the acoustic environment are reflected through the environmental sensitivity coefficient independent of the noise intensity. The environmental sensitivity coefficient is preset according to the device type and application scenario. For example, vehicles may be set with a higher sensitivity due to the need to adapt to rapidly changing road noise, while home audio systems may adopt a lower sensitivity to maintain stability.

[0055] Further, by traversing the matching degree differences between all adjacent candidate tracks in the pre-matching degree sequence, the maximum fluctuation amplitude is captured. For example, in a KTV scenario, when a user requests "FRTY", background vocal interference may cause the matching degree sequence to show a distribution of "FRTY" 0.75, "RTY" 0.72, "FRT" 0.70. At this time, the maximum adjacent difference of 0.05 is identified as the basic parameter for subsequent threshold calculation.

[0056] In S32, by fusing the environmental factor, the sensitivity coefficient, and the maximum fluctuation value, a dynamic threshold generation model is constructed. This model makes the threshold decrease as the noise intensity increases, and is simultaneously adjusted by the sensitivity coefficient: high-sensitivity devices trigger a larger threshold reduction amplitude under the same noise, so as to accurately identify the compressed matching fault in the noise interference. This dual adjustment mechanism not only maintains the environmental adaptability of the threshold adjustment but also takes into account the characteristic differences of different devices.

[0057] In one implementation, the difference threshold obtained according to the influencing factor in S3 is expressed as:

[0058] ; where

[0059] is the difference threshold, is the matching difference between the j-th group of adjacent matching degrees in the pre-matching degree sequence, is the number of different adjacent matching degrees in the pre-matching degree sequence, is the influencing factor, is the environmental sensitivity coefficient.

[0060] In this implementation, it should be noted that in the entire expression, is for dynamic reference value extraction, capturing the maximum difference amplitude of the current matching degree sequence, and using this as the reference for fault judgment; in a quiet environment, semantic matching is clear, and the matching degree difference between adjacent tracks is large (such as 0.4); while in a high-noise scenario, noise interference will cause the matching degree difference to be compressed (such as 0.15); by taking the maximum value, it is ensured that the threshold reference always reflects the actual matching fluctuation level in the current environment.

[0061] Further, is the noise attenuation term, which dynamically reduces the threshold according to the noise intensity; where controls the noise intensity, the greater the noise, is closer to 1, resulting in being smaller, being larger; adjusts the sensitivity, high-sensitivity devices (such as in-vehicle, ) under the same noise, is smaller, is also smaller, the reduction of the difference threshold is greater, and the fault is more stringent; for low-sensitivity devices (such as household devices, ), under the same noise, is larger, is also larger, the reduction of the difference threshold is smaller, and the fault is more lenient.

[0062] Furthermore, when is closer to 0 (no noise), is closer to 0, the difference threshold approaches 0, and a very small difference in matching degree is required to trigger a fault, which means that only the track corresponding to the highest matching degree needs to be pushed. When is closer to 1 (extreme noise), and at the same time , the difference threshold , and the difference threshold is compressed less.

[0063] As Figure 5 shown, in one embodiment, obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the music library in S2 includes:

[0064] S21. Obtain the maximum fuzzy increment coefficient;

[0065] S22. Obtain the matching degree corresponding to each track according to the maximum fuzzy increment coefficient, the influence factor, the semantic word, and the word segmentation of multiple tracks in the music library.

[0066] In this embodiment, it should be noted that in S21, the boundary constraint of semantic expansion is set, and the upper limit of error tolerance for matching degree calculation in a noisy environment is controlled by the maximum fuzzy increment coefficient. This coefficient is preset according to the device type and application scenario. For example, in-vehicle devices need to cope with sudden road noise, so a higher maximum fuzzy increment coefficient (such as 1.3) is set to allow a wider semantic association; while household devices use a lower maximum fuzzy increment coefficient (such as 1.1) to prevent misjudgment caused by excessive fuzzy matching. This preset mechanism ensures that voice recognition errors can be compensated in different scenarios while avoiding out-of-control semantic spans. For example, in a concert scenario, the user's true intention "play XYZGH" may be recognized as "play KLO of XYZ". By dynamically adjusting the weight range of near-sound words such as "GH" and "KLO" in the lyrics through the maximum fuzzy increment coefficient value, the correct track can still enter the candidate set.

[0067] In S22, the dynamic coupling of noise intensity and semantic expansion is achieved through the synergistic effect of the impact factor and the maximum fuzzy increment coefficient. In low-noise scenarios, the matching calculation focuses on precise text overlap, and only obtains a high weight when the semantic word and the song segmentation are completely matched; as the noise increases, the fuzzy increment is gradually activated, allowing some matching items to improve their scores through semantic relevance. For example, in a car scenario, the user says "play ABC" and is identified as "play ABD" by noise interference. At this time, the high impact factor value triggers the fuzzy increment, so that the songs in the music library that contain both "ABC" lyrics and "ABD" homophonic titles are all matched. The maximum fuzzy increment coefficient is then used to constrain the larger matching increase, successfully achieving full recommendation of multiple songs, ensuring that the songs that are truly needed are not eliminated.

[0068] In one embodiment, the matching degree corresponding to each track is obtained in S2 based on the impact factor, semantic words, and word segmentation of multiple tracks in the music library as follows:

[0069] ;in,

[0070] is the matching degree corresponding to the i-th track, is the maximum fuzzy increment coefficient, is the impact factor, For semantic words, is the word segment corresponding to the i-th track, is the number of words in the semantic word, is the number of words in the word segmentation corresponding to the i-th track.

[0071] In this embodiment, it should be noted that, in the entire expression, This is the basic matching calculation item, which measures the text overlap ratio between user semantic words and song segmentation words, to avoid inflated matching scores for long text tracks due to the large number of segmentation words. Example: In a noisy environment, the user's song request voice signal is converted into the voice words "play XYZ's BCD" (A={play, um, XYZ, B, C, D}, |A|=6), and the title of track 1 is "BCE"-XYZ ( , ),but , the basic matching degree is Track 2 is titled "OABT" - QWE ( , ),but , the basic matching degree is

[0072] Further, is the fuzzy incremental compensation term. The greater the noise, the greater the compensation value The larger (such as at 1), so as to achieve that the greater the noise, the more extra points are allowed for the tracks with incomplete semantic matching; for example: when the user says "Play DEF" (recognized as "Play DEG"), the matching degree of the track "DEF" is compensated. Among them, in, the basic matching degree is an objective quantitative index of semantic association, which is used to measure the exact coincidence degree between the user's semantic words and the track word segmentation, and the whole incorporating the min function is to retain the weight of the original matching information during the noise compensation process, and avoid the matching degree being too high or too low due to complete dependence on the compensation term; for example, assume that the user says "ABC" and is recognized as "ABE", and the basic matching degree of the real expected track "ABC" is 0.67. If the compensation term (such as 0.8) is directly used and the basic matching degree is ignored, the matching degree of "ABC" will be lower, and may be much lower than other partially matched tracks. However, through the formula design, the final matching degree is: , in this way, the matching degree of this track can be significantly improved. Among them, and the maximum fuzzy increment coefficient are set to reflect the rigid ceiling of the matching degree increase, limit the maximum allowable range of noise compensation, and prevent the semantic expansion from getting out of control due to too strong environmental noise; just like in the above example, if there is no limitation, then the directly calculated result is , resulting in the matching degree of a track with a low matching degree being immediately over-increased, leading to out-of-control matching.

[0073] In summary, semantic fault tolerance in a noisy environment is achieved. Even if noise causes speech recognition errors (such as "ABC" being misrecognized as "ABE"), this embodiment can still raise the matching degree of the track with the true intention. Further, it also achieves the prevention of excessive fuzzy matching; excessive compensation may lead to irrelevant tracks being recommended due to a single-word coincidence (such as "ABEF" being misrecommended because it contains "ABE"). Further, it also achieves dynamic weight and device adaptation. For example, in the vehicle-mounted device in device differentiation, set , and in the home device, set and so on.

[0074] It also provides a voice song-request interaction data processing system based on multimodal fusion. The system includes:

[0075] A data interaction module, which is used to receive the song-request voice signal, obtain the processing time point when the song-request voice signal is received, obtain the semantic word according to the song-request voice signal, obtain the processing time period with the end time point being the processing time point, obtain the environmental sound signal recorded within the processing time period, and obtain the influence factor according to the environmental sound signal;

[0076] A data processing module, configured to obtain the matching degree corresponding to each track according to the impact factor, semantic words, and word segmentation of multiple tracks in the music library, obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order, and form a pre-matching degree sequence;

[0077] A data screening module, configured to obtain the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtain a difference threshold according to the impact factor, and obtain a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence;

[0078] A data pushing module, configured to obtain an interactive push sequence according to the tracks corresponding to the target matching degree sequence, and push the tracks according to the interactive push sequence.

[0079] In one embodiment, the data screening module is further configured to: obtain adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; divide the pre-matching degree sequence according to one or more division positions to form multiple sub-matching degree sequences; obtain the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.

[0080] In one embodiment, the data interaction module is further configured to: preset multiple noise intensity ranges, and pre-define that different noise intensity ranges correspond to different impact factors; obtain the ambient intensity according to the ambient sound signal, and obtain the corresponding impact factor according to the noise intensity range in which the ambient intensity falls.

[0081] In this embodiment, it should be noted that regarding the above-mentioned voice song ordering interaction data processing system based on multi-modal fusion, the specific manner of performing operations has been described in detail in the embodiments of the voice song ordering interaction data processing method based on multi-modal fusion, and will not be elaborated here.

[0082] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0083] In addition, it should be noted that in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.

[0084] In addition, it should be noted that in the above embodiments, the description of the consecutive use of one or more of "A, B, C, D, E, F, G, H, I, J, K, L, M, N, O, P, Q, R, S, T, U, V, W, X, Y, Z" is used to refer to one or more characters in the current song track name and one or more characters in the singer name; at the same time, the same letters between the two embodiments may have no association.

[0085] In addition, any combination can be made between various different embodiments of the present disclosure, as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

[0086] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. A method for processing voice song-request interaction data based on multimodal fusion, characterized in that including: Receiving a song-request voice signal, obtaining the processing time point when the song-request voice signal is received, obtaining the semantic word according to the song-request voice signal, obtaining the processing time period with the end time point being the processing time point, obtaining the ambient sound signal recorded within the processing time period, and obtaining the influence factor according to the ambient sound signal; Obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library, obtaining multiple matching degrees exceeding the matching threshold, arranging the multiple matching degrees in descending order and forming a pre-matching degree sequence; Obtaining the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtaining the difference threshold according to the influence factor, and obtaining the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; Obtaining an interactive push sequence according to the tracks corresponding to the target matching degree sequence, and pushing the tracks according to the interactive push sequence.

2. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, wherein The obtaining the target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence includes: Obtaining adjacent matching degrees in the pre-matching degree sequence where the matching difference is greater than the difference threshold as the division positions; Dividing the pre-matching degree sequence according to one or more division positions and forming multiple sub-matching degree sequences; Obtaining the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.

3. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, characterized in that The obtaining the influence factor according to the ambient sound signal includes: Presetting multiple noise intensity ranges, and predefined that multiple noise intensity ranges respectively correspond to different influence factors; Obtaining the ambient intensity according to the ambient sound signal, and obtaining the corresponding influence factor according to the noise intensity range into which the ambient intensity falls.

4. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, wherein The obtaining the difference threshold according to the influence factor includes: Obtaining the environmental sensitivity coefficient, and obtaining the maximum value of the matching differences between adjacent matching degrees in the pre-matching degree sequence; Obtaining the difference threshold according to the influence factor, the environmental sensitivity coefficient, and the maximum value of the matching differences between adjacent matching degrees in the pre-matching degree sequence.

5. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, wherein The obtaining the difference threshold according to the influence factor is expressed as: ; wherein, is the difference threshold, is the matching difference between the j-th group of adjacent matching degrees in the pre-matching degree sequence, is the number of different adjacent matching degrees in the pre-matching degree sequence, is the influence factor, is the environmental sensitivity coefficient.

6. The method for processing voice karaoke interaction data based on multimodal fusion according to claim 1, wherein The obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library includes: Obtaining the maximum fuzzy increment coefficient; Obtaining the matching degree corresponding to each track according to the maximum fuzzy increment coefficient, the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library.

7. The method for processing voice song-request interaction data based on multimodal fusion according to claim 1, wherein The obtaining the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library is expressed as: ; wherein, is the matching degree corresponding to the i-th track, is the maximum fuzzy increment coefficient, is the influence factor, is the semantic word, is the word segmentation corresponding to the i-th track, is the number of words in the semantic word, is the number of words in the word segmentation corresponding to the i-th track.

8. A voice song-request interaction data processing system based on multimodal fusion, characterized in that The system includes: A data interaction module, configured to receive a song-request voice signal, obtain the processing time point when the song-request voice signal is received, obtain the semantic word according to the song-request voice signal, obtain the processing time period with the end time point being the processing time point, obtain the ambient sound signal recorded within the processing time period, and obtain the influence factor according to the ambient sound signal; A data processing module, configured to obtain the matching degree corresponding to each track according to the influence factor, the semantic word, and the word segmentation of multiple tracks in the song library, obtain multiple matching degrees exceeding the matching threshold, arrange the multiple matching degrees in descending order and form a pre-matching degree sequence; A data screening module, which is used to obtain the matching difference between adjacent matching degrees in the pre-matching degree sequence, obtain a difference threshold according to an influence factor, and obtain a target matching degree sequence according to the difference threshold and the matching difference between adjacent matching degrees in the pre-matching degree sequence; A data push module, which is used to obtain an interactive push sequence according to the track corresponding to the target matching degree sequence, and push the track according to the interactive push sequence.

9. The voice-based song ordering interactive data processing system based on multimodal fusion according to claim 8, wherein The data screening module is further used for: Obtaining adjacent matching degrees in the pre-matching degree sequence whose matching difference is greater than the difference threshold as the division positions; Dividing the pre-matching degree sequence according to one or more division positions to form multiple sub-matching degree sequences; Obtaining the sub-matching degree sequence containing the maximum matching degree as the target matching degree sequence.

10. The voice-based song ordering interactive data processing system based on multimodal fusion according to claim 9, characterized in that, The data interaction module is further used for: Presetting multiple noise intensity ranges, and predefined that multiple noise intensity ranges respectively correspond to different influence factors; Obtaining the environmental intensity according to the environmental sound signal, and obtaining the corresponding influence factor according to the noise intensity range in which the environmental intensity falls.

Citation Information

Patent Citations

  • Voice control method of intelligent household electrical appliance and intelligent household electrical appliance

    CN119993146A

  • Fuzzy instruction analysis method and system based on dual-channel noise reduction and dynamic semantic map

    CN120048269A