A voice content association determination method and related device

By dynamically adjusting the association threshold and association score, and combining the correction value for voice interruption events, the problem of insufficient accuracy in associating adjacent voice content in vehicles was solved, achieving higher accuracy and real-time performance.

CN122135734APending Publication Date: 2026-06-02ZHEJIANG GEELY HLDG GRP CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of determining the association between adjacent voice content in a vehicle is low due to real-time changes in the driver's voice dialogue, and fixed thresholds cannot adapt to dynamic voice scenarios, resulting in misjudgments and insufficient accuracy.

Method used

By acquiring data such as vehicle driving scenarios, driver's historical voice data, and driving status data, thresholds are determined, and association thresholds and association scores are dynamically adjusted. Combined with correction values ​​for voice interruption events, the content association of adjacent voices is accurately determined.

Benefits of technology

It improves the accuracy of determining the association between adjacent voice content, reduces the false judgment rate, and meets the real-time requirements of vehicle assisted driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135734A_ABST
    Figure CN122135734A_ABST
Patent Text Reader

Abstract

The application provides a speech content association determination method and related equipment, the speech content association determination method comprises: acquiring a first speech currently collected by a vehicle and a second speech collected last time, and acquiring threshold determination data of the vehicle; determining an association threshold according to the threshold determination data, and determining an association score of the first speech and the second speech; determining whether the first speech is associated with the second speech in content according to a comparison result between the association score and the association threshold. In the application, the determination accuracy of the association between adjacent speeches in content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a method and related equipment for determining speech content association. Background Technology

[0002] With the development of vehicle intelligence, voice capture functions that collect driver voice data for driver assistance have become common intelligent features. In multi-turn voice dialogues with the driver within the vehicle, it is necessary to determine whether the content of the currently collected voice data is related to that of historical dialogues, and to provide driver assistance based on the determination result.

[0003] In an exemplary technique, the similarity between adjacent speech segments is compared with a preset threshold to determine whether adjacent speech segments are related in content.

[0004] However, the voice dialogue between the driver and the vehicle changes in real time, while the preset threshold is fixed. This means that the dynamic voice scenario of the vehicle does not match the preset threshold, resulting in low accuracy in determining the content association of adjacent voices. Summary of the Invention

[0005] Based on the aforementioned technological status, this application provides a method and related equipment for determining the association of speech content, which addresses the problem of low accuracy in determining the content association of adjacent speech items.

[0006] To achieve the above-mentioned technical objectives, this application proposes the following technical solution: Firstly, this application provides a method for determining the association of voice content, including: The system acquires the first voice currently collected by the vehicle and the second voice collected previously, and acquires threshold determination data for the vehicle. The threshold determination data includes at least one of the vehicle's driving scenario, the historical voice data of the user driving the vehicle, and the user's driving status data. Based on the threshold, the association threshold is determined, and the association score between the first speech and the second speech is determined. Based on the comparison between the association score and the association threshold, it is determined whether the first speech and the second speech are associated in terms of content.

[0007] In some implementations, determining the association threshold based on the threshold-determined data includes: Obtain the coefficients and first correction values ​​corresponding to the sub-data in the threshold determination data, wherein each sub-data is the driving scenario, the historical voice data, and the driving state data; Based on the coefficients corresponding to the sub-data and the first correction value, determine the target correction value corresponding to the sub-data; The associated threshold is obtained by correcting the preset threshold based on the target correction value of each of the sub-data.

[0008] In some implementations, the step of correcting a preset threshold based on the target correction value of each of the sub-data to obtain the association threshold includes: In response to the detection of a voice interruption event, the preset threshold is corrected according to the target correction value of each of the sub-data to obtain an intermediate threshold; A second correction value is determined based on the voice interruption event; The association threshold is obtained by reducing the intermediate threshold based on the second correction value.

[0009] In some implementations, determining the second correction value based on the voice interruption event includes: Determine the initial correction value based on the voice interruption event; A compensation coefficient is determined based on the interruption duration of the voice interruption event, and the compensation coefficient is negatively correlated with the interruption duration. The second correction value is determined based on the compensation coefficient and the initial correction value.

[0010] In some implementations, determining the association score between the first speech and the second speech includes: Obtain content association parameters between the first speech and the second speech, wherein the content association parameters include at least one of the following: scene-aware time decay value, semantic similarity, intent continuity, and content overlap between the first speech and the second speech; The auxiliary correlation parameters are determined based on the user's vehicle driving environment data. The auxiliary correlation parameters include at least one of the following: user attention, driving scene stability, impact of voice interruption events, spatial correlation, and user interaction correlation. The association score between the first speech and the second speech is determined based on the content association parameters, the corresponding weights, the auxiliary association parameters, and the corresponding weights.

[0011] In some implementations, determining the association score between the first speech and the second speech includes: In response to the detection of a voice interruption event, an intermediate score between the first voice and the second voice is determined; A third correction value is determined based on the voice interruption event, and the intermediate score is increased based on the third correction value to obtain the associated score.

[0012] In some implementations, determining the third correction value based on the voice interruption event includes: Obtain the interruption type and interruption duration corresponding to the voice interruption event; The third correction value is determined based on the interruption type and the interruption duration.

[0013] In some implementations, obtaining the content association parameters between the first speech and the second speech includes: The scene decay duration of the vehicle is determined based on the vehicle's speed. In response to the scenario attenuation duration being greater than or equal to the attenuation threshold, the content association parameters between the first speech and the second speech are obtained.

[0014] In some implementations, after determining the scene decay duration of the vehicle based on its speed, the method further includes: In response to the scenario attenuation duration being less than the attenuation threshold, it is determined whether the first speech and the second speech are related in content based on the acquisition interval duration between the first speech and the second speech.

[0015] In some implementations, determining the association threshold based on the threshold determination data and determining the association score between the first speech and the second speech includes: In response to the detection of a voice interruption event, an association threshold is determined based on the threshold determined data and the voice interruption event, and an association score between the first voice and the second voice is determined based on the voice interruption event.

[0016] Secondly, this application provides a speech content association determination device, comprising: The acquisition module is used to acquire the first voice currently collected by the vehicle and the second voice collected previously, and to acquire threshold determination data of the vehicle. The threshold determination data includes at least one of the vehicle's driving scenario, the historical voice data of the user driving the vehicle, and the user's driving status data. The first determining module is used to determine the association threshold based on the threshold determining data, and to determine the association score between the first speech and the second speech; The second determining module is used to determine whether the first speech is related to the second speech in terms of content based on the comparison result between the association score and the association threshold.

[0017] Thirdly, this application provides a speech content association determination device, including a memory and a processor, wherein, The memory is connected to the processor and is used to store programs; The processor is used to implement the voice content association determination method as described in the first aspect or any implementation thereof by running a program in the memory.

[0018] Fourthly, this application provides a vehicle, the vehicle including a voice content association determination device, the voice content association determination device implementing the voice content association determination method as described in the first aspect or any implementation thereof.

[0019] Fifthly, this application provides a computer program product, which, when executed by a processor, implements the voice content association determination method as described in the first aspect or any implementation thereof.

[0020] In a sixth aspect, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the voice content association determination method as described in the first aspect or any implementation thereof.

[0021] This application provides a method and related equipment for determining voice content association. It acquires a first voice currently collected by a vehicle and a second voice collected previously, and also acquires threshold determination data such as the vehicle's driving scenario, the driver's historical voice data, and the user's driving status data. An association threshold is determined using this threshold determination data, and an association score is determined between the first and second voices. By comparing the association score with the association threshold, it is determined whether the first and second voices are content-related. In this application, the association threshold is dynamically determined based on data such as the vehicle's driving scenario, the driver's historical voice data, and the user's driving status data. Therefore, based on the comparison result of the association score representing whether the first and second voices are related, the method accurately determines whether the first and second voices are content-related, improving the accuracy of determining the content-related association of adjacent voices. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0023] Figure 1 A flowchart of a voice content association determination method provided in this application embodiment Figure 1 .

[0024] Figure 2 A flowchart of a voice content association determination method provided in this application embodiment Figure 2 .

[0025] Figure 3 A flowchart of a voice content association determination method provided in this application embodiment Figure 3 .

[0026] Figure 4 A flowchart of a voice content association determination method provided in this application embodiment Figure 4 .

[0027] Figure 5 A flowchart of a voice content association determination method provided in this application embodiment Figure 5 .

[0028] Figure 6 This is a schematic diagram of the functional modules of a voice content association determination device provided in an embodiment of this application.

[0029] Figure 7 This is a structural schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0031] It should be noted that the user information (including but not limited to electrical equipment information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0032] With the development of vehicle intelligence, voice capture functions that collect driver voice data for driver assistance have become common intelligent features. In multi-turn voice dialogues with the driver within the vehicle, it is necessary to determine whether the content of the currently collected voice data is related to that of historical dialogues, and to provide driver assistance based on the determination result.

[0033] In an exemplary technique, the similarity between adjacent speech segments is compared with a preset threshold to determine whether adjacent speech segments are related in content.

[0034] However, the voice dialogue between the driver and the vehicle changes in real time, while the preset threshold is fixed. This means that the dynamic voice scenario of the vehicle does not match the preset threshold, resulting in low accuracy in determining the content association of adjacent voices.

[0035] Furthermore, in the exemplary technology, the similarity of adjacent voices is calculated solely based on semantic similarity. The single-dimensional score cannot accurately reflect the true relevance and cannot handle voice scenarios with ambiguous intentions such as "this is the place" or "open this."

[0036] In addition, the exemplary technology first fixes the threshold and then optimizes the similarity of adjacent voices, or first fixes the similarity and then adjusts the threshold. When the vehicle scene changes (such as switching from a highway to a parking spot), even if the similarity is accurate, the fixed threshold will still misjudge.

[0037] To address the aforementioned technical problems, this application proposes a method for determining voice content association. The following detailed description of this method through various embodiments is provided.

[0038] Reference Figure 1 , Figure 1 A flowchart of a voice content association determination method provided in this application embodiment Figure 1 .like Figure 1 As shown, the voice content association determination method provided in this embodiment includes: Step S101: Obtain the first voice currently collected by the vehicle and the second voice collected previously, and obtain the threshold determination data of the vehicle. The threshold determination data includes at least one of the following: the vehicle's driving scenario, the historical voice data of the user driving the vehicle, and the user's driving status data.

[0039] In this embodiment, the executing entity is a voice content association determination device. For ease of description, the term "device" will be used to refer to the voice content association determination device below. The device can be a component in the vehicle or the vehicle itself.

[0040] The driver, while driving, interacts with the vehicle via voice commands. These commands can include playing music, providing navigation, etc. Examples include commands like "Play audio A," "Go to the nearest shopping mall," and "Reach the nearest gas station." The device needs to determine the user's intent by comparing the currently captured voice commands with those captured previously. For instance, if the previously captured command was "Go to the highway exit," and the currently captured command is "Go to the nearest gas station," then the two commands are related, meaning the gas station being sought is one along the route to the highway exit.

[0041] If the vehicle captures voice data, it then acquires the previously captured voice data, which is defined as the second voice data, and the currently captured voice data is defined as the first voice data. In addition, the device acquires threshold determination data for the vehicle, which includes at least one of the following: the vehicle's driving scenario, the user's historical voice data, and the user's driving status data.

[0042] The driving scenarios of the vehicle include, but are not limited to: high-speed driving, urban road driving, parking, and complex maneuvering. High-speed driving refers to the vehicle speed being greater than a first speed threshold, such as 80 km / h; urban road driving refers to the vehicle speed being less than or equal to the first speed threshold and greater than a second speed threshold, such as 20 km / h; parking refers to the vehicle speed being less than or equal to a third speed threshold, such as 5 km / h; and complex maneuvering refers to the vehicle making a sharp turn or braking.

[0043] Historical voice data refers to the historical voice messages that users have made to the vehicle.

[0044] Driving status data includes, but is not limited to, the user's attention and fatigue level while driving. Attention is assessed by whether the user checks their phone, takes their hands off the steering wheel, or keeps their eyes on the road ahead. If any of these occur, points are deducted from the maximum score for that activity. The final score represents the user's attention level. Fatigue level is determined based on the user's driving duration and whether the time falls during rest periods. For example, if the current time is during lunch break or nighttime sleep, the fatigue level is 0.8 times the maximum score. The longer the driving duration, the more points are deducted from the maximum score. This is how the fatigue level is calculated.

[0045] Step S102: Determine the association threshold based on the threshold determination data, and determine the association score between the first speech and the second speech.

[0046] After determining the threshold and the data, the associated threshold is determined using the threshold-determined data.

[0047] For example, the device sets a preset threshold, which can be any suitable value, such as 0.6. The device adjusts the preset threshold using threshold determination data to obtain the threshold value.

[0048] In one example, the threshold determination data is a driving scenario. In a high-speed driving scenario, since the road conditions on highways are simple and not complicated, the interval between the user's voice conversation with the vehicle is relatively long. That is, there are two voice conversations with a long interval, but the two voice conversations may still be related in content. Therefore, the preset threshold can be reduced to obtain the association threshold.

[0049] In another example, the threshold determination data is historical voice data. By analyzing the historical voice data, the user's voice patterns are determined. The voice patterns reveal that the interval between the user's conversation with the vehicle is relatively short. Since the interval is short, the probability that adjacent voices are related in content is relatively high. Therefore, the preset threshold is lowered to obtain the association threshold. Conversely, if the voice patterns reveal that the interval between the user's conversation with the vehicle is relatively long, the probability that adjacent voices are related in content is relatively low. Therefore, the preset threshold is increased to obtain the association threshold.

[0050] In another example, the threshold is determined using user state data. When user state data shows low attention or high fatigue, the probability of adjacent speech items being content-related is low, so the preset threshold is increased to obtain the association threshold.

[0051] The device then determines the association score between the first speech and the second speech. For example, the device can convert the text corresponding to the first speech into a first feature vector and the text corresponding to the second speech into a second feature vector, calculate the similarity between the first feature vector and the second feature vector, and then determine the association score based on the similarity. The greater the similarity, the greater the association score.

[0052] Step S103: Based on the comparison result between the association score and the association threshold, determine whether the first speech is related to the second speech in terms of content.

[0053] After determining the association threshold and the association score, the device can obtain the comparison result by comparing the association threshold and the association score. Based on the comparison result, it can determine whether the first speech and the second speech are related in terms of content, that is, whether the first speech and the second speech are related.

[0054] For example, when the association score is greater than or equal to the association threshold, the first speech and the second speech are associated in terms of content, that is, the first speech and the second speech are related; when the association score is less than the association threshold, the first speech and the second speech are not associated in terms of content, that is, the first speech and the second speech are not related.

[0055] In this embodiment, the system acquires the first voice currently collected by the vehicle and the second voice collected previously. It also acquires threshold determination data, including the vehicle's driving scenario, the driver's historical voice data, and the user's driving status data. A correlation threshold is determined using this threshold determination data, and a correlation score is calculated between the first and second voices. The comparison between the correlation score and the correlation threshold determines whether the first and second voices are content-related. This embodiment dynamically determines the correlation threshold based on data such as the vehicle's driving scenario, the driver's historical voice data, and the user's driving status data. Therefore, based on the comparison between the correlation score (indicating whether the first and second voices are related) and the correlation threshold, it accurately determines whether the first and second voices are content-related, improving the accuracy of determining the content-related association between adjacent voices.

[0056] Reference Figure 2 , Figure 2 A flowchart of a voice content association determination method provided in this application embodiment Figure 2 ,based on Figure 1 In the embodiment shown, step S102 includes: Step S201: Obtain the coefficients and first correction values ​​corresponding to the sub-data in the threshold determination data. Each sub-data is a driving scenario, historical voice data, and driving status data.

[0057] Step S202: Determine the target correction value corresponding to the sub-data based on the coefficients corresponding to the sub-data and the first correction value.

[0058] In this embodiment, the device acquires the coefficients and first correction values ​​corresponding to the sub-data in the threshold determination data, and then determines the target correction value through the coefficients and the first correction value. Each sub-data is the driving scenario, historical voice data, and driving status data.

[0059] In one example, when the sub-data is a driving scene, the target correction value is β × f_scene_factor, where β is a coefficient and can be a preset value, and f_scene_factor is the first correction value, for example: High-speed driving: f_scene_factor=-0.15 (reduce threshold, tolerate long intervals); City roads: f_scene_factor=0 (standard threshold); Parking status: f_scene_factor=+0.15 (increase the threshold for faster switching); Complex handling (sharp turns / braking): f_scene_factor=-0.08.

[0060] In another example, the target correction value is α×(ΔT_user_avg-ΔT_global_avg), where α is a coefficient, for example, 0.001, used to control the degree of influence of user voice habits on the preset threshold; (ΔT_user_avg-ΔT_global_avg) is the first correction value, ΔT_user_avg is the average interval duration of user (driver) voice dialogue in historical voice data, and ΔT_global_avg is the average dialogue interval duration of all users, which is a baseline value.

[0061] In another example, the target correction value is γ×(1-attention_t)+δ×fatigue_t, where γ and δ are coefficients, attention_t represents the score corresponding to attention, and fatigue_t represents the score corresponding to fatigue.

[0062] Step S203: Based on the target correction value of each sub-data, the preset threshold is corrected to obtain the associated threshold.

[0063] Once the target correction value is obtained, the associated threshold can be obtained by correcting the preset threshold using the target correction value.

[0064] For example, the association threshold θ_final = θ_base + β·f_scene_factor + α×(ΔT_user_avg-ΔT_global_avg) + γ×(1-attention_t) + δ×fatigue_t, where θ_base is a preset threshold.

[0065] In addition, the correlation threshold can also be determined by one or two of the target correction values ​​corresponding to the three types of sub-data mentioned above.

[0066] It should be noted that the association threshold needs to be within a certain range, for example, [0.35, 0.85]. Specifically, statistics show that when the association threshold is less than 0.35, the true correlation probability between adjacent speech is less than 5%. Setting the lower limit of the threshold range to less than 0.35 will increase the false positive rate. Conversely, when the association threshold is greater than 0.85, the true correlation rate between adjacent speech is greater than 98%. Setting the upper limit of the threshold range to greater than 0.85 will not significantly increase the correlation rate, but will instead increase the false positive rate.

[0067] Based on the correlation thresholds determined using the aforementioned sub-data, the following technical effects are observed statistically: Scene accuracy variance decreased from 8-12% to 2-3% (a reduction of 75-88%), while accuracy remained stable during scene switching; Accuracy in high-speed scenarios improved from 68% to 92% (+24%). Accuracy in parking scenarios improved from 75% to 91% (+16%). Example of dynamic threshold range: High speed + Focus = 0.48, Parking + Distraction = 0.81, City + Standard θ = 0.60.

[0068] In this embodiment, the coefficients and correction values ​​of sub-data in the data are determined by threshold, and the association threshold is dynamically determined.

[0069] Figure 3 A flowchart of a voice content association determination method provided in this application embodiment Figure 3 ,based on Figure 1 or Figure 2 In the embodiment shown, step S102 includes: Step S301: In response to the detection of a voice interruption event, determine the data based on the threshold and determine the intermediate threshold.

[0070] In this embodiment, during a user's voice conversation with the vehicle, there may be moments when the voice is interrupted. For example, when the user speaks to the vehicle, the vehicle's in-vehicle terminal may display an incoming call, broadcast navigation, issue a warning, or output a prompt tone. These events that interrupt the user's voice are defined as voice interruption events. Such voice interruption events can interfere with the correlation of adjacent voice messages; therefore, it is necessary to reduce the impact of voice interruption events on voice correlation determination.

[0071] In response, if the device detects voice interruption events such as the vehicle terminal displaying an incoming call, the vehicle terminal broadcasting navigation, the vehicle issuing an alarm, or the vehicle outputting a prompt tone, it first determines an intermediate threshold through threshold determination data. The intermediate threshold can be θ_final in the above embodiment.

[0072] Step S302: Determine the second correction value based on the voice interruption event.

[0073] The device determines a correction value based on the voice interruption event, and this correction value is defined as the second correction value.

[0074] In one example, the second correction value is related to the voice interruption event. For instance, if the voice interruption event is an incoming call interruption (a strong interruption), the second correction value is the largest first setting value, which is, for example, 0.15. If the voice interruption event is a vehicle warning tone or other alert, this is a medium-strong interruption, and the second correction value is the second setting value, which is smaller than the first setting value, for example, 0.12. If the voice interruption event is a navigation broadcast, this is a medium interruption, and the second correction value is the smaller third setting value, which is smaller than the second setting value, for example, 0.105. It is understood that the second correction value is related to the severity of the voice interruption event; the higher the severity, the larger the second correction value.

[0075] In another example, the effect of the voice interruption event compensation threshold decays exponentially over time to prevent misjudgments caused by excessively long voice interruptions. To address this, the device determines an initial correction value based on the voice interruption event, and then determines a compensation coefficient based on the interruption duration. The compensation coefficient is negatively correlated with the interruption duration; that is, the longer the interruption duration, the smaller the compensation coefficient. A second correction value is then determined based on the compensation coefficient and the initial correction value. For example, the second correction value = a × w_interrupt, where a is the compensation coefficient and w_interrupt is the initial correction value.

[0076] Step S303: Based on the second correction value, reduce the intermediate threshold to obtain the associated threshold.

[0077] After obtaining the second correction value, the intermediate threshold is reduced based on the second correction value to obtain the association threshold. For example, the association threshold = θ_adaptive - a × w_interrupt, where θ_adaptive is the intermediate threshold and a × w_interrupt is the second correction value.

[0078] In this embodiment, when a voice interruption event is detected, the threshold is corrected based on the voice interruption event to obtain the associated threshold, thereby avoiding the problem of excessively high voice-related misjudgment rate caused by the voice interruption event.

[0079] Figure 4 A flowchart of a voice content association determination method provided in this application embodiment Figure 4 .based on Figures 1 to 3 In any of the embodiments shown, step S102 includes: Step S401: Obtain content association parameters between the first speech and the second speech. The content association parameters include at least one of the following: scene-aware time decay value, semantic similarity, intent continuity, and content overlap between the first speech and the second speech.

[0080] In this embodiment, the device acquires content association parameters between the first speech and the second speech. The content association parameters include at least one of the following: scene-aware time decay value, semantic similarity, intent continuity, and content overlap between the first speech and the second speech.

[0081] The scene-aware time decay value refers to the degree of relevance of the driving scene to the user's multi-turn speech, and can be determined based on the driving scene. For example, the scene-aware time decay value f_time = exp(-λ_scene × ΔT), where λ_scene is the decay parameter and ΔT is the average speech interval duration of the user. λ_scene can be determined based on the driving scene: 0.025 for highway driving, 0.05 for city driving, and 0.075 for parked driving.

[0082] Semantic similarity refers to the similarity between the textual feature vectors of two speech words.

[0083] Intent continuity refers to the parameter determined by the device after recognizing the intent and performing a lookup table. For example, if the second voice contains "navigation" and the first voice contains "confirm", the intent continuity obtained by looking up the table is 0.9; if the first voice contains "play music", "navigation" and "play music" are obviously mismatched, the intent continuity obtained by looking up the table is 0.1.

[0084] Content overlap refers to entity overlap, where entities are such as specific locations or names. It is determined by the ratio between the number of entity intersections and the number of entity unions. For example, if the second speech contains "County A" and "Shopping Mall C", and the first speech contains "County A" and "Gas Station B", then the entity intersection is "County A", and the entity union is "County A", "Gas Station B", and "Shopping Mall C".

[0085] Step S402: Determine auxiliary correlation parameters based on the user's vehicle driving environment data. The auxiliary correlation parameters include at least one of the following: user attention, driving scene stability, impact of voice interruption events, spatial correlation, and user interaction correlation.

[0086] The device then determines auxiliary correlation parameters based on the user's vehicle driving environment data. These auxiliary correlation parameters include at least one of the following: user attention, driving scene stability, impact of voice interruption events, spatial correlation, and user interaction correlation.

[0087] User attention can be assessed by whether the user checks their phone while driving, whether their hands are off the steering wheel, and whether their gaze is fixed on the road ahead. If any of these occur, points are deducted from the maximum score. The final score represents the user's attention level.

[0088] Driving scenario stability characterizes the degree of vehicle driving stability and can be determined by the rate of change of vehicle speed and the rate of change of steering wheel speed. For example, driving scenario stability = (rate of change of vehicle speed + rate of change of steering wheel speed) / 2.

[0089] The impact of voice interruption events can be determined by looking up a table. Voice interruption events include incoming calls, prompts, and navigation announcements. If the voice interruption event is an incoming call, the impact is 1; if it is a prompt, the impact is 0.8; and if it is a navigation announcement, the impact is 0.7.

[0090] Spatial relevance is determined by whether the vehicle is located in the region of interest and whether the specified word is present. For example, if the distance between the vehicle and the region of interest is less than a distance threshold and the first speech contains the specified word, the spatial relevance is 1; if the distance between the vehicle and the region of interest is greater than or equal to the distance threshold, and / or the first speech does not contain the specified word, the spatial relevance is 0. The specified word can be an entity in the second speech.

[0091] User interaction relevance refers to the numerical value derived from the interaction between the user and the vehicle. For example, a user says "Recommend nearby Sichuan restaurants" (second voice), the vehicle screen displays a list of restaurants, the user taps the list to browse, and 3 seconds later says "Second" (first voice). In this case, the two voices are interactive and related, therefore, the user interaction relevance is 1. As another example, a user says "Recommend nearby Sichuan restaurants" (second voice), the screen displays a list, but the user doesn't touch the screen, and 10 seconds later says "Second" (first voice). In this case, the user saying the second voice timed out and didn't interact with the list, therefore, the user interaction relevance is 0.

[0092] Step S403: Determine the association score between the first speech and the second speech based on the content association parameters, the corresponding weights, the auxiliary association parameters, and the corresponding weights.

[0093] After determining the content association parameters and auxiliary association parameters, a weighted calculation is performed based on the content association parameters, their corresponding weights, the auxiliary association parameters, and their corresponding weights to obtain the association score. The association score is expressed as: Score=Σ(w_i×f_i), where each f_i is an auxiliary association parameter and a content association parameter, and w_i is a weight.

[0094] Furthermore, the device can be configured with hierarchical judgment rules, including rapid judgment and in-depth evaluation. In-depth evaluation refers to determining the association score through content association parameters and auxiliary association parameters, and determining whether the first speech is related to the second speech by comparing the association score with the association threshold. This is suitable for complex scenarios in vehicles. Rapid judgment is suitable for simple scenarios in vehicles. The distinction between simple and complex scenarios is determined by the decay duration of the scenario. A longer decay duration indicates a longer duration and thus a complex scenario; a shorter decay duration indicates a shorter duration and thus a simple scenario.

[0095] In response, the device determines the scene decay time based on the vehicle's speed. For example, when the vehicle speed is high speed, λ_scene is 0.025, and the scene decay time = ln(2) / 0.025 = 0.693 / 0.025 = 27.7 milliseconds; when the vehicle speed is urban driving speed, λ_scene is 0.05, and the scene decay time = ln(2) / 0.05 = 0.693 / 0.05 = 13.9 milliseconds; when the vehicle speed is stationary, λ_scene is 0.075, and the scene decay time = ln(2) / 0.075 = 0.693 / 0.075 = 9.2 milliseconds.

[0096] If the scene decay duration is greater than or equal to the decay threshold, a deep evaluation is performed, that is, the content association parameters between the first speech and the second speech are obtained. In this case, steps S401 to S403 are executed. The decay threshold can be any suitable value, for example, a decay threshold of 10 milliseconds.

[0097] If the scene decay duration is less than the decay threshold, a rapid judgment is made, that is, based on the acquisition interval between the first and second speech, to determine whether the first and second speech are related in content. For example, when the acquisition interval is less than a first preset interval and the intent of the first and second speech is the same, the first and second speech are related. The preset interval can be any suitable value, such as 2 seconds. When the acquisition interval is longer than a second preset interval and the first speech is not interrupted, the first and second speech are unrelated. The second preset interval is longer than the first speech interval, for example, the second preset interval is 60 seconds. In addition, if a transition word is detected in the first speech, the first and second speech are unrelated. In cases other than the above three, the correlation score is determined by the content correlation parameter between the first and second speech, and then the correlation score is compared with the correlation threshold to determine whether the first and second speech are related.

[0098] In this embodiment, statistical analysis shows that determining the association score through content association parameters and auxiliary association parameters has the following technical effects: The accuracy of the correlation score was improved by 96%; the accuracy of multimodal scenarios increased from 52% to 88%; the accuracy of GPS location-dependent scenarios for vehicles increased from 45% to 88% (+43%); and the average latency of responding to user voice in vehicles decreased from 80-120ms to 22ms, meeting real-time requirements.

[0099] Figure 5 A flowchart of a voice content association determination method provided in this application embodiment Figure 5 .based on Figures 1 to 4 In any of the embodiments shown, step S102 includes: Step S501: In response to the detection of a voice interruption event, determine the intermediate score between the first voice and the second voice.

[0100] In this embodiment, when the vehicle detects a voice interruption event targeting the first voice, such as an incoming call, vehicle notification sound, or navigation announcement, compensation is needed to adjust the correlation score between the first and second voices. To this end, the device first determines the score between the first and second voices as an intermediate score. This intermediate score can be... Figure 4 The scores determined by the content association parameters and auxiliary association parameters in the illustrated embodiment can also be... Figure 1 The associated score determined in the illustrated embodiment.

[0101] Step S502: Determine the third correction value based on the voice interruption event, and increase the intermediate score based on the third correction value to obtain the associated score.

[0102] The device determines the correction value corresponding to the voice interruption event, and this correction value is defined as the third correction value.

[0103] For example, the third correction value is determined by the voice interruption event. If the voice interruption event is an incoming call which is a strong interruption, the third correction value is a larger first value. If the voice interruption event is a prompt tone which is a medium to strong interruption, the third correction value is a second value, which is smaller than the first value. If the voice interruption event is a navigation broadcast which is a medium interruption, the third correction value is a smaller third value, which is smaller than the second value.

[0104] In another example, the device determines a third correction value based on the interruption type and duration of the voice interruption event. For example, the third correction value is: 0.20×w_interrupt×exp((-t_since_interrupt) / T_protect).

[0105] Where w_interrupt is the compensation coefficient corresponding to the interruption type, t_since_interrupt is the interruption duration, T_protect is the interruption duration, and t_since_interrupt is the parameter corresponding to the level of the interruption type.

[0106] For example, when the voice interruption event is an incoming call, which is a strong interruption, then w_interrupt=1.0 and T_protect=300 seconds; When the voice interruption event is a notification tone, it is considered a medium-strong interruption, with w_interrupt=0.8 and T_protect=180 seconds; When the voice interruption event is a navigation broadcast, it is considered a medium interruption, with w_interrupt=0.7 and T_protect=120 seconds.

[0107] After obtaining the third correction value, the intermediate score can be increased by the third correction value to obtain the correlation score, which is: score_base+0.20×w_interrupt×exp((-t_since_interrupt) / T_protect).

[0108] Here, score_base is the median score.

[0109] In this embodiment, when a voice interruption event is detected, the correlation score between the first voice and the second voice is compensated based on the voice interruption event, thereby improving the accuracy of the judgment on whether the first voice and the second voice are related.

[0110] In one embodiment, upon detecting a speech interruption event, a correlation threshold is determined based on threshold data and the speech interruption event, and a correlation score between the first speech and the second speech is determined based on the speech interruption event. It is understood that upon detecting a speech interruption event, both the correlation score and the correlation threshold are compensated, i.e., the correlation score is increased while the correlation threshold is decreased. The specific details of determining the correlation threshold based on threshold data and the speech interruption event, and determining the correlation score based on the speech interruption event, are as described above and will not be repeated here.

[0111] It should be noted that in the process of calculating the association threshold and association score, the weights {w_i} and coefficients {α,β,γ,δ} will be optimized simultaneously. These parameters can be optimized simultaneously through the model.

[0112] The joint optimization objective of the model is: L_total = L_classification + λ_stability·L_threshold_stability, ensuring that the distribution of association scores changes in tandem with the distribution of association thresholds.

[0113] Where L_total is the total loss function, the overall objective function for joint optimization, used to simultaneously optimize classification accuracy and threshold stability; L_classification is the classification loss, used for the accuracy loss in context-based judgment. It is calculated as: Cross-entropy loss = -Σ[y·log( )+(1-y)log(1- )], where y is the real label (1 = relevant, 0 = irrelevant). It is the prediction result; the role of classification loss is to ensure high accuracy in judgment.

[0114] λ_stability is a balancing coefficient that controls the weight balance of the two loss terms. Its value ranges from 0.1 to 0.3 and can adjust the accuracy and stability.

[0115] L_threshold_stability is the threshold stability loss, a constraint term that ensures the distribution of the associated threshold score changes in tandem with the distribution of the associated threshold. Calculation method: L_threshold_stability = Var(Score - θ) + |E[Score_correlation] - E[θ] - margin|, where Var(Score - θ) represents the variance of the difference between the correlation score and the correlation threshold θ (the smaller the variance, the more stable the correlation), E[Score_correlation] refers to the average score of the correlated samples, E[θ] is the average threshold, and margin is the expected safety margin (usually 0.15-0.25). L_threshold_stability can prevent systematic misjudgments caused by the mismatch between the distribution of the correlation score and the correlation threshold θ.

[0116] In this embodiment, statistics show that the accuracy rate of voice interruption recovery increased from 45% to 87%, and the accuracy rate of recovery after incoming call increased from 35% to 88%.

[0117] Corresponding to the above-described method for determining the association of voice content, this application also provides a device for determining the association of voice content. Figure 6 This is a schematic diagram of a speech content association determination device provided in an embodiment of this application. The speech content association determination device 600 provided in this embodiment includes: The acquisition module 610 is used to acquire the first voice currently collected by the vehicle and the second voice collected previously, and to acquire the threshold determination data of the vehicle. The threshold determination data includes at least one of the following: the vehicle's driving scenario, the historical voice data of the user driving the vehicle, and the user's driving status data. The first determining module 620 is used to determine the association threshold based on the threshold determining data, and to determine the association score between the first speech and the second speech. The second determining module 630 is used to determine whether the first speech is related to the second speech in terms of content based on the comparison result between the association score and the association threshold.

[0118] In some implementations, the voice content association determination device 600 is also used for: Obtain the coefficients and first correction values ​​corresponding to the sub-data in the threshold determination data. Each sub-data includes driving scenario, historical voice data, and driving status data. Based on the coefficients corresponding to the sub-data and the first correction value, determine the target correction value corresponding to the sub-data; The associated threshold is obtained by correcting the preset threshold based on the target correction value of each sub-data.

[0119] In some implementations, the voice content association determination device 600 is also used for: In response to the detection of a voice interruption event, the preset threshold is corrected according to the target correction value of each sub-data to obtain an intermediate threshold; Determine the second correction value based on the voice interruption event; Based on the second correction value, the intermediate threshold is reduced to obtain the associated threshold.

[0120] In some implementations, the voice content association determination device 600 is also used for: Determine the initial correction value based on the voice interruption event; The compensation coefficient is determined based on the interruption duration of the voice interruption event, and the compensation coefficient is negatively correlated with the interruption duration. The second correction value is determined based on the compensation coefficient and the initial correction value.

[0121] In some implementations, the voice content association determination device 600 is also used for: Obtain content association parameters between the first speech and the second speech. The content association parameters include at least one of the following: scene-aware time decay value, semantic similarity, intent continuity, and content overlap between the first speech and the second speech. The auxiliary correlation parameters are determined based on the user's vehicle driving environment data. The auxiliary correlation parameters include at least one of the following: user attention, driving scenario stability, impact of voice interruption events, spatial correlation, and user interaction correlation. The association score between the first speech and the second speech is determined based on the content association parameters, their corresponding weights, auxiliary association parameters, and their corresponding weights.

[0122] In some implementations, the voice content association determination device 600 is also used for: In response to the detection of a voice interruption event, an intermediate score between the first and second voices is determined; The third correction value is determined based on the voice interruption event, and the intermediate score is increased based on the third correction value to obtain the correlation score.

[0123] In some implementations, the voice content association determination device 600 is also used for: Get the interruption type and interruption duration corresponding to the voice interruption event; The third correction value is determined based on the interruption type and interruption duration.

[0124] In some implementations, the voice content association determination device 600 is also used for: The scene decay duration of the vehicle is determined based on the vehicle's speed. In response to a scene decay duration greater than or equal to a decay threshold, the content association parameters between the first speech and the second speech are obtained.

[0125] In some implementations, the voice content association determination device 600 is also used for: In response to the scene decay duration being less than the decay threshold, the system determines whether the first and second voice recordings are related in content based on the acquisition interval between them.

[0126] In some implementations, the voice content association determination device 600 is also used for: In response to the detection of a voice interruption event, a correlation threshold is determined based on the threshold data and the voice interruption event, and the correlation score between the first voice and the second voice is determined based on the voice interruption event.

[0127] The speech content association determination device and the speech content association determination method provided in the above embodiments of this application belong to the same application concept. They can execute the speech content association determination method provided in any of the above embodiments of this application, and have the corresponding functional modules and beneficial effects for executing the speech content association determination method. Technical details not described in detail in this embodiment can be found in the specific processing content of the speech content association determination method provided in the above embodiments of this application, and will not be repeated here.

[0128] The functions implemented by each module in the voice content association determination device can be implemented by the same or different processors, and this application embodiment does not limit this.

[0129] It should be understood that the modules in the above-mentioned voice content association determination device can be implemented in the form of processor calling firmware. For example, the system includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each module of the device. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal to the device or external to the system. Alternatively, the modules in the system can be implemented in the form of hardware circuits. By designing the hardware circuits, some or all of the module functions can be implemented. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above modules are implemented by designing the logical relationships of the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files to implement the functions of some or all of the above modules. All modules of the above-mentioned voice content association determination device can be implemented entirely by processor calling firmware, or entirely by hardware circuits, or partially by processor calling firmware with the remaining parts implemented by hardware circuits.

[0130] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.

[0131] As can be seen, each module in the above voice content association determination device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor types.

[0132] Furthermore, the modules in the above-mentioned voice content association determination device can be integrated in whole or in part, or they can be implemented independently. In one implementation, these modules are integrated together and implemented in the form of a System-on-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the modules of the device. The at least one processor can be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.

[0133] This application provides a structural schematic diagram of a vehicle, see [link / reference] Figure 7 As shown, the vehicle includes a memory 700 and a processor 710; wherein the memory 700 is connected to the processor 710 and is used to store programs; the processor 710 is used to implement the voice content association determination method disclosed in any of the above embodiments by running the programs stored in the memory 700.

[0134] Specifically, the vehicle may also include: a bus, a communication interface 720, an input device 730, an output device 740, and a voice content association determination device 750. The vehicle may also include a data transceiver module, an image monitoring module, and a signal monitoring module.

[0135] The processor 710, memory 700, communication interface 720, input device 730, output device 740, and voice content association determination device 750 are interconnected via a bus. Among them: A bus can include a pathway for transmitting information between various components in a vehicle.

[0136] The processor 710 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0137] The processor 710 may include a main processor, as well as a baseband chip, modem, etc.

[0138] The memory 700 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 700 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0139] Input device 730 may include a device for receiving data and information input by a user, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0140] Output device 740 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0141] The communication interface 720 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0142] The processor 710 executes the program stored in the memory 700 and calls other devices, which can be used to implement each step of any of the voice content association determination methods provided in the above embodiments of this application.

[0143] It should be noted that the vehicle can be an in-vehicle terminal, mobile phone, wearable device or server, etc.; or it can be a vehicle that includes an in-vehicle terminal, etc.

[0144] This application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in the memory through the data interface to execute the speech content association determination method described in any of the above embodiments. For the specific processing procedure and its beneficial effects, please refer to the embodiments of the speech content association determination method described above.

[0145] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the speech content association determination method according to various embodiments of this application as described in any of the above embodiments of this specification.

[0146] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the power device, as a standalone firmware package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0147] Furthermore, embodiments of this application may also be storage media storing computer programs, which are executed by a processor in the speech content association determination method according to various embodiments of this application described in any of the above embodiments of this specification, specifically implementing the steps of the speech content association determination method described above.

[0148] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0149] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0150] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0151] The units of the apparatus in the various embodiments of this application can be merged, divided, and deleted according to actual needs.

[0152] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0153] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0154] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or as firmware functional modules or sub-modules.

[0155] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer firmware, or a combination of both. To clearly illustrate the interchangeability of hardware and firmware, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or firmware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0156] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly using hardware, firmware units executed by a processor, or a combination of both. The firmware unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0157] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0158] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for determining the association of speech content, characterized in that, include: The system acquires the first voice currently collected by the vehicle and the second voice collected previously, and acquires threshold determination data for the vehicle. The threshold determination data includes at least one of the vehicle's driving scenario, the historical voice data of the user driving the vehicle, and the user's driving status data. Based on the threshold, the association threshold is determined, and the association score between the first speech and the second speech is determined. Based on the comparison between the association score and the association threshold, it is determined whether the first speech and the second speech are associated in terms of content.

2. The method for determining speech content association according to claim 1, characterized in that, The step of determining the association threshold based on the threshold data includes: Obtain the coefficients and first correction values ​​corresponding to the sub-data in the threshold determination data, wherein each sub-data is the driving scenario, the historical voice data, and the driving state data; Based on the coefficients corresponding to the sub-data and the first correction value, determine the target correction value corresponding to the sub-data; The associated threshold is obtained by correcting the preset threshold based on the target correction value of each of the sub-data.

3. The speech content association method according to claim 1, characterized in that, The step of determining the association threshold based on the threshold data includes: In response to the detection of a voice interruption event, data is determined based on the threshold, and an intermediate threshold is determined. A second correction value is determined based on the voice interruption event; The association threshold is obtained by reducing the intermediate threshold based on the second correction value.

4. The speech content association method according to claim 3, characterized in that, determining the second correction value based on the speech interruption event includes: Determine the initial correction value based on the voice interruption event; A compensation coefficient is determined based on the interruption duration of the voice interruption event, and the compensation coefficient is negatively correlated with the interruption duration. The second correction value is determined based on the compensation coefficient and the initial correction value.

5. The method for determining speech content association according to claim 1, characterized in that, Determining the association score between the first speech and the second speech includes: Obtain content association parameters between the first speech and the second speech, wherein the content association parameters include at least one of the following: scene-aware time decay value, semantic similarity, intent continuity, and content overlap between the first speech and the second speech; The auxiliary correlation parameters are determined based on the user's vehicle driving environment data. The auxiliary correlation parameters include at least one of the following: user attention, driving scene stability, impact of voice interruption events, spatial correlation, and user interaction correlation. The association score between the first speech and the second speech is determined based on the content association parameters, the corresponding weights, the auxiliary association parameters, and the corresponding weights.

6. The method for determining speech content association according to claim 1, characterized in that, Determining the association score between the first speech and the second speech includes: In response to the detection of a voice interruption event, an intermediate score between the first voice and the second voice is determined; A third correction value is determined based on the voice interruption event, and the intermediate score is increased based on the third correction value to obtain the associated score.

7. The method for determining speech content association according to claim 6, characterized in that, Determining the third correction value based on the voice interruption event includes: Obtain the interruption type and interruption duration corresponding to the voice interruption event; The third correction value is determined based on the interruption type and the interruption duration.

8. The method for determining speech content association according to claim 5, characterized in that, The step of obtaining the content association parameters between the first speech and the second speech includes: The scene decay duration of the vehicle is determined based on the vehicle's speed. In response to the scenario attenuation duration being greater than or equal to the attenuation threshold, the content association parameters between the first speech and the second speech are obtained.

9. The method for determining speech content association according to claim 8, characterized in that, After determining the scene decay duration of the vehicle based on the vehicle speed, the method further includes: In response to the scenario attenuation duration being less than the attenuation threshold, it is determined whether the first speech and the second speech are related in content based on the acquisition interval duration between the first speech and the second speech.

10. The method for determining speech content association according to claim 1, characterized in that, The step of determining the association threshold based on the threshold data and determining the association score between the first speech and the second speech includes: In response to the detection of a voice interruption event, an association threshold is determined based on the threshold determined data and the voice interruption event, and an association score between the first voice and the second voice is determined based on the voice interruption event.

11. A device for determining the association of voice content, characterized in that, Including memory and processor, among which, The memory is connected to the processor and is used to store programs; The processor is used to implement the voice content association determination method as described in any one of claims 1-10 by running the program in the memory.

12. A vehicle, characterized in that, The vehicle includes a voice content association determination device, which implements the voice content association determination method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the voice content association determination method as described in any one of claims 1-10.

14. A computer program product, wherein when executed by a processor, the computer program implements the speech content association determination method as described in any one of claims 1-10.