Parameter management method of vehicle-mounted voice interaction system and electronic equipment
By dynamically adjusting the end time of voice activity detection, based on the vehicle driving state and user status, the problem of premature voice interaction ending in the vehicle voice interaction system is solved, and the accuracy and user experience of the interaction are improved.
Patent Information
- Application Number
- CN202510365099.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-06
AI Technical Summary
The fixed end time of voice activity detection in the vehicle voice interaction system may lead to the premature end of voice interaction and affect the user experience.
By obtaining the vehicle's driving status information and/or the user's user status information, the voice activity detection end time is dynamically adjusted to ensure that the user can complete the voice input.
Reduce or avoid the situation where voice interaction ends prematurely, improve the accuracy and user experience of voice interaction, improve the timeliness and accuracy of vehicle voice control, and enhance driving safety.
Smart Images

Figure CN120108435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle technology, and in particular to a parameter management method and electronic equipment for a vehicle-mounted voice interaction system. Background Art
[0002] Voice Activity Detection (VAD) technology is a technology used for voice processing, the purpose of which is to detect whether a voice signal exists. Currently, the parameter of the VAD end time of the in-vehicle voice interaction system is fixed. Therefore, the current in-vehicle voice interaction system adopts a fixed post-voice interaction endpoint strategy to end the voice input. That is, if the in-vehicle voice interaction system does not detect new user voice within the fixed VAD end time, the voice interaction will end. However, in actual driving, there are often situations where users fail to complete the input of voice commands in time. Therefore, based on the fixed VAD end time, that is, the fixed post-voice interaction endpoint, ending the voice interaction may result in the voice interaction being ended prematurely, affecting the user experience. Summary of the invention
[0003] The purpose of the present invention is to solve the problem that when a fixed VAD end time, that is, a fixed endpoint strategy after voice interaction, is used in a vehicle to end the voice interaction, the voice interaction may be ended prematurely, thereby affecting the user experience.
[0004] To solve the above technical problems, in the first aspect, an embodiment of the present invention discloses a parameter management method for an in-vehicle voice interaction system, the method comprising: obtaining target information, the target information including driving status information of the vehicle and / or user status information of users in the vehicle; and adjusting the voice activity detection end time of the in-vehicle voice interaction system according to the target information.
[0005] When the above method is used, the end time of the voice activity detection of the vehicle-mounted voice interaction system can be adjusted based on the obtained driving state information of the vehicle and / or the user state information of the user in the vehicle. By adjusting the end time of the voice activity detection of the vehicle-mounted voice interaction system based on the driving state information of the vehicle and / or the user state information of the user, the end time of the voice activity detection of the vehicle-mounted voice interaction system can be dynamically adjusted, which can reduce or avoid the situation where the voice interaction is terminated prematurely, thereby ensuring that the user can fully complete the voice input, improve the accuracy of the voice interaction, and enhance the user experience. In addition, the timeliness and accuracy of the vehicle voice control can be improved, and the vehicle driving safety can be improved.
[0006] According to the above claims and related contents, in a parameter management method of an in-vehicle voice interaction system disclosed in an embodiment of the present invention, when the target information includes driving status information, the voice activity detection end time of the in-vehicle voice interaction system is adjusted according to the target information, including: inputting the driving status information into a driving scene recognition model to obtain the driving scene of the vehicle; and determining the voice activity detection end time of the in-vehicle voice interaction system according to the driving scene of the vehicle.
[0007] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, the voice activity detection end time is determined according to the driving scenario of the vehicle, so that the voice activity detection end time corresponds to the driving scenario, thereby having corresponding voice activity detection end times in different driving scenarios. This can effectively ensure that when the vehicle is in different driving scenarios, the user can still complete the voice input completely, thereby improving the accuracy of voice interaction and improving the user experience.
[0008] According to the above claims and related contents, in a parameter management method of an in-vehicle voice interaction system disclosed in an embodiment of the present invention, when the target information includes user status information, the end time of the voice activity detection of the in-vehicle voice interaction system is adjusted according to the target information, including: inputting the user status information into a user status recognition model to obtain the user status; and determining the end time of the voice activity detection of the in-vehicle voice interaction system according to the user status.
[0009] When the parameter management method of the vehicle-mounted voice interaction system is adopted, the voice activity detection end time is determined according to the user status, that is, the voice activity detection end time can be changed based on the user status. Therefore, when the user status is different, the voice activity detection end time of the voice interaction system is also different, so as to ensure that the user can still complete the voice input completely when the user status is different, improve the accuracy of voice interaction, and improve the user experience.
[0010] According to the above claims and related contents, in a parameter management method of an in-vehicle voice interaction system disclosed in an embodiment of the present invention, when the target information includes driving status information and user status information, the voice activity detection end time of the in-vehicle voice interaction system is adjusted according to the target information, including: inputting the driving status information into a driving scene recognition model to obtain the driving scene of the vehicle; inputting the user status information into a user status recognition model to obtain the user status; and determining the voice activity detection end time of the in-vehicle voice interaction system according to the driving scene and the user status.
[0011] When the parameter management method of the in-vehicle voice interaction system is adopted, the end time of the voice activity detection is determined according to the driving scenario and the user status, that is, the driving scenario and the user status are combined to determine the end time of the voice activity detection, so that the end time of the voice activity detection of the in-vehicle voice interaction system is more accurate.
[0012] According to the above claims and related contents, in a parameter management method of an in-vehicle voice interaction system disclosed in an embodiment of the present invention, the voice activity detection end time of the in-vehicle voice interaction system is adjusted according to the driving scenario and the user status, including: determining the initial voice activity detection end time according to the driving scenario; adjusting the initial voice activity detection end time according to the user status to obtain the voice activity detection end time.
[0013] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, after determining the end time of the initial voice activity detection according to the driving scenario, the end time of the initial voice activity detection can be further adjusted according to the user status to determine the end time of the voice activity detection of the current in-vehicle voice interaction system, so that the obtained voice activity detection end time is more accurate.
[0014] According to the above claims and related contents, in a parameter management method of an in-vehicle voice interaction system disclosed in an embodiment of the present invention, the user state includes the user attention state, and the voice activity detection end time in the in-vehicle voice interaction system is determined according to the driving scene and the user state, including determining by the following formula:
[0015]
[0016] in, is the end time of voice activity detection, T base is the preset end time of basic voice activity detection, is the sensitivity coefficient corresponding to the driving scene, Z is F attention or A score , F attention is the user attention coefficient corresponding to the user attention state, A score is the user attention score corresponding to the user attention state, and i is the number of the driving scenes.
[0017] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, the voice activity detection end time corresponding to the current driving scene of the vehicle and the current user state of the user can be accurately calculated based on the preset basic voice activity detection end time, the sensitivity coefficient corresponding to the driving scene and the user attention coefficient or user attention score corresponding to the user attention state, so as to ensure that the user can complete the voice input completely and improve the accuracy of voice interaction.
[0018] According to the above claims and related contents, a parameter management method of a vehicle voice interaction system disclosed in an embodiment of the present invention also includes obtaining a user attention score A by the following formula: score :
[0019]
[0020] Wherein, W is a preset weight, b is a preset bias parameter, and X is a feature vector calculated based on the user state.
[0021] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, the user attention score can be accurately calculated based on the preset weight and bias parameters and the feature vector calculated based on the user state information, and the end time of the voice activity detection can be adjusted based on the user attention score, so that the end time of the voice activity detection is more in line with the time required for voice interaction in the current state of the user.
[0022] According to the above claims and related contents, in a parameter management method of an in-vehicle voice interaction system disclosed in an embodiment of the present invention, driving status information includes vehicle motion status information collected by a vehicle motion sensor, and the vehicle motion sensor includes an accelerometer, a gyroscope and a positioning speed sensor; user status information includes user facial status information collected by a vehicle camera, and the user facial status information includes facial expression information and / or eye movement information.
[0023] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, multiple vehicle motion state information of the vehicle can be collected based on the accelerometer, gyroscope and positioning speed sensor to better determine the vehicle's driving state information and accurately calculate the voice activity detection end time corresponding to the vehicle's current driving state.
[0024] When the parameter management method of the in-vehicle voice interaction system is adopted, the user status information can be accurately obtained based on the user facial status information collected by the vehicle camera to better adjust the end time of the voice activity detection.
[0025] According to the above claims and related contents, a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present invention also includes: obtaining a speech completion rate corresponding to the in-vehicle voice interaction system; when the ratio of the speech completion rate to the voice activity detection end time is lower than a preset threshold, adjusting the voice activity detection end time of the in-vehicle voice interaction system according to the target information.
[0026] When adopting the parameter management method of the above-mentioned in-vehicle voice interaction system, the end time of the voice activity detection in the in-vehicle voice interaction system can be further adjusted based on the comparison of the ratio of the voice completeness rate corresponding to the in-vehicle voice interaction system to the end time of the voice activity detection with a preset threshold, so that the end time of the voice activity detection in the in-vehicle voice interaction system can ensure that the time when the user inputs voice is within the end time of the voice activity detection under the current driving state of the vehicle and the current user state of the user in the vehicle.
[0027] In a second aspect, an embodiment of the present invention further discloses an electronic device, which includes: a processor, a memory communicatively connected to the processor, the memory storing computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the electronic device implements a parameter management method for an in-vehicle voice interaction system such as any one of the above-mentioned items.
[0028] When using the above-mentioned electronic device, the processor can adjust the end time of the voice activity detection in the vehicle-mounted voice interaction system based on the vehicle's driving status information and / or the user status information of the user in the vehicle to avoid the voice interaction process from ending prematurely, thereby ensuring that the user can fully complete the voice input, improve the accuracy of the voice interaction, and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A flow chart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application;
[0030] Figure 2 Another flowchart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application;
[0031] Figure 3 Another flowchart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application;
[0032] Figure 4 Another flowchart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application;
[0033] Figure 5 Another flowchart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application;
[0034] Figure 6 Another flowchart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application;
[0035] Figure 7 Another flowchart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0036] In the voice interaction system, the VAD end time is fixed. Therefore, the voice interaction system uses a fixed post-voice interaction endpoint strategy to end the voice input. The post-voice interaction endpoint strategy is used to determine the end position of the voice signal in the voice interaction system. However, during the voice interaction process, unexpected situations may occur that cause the user to be unable to complete the input of the voice command in time, affecting the user experience.
[0037] Furthermore, the in-vehicle voice interaction system is used for human-vehicle interaction as an example for explanation. When the user is driving the vehicle, if the driving scene suddenly changes (for example, sudden braking or sharp turns), the user's attention will be focused on driving the vehicle. At this time, the user may not be able to input voice commands in time, but the in-vehicle voice interaction system uses a fixed post-voice interaction endpoint strategy to end the voice, that is, if the in-vehicle voice interaction system does not detect new user voice within the fixed VAD end time, it will end the voice interaction. Therefore, it is easy to cause the voice interaction to be terminated prematurely, making the voice interaction inaccurate, thereby affecting the user experience. Furthermore, if the user controls the automatic driving of the vehicle based on voice commands, if the user fails to complete the input of the voice command in time, the voice interaction will be terminated prematurely, which is prone to driving safety problems.
[0038] Based on the above problems, the present application proposes a parameter management method and electronic device for an in-vehicle voice interaction system. In the parameter management method for the in-vehicle voice interaction system, the voice activity detection end time of the in-vehicle voice interaction system can be adjusted based on the obtained driving status information of the vehicle and / or the user status information of the user in the vehicle, so that the voice activity detection end time of the in-vehicle voice interaction system can be dynamically adjusted, which can reduce or avoid the situation where the voice interaction is terminated prematurely, thereby ensuring that the user can complete the voice input completely, improving the accuracy of the voice interaction, and enhancing the user experience.
[0039] refer to Figure 1 , Figure 1 A flow chart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application. In a first aspect, an embodiment of the present invention discloses a parameter management method for an in-vehicle voice interaction system, the method comprising:
[0040] Step S100: Acquire target information, where the target information includes driving status information of the vehicle and / or user status information of users in the vehicle.
[0041] Step S200: adjusting the voice activity detection end time of the in-vehicle voice interaction system according to the target information.
[0042] By adopting the above method, the end time of the voice activity detection of the vehicle-mounted voice interaction system can be adjusted based on the obtained driving state information of the vehicle and / or the user state information of the user in the vehicle. By adjusting the end time of the voice activity detection of the vehicle-mounted voice interaction system based on the driving state information of the vehicle and / or the user state information of the user, the end time of the voice activity detection of the vehicle-mounted voice interaction system can be dynamically adjusted, which can reduce or avoid the situation where the voice interaction is terminated prematurely, thereby ensuring that the user can fully complete the voice input, improve the accuracy of the voice interaction, and enhance the user experience. In addition, the timeliness and accuracy of the vehicle voice control can also be improved, and the vehicle driving safety can be improved.
[0043] The parameter management method of the in-vehicle voice interaction system disclosed in the embodiment of the present invention can be applied to the voice interaction system in the server of the vehicle. In the case where the parameter management method of the in-vehicle voice interaction system is applied to the voice interaction system, the information collection module in the vehicle can collect the information of the users in the vehicle in real time, and determine the user status information based on the collected user information. In addition, the voice interaction system can collect the driving status information of the vehicle in real time. In the case where the vehicle obtains the user status information and / or the driving status information of the vehicle, the end time of the voice activity detection of the in-vehicle voice interaction system is determined based on the user status information and / or the driving status information of the vehicle.
[0044] In possible implementations of the parameter management method of the in-vehicle voice interaction system disclosed in the present invention, the voice interaction system may adjust the end time of the voice activity detection only based on the driving status information of the vehicle; the voice interaction system may also adjust the end time of the voice activity detection only based on the user status information of the user in the vehicle; the voice interaction system may also adjust the end time of the voice activity detection in combination with the driving status information of the vehicle and the user status information of the user in the vehicle.
[0045] Furthermore, the vehicle's driving state information may include: XYZ axis velocity, XYZ axis acceleration, steering angular velocity and angular acceleration. The user state information may include: user sitting posture.
[0046] refer to Figure 2 , Figure 2 Another flow chart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application. In a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present invention, when the target information includes driving state information, adjusting the end time of the voice activity detection in the in-vehicle voice interaction system according to the target information includes: step S211: inputting the driving state information into a driving scene recognition model to obtain the driving scene of the vehicle; step S212: determining the end time of the voice activity detection in the in-vehicle voice interaction system according to the driving scene of the vehicle.
[0047] In another specific embodiment of the present invention, the target information is the driving state information of the vehicle. The specific method for determining the end time of voice activity detection based on the driving state information includes step S211 and step S212.
[0048] Specifically, the specific steps of step S211 may include: inputting the driving status information into a preset driving scene recognition model, wherein the preset driving scene recognition model is provided with all driving scenes of the vehicle in different states, each driving scene has corresponding driving status information, and the driving scene recognition model may be provided in an in-vehicle voice interaction system or a server in the vehicle; based on the driving status information of the current vehicle, searching in the driving scene recognition model to determine the driving scene corresponding to the current driving status information.
[0049] Specifically, the specific steps of step S212 may include: according to the determined driving scene, the parameters corresponding to the driving scene may be determined, and based on the parameters corresponding to the driving scene, the end time of the voice activity detection in the vehicle-mounted voice interaction system may be calculated and determined. In addition, based on the correspondence between the driving scene and the activity detection time pre-set in the vehicle-mounted voice interaction system or in the server, when the driving scene is determined, the correspondence between the driving scene and the end time of the voice activity detection is called to determine the end time of the voice activity detection corresponding to the current driving scene. Among them, the correspondence between the driving scene and the end time of the voice activity detection may be the relationship between all driving scenes and the corresponding voice activity detection end time in the driving scene recognition model set in the vehicle-mounted voice interaction system or in the server. The end time of the voice activity detection corresponding to the driving scene can be determined based on multiple tests.
[0050] Furthermore, the driving scene recognition model is based on pre-obtaining all motion state data of the vehicle in different driving scenarios, analyzing and calculating all the obtained motion state data based on a machine learning algorithm, determining the vehicle's driving state information, and thereby setting the relationship between the driving scene and the corresponding driving state information. Based on the motion state data and the relationship between the driving scene and the corresponding driving state information, a corresponding neural network classifier is trained to identify different driving scenarios.
[0051] Furthermore, the driving scene recognition model can also be set in the security operation center in the vehicle.
[0052] Furthermore, the driving scenarios in the driving scenario recognition model include, but are not limited to: steady vehicle driving, sudden deceleration of the vehicle in a straight line, sudden steering of the vehicle, sudden acceleration of the vehicle, large arc turning of the vehicle, sudden braking of the vehicle and sharp turning of the vehicle.
[0053] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, the voice activity detection end time is determined according to the driving scenario of the vehicle, so that the voice activity detection end time corresponds to the driving scenario, thereby having corresponding voice activity detection end times in different driving scenarios. This can effectively ensure that when the vehicle is in different driving scenarios, the user can still complete the voice input completely, thereby improving the accuracy of voice interaction and enhancing the user experience.
[0054] refer to Figure 3 , Figure 3 Another flow chart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application. In a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present invention, when the target information includes user status information, adjusting the end time of the voice activity detection of the in-vehicle voice interaction system according to the target information includes: step S221: inputting the user status information into a user status recognition model to obtain the user status; step S222: determining the end time of the voice activity detection of the in-vehicle voice interaction system according to the user status.
[0055] In a parameter management method of an in-vehicle voice interaction system disclosed in another specific embodiment of the present invention, a specific method for determining the end time of voice activity detection according to user status information includes step S221 and step S222.
[0056] Specifically, the specific steps of step S221 may include: inputting the obtained user status information into a preset user status recognition model, wherein the preset user status recognition model is set with all the user statuses under different driving conditions, each user status has corresponding user status information, and the user status recognition model can be set in the vehicle-mounted voice interaction system or the server in the vehicle; searching in the user status recognition model based on the user status information to determine the user status corresponding to the current user status information.
[0057] Specifically, the specific steps of step S222 may include: according to the determined user state, the parameters corresponding to the user state may be determined, and based on the parameters corresponding to the user state, the end time of the voice activity detection of the in-vehicle voice interaction system may be calculated and determined. In addition, based on the correspondence between the user state and the activity detection time pre-set in the in-vehicle voice interaction system or in the server, when the user state is determined, the correspondence between the user state and the end time of the voice activity detection is called to determine the end time of the voice activity detection corresponding to the current user state. Among them, the correspondence between the user state and the end time of the voice activity detection may be the relationship between all user states and the corresponding end time of the voice activity detection in the user state recognition model set in the in-vehicle voice interaction system or in the server. The end time of the voice activity detection corresponding to the user state may be determined based on multiple experiments.
[0058] Furthermore, the user state recognition model is based on obtaining all state data of the user in different driving scenarios in advance, analyzing and calculating all the obtained user state data based on a preset algorithm, determining the user's state information, and thus setting the relationship between the user state and the corresponding user state information.
[0059] When the parameter management method of the vehicle-mounted voice interaction system is adopted, the voice activity detection end time is determined according to the user status, that is, the voice activity detection end time can be changed based on the user status. Therefore, when the user status is different, the voice activity detection end time in the voice interaction system is also different, so as to ensure that the user can still complete the voice input completely when the user status is different, improve the accuracy of voice interaction, and enhance the user experience.
[0060] refer to Figure 4 , Figure 4 Another flow chart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application. In a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present invention, when the target information includes driving state information and user state information, adjusting the voice activity detection end time of the in-vehicle voice interaction system according to the target information includes:
[0061] Step S231: input the driving state information into the driving scene recognition model to obtain the driving scene of the vehicle; input the user state information into the user state recognition model to obtain the user state.
[0062] Step S232: Determine the end time of the voice activity detection in the in-vehicle voice interaction system according to the driving scenario and the user status.
[0063] In a parameter management method for an in-vehicle voice interaction system disclosed in another specific embodiment of the present invention, when the driving scene of the vehicle is obtained based on the driving state information and the user state is obtained based on the user state information, the end time of the voice activity detection is determined according to the driving scene and the user state. There is no order of steps between obtaining the driving scene of the vehicle based on the driving state information and obtaining the user state based on the user state information. The driving scene of the vehicle can be obtained based on the driving state information first, and then the user state can be obtained based on the user state information; or the user state can be obtained based on the user state information first, and then the driving scene of the vehicle can be obtained based on the driving state information; or the driving scene of the vehicle can be obtained based on the driving state information and the user state can be obtained based on the user state information at the same time.
[0064] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, the end time of the voice activity detection is determined according to the driving scenario and the user status, that is, the driving scenario and the user status are combined for analysis and calculation to determine the end time of the voice activity detection, so that the end time of the voice activity detection in the in-vehicle voice interaction system is more accurate.
[0065] refer to Figure 5 , Figure 5 Another flow chart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application. In a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present invention, adjusting the voice activity detection end time of the in-vehicle voice interaction system according to the driving scene and the user status includes:
[0066] Step S241: Determine the end time of the initial voice activity detection according to the driving scenario.
[0067] Step S242: adjusting the initial voice activity detection end time according to the user status to obtain the voice activity detection end time.
[0068] In a parameter management method of an in-vehicle voice interaction system disclosed in another specific embodiment of the present invention, when the driving scenario of the vehicle is determined, the end time of the initial voice activity detection is determined based on step S241, and when the user status is determined, the end time of the voice activity detection is determined based on step S242.
[0069] After determining the end time of the initial voice activity detection based on the driving scenario of the vehicle, the end time of the initial voice activity detection can be adjusted based on the user state of the user in the vehicle to obtain the end time of the voice activity detection. When adjusting the end time of the initial voice activity detection based on the user state of the user, a user state influence factor is calculated based on a user state parameter included in the user state, and the end time of the initial voice activity detection is adjusted based on the user state influence factor to obtain the end time of the voice activity detection.
[0070] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, after determining the initial voice activity detection end time according to the driving scenario, the initial voice activity detection end time can be further adjusted according to the user status to determine the voice activity detection end time in the current in-vehicle voice interaction system, so that the obtained voice activity detection end time is more accurate.
[0071] According to another specific embodiment of the present invention, in a parameter management method of an in-vehicle voice interaction system disclosed in an embodiment of the present invention, the user state includes the user attention state, and the voice activity detection end time in the in-vehicle voice interaction system is determined according to the driving scene and the user state, including determining by the following formula:
[0072]
[0073] in, is the end time of voice activity detection, T base is the preset end time of basic voice activity detection, is the sensitivity coefficient corresponding to the driving scene, Z is F attention or A score , F attention is the user attention coefficient corresponding to the user attention state, A score is the user attention score corresponding to the user attention state, and i is the number of the driving scenes.
[0074] In a parameter management method for an in-vehicle voice interaction system disclosed in another specific embodiment of the present invention, the voice activity detection end time can be calculated based on the driving scene and the user state. Further, the voice activity detection end time can be calculated based on the sensitivity coefficient corresponding to the driving scene and the user attention coefficient or user attention score corresponding to the user attention state included in the user state.
[0075] Specifically, based on the sensitivity coefficient corresponding to the driving scenario and the user attention coefficient or user attention score corresponding to the user attention state included in the user state, the specific method for calculating the end time of the voice activity detection is as follows: based on formula (1), the voice activity detection end time can be calculated, and the basic voice activity detection end time preset in formula (1) can be the system default basic voice activity detection end time. The basic voice activity detection end time can be based on the voice activity detection end time required for user voice interaction when the vehicle is driving smoothly, and can also be determined based on the historical voice activity detection end time. The specific setting of the basic voice activity detection end time is not limited here, and can be set based on specific circumstances.
[0076] Further, The i in is the driving scene identifier. is the sensitivity coefficient corresponding to the driving scenario, which can be determined based on the driving scenario. The higher the user's attention required for the driving scenario, the higher the value of the sensitivity coefficient. For example, in the case of sudden braking of the vehicle, the sensitivity coefficient corresponding to the driving scenario can be set to 1.5; in the case of stable driving of the vehicle, the sensitivity coefficient corresponding to the driving scenario can be set to 0.2.
[0077] Furthermore, the user attention coefficient F attention Can be a fixed value, and F attention The value of can be determined based on historical user status data. The value range of can be 0 to 1. Among them, in the user attention coefficient Fattention When the user attention coefficient F is 0, it can be said that the user's attention is completely distracted. attention When it is 1, it can be said that the user's attention is fully focused.
[0078] Furthermore, the user attention coefficient and the user attention score are different parameters corresponding to the user attention state. In different implementations, the voice activity detection end time can be calculated based on the user attention coefficient or the user attention score. When calculating the voice activity detection end time, there is no fixed requirement for the selection of the user attention coefficient and the user attention score, and the user can determine it based on specific needs.
[0079] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, the voice activity detection end time corresponding to the current driving scene of the vehicle and the current user state of the user can be accurately calculated based on the preset basic voice activity detection end time, the sensitivity coefficient corresponding to the driving scene and the user attention coefficient or the user attention state score corresponding to the user attention state, so as to ensure that the user can complete the voice input completely and improve the accuracy of voice interaction.
[0080] According to another specific embodiment of the present invention, a parameter management method of a vehicle voice interaction system disclosed in an embodiment of the present invention further includes obtaining a user attention score A by the following formula: score :
[0081]
[0082] Wherein, W is a preset weight, b is a preset bias parameter, and X is a feature vector calculated based on the user state.
[0083] In another specific embodiment of the present invention, the parameter management method of the vehicle voice interaction system disclosed in the present invention, the user attention coefficient A score It can be calculated based on formula (2). In formula (2), the preset weight and the preset bias parameter are calculated based on the historical state data. X is a feature vector calculated based on the user state.
[0084] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, the user attention score can be accurately calculated based on the preset weight and bias parameters and the feature vector calculated based on the user state information, and the end time of the voice activity detection can be adjusted based on the user attention score, so that the end time of the voice activity detection is more in line with the time required for voice interaction in the current state of the user.
[0085] According to another specific embodiment of the present invention, in a parameter management method of a vehicle-mounted voice interaction system disclosed in an embodiment of the present invention, the driving state information includes vehicle motion state information collected by a vehicle motion sensor, and the vehicle motion sensor includes an accelerometer, a gyroscope, and a positioning speed sensor;
[0086] The user status information includes user facial status information collected by the vehicle camera, and the user facial status information includes facial expression information and / or eye movement information.
[0087] In a parameter management method of a vehicle-mounted voice interaction system disclosed in another specific embodiment of the present invention, driving status information is collected based on a vehicle motion sensor. Multiple vehicle motion sensors can be set in the vehicle, and the multiple motion sensors work together to collect different vehicle motion status data. While ensuring that the timestamps of all vehicle motion status data are the same, all vehicle motion status data are sent to a server or a vehicle-mounted voice interaction system based on the vehicle-mounted communication module, wherein the vehicle motion sensor may include: the vehicle motion sensor includes an accelerometer, a gyroscope and a positioning (Global Positioning System, GPS) speed sensor.
[0088] When the parameter management method of the above-mentioned in-vehicle voice interaction system is adopted, multiple vehicle motion state information of the vehicle can be collected based on the accelerometer, gyroscope and positioning speed sensor to better determine the vehicle's driving state information and accurately calculate the voice activity detection end time corresponding to the vehicle's current driving state.
[0089] In a parameter management method for an in-vehicle voice interaction system disclosed in another specific embodiment of the present invention, the user status information may include user facial expression information and / or eye movement information collected by a vehicle camera. Based on the user facial expression information and / or eye movement information, the user's attention status can be analyzed to obtain the user status.
[0090] When the parameter management method of the in-vehicle voice interaction system is adopted, the user status information can be accurately obtained based on the user facial status information collected by the vehicle camera to better adjust the end time of the voice activity detection.
[0091] refer to Figure 6 , Figure 6 Another flow chart of a parameter management method of a vehicle-mounted voice interaction system disclosed in an embodiment of the present application. A parameter management method of a vehicle-mounted voice interaction system disclosed in an embodiment of the present invention further includes: Step S300: Obtaining a speech integrity rate corresponding to the vehicle-mounted voice interaction system.
[0092] Step S400: When the ratio of the speech completeness rate to the speech activity detection end time is lower than a preset threshold, adjusting the speech activity detection end time of the in-vehicle voice interaction system according to the target information.
[0093] refer to Figure 6 After determining the end time of voice activity detection, the historical voice interaction data of the vehicle can be automatically monitored based on the large language model to obtain the voice completeness rate C. And the voice completeness rate of the current vehicle voice interaction system can be evaluated based on the voice completeness rate C. The voice completeness rate evaluation of the current vehicle voice interaction system is to compare the obtained voice completeness rate C with the calculated voice activity detection end time If the speech integrity rate C is compared with the calculated speech activity detection end time The larger the ratio is, the better the voice interaction effect of the voice interaction system is. If the voice completeness rate C is equal to the calculated voice activity detection end time If the ratio of is lower than a preset threshold, it is necessary to adjust the voice activity detection end time in the in-vehicle voice interaction system according to the target information, wherein the preset threshold can be determined based on the specific situation.
[0094] refer to Figure 7 , Figure 7 Another flow chart of a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present application. Further, a parameter management method for an in-vehicle voice interaction system disclosed in an embodiment of the present invention further includes: iterating a driving scene recognition model and / or a user state recognition model according to a target message.
[0095] refer to Figure 7 When the vehicle starts to drive, the vehicle motion sensor collects the vehicle motion state data and sends the collected vehicle motion state data to the server or the in-vehicle voice interaction system. The server or the in-vehicle voice interaction system further optimizes the system to identify the vehicle driving scene based on the vehicle motion data, and dynamically adjusts the VAD time based on the user's determination of the voice activity detection end time formula (1) in the in-vehicle voice interaction system, i.e., the VAD time adjustment strategy and the driving scene and the user's attention coefficient or user attention state score corresponding to the user's attention state. After dynamically adjusting the VAD time, the voice completeness rate C of the voice interaction system is obtained, and the effect of the voice interaction system is evaluated based on the voice completeness rate and the voice activity detection end time. If the effect of the voice interaction system is lower than the preset effect, the driving scene recognition model used for driving scene recognition is iteratively optimized, i.e., the iterative driving scene recognition model is iterated to further optimize the system.
[0096] When adopting the parameter management method of the above-mentioned in-vehicle voice interaction system, the end time of the voice activity detection in the in-vehicle voice interaction system can be further adjusted based on the comparison of the ratio of the voice completeness rate corresponding to the in-vehicle voice interaction system to the end time of the voice activity detection with a preset threshold, so that the end time of the voice activity detection in the in-vehicle voice interaction system can ensure that the time when the user inputs voice is within the end time of the voice activity detection under the current driving state of the vehicle and the current user state of the user in the vehicle.
[0097] In a second aspect, an embodiment of the present invention further discloses an electronic device, which includes: a processor, a memory communicatively connected to the processor, the memory storing computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the electronic device implements a parameter management method for an in-vehicle voice interaction system such as any one of the above-mentioned items.
[0098] In a second aspect, an embodiment of the present invention further discloses an electronic device, based on which the parameter management method of the above-mentioned in-vehicle voice interaction system can be implemented to adjust the end time of the voice activity detection in the in-vehicle voice interaction system.
[0099] When the above electronic device is used, the processor can adjust the end time of the voice activity detection in the vehicle voice interaction system based on the driving state information of the vehicle and / or the user state information of the user in the vehicle to avoid the voice interaction process from ending prematurely, thereby ensuring that the user can complete the voice input completely, improving the accuracy of the voice interaction, and improving the user experience. In addition, the timeliness and accuracy of the vehicle voice control can also be improved, and the vehicle driving safety can be improved.
[0100] It should be noted that, in addition to the implementation methods of the present invention described in the above-mentioned specific embodiments, those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Although the description of the present invention is introduced in conjunction with the preferred embodiment, this does not mean that the features of this invention are limited to this implementation method. On the contrary, the purpose of introducing the invention in conjunction with the implementation method is to cover other options or modifications that may be extended based on the claims of the present invention. In order to provide a deep understanding of the present invention, the above description contains many specific details, and the present invention can also be implemented without using these details. In addition, in order to avoid confusion or blurring the focus of the present invention, some specific details will be omitted in the description. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0101] In the description of this embodiment, it should be noted that the terms "upper", "lower", "inner", "bottom", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or are the orientations or positional relationships in which the inventive product is usually placed when used. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0102] The terms “first”, “second”, etc. are only used for distinguishing descriptions and should not be understood as indicating or implying relative importance.
[0103] In the description of this embodiment, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this embodiment can be understood according to specific circumstances.
[0104] Although the present invention has been illustrated and described with reference to certain preferred embodiments of the present invention, it should be understood by those skilled in the art that the above is a further detailed description of the present invention in conjunction with specific embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. Those skilled in the art may make various changes in form and details, including making several simple deductions or substitutions, without departing from the spirit and scope of the present invention.
Claims
1. A parameter management method for a vehicle-mounted voice interaction system, characterized in that: The parameter management method of the vehicle-mounted voice interaction system includes: Acquiring target information, wherein the target information includes driving state information of the vehicle and / or user state information of a user in the vehicle; According to the target information, the voice activity detection end time of the in-vehicle voice interaction system is adjusted.
2. The parameter management method according to claim 1, characterized in that: In a case where the target information includes the driving state information, adjusting the voice activity detection end time in the in-vehicle voice interaction system according to the target information includes: Inputting the driving state information into a driving scene recognition model to obtain a driving scene of the vehicle; According to the driving scenario, the voice activity detection end time of the in-vehicle voice interaction system is determined.
3. The parameter management method according to claim 1, characterized in that: In a case where the target information includes the user status information, adjusting the voice activity detection end time in the in-vehicle voice interaction system according to the target information includes: Inputting the user status information into a user status recognition model to obtain the user status; Determine an end time of the voice activity detection of the in-vehicle voice interaction system according to the user status.
4. The parameter management method according to claim 1, characterized in that: In a case where the target information includes the driving state information and the user state information, adjusting the voice activity detection end time of the in-vehicle voice interaction system according to the target information includes: Inputting the driving state information into a driving scene recognition model to obtain a driving scene of the vehicle; Inputting the user status information into a user status recognition model to obtain the user status; Determine an end time of voice activity detection of the in-vehicle voice interaction system according to the driving scenario and the user status.
5. The parameter management method according to claim 4, characterized in that: Determining the end time of the voice activity detection of the in-vehicle voice interaction system according to the driving scenario and the user state includes: Determining an initial voice activity detection end time according to the driving scenario; The initial voice activity detection end time is adjusted according to the user status to obtain the voice activity detection end time.
6. The parameter management method according to claim 4, characterized in that: The user state includes a user attention state, and determining the end time of the voice activity detection in the in-vehicle voice interaction system according to the driving scenario and the user state includes determining by the following formula: in, is the end time of the voice activity detection, T base is the preset end time of basic voice activity detection, is the sensitivity coefficient corresponding to the driving scenario, Z is F attention or A score , F attention is the user attention coefficient corresponding to the user attention state, A score is the user attention score corresponding to the user attention state, and i is the number of the driving scenes.
7. The parameter management method according to claim 6, characterized in that: The method also includes obtaining the user attention score A by the following formula: score : Wherein, W is a preset weight, b is a preset bias parameter, and X is a feature vector calculated based on the user state.
8. The parameter management method according to any one of claims 1 to 7, characterized in that: The driving state information includes vehicle motion state information collected by a vehicle motion sensor, wherein the vehicle motion sensor includes an accelerometer, a gyroscope and a positioning speed sensor; The user status information includes user facial status information collected by a vehicle camera, and the user facial status information includes facial expression information and / or eye movement information.
9. The parameter management method according to any one of claims 1 to 7, characterized in that: The method further comprises: Obtaining a speech completeness rate corresponding to the in-vehicle speech interaction system; When the ratio of the speech completeness rate to the speech activity detection end time is lower than a preset threshold, the speech activity detection end time in the in-vehicle voice interaction system is adjusted according to the target information.
10. An electronic device, characterized in that: The electronic device comprises: a processor, a memory communicatively coupled to the processor, the memory storing computer executable instructions; The processor executes the computer-executable instructions stored in the memory so that the electronic device implements the parameter management method of the in-vehicle voice interaction system as described in any one of claims 1-9.