Vehicle-mounted voice broadcast method and related device
By generating and distributing an SDK package containing the voice ID of the in-vehicle voice AI assistant, the problem of inconsistent voice timbres in in-vehicle voice systems has been solved, achieving voice uniformity and enhancing the passenger's auditory experience and the premium feel of the cabin.
Patent Information
- Application Number
- CN202310099403.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-01
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-02-01
AI Technical Summary
The inconsistent timbre of different applications in existing in-vehicle voice systems results in a poor auditory experience for passengers, lacks uniformity, and diminishes the sense of luxury in the cabin.
By generating an SDK package containing the voice ID of the in-vehicle voice AI assistant and distributing it to the target in-vehicle application, the voice ID is unified, enabling the target in-vehicle application to use the current voice ID of the in-vehicle voice AI assistant when making broadcasts, thus achieving voice uniformity.
It solves the problem of inconsistent voice quality in vehicles, improves the passenger's auditory experience, and enhances the sense of luxury in the cabin.
Smart Images

Figure CN116229934B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automotive in-vehicle systems, and more particularly to an in-vehicle voice broadcasting method and related equipment. Background Technology
[0002] With people's aspirations for a better life and the development of automotive intelligent voice technology, people have more demands and higher requirements for the announcements in in-vehicle intelligent cockpits. In the existing in-vehicle infotainment systems, the in-vehicle voice AI assistant has one announcement tone, the navigation has another, and parking and driving-related reminders have yet another. The coexistence of different tones affects the passenger's experience, and the lack of uniformity among different voice announcement software reduces the sense of luxury in the cockpit. Summary of the Invention
[0003] In view of the above problems, embodiments of the present invention provide a vehicle-mounted voice broadcasting method and related equipment with a more user-friendly and consistent tone across multiple applications.
[0004] To address the aforementioned problems, a first aspect of this invention provides a method for in-vehicle voice broadcasting, the method comprising:
[0005] An SDK package is generated based on the TTS function of the in-vehicle voice AI assistant. The SDK package includes the voice ID currently used by the in-vehicle voice AI assistant.
[0006] The SDK package is distributed to at least one target in-vehicle application of the vehicle to which the in-vehicle voice AI assistant belongs, wherein the target in-vehicle application includes a voice broadcast function;
[0007] When the aforementioned target in-vehicle application performs voice broadcasts, the voice ID currently used by the aforementioned in-vehicle voice AI assistant will be adopted.
[0008] In one feasible implementation, the above-described vehicle-mounted voice broadcasting method includes:
[0009] In response to the user's command to switch the voice tone of the in-vehicle voice AI assistant, the voice tone of the in-vehicle voice AI assistant is switched.
[0010] Send a tone ID modification instruction to at least one target in-vehicle application of the vehicle to which the aforementioned in-vehicle voice AI assistant belongs. The tone ID modification instruction includes switching tone IDs so that the currently used tone ID in the SDK package of the target in-vehicle application that receives the tone ID modification instruction is replaced with the switched tone ID.
[0011] In one feasible implementation, the SDK package further includes: at least one audio stream for broadcasting, wherein the audio stream for broadcasting includes the audio stream associated with the timbre ID currently used by the in-vehicle voice AI assistant;
[0012] In the case of voice broadcasting in the aforementioned target in-vehicle application, the voice ID currently used by the aforementioned in-vehicle voice AI assistant is adopted, including:
[0013] When the aforementioned target in-vehicle application performs voice broadcast, the audio stream associated with the timbre ID currently used by the in-vehicle voice AI assistant is retrieved from the SDK package of the aforementioned target in-vehicle application based on the application broadcast method of the aforementioned target in-vehicle application itself, and the broadcast is performed using the audio stream.
[0014] In one feasible implementation, the method for vehicle-mounted voice broadcasting described above further includes:
[0015] When at least two of the aforementioned target in-vehicle applications need to perform voice broadcasts simultaneously, obtain the attribute levels of at least two of the aforementioned target in-vehicle applications respectively.
[0016] Retrieve the audio stream used for broadcasting that is associated with the timbre ID currently used by the aforementioned in-vehicle voice AI assistant from the SDK package of the aforementioned target in-vehicle application;
[0017] Voice broadcasts are only performed based on the application broadcasting method of the target in-vehicle application with the highest attribute level.
[0018] In one feasible implementation, the method for vehicle-mounted voice broadcasting described above further includes:
[0019] When at least two of the aforementioned target in-vehicle applications need to perform voice broadcasts simultaneously, obtain the attribute levels of at least two of the aforementioned target in-vehicle applications respectively.
[0020] Retrieve the audio stream used for broadcasting that is associated with the timbre ID currently used by the aforementioned in-vehicle voice AI assistant from the SDK package of the aforementioned target in-vehicle application;
[0021] Voice broadcasts are performed simultaneously based on the application broadcasting methods of at least two of the aforementioned target in-vehicle applications, wherein the voice broadcasting volume of the target in-vehicle application with the higher attribute level is greater than the voice broadcasting volume of the target in-vehicle application with the lower attribute level.
[0022] The volume difference between the voice broadcast volume of the target in-vehicle application with the higher attribute level and the voice broadcast volume of the target in-vehicle application with the lower attribute level is greater than one-fifth of the vehicle's volume range, and the voice broadcast volume of the target in-vehicle application with the lower attribute level meets the minimum value of the user's hearing range.
[0023] In one feasible implementation, when at least two of the aforementioned target in-vehicle applications need to perform voice broadcasts simultaneously, the in-vehicle target user associated with each of the aforementioned target in-vehicle applications is obtained respectively.
[0024] Based on the seating distribution of the target users in the vehicle, the vehicle speaker unit near each target user is selected to broadcast the voice announcement of the target vehicle application associated with that target user.
[0025] A second aspect of the present invention provides an in-vehicle voice broadcasting device, the device comprising:
[0026] The generation module is used to generate an SDK package based on the TTS function of the in-vehicle voice AI assistant. The SDK package includes the voice ID currently used by the in-vehicle voice AI assistant.
[0027] The distribution module is used to distribute the SDK package to at least one target in-vehicle application of the vehicle to which the in-vehicle voice AI assistant belongs, wherein the target in-vehicle application includes a voice broadcast function.
[0028] The broadcast module is used to use the voice ID currently used by the in-vehicle voice AI assistant when the target in-vehicle application broadcasts voice messages.
[0029] A third aspect of the present invention provides an electronic device, which includes at least one processor and at least one memory connected to the processor, wherein the processor is configured to call program instructions in the memory to execute the vehicle voice broadcasting method described in any of the first aspects.
[0030] A fourth aspect of the present invention provides a storage medium including a stored program, wherein, when the program is executed, the device on which the storage medium is located executes the in-vehicle voice broadcasting method described in any of the first aspects.
[0031] Compared with the prior art, the present invention has at least the following beneficial effects: This application provides a method for in-vehicle voice broadcasting, the method comprising: generating an SDK package based on the TTS function of an in-vehicle voice AI assistant, wherein the SDK package includes the voice ID currently used by the in-vehicle voice AI assistant; distributing the SDK package to at least one target in-vehicle application of the vehicle to which the in-vehicle voice AI assistant belongs, wherein the target in-vehicle application includes a voice broadcasting function; and using the voice ID currently used by the in-vehicle voice AI assistant when the target in-vehicle application performs voice broadcasting. The in-vehicle voice broadcasting method provided in this solution generates an SDK package based on the TTS function of the in-vehicle voice AI assistant. This SDK package contains the currently used voice ID of the in-vehicle voice AI assistant. The SDK package is distributed to at least one target in-vehicle application with voice broadcasting function in the vehicle to which the in-vehicle voice AI assistant belongs. When the target in-vehicle application broadcasts, it will use the current voice ID of the in-vehicle voice AI assistant, thereby unifying the voice ID of the target in-vehicle application with the current voice ID of the in-vehicle voice AI assistant. This solution reduces the interference of different in-vehicle voice timbres, and the voice timbres of the aforementioned in-vehicle voice AI assistant are more in line with the passengers' hearing habits, improving the passengers' auditory experience and enhancing the sense of luxury in the cabin. Attached Figure Description
[0032] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0033] Figure 1 A schematic flowchart illustrating an in-vehicle voice broadcasting method provided in this application embodiment;
[0034] Figure 2 A schematic diagram illustrating the arrangement of vehicle voice attribute levels provided in this application embodiment;
[0035] Figure 3 This is a schematic diagram of an audio conflict strategy type provided in an embodiment of this application;
[0036] Figure 4 A schematic structural block diagram of the hardware and software architecture of an in-vehicle voice broadcasting device provided in this application embodiment;
[0037] Figure 5 A schematic structural block diagram of an in-vehicle voice broadcasting device provided in an embodiment of this application;
[0038] Figure 6 This is a schematic structural block diagram of an electronic device provided in an embodiment of this application.
[0039] Figure 4 40 is the in-vehicle central computing unit, 401 is the operating system, 402 is the in-vehicle voice AI assistant, 403 is the navigation in-vehicle application, 404 is the multimedia in-vehicle application, 405 is the application broadcast in-vehicle application, 406 is the front left speaker, 407 is the front right speaker, 408 is the front camera, 409 is the front voice interaction device, 410 is the rear voice interaction device, 411 is the rear camera, 412 is the rear left speaker, and 413 is the rear right speaker. Specific Implementation
[0040] This invention provides a vehicle-mounted voice broadcasting method and related equipment, which solves the problems of inconsistent voice timbre and chaotic broadcasting logic in the prior art.
[0041] To better understand the technical solution of the present invention, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention. Furthermore, it should be noted that, for ease of description, only the parts related to the present invention are shown in the drawings, not the entire structure.
[0042] The first aspect of this application provides a vehicle-mounted voice broadcasting method, such as... Figure 1 As shown, the method may include:
[0043] Step S110: Generate an SDK package based on the TTS function of the in-vehicle voice AI assistant, wherein the SDK package includes the voice ID currently used by the in-vehicle voice AI assistant.
[0044] The aforementioned in-vehicle voice broadcasting method can be implemented through an internet connection between a cloud server and the in-vehicle terminal, or through a controller within the intelligent in-vehicle terminal. This controller connects to an operating system to issue commands to various application software. The operating system is a set of interconnected system software programs that control computer operation, utilize and run hardware and software resources, and provide public services to organize user interaction. In this application, it can be described as an operating system based on vehicle applications. The aforementioned intelligent in-vehicle terminal can be a front-end device for a vehicle monitoring and management system. , For modern management of transport vehicles, this application provides a controller that implements the above functions, including GPS networking, positioning, voice interaction, and a black box.
[0045] Specifically, the aforementioned in-vehicle voice AI assistant software package utilizes existing TTS (Text To Speech) capabilities, which can be understood as a function of converting text into speech and is part of human-computer dialogue. It generates a JAR file (based on pre-written Java classes packaged together) using this TTS capability. This JAR file is integrated into the aforementioned in-vehicle voice AI assistant software's SDK (Software Development Kit). The SDK package can use the JAR file to implement the in-vehicle voice AI assistant's TTS function. The SDK package also includes an API (Application Programming Interface), which enables interoperability between different programs. The JAR file includes the in-vehicle voice AI assistant's TTS class, TTS class attributes, and TTS playback methods. The TTS class attributes include the user object, timbre ID, speed, background sound, playback scene, authorization code, and the audio stream used for playback, etc., and the timbre ID is the timbre ID currently used by the in-vehicle voice AI assistant.
[0046] Through the various attributes and TTS broadcasting methods in the aforementioned in-vehicle voice AI assistant SDK package, the in-vehicle voice AI assistant includes, but is not limited to, functions such as adjusting timbre, adjusting volume, adjusting speech rate, adjusting background noise, and adjusting broadcasting scenarios.
[0047] Step S120: Distribute the SDK package to at least one target in-vehicle application of the vehicle to which the in-vehicle voice AI assistant belongs, wherein the target in-vehicle application includes a voice broadcast function.
[0048] It is understandable that the aforementioned in-vehicle applications with voice broadcasting capabilities are not limited to application software such as parking, navigation, and multimedia, but also include TTS-type voice broadcasting information such as emergency assistance communications, instrument driving information reminders, hardware broadcasting, and vehicle system prompts.
[0049] Specifically, the in-vehicle voice AI assistant can distribute the aforementioned SDK package to at least one target in-vehicle application on the vehicle to which the in-vehicle voice AI assistant belongs. This target in-vehicle application includes voice playback functionality. The JAR file within the SDK package contains the voice ID currently used by the in-vehicle voice AI assistant. Distribution can be achieved through access between the API of the target in-vehicle application and the API within the SDK package. By introducing the SDK package into the at least one target in-vehicle application, the target in-vehicle application can directly call the TTS class attributes and TTS playback methods in the aforementioned JAR file within the SDK package. By distributing the in-vehicle voice AI assistant SDK package to at least one target in-vehicle application on the vehicle to which the in-vehicle voice AI assistant belongs, the TTS class attribute parameters of the target in-vehicle application and the in-vehicle voice AI assistant can be unified, allowing users of the vehicle to control the target in-vehicle application by controlling certain parameters of the in-vehicle voice AI assistant.
[0050] Step S130: When the target in-vehicle application makes a voice broadcast, the voice ID currently used by the in-vehicle voice AI assistant is adopted.
[0051] This can be explained as follows: When the target in-vehicle application needs to perform a voice broadcast task, it first uses TTS capabilities to generate an SDK package containing a JAR file, also known as a JDK package. To prevent exceptions during the call, the JDK package first initializes this JAR file. This JDK package contains the TTS class attributes and TTS broadcast methods from the in-vehicle voice AI assistant JAR file. The TTS class attributes include a voice ID. When the target in-vehicle application calls the TTS broadcast method in the JDK package, the TTS broadcast method uses the currently used voice ID attribute from the in-vehicle voice AI assistant SDK package, thus completing the voice broadcast task using the currently used voice ID of the in-vehicle voice AI assistant. By calling the TTS class attribute within the in-vehicle voice AI assistant SDK package from the target in-vehicle application, this solution addresses the issue of inconsistent in-vehicle voice timbres in existing technologies. This improves the user's auditory experience and enhances the premium feel of the smart cockpit. The target in-vehicle application can then directly use the same timbre ID when performing voice broadcasts.
[0052] For example, when the aforementioned in-vehicle voice AI assistant's JAR file contains TTS-type attribute parameters that can provide background sound and broadcast scene services, such as selecting "soothing" for the background sound and "cinema" for the broadcast scene, the target in-vehicle application can display soothing background sound effects and create a "cinema" atmosphere when performing voice broadcast tasks. Through this example, the functionality of the target in-vehicle application during voice broadcast can be extended, satisfying users' pursuit of auditory aesthetics.
[0053] In this solution, the controller controls the aforementioned in-vehicle voice AI assistant to generate a JAR file based on its TTS (Text-to-Speech) function. This JAR file contains TTS class attribute parameters, including a voice ID. This JAR file is included within the in-vehicle voice AI assistant's SDK package. The SDK package can be distributed to the target in-vehicle application via API communication. When performing voice broadcast tasks, the target in-vehicle application can call the TTS attribute parameters of the in-vehicle voice AI assistant, which include the current voice ID of the in-vehicle voice AI assistant. This solution unifies the voice IDs of the target in-vehicle applications with the current voice ID of the in-vehicle voice AI assistant, ensuring consistency across different target in-vehicle applications. This enhances the auditory experience for users in vehicles using the in-vehicle voice AI assistant, improving the premium feel of the smart cockpit.
[0054] In one possible embodiment, the aforementioned in-vehicle voice broadcasting method may further include: switching the voice of the in-vehicle voice AI assistant in response to a user's voice timbre switching command.
[0055] Send a tone ID modification instruction to at least one target in-vehicle application of the vehicle to which the aforementioned in-vehicle voice AI assistant belongs. The tone ID modification instruction includes switching tone IDs so that the currently used tone ID in the SDK package of the target in-vehicle application that receives the tone ID modification instruction is replaced with the switched tone ID.
[0056] Specifically, users of the aforementioned vehicles can issue voice switching commands through the operating interface of the in-vehicle voice AI assistant. These commands allow users to change the voice of the in-vehicle voice assistant. The operating interface refers to the human-computer interaction interface of the in-vehicle voice AI assistant, a two-way information exchange platform between humans and computers. The voice of the in-vehicle voice assistant refers to the voice type currently used by the assistant during voice broadcasting, based on the current voice ID. The controller can switch the voice IDs within the SDK package of the in-vehicle voice AI assistant; the switched voice ID is generated based on the voice switching command. This solution provides users with multiple choices of voice IDs, improving the user experience of the in-vehicle voice AI assistant and enhancing the premium feel of the smart cockpit.
[0057] The aforementioned in-vehicle voice AI assistant sends a voice ID modification instruction to at least one target in-vehicle application of the vehicle to which it belongs. The voice ID modification instruction is generated based on the aforementioned voice switching instruction. The voice ID modification instruction includes switching the voice ID, so that the current voice ID in the SDK package of the aforementioned target in-vehicle application is switched to the aforementioned switched voice ID. The SDK package comes from the aforementioned in-vehicle voice AI assistant.
[0058] This can be explained as follows: After a user of the vehicle to which the aforementioned in-vehicle voice AI assistant belongs issues a voice tone switching command through the assistant's interface, the controller switches the current voice tone ID of the in-vehicle voice AI assistant. The switched voice tone ID is based on the user's voice tone switching command. Based on this voice tone modification command, the controller adjusts the voice tone ID parameter in the TTS class attribute parameters within the target in-vehicle application SDK package. The adjusted voice tone ID parameter is the switched voice tone ID parameter.
[0059] In this solution, the user issues a voice ID switching command to the aforementioned in-vehicle voice AI assistant. The controller then distributes the switched voice ID of the in-vehicle voice AI assistant to at least one target in-vehicle application within the vehicle to which the in-vehicle voice AI assistant belongs. The voice ID of the target in-vehicle application is then replaced with the switched voice ID of the in-vehicle voice AI assistant. Users of the vehicle to which the in-vehicle voice AI assistant belongs can switch the voice ID of at least one of the target in-vehicle applications within the vehicle by switching the voice ID of the in-vehicle voice AI assistant. This solution enriches the user's voice selection, improves the convenience of controlling different target in-vehicle applications within the vehicle, strengthens the consistency between different target in-vehicle applications, and enhances the premium feel of the smart cockpit.
[0060] In one possible embodiment, the SDK package further includes: at least one audio stream used for broadcasting, the audio stream used for broadcasting including the audio stream associated with the timbre ID currently used by the in-vehicle voice AI assistant;
[0061] In the case of voice broadcasting in the aforementioned target in-vehicle application, the voice ID currently used by the aforementioned in-vehicle voice AI assistant is adopted, including:
[0062] When the aforementioned target in-vehicle application performs voice broadcast, the audio stream associated with the timbre ID currently used by the in-vehicle voice AI assistant is retrieved from the SDK package of the aforementioned target in-vehicle application based on the application broadcast method of the aforementioned target in-vehicle application itself, and the broadcast is performed using the audio stream.
[0063] Specifically, the aforementioned audio stream can be considered as a data stream of audio signals. The controller executes the voice broadcasting task through the audio channel opened by the operating system, and the audio channel carries the aforementioned audio stream. The aforementioned audio signal data stream contains audio information, such as sampling rate, harmonics, number of channels, bit rate, etc. This data can be used to determine the playback device of the vehicle to which the aforementioned in-vehicle voice AI assistant belongs, such as the vehicle speaker. Different timbres correspond to different sound waveforms, and the sound waveform is mainly determined based on harmonics. To further explain, in this embodiment, different timbres correspond to different timbre IDs, and different timbre IDs correspond to different data streams, i.e., the aforementioned audio stream. Different timbre IDs within the aforementioned in-vehicle voice AI assistant SDK package correspond to different audio streams. The aforementioned audio stream used for broadcasting refers to the audio stream used when currently executing the voice broadcasting task.
[0064] The aforementioned target in-vehicle applications may include different audio stream types, such as system prompts, error alarms, and voice-announced task sounds. Different broadcast strategies can be determined based on different audio stream types, and these broadcast strategies can generate different application broadcast methods. Therefore, different application broadcast methods include different audio streams used for broadcasting.
[0065] When the aforementioned target in-vehicle application needs to perform a voice broadcast task, the type of audio stream used for broadcasting is determined based on the application's required broadcasting method. Then, an audio stream associated with the current voice ID of the in-vehicle voice AI assistant is selected from the eligible audio stream types for broadcasting. This solution allows the use of the in-vehicle voice AI assistant's voice ID without altering the target in-vehicle application's own broadcasting method, thus meeting the needs of the users of the vehicle to obtain the target in-vehicle application's voice broadcasting information. Furthermore, adjusting the target in-vehicle application's broadcast voice ID to the in-vehicle voice AI assistant's voice ID better suits the users' auditory habits. The unified voice ID across different TTS-type voice broadcasting software improves the users' auditory experience and enhances the premium feel of the smart cockpit.
[0066] In one possible embodiment, the aforementioned in-vehicle voice broadcasting method further includes:
[0067] When at least two of the aforementioned target in-vehicle applications need to perform voice broadcasts simultaneously, obtain the attribute levels of at least two of the aforementioned target in-vehicle applications respectively.
[0068] Retrieve the audio stream used for broadcasting that is associated with the timbre ID currently used by the aforementioned in-vehicle voice AI assistant from the SDK package of the aforementioned target in-vehicle application;
[0069] Voice broadcasts are only performed based on the application broadcasting method of the target in-vehicle application with the highest attribute level.
[0070] Specifically, when at least two target in-vehicle applications within the vehicle to which the aforementioned in-vehicle voice AI assistant belongs need to perform voice broadcasts simultaneously, the need for simultaneous voice broadcasts means that their own broadcast conditions are met simultaneously. These broadcast conditions are based on the broadcasting method of the aforementioned target in-vehicle applications, and the attribute level of the aforementioned target in-vehicle applications is based on the broadcasting strategy and audio stream type in the audio stream of the aforementioned target in-vehicle applications.
[0071] To implement this solution, this embodiment provides an audio policy program that can determine the attribute levels of different target in-vehicle applications. This audio policy program is based on the aforementioned operating system and can be written in programming languages such as C or C++. Further, the audio policy program determines the audio stream of the target in-vehicle application with the highest attribute level based on the attribute levels of at least two target in-vehicle applications that need to be broadcast simultaneously. The audio policy program only allows the audio channel carrying the audio stream of the target in-vehicle application with the highest attribute level, and the controller only broadcasts the voice announcements of the target in-vehicle application with the highest attribute level.
[0072] For example, the above audio strategy determines the attribute level of the target in-vehicle application based on its different types. Specifically, this can be achieved through a preset configuration file, which includes a specific arrangement of the target in-vehicle application attribute levels. An embodiment of this application provides a schematic diagram of the arrangement of in-vehicle voice attribute levels, as shown below. Figure 2 As shown, Level 1 is the highest attribute level, and Level 10 is the lowest attribute level. The configuration file also includes audio conflict strategy types. These audio conflict strategy types are proposed based on the different needs of target in-vehicle applications under different circumstances at different attribute levels. An embodiment of this application provides a schematic diagram of an audio conflict strategy type, as shown below. Figure 3 As shown, C1 represents the currently playing target in-vehicle application, and C2 represents the inserted target in-vehicle application. It should be noted that the ranking of target in-vehicle application attributes and audio conflict strategy types listed above are not unique; users can adapt them according to their specific settings and personal habits. The examples provided in this application are for illustrative purposes only.
[0073] Meanwhile, the aforementioned configuration file also includes a specific scenario conflict design hierarchy table based on the arrangement of the target in-vehicle application attribute levels and the audio playback conflict strategy, as shown in Table 1. This scenario conflict design hierarchy table is proposed based on the different broadcast levels between the target in-vehicle applications with different attribute levels. The broadcast level refers to the broadcast level of different target in-vehicle applications when the broadcast requirements are met simultaneously. In this table, column C1 represents the currently playing target in-vehicle application, column C2 represents the inserted target in-vehicle application, and L1 to L10 correspond to... Figure 2 The target in-vehicle application types range from Level 1 to Level 10. It should be noted that the specific strategies in this scenario conflict design table only provide one approach and are not unique.
[0074]
[0075] Table 1
[0076] When there are at least two target in-vehicle applications requiring voice broadcasts, the aforementioned audio policy program determines the corresponding attribute level based on the type of the at least two target in-vehicle applications, and controls the aforementioned operating system audio channel to allow the audio stream corresponding to the target in-vehicle application with the higher attribute level to be broadcast. The aforementioned controller only broadcasts the voice broadcast of the target in-vehicle application with the highest attribute level.
[0077] Here's a possible example: when the user of the vehicle with the aforementioned in-vehicle voice AI assistant is using an app that broadcasts news, such as a "news broadcast," the vehicle might issue a low tire pressure warning, i.e., a dashboard driving information alert. Figure 2As shown, application broadcast software belongs to Level 9, and instrument panel driving information reminders belong to Level 2. According to Table 1, the conflict strategy type is C. Figure 3 As shown, in the above conflict strategy type C, that is, C2 interrupts the broadcast without restoring C1, and C1 stops broadcasting, C2 is equivalent to the low tire pressure alarm message, and C1 is equivalent to the news broadcast software. Through the above audio strategy program, the program completely interrupts the news broadcast software broadcast according to the above preset configuration file, and only plays the low tire pressure alarm message.
[0078] Through the above embodiments, this solution provides a processing method when at least two target in-vehicle applications need to be broadcast simultaneously. By using a preset configuration file to determine the attribute level of the at least two target in-vehicle applications, voice broadcast is only performed on the target in-vehicle applications with higher attribute levels. This improves the problem of logical confusion among multiple target in-vehicle applications, enhances the ability of users of the vehicles, especially drivers, to distinguish the current voice information, improves the user's auditory experience, and also helps to improve driving safety.
[0079] In one possible embodiment, the aforementioned in-vehicle voice broadcasting method further includes:
[0080] If at least two of the aforementioned target in-vehicle applications need to perform voice broadcasts simultaneously, obtain the attribute levels of at least two of the aforementioned target in-vehicle applications respectively.
[0081] Retrieve the audio stream used for broadcasting that is associated with the timbre ID currently used by the aforementioned in-vehicle voice AI assistant from the SDK package of the aforementioned target in-vehicle application;
[0082] Voice broadcasts are performed simultaneously based on the application broadcasting methods of at least two of the aforementioned target in-vehicle applications, wherein the voice broadcasting volume of the target in-vehicle application with the higher attribute level is greater than the voice broadcasting volume of the target in-vehicle application with the lower attribute level.
[0083] Specifically, the aforementioned audio strategy program can determine the attribute level of the target in-vehicle application based on the priority strategy and scene conflict design table in the aforementioned preset configuration file. Based on the attribute level of the target in-vehicle application, it determines that the target in-vehicle applications can be played simultaneously. This playback method can be called mixed playback. To further explain, the aforementioned audio strategy program controls the voice playback volume of the target in-vehicle application based on the attribute level of the target in-vehicle application. The voice playback volume parameter of the target in-vehicle application with a higher attribute level is higher than that of the target in-vehicle application with a lower attribute level.
[0084] For example, when the user of the vehicle to which the aforementioned in-vehicle voice AI assistant belongs is using navigation software, such as "Gaode Navigation," the navigation software... Figure 2The indicated level is Level 8. At this point, the user is using voice interaction software, such as making a voice call. According to Table 1, the conflict strategy type is D, which means that C1 and C2 coexist, and C1 is weakened. At this time, C1 is a consonant and C2 is the dominant sound, which means reducing the volume of the voice broadcast of the navigation software.
[0085] This solution determines the attribute levels of different target in-vehicle applications that need to be played simultaneously based on the aforementioned audio strategy procedure. It then determines the voice playback volume parameters of these applications based on their attribute levels, with higher-level applications receiving louder voice playback than lower-level applications. This solution allows users of the same vehicle to simultaneously receive voice information from different target in-vehicle applications. Furthermore, by controlling the voice playback volume, it differentiates the importance of the voice information from different applications, improving the logical inconsistencies that arise when different in-vehicle voices are played simultaneously and enhancing the competitiveness of smart cockpit products.
[0086] In one possible embodiment, the volume difference between the voice broadcast volume of the target in-vehicle application with the higher attribute level and the voice broadcast volume of the target in-vehicle application with the lower attribute level is greater than one-fifth of the vehicle's volume range, and the voice broadcast volume of the target in-vehicle application with the lower attribute level meets the minimum value of the user's hearing range.
[0087] Specifically, the aforementioned vehicle volume range can be either the volume range of the in-vehicle voice AI assistant or the volume range of the in-vehicle voice terminal. For ease of explanation, this embodiment uses the volume range of the in-vehicle voice AI assistant for illustration.
[0088] For example, the volume adjustment range should be synchronized with the adjustable volume range of the in-vehicle voice AI assistant. If the volume adjustment range is (0-100), and the volume is adjusted to 1, which represents the lowest value within the user's hearing range, the voice playback volume parameter of the target in-vehicle application with the higher attribute level before mixing is set to V primary, and the voice playback volume parameter of the target in-vehicle application with the lower attribute level before mixing is set to V secondary. During mixing, the voice playback volume parameter of the target in-vehicle application with the lower attribute level is set to V secondary mix. The aforementioned mixing refers to the situation where the different target in-vehicle applications need to play simultaneously.
[0089] When both the target in-vehicle application with the higher attribute level and the target in-vehicle application with the lower attribute level have playback needs and the audio strategy program determines that mixed audio playback is possible, the audio strategy program will assign the current voice playback volume parameter of the in-vehicle voice AI assistant software, such as 60, to the target in-vehicle application with the higher attribute level, i.e., V primary is 60, and at the same time adjust the volume of V secondary mixed audio playback to less than 40, such as 35.
[0090] Another example is that if the current voice broadcast volume parameter of the aforementioned in-vehicle voice AI assistant software is 20, then this parameter will be assigned to the target in-vehicle application with the highest current attribute level, i.e., V_main = 20, while V_auxiliary will be adjusted to 1, i.e., V_auxiliary_mix = 1.
[0091] Another exemplary embodiment of this application may further include, when the target in-vehicle application with a lower attribute level enters and exits the mixed audio playback process, the audio strategy program needs to perform a fade-in / fade-out effect of 100 milliseconds on the target in-vehicle application audio source (i.e., consonant audio source) with the lower attribute level. For example, 100 milliseconds before the consonant audio source enters the playback, the audio playback volume parameter of the audio source will gradually increase to the desired audio playback volume parameter; 100 milliseconds before the consonant audio source needs to exit the playback, the audio playback volume parameter of the audio source will gradually decrease from the current audio playback volume parameter to zero. The above processing method improves the sophistication of the target in-vehicle application's mixed audio playback.
[0092] This solution creates a sense of hierarchy in the voice broadcasting of different target in-vehicle applications by setting the voice broadcasting volume of different attribute levels during mixed audio broadcasting. It reasonably classifies the importance and urgency of the broadcasting content of different target in-vehicle applications under different situations. By controlling the voice broadcasting volume parameters of different target in-vehicle applications, it improves the user's, especially the driver's, ability to recognize different in-vehicle voice information at the same time, improves the user experience, and also makes an important contribution to improving driving safety and reliability.
[0093] In one possible embodiment, the aforementioned in-vehicle voice broadcasting method further includes:
[0094] When at least two of the aforementioned target in-vehicle applications need to perform voice broadcasts simultaneously, obtain the in-vehicle target users associated with each of the aforementioned target in-vehicle applications;
[0095] Based on the seating distribution of the target users in the vehicle, the vehicle speaker unit near each target user is selected to broadcast the voice announcement of the target vehicle application associated with that target user.
[0096] It can be explained that the vehicle to which the aforementioned in-vehicle voice AI assistant belongs needs to collect the playback needs of users in different positions within the vehicle. To achieve this, the vehicle may be equipped with an in-vehicle camera (with image acquisition function), intelligent in-vehicle terminals for the front and rear rows respectively, and independent in-vehicle speakers for the front and rear rows. The intelligent in-vehicle terminal should also be equipped with a voice interaction device. This application provides a schematic structural block diagram of the hardware and software architecture of an in-vehicle voice broadcasting device, as shown below. Figure 4 As shown.
[0097] Specifically, Figure 4 The in-vehicle central computing unit 40 is the core of computing and control for the terminal of the vehicle to which the aforementioned in-vehicle voice AI assistant belongs. It is the final execution unit for information processing and program execution, such as a processor. The in-vehicle central computing unit handles the interaction of information, functional linkage, and effect realization between the electrical, hardware, and software layers of the vehicle. At the software layer, the in-vehicle central computing unit 40, through an operating system 401 (such as an Android-based in-vehicle operating system), carries target in-vehicle applications such as navigation applications 403, multimedia applications 404, and application broadcasting applications 405. At the electrical layer, the in-vehicle central computing unit also manages some intelligent in-vehicle devices within the vehicle. These intelligent in-vehicle devices refer to devices, instruments, or machines with computing capabilities used in the vehicle, such as… Figure 4 As shown, the front left speaker 406, front right speaker 407, front camera 408, front voice interaction device 409, rear voice interaction device 410, rear camera 411, rear left speaker 412, and rear right speaker 413 collect information and process it to control the aforementioned devices to achieve the functions of the target in-vehicle application. At the hardware level, the aforementioned in-vehicle central computing unit is connected to memory, device interfaces, and other devices via an I / O bus. It should be noted that... Figure 4 The listed target vehicle application types and intelligent vehicle devices are only examples and do not represent all of them. The connection relationships between the parts are for illustrative purposes only.
[0098] For example, when there are multiple passengers in both the front and rear seats of the aforementioned vehicle, and two different target in-vehicle applications need to be broadcast simultaneously, such as "navigation" software 403 and "multimedia application" software 404, both target in-vehicle applications are broadcast through the aforementioned in-vehicle voice AI assistant 402. The aforementioned front-seat camera 408 and rear-seat camera 411 in the aforementioned vehicle will respectively collect information on the upper body of the passengers in the vehicle. The aforementioned in-vehicle central computing unit will make a preliminary judgment on the collected information. If it is recognized that there is a user sitting in the driver's seat, it will control the front left speaker 406 to play the audio source of the "navigation" software 403 through the in-vehicle central computing unit, and other speakers will not play this audio source. The aforementioned rear-seat camera 411 in the aforementioned vehicle recognizes that there is a passenger in the left rear position, and will control the rear left speaker 412 to play the audio source of the "multimedia application" software through the in-vehicle central computing unit, and other speakers will not play this audio source. In the above embodiment, if the passenger in the front seat of the vehicle asks "What will the weather be like tomorrow?" through the front voice interaction device 409, the vehicle central computing unit 40 will judge the question through the vehicle's built-in voice collection device and recognition function, and then control the "application broadcasting" software 405 in the vehicle operating system to broadcast the voice, and control the vehicle speaker on the passenger side to broadcast the voice, while other speakers will not play this sound source.
[0099] Another example is that the aforementioned in-vehicle central computing unit can perform predictive control in advance through historical information records. These historical information records include boarding time, passenger's usual seating position, frequently opened applications and corresponding time periods, and image information collection. For ease of explanation, this application provides an example: if a passenger in the left rear seat of the aforementioned vehicle frequently boards between 5:20 PM and 5:30 PM and frequently uses the "Anime Channel" application between 5:30 PM and 6:10 PM, the in-vehicle central computing unit will sense the passenger boarding based on the aforementioned rear-seat in-vehicle camera and record the boarding time. At the same time, it will record the passenger's body shape based on the image acquisition function. When the passenger operates the in-vehicle operating system and launches the selected target in-vehicle application, the aforementioned in-vehicle central computing unit will record the time and the running application. After multiple operations that match the above behavioral characteristics, the aforementioned in-vehicle central computing unit will perform a profile operation on the passenger. When a passenger matching the above body shape boards the vehicle, the in-vehicle central computing unit will control the in-vehicle operating system to automatically open the "Anime Channel" application for broadcast.
[0100] The above embodiments, by utilizing the information interaction processing between the intelligent in-vehicle equipment in the vehicle and the target in-vehicle application and voice AI assistant, meet the different needs of different customers in different situations, profile users and predict their behavior in advance, enhance the interaction between passengers and the intelligent cockpit, and improve the user's auditory experience by controlling the in-vehicle speakers in different locations to reduce the influence of different target in-vehicle application sound sources in the vehicle, thereby enhancing the premium feel of the intelligent cockpit.
[0101] A second aspect of this application provides an in-vehicle voice broadcasting device for performing the steps of the in-vehicle voice method as described in any of the first aspects above, such as... Figure 5 As shown, the in-vehicle voice broadcasting device 50 includes a generation module 501, a distribution module 502, and a broadcasting module 503, wherein:
[0102] The generation module 501 is used to generate an SDK package based on the TTS function of the in-vehicle voice AI assistant, wherein the SDK package includes the voice ID currently used by the in-vehicle voice AI assistant;
[0103] The distribution module 502 is used to distribute the SDK package to at least one target vehicle application of the vehicle to which the vehicle voice AI assistant belongs, wherein the target vehicle application includes a voice broadcast function.
[0104] The broadcast module 503 is used to use the voice ID currently used by the in-vehicle voice AI assistant when the target in-vehicle application performs voice broadcast.
[0105] A third aspect of this application provides an electronic device 60, such as... Figure 6 As shown, the electronic device includes at least one processor 601 and at least one memory 602 connected to the processor, wherein the processor 601 is used to call program instructions in the memory 602 to perform the steps of the in-vehicle voice method as proposed in any of the first aspects above.
[0106] The fourth aspect of this application provides a storage medium having a program stored thereon, which, when executed, controls the device on which the storage medium is located to perform the steps of the in-vehicle voice broadcasting method as described in any of the first aspects above.
[0107] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0108] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0109] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0110] This application also provides a computer program product, which includes computer software instructions that, when executed on a processing device, cause the processing device to perform actions such as... Figure 1 The process of the in-vehicle voice broadcasting method in the corresponding embodiment.
[0111] The aforementioned computer program product includes one or more computer instructions. When the aforementioned computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The aforementioned computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The aforementioned computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0113] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.
[0114] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0115] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0116] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0117] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A vehicle-mounted voice broadcasting method, characterized in that, include: An SDK package is generated based on the TTS function of the in-vehicle voice AI assistant, wherein the SDK package includes the voice ID currently used by the in-vehicle voice AI assistant; The SDK package is distributed to at least one target in-vehicle application of the vehicle to which the in-vehicle voice AI assistant belongs, wherein the target in-vehicle application includes a voice broadcast function; When the target in-vehicle application performs voice broadcast, the voice ID currently used by the in-vehicle voice AI assistant is used, including: when the target in-vehicle application performs a voice broadcast task, the target in-vehicle application calls the TTS class attribute in the SDK package of the in-vehicle voice AI assistant, the TTS class attribute includes the voice ID currently used by the in-vehicle voice AI assistant, and the voice broadcast task is completed using the voice ID currently used by the in-vehicle voice AI assistant.
2. The vehicle-mounted voice broadcasting method according to claim 1, characterized in that, Also includes: In response to the user's command to switch the voice tone of the in-vehicle voice AI assistant, the voice tone of the in-vehicle voice AI assistant is switched. Send a tone ID modification instruction to at least one target in-vehicle application of the vehicle to which the in-vehicle voice AI assistant belongs. The tone ID modification instruction includes switching tone IDs so that the currently used tone ID in the SDK package of the target in-vehicle application that receives the tone ID modification instruction is replaced with the switched tone ID.
3. The vehicle-mounted voice broadcasting method according to claim 1, characterized in that, The SDK package also includes: at least one audio stream for broadcasting, the audio stream for broadcasting including the audio stream associated with the timbre ID currently used by the in-vehicle voice AI assistant; When the target in-vehicle application performs voice broadcast, the voice ID currently used by the in-vehicle voice AI assistant is adopted, including: When the target in-vehicle application performs voice broadcast, the audio stream associated with the timbre ID currently used by the in-vehicle voice AI assistant is retrieved from the SDK package of the target in-vehicle application based on the application broadcast method of the target in-vehicle application, and the broadcast is performed using the audio stream.
4. The vehicle-mounted voice broadcasting method according to claim 1, characterized in that, Also includes: When at least two of the target vehicle applications need to perform voice broadcasts simultaneously, the attribute levels of at least two of the target vehicle applications are obtained respectively. Retrieve the audio stream used for broadcasting that is associated with the timbre ID currently used by the in-vehicle voice AI assistant from the SDK package of the target in-vehicle application; Voice broadcasting is performed only based on the application broadcasting method of the target in-vehicle application with the highest attribute level.
5. The vehicle-mounted voice broadcasting method according to claim 1, characterized in that, Also includes: When at least two of the target vehicle applications need to perform voice broadcasts simultaneously, the attribute levels of at least two of the target vehicle applications are obtained respectively. Retrieve the audio stream used for broadcasting that is associated with the timbre ID currently used by the in-vehicle voice AI assistant from the SDK package of the target in-vehicle application; Voice broadcasts are performed simultaneously based on the application broadcasting methods of at least two of the target in-vehicle applications, wherein the voice broadcasting volume of the target in-vehicle application with the higher attribute level is greater than the voice broadcasting volume of the target in-vehicle application with the lower attribute level.
6. The vehicle-mounted voice broadcasting method according to claim 5, characterized in that, The volume difference between the voice broadcast volume of the target in-vehicle application with a higher attribute level and the voice broadcast volume of the target in-vehicle application with a lower attribute level is greater than one-fifth of the vehicle's volume range, and the voice broadcast volume of the target in-vehicle application with a lower attribute level meets the minimum value of the user's hearing range.
7. The vehicle-mounted voice broadcasting method according to claim 1, characterized in that, Also includes: When at least two of the target in-vehicle applications need to perform voice broadcasts simultaneously, obtain the in-vehicle target user associated with each of the target in-vehicle applications; Based on the seating distribution of the target users in the vehicle, the vehicle speaker unit near each target user is selected to broadcast the voice announcement of the target vehicle application associated with that target user.
8. A vehicle-mounted voice broadcasting device, characterized in that, include: A generation module is used to generate an SDK package based on the TTS function of the in-vehicle voice AI assistant, wherein the SDK package includes the voice ID currently used by the in-vehicle voice AI assistant; The distribution module is used to distribute the SDK package to at least one target in-vehicle application of the vehicle to which the in-vehicle voice AI assistant belongs, wherein the target in-vehicle application includes a voice broadcast function; The broadcast module is used to use the voice ID currently used by the in-vehicle voice AI assistant when the target in-vehicle application performs voice broadcast. This includes: when the target in-vehicle application performs a voice broadcast task, the target in-vehicle application calls the TTS class attribute in the SDK package of the in-vehicle voice AI assistant. The TTS class attribute includes the voice ID currently used by the in-vehicle voice AI assistant, and the voice broadcast task is completed using the voice ID currently used by the in-vehicle voice AI assistant.
9. An electronic device, characterized in that, The electronic device includes at least one processor and at least one memory connected to the processor, wherein the processor is used to call program instructions in the memory to execute the vehicle voice broadcasting method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the in-vehicle voice broadcasting method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Voice broadcasting method and device, electronic equipment and storage medium
CN113449141A
Voice playing method and device, equipment and storage medium
CN114121028A
System and method for providing personalized virtual personal assistant
CN115428067A
Video call method and device
WO2022135411A1