Voice playing method and device, vehicle terminal and storage medium
By generating text content based on trigger commands and timbre in the vehicle terminal and converting it into speech, the problem of rigid voice interaction functions in the existing technology is solved, and a more personalized interactive experience with more adapted and harmonious voice content and timbre is achieved.
Patent Information
- Application Number
- CN202410660356.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2044-05-27
AI Technical Summary
The voice interaction functions in existing technologies produce monotonous, inflexible, and unpersonalized voices.
Text content is generated based on trigger commands and the voice of the in-vehicle voice assistant. Text-to-speech synthesis technology is used to convert text into speech and play it in the in-vehicle terminal. It supports multiple voice selections and personalized settings, including voice display, selection, and voice feature matching for text content generation.
It improves the flexibility and adaptability of voice interaction functions, making the played voice content and tone more harmonious and enhancing the user experience.
Smart Images

Figure CN118711566B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of human-computer interaction, and in particular to a voice playing method and device, a vehicle terminal, and a storage medium. BACKGROUND
[0002] With the development of human-computer interaction technology, voice interaction function is born.
[0003] In the related art, the voice interaction function only directly converts preset text content into voice to play, which makes the voice played in the voice interaction process appear to be rather stiff and not flexible enough. SUMMARY
[0004] Embodiments of the present application provide a voice playing method and device, a vehicle terminal, and a storage medium, which can improve the flexibility of voice interaction function. The technical solutions provided by embodiments of the present application are as follows:
[0005] According to an aspect of embodiments of the present application, a voice playing method is provided, which includes:
[0006] In the case of receiving a trigger instruction for triggering a vehicle voice assistant in the vehicle terminal, generating text content based on the trigger instruction and the timbre of the vehicle voice assistant;
[0007] Generating voice corresponding to the text content based on the timbre of the vehicle voice assistant and the text content;
[0008] Playing the voice.
[0009] According to an aspect of embodiments of the present application, a voice playing device is provided, which includes:
[0010] A text generation module configured to, in the case of receiving a trigger instruction for triggering a vehicle voice assistant in the vehicle terminal, generate text content based on the trigger instruction and the timbre of the vehicle voice assistant;
[0011] A voice generation module configured to generate voice corresponding to the text content based on the timbre of the vehicle voice assistant and the text content;
[0012] A voice playing module configured to play the voice.
[0013] According to an aspect of embodiments of the present application, a computer device is provided, which includes a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the voice playing method described above.
[0014] According to an aspect of the embodiments of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the voice playing method.
[0015] According to an aspect of the embodiments of the present application, a computer program product is provided, and the computer program product is loaded and executed by a processor to implement the voice playing method.
[0016] The technical solutions provided by the embodiments of the present application can have the following beneficial effects.
[0017] By generating the text content based on the trigger instruction and the specific tone, the generated text content takes into account the tone characteristics of the vehicle-mounted voice assistant, so that the played voice content and the tone are more adaptive and more harmonious, thereby improving the flexibility of the voice interaction function.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 is a flowchart of the voice playing method provided by an embodiment of the present application;
[0021] Figure 2 is a block diagram of a vehicle-mounted terminal provided by another embodiment of the present application;
[0022] Figure 3 is a block diagram of a voice playing device provided by an embodiment of the present application;
[0023] Figure 4 is a block diagram of a vehicle-mounted terminal provided by an embodiment of the present application. DETAILED DESCRIPTION
[0024] The exemplary embodiments will be described in detail herein with reference to the drawings. Unless otherwise indicated, the same numbers on the different drawings indicate the same or similar elements. The embodiments described in the following exemplary embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of methods consistent with some aspects of the present application as detailed in the appended claims.
[0025] The execution subject of each step of the method provided in the embodiments of the present application can be a vehicle-mounted terminal, which refers to an electronic device integrated in a vehicle and having data calculation, processing and storage capabilities. In the following, the technical solutions of the present application are introduced and described through several embodiments.
[0026] Please refer to Figure 1 which shows a flowchart of a voice playing method provided in an embodiment of the present application. In the present embodiment, the method is mainly exemplified by being applied to the vehicle-mounted terminal introduced above. The method can include at least one of the following steps 110-130.
[0027] Step 110, in the case of receiving a trigger instruction for triggering a vehicle-mounted voice assistant in the vehicle-mounted terminal, generating text content based on the trigger instruction and the timbre of the vehicle-mounted voice assistant.
[0028] In some embodiments, a vehicle-mounted terminal is installed in a vehicle, which is a front-end device of a vehicle monitoring and management system and can also be referred to as a vehicle transmission control unit (TCU). In some embodiments, the vehicle-mounted terminal can be integrated with positioning, communication, navigation, vehicle driving information recording and other functions. In some embodiments, the vehicle-mounted terminal can also support telephone calling, voice interaction and other functions. In some embodiments, the vehicle-mounted voice assistant can be a computer program installed in the vehicle-mounted terminal for realizing voice interaction function. In some embodiments, the vehicle-mounted voice assistant can issue prompt information to people inside or outside the vehicle through voice interaction. In some embodiments, the vehicle-mounted voice assistant can also have a conversation with the user of the vehicle through voice interaction.
[0029] In some embodiments, the trigger instruction can be an instruction generated or activated by the vehicle-mounted terminal itself, such as a seat belt prompt instruction, an overspeed prompt instruction, a door closing prompt instruction, an oil consumption prompt instruction, etc. For example, during the starting or driving of the vehicle, if the vehicle-mounted terminal does not detect that the user has fastened the seat belt, it will automatically generate a seat belt prompt instruction. Based on the seat belt prompt instruction, text content for prompting the user to fasten the seat belt as soon as possible can be generated, such as “seat belt not fastened” and “please fasten the seat belt”. For another example, during the driving of the vehicle, if the vehicle-mounted terminal detects that the vehicle speed exceeds the specified speed, the vehicle-mounted terminal can generate an overspeed prompt instruction. Based on the overspeed prompt instruction, text content for prompting the user that the vehicle has overspeed can be generated, such as “vehicle has overspeed” and “please reduce the vehicle speed”. For another example, during the starting or driving of the vehicle, if a certain door is not closed, the vehicle-mounted terminal can generate a door closing prompt instruction. Based on the door closing prompt instruction, text content for prompting the driver and / or passenger in the vehicle to close the door can be generated, such as “door not closed” and “please close the door”.
[0030] In some embodiments, the trigger instruction can also be based on the behavior of the driver or passenger in the vehicle to trigger the generated instruction. For example, after the driver of the vehicle starts the vehicle, an instruction can be generated to prompt that the vehicle has been started, based on which corresponding text content such as "the vehicle has been started", "the vehicle has been successfully started" and the like can be generated. In some embodiments, the trigger instruction can also be a voice instruction. For example, if the driver or passenger in the vehicle issues a voice instruction "what is the weather today", the vehicle terminal can generate corresponding text content according to the weather information of the day after obtaining the weather information of the day.
[0031] In some embodiments, the text content generated in the present application based on the trigger instruction is adapted to the timbre of the vehicle voice assistant. For example, if the timbre of the vehicle voice assistant is the timbre of a young woman, the text content generated based on the seat belt prompt instruction to prompt the user to fasten the seat belt as soon as possible can be "fasten the seat belt as soon as possible, be careful~"; if the timbre of the vehicle voice assistant is the timbre of an older woman, the text content generated based on the seat belt prompt instruction to prompt the user to fasten the seat belt as soon as possible can be "fasten the seat belt as soon as possible, child"; if the timbre of the vehicle voice assistant is the timbre of a young man, the text content generated based on the seat belt prompt instruction to prompt the user to fasten the seat belt as soon as possible can be "fasten the seat belt as soon as possible, brother".
[0032] Step 120, generating the voice corresponding to the text content based on the timbre of the vehicle voice assistant and the text content.
[0033] In some embodiments, the text content is converted into corresponding voice by using the timbre of the vehicle voice assistant, such as converting the text content into a corresponding voice file, so as to obtain the voice to be played. This step of converting text into voice involves text-to-speech (TTS, Text To Speech) technology.
[0034] Step 130, playing the voice.
[0035] In some embodiments, the voice generated in step 120 above can be played through the loudspeaker of the vehicle terminal or the loudspeaker connected to the vehicle terminal.
[0036] In some embodiments, the application supporting the vehicle voice assistant can also support the user to turn off the corresponding vehicle voice assistant function at any time. That is, the user can arbitrarily turn on and off the voice assistant function of each application program or each function in the vehicle, so as to make the vehicle voice assistant more personalized and improve the setting flexibility of the vehicle voice assistant.
[0037] In summary, the technical scheme provided by the embodiments of the present application generates text content based on a trigger instruction and a specific timbre, so that the generated text content takes into account the timbre characteristics of the vehicle-mounted voice assistant, so that the played voice content and the timbre are more suitable and more harmonious, thereby improving the flexibility of voice interaction function.
[0038] In some possible implementations, the vehicle-mounted voice assistant corresponds to multiple candidate timbres; and the method further includes at least one of the following steps:
[0039] 1. displaying multiple candidate timbres;
[0040] 2. in response to a selection operation on a first timbre of the multiple candidate timbres, determining the first timbre as the timbre of the vehicle-mounted voice assistant.
[0041] In some embodiments, the vehicle-mounted voice assistant can have multiple candidate timbres, and the vehicle-mounted terminal can display the multiple candidate timbres to the user, and determine one of the timbres, such as the first timbre, as the timbre of the vehicle-mounted voice assistant. Thereafter, the first timbre is the default timbre of the vehicle-mounted voice assistant before the user reselects the timbre, thereby saving the time cost of the user selecting the timbre each time and improving the convenience of user operation.
[0042] In some embodiments, displaying multiple candidate timbres can mean that the display screen of the vehicle-mounted terminal displays multiple candidate timbres each with an option (such as the name of each timbre, the avatar of the corresponding character / role, etc.), and the user can determine the first timbre as the timbre of the vehicle-mounted voice assistant through a selection operation (such as a click operation) on the option corresponding to the first timbre.
[0043] In some embodiments, displaying multiple candidate timbres can also mean playing multiple candidate timbres each with an example voice corresponding thereto, so as to facilitate the user to make a more suitable selection. In some embodiments, the multiple candidate timbres can include a main accompanying timbre, such as the timbre of the user's relatives or friends.
[0044] In some embodiments, the multiple candidate timbres include at least one of the following: the timbre of an IP character; and the timbre of a user-provided voice.
[0045] In some embodiments, the candidate timbre can be the timbre of a character in game, movie, TV series, or other media content. In some embodiments, the timbre of an IP character can be pre-stored in the vehicle-mounted terminal, or can be downloaded by the user and stored in the vehicle-mounted terminal. In some embodiments, the candidate timbre can be the timbre of a user-provided voice.
[0046] In some embodiments, the process of obtaining the candidate timbre can include the following steps:
[0047] 1. Tone collection
[0048] In some embodiments, a plurality of different sound samples are recorded by a developer or a user of the in-vehicle voice assistant. These sound samples can be the voices of IP characters in respective game, movie, TV series, etc. media content, or the voices of the user of the in-vehicle voice assistant or the user's family and friends, or the voices of virtual human voices generated by voice synthesis technology. In some embodiments, these sound samples can be the voices of different groups of people such as men, women, children, and the elderly.
[0049] 2. Processing and storage of tone
[0050] The collected sound samples need to be processed and stored. The processing of the sound samples can include noise removal, tone adjustment, addition of audio effects, etc. After that, the sound samples are stored in the audio library of the in-vehicle terminal or the background server of the in-vehicle terminal, ready for user selection and playback.
[0051] In some embodiments, the text content can be synthesized into voice to be played by a synthesis engine, which can use mixing, pitch adjustment, etc. to ensure that the output voice is consistent with the tone of the in-vehicle voice assistant selected by the user. In some embodiments, once the user selects a certain tone (such as the first tone), the in-vehicle terminal can set and store it as the user's personal preference setting, and use it as the default tone of the in-vehicle voice assistant in the future.
[0052] In some embodiments, the in-vehicle terminal can include the following modules: a signal processing module 10, an ASR (Automatic Speech Recognition) module 20, an NLP (Natural Language Processing) module 30, and a TTS module 40. Among them, the signal processing module 10 can include the following units: a microphone array, and a VAD (Voice Activity Detection); the ASR module 20 can include the following units: an acoustic model, a language model; the NLP module 30 can include the following units: NLU (Natural Language Understanding), DM (Data Mining), NLG (Natural Language Generation); the TTS module 40 can include the following units: a linguistic front end, a vocoder.
[0053] In the implementation manners described above, the user can select a tone that meets his / her own intention from multiple candidate tones, thereby enriching the tone selection of the in-vehicle voice assistant and making the tone of the in-vehicle voice assistant more personalized.
[0054] In some possible implementation manners, the text content is generated based on the trigger instruction and the tone of the in-vehicle voice assistant, and includes at least one of the following steps:
[0055] 1. Generating standard text content corresponding to the trigger instruction based on the trigger instruction.
[0056] 2. Adjusting the standard text content based on the tone of the in-vehicle voice assistant to generate the text content.
[0057] In some embodiments, the standard text content generated based on the trigger instruction is the same regardless of the tone of the in-vehicle voice assistant. Then, the standard text content can be adjusted based on the tone characteristics of the in-vehicle voice assistant, so that the finally generated text content is more in line with the impression of the tone of the in-vehicle voice assistant to the user. For example, for a seat belt prompt instruction, the standard text content can be "Please fasten your seat belt", if the tone of the in-vehicle voice assistant is the tone of a certain general role, the text content "Listen to orders! Fasten your seat belt!" can be generated based on the standard text content; if the tone of the in-vehicle voice assistant is the tone of a certain doctor role, the text content "The seat belt must be fastened, don't become my patient, please" can be generated based on the standard text content; if the tone of the in-vehicle voice assistant is the tone of the user's partner, the text content "My love, fasten your seat belt, wait for you to come back safely" can be generated based on the standard text content; if the tone of the in-vehicle voice assistant is the tone of the user's child, the text content "Dad / Mom, don't forget the seat belt, wait for you to come home safely" can be generated based on the standard text content.
[0058] In some embodiments, adjusting the standard text content based on the tone of the in-vehicle voice assistant to generate the text content can be implemented as at least one of the following steps:
[0059] 1. Obtaining a common phrase corresponding to the tone of the in-vehicle voice assistant; and adding the common phrase into the standard text content to generate the text content.
[0060] In some embodiments, the commonly used words can also be referred to as a catchphrase. According to the statistical analysis of the voice samples of the character or person corresponding to the voice tone of the in-vehicle voice assistant, the words or sentences with higher frequency of occurrence are determined as the commonly used words of the in-vehicle voice assistant corresponding to the voice tone, and the commonly used words are added to the standard text content to generate the text content corresponding to the voice tone of the in-vehicle voice assistant. For example, if the statistical analysis of the voice samples of the character or person corresponding to the voice tone of the in-vehicle voice assistant shows that the commonly used words are "good", the text content obtained after adding the corresponding commonly used words to the standard text content "vehicle start" can be "good! Vehicle start".
[0061] 2. Obtain the address of the user by the character or person corresponding to the voice tone of the in-vehicle voice assistant; add the address to the standard text content to generate the text content.
[0062] In some embodiments, the user can input or select the address of the user by the character or person corresponding to the voice tone of the in-vehicle voice assistant in advance, such as "dad", "mom", "son", "daughter", "Xiaoming", "Lili", "Dad", "Mom", "girlfriend", "boyfriend", "best friend", "dog", "baby" and the like, and store the address in the in-vehicle terminal. When the in-vehicle voice assistant plays the voice, the address is added to the standard text content to generate the text content corresponding to the voice.
[0063] For example, if the voice tone of the in-vehicle voice assistant is the voice tone of the user's child, the user can determine the address as "dad" or "mom", and add "dad" or "mom" to the text content when generating the text content.
[0064] 3. Obtain the self-address of the IP character or person corresponding to the voice tone of the in-vehicle voice assistant; add the self-address to the standard text content to generate the text content.
[0065] In some embodiments, the IP character or person corresponding to the voice tone of the in-vehicle voice assistant can have a commonly used self-address, such as "Xiaoqi", "Xiao Q", "Jiangjiang", "Zai Xia", "Ben Shuai", "Ben Girl", "Old Monk", "Old Nun", "Poor Monk" and the like. The self-address corresponding to each candidate voice tone can be selected and determined by a related technical person or the user himself / herself, and the self-address is added when generating the text content.
[0066] In some embodiments, the self-address is directly added in the case where there is no first person in the standard text content. For example, if the standard text content is "the air conditioner in the vehicle has been adjusted by 5 degrees Celsius", then after adding the self-address, the text content can be "Xiao Q has adjusted the air conditioner in the vehicle by 5 degrees Celsius".
[0067] In some embodiments, in the case that there is a first person pronoun (such as "I") in the standard text content, the first person pronoun can be replaced with a corresponding self-reference. For example, if the standard text content is "I have turned down the air conditioner in the car by 5 degrees Celsius", after replacing "I" with the self-reference "Little Q", the text content can be "Little Q has turned down the air conditioner in the car by 5 degrees Celsius".
[0068] Of course, the content adjustment manner for the standard text content is not limited to the above examples, and other adjustment manners can also exist, and the embodiments of the present application do not make specific limitations thereto.
[0069] In this embodiment, by adding common phrases, addresses to the user, self-references and other content, the final obtained text content is better matched and more harmonious with the timbre of the vehicle-mounted voice assistant, and the final played voice is more vivid and lifelike.
[0070] In the above implementation manner, based on the trigger instruction, the standard text content generated uniformly, and the text content corresponding to the final played voice is obtained based on the standard text content, so that the text content conforms to the timbre of the vehicle-mounted voice assistant to the user's impression, while also trying to ensure that the semantic of the text content is consistent with the trigger instruction.
[0071] In some possible implementation manners, in the case that there is no passenger in the co-pilot seat of the vehicle where the vehicle-mounted terminal is located, the first loudspeaker is used to play the voice; wherein the first loudspeaker is the loudspeaker corresponding to the co-pilot seat.
[0072] In some embodiments, in the case that there is no passenger in the co-pilot seat of the vehicle where the vehicle-mounted terminal is located, or in the case that there is only the driver and no other person in the vehicle, the first loudspeaker corresponding to the co-pilot seat is used to play the voice. Thus, the driver hears the sound of the vehicle-mounted voice assistant coming from the co-pilot position, thereby simulating the passenger in the co-pilot position talking to the driver, and improving the feeling and experience of the driver in voice interaction with the vehicle-mounted voice assistant.
[0073] In some embodiments, in the case that there is a passenger in the co-pilot seat, other loudspeakers other than the co-pilot seat are used to play the voice, such as the loudspeaker corresponding to the driver seat, the loudspeaker between the driver seat and the co-pilot seat, the roof loudspeaker, etc.
[0074] In some embodiments, whether there is a person in each seat in the vehicle can be identified by a pressure detection identification manner. For example, if the pressure of the co-pilot seat is less than or equal to a first threshold value, it indicates that there is no passenger in the co-pilot seat. The specific value of the first threshold value can be determined by relevant technical personnel, and the embodiments of the present application do not make specific limitations thereto.
[0075] In the implementation manners, the passenger in the co-pilot position is made to speak to the driver in simulation of the hearing of the driver, so that the driver has a real feeling of talking to a real person, and the experience of the driver in voice interaction with the voice assistant of the vehicle is improved.
[0076] The following is a device embodiment of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0077] Please refer to Figure 3 which shows a block diagram of a voice playing device according to an embodiment of the present application. The device has the functions of the voice playing method examples described above, which can be implemented by hardware or corresponding software executed by hardware. The device can be the vehicle terminal introduced above, or can be arranged on the vehicle terminal. The device 300 can include a text generation module 310, a voice generation module 320, and a voice playing module 330.
[0078] The text generation module 310 is configured to, in a case where a trigger instruction for triggering a voice assistant of the vehicle terminal is received, generate text content based on the trigger instruction and a timbre of the voice assistant of the vehicle.
[0079] The voice generation module 320 is configured to generate a voice corresponding to the text content based on the timbre of the voice assistant of the vehicle and the text content.
[0080] The voice playing module 330 is configured to play the voice.
[0081] In some embodiments, the device 300 further includes a timbre display module and a timbre determination module.
[0082] The timbre display module is configured to display the plurality of candidate timbres.
[0083] The timbre determination module is configured to, in response to a selection operation on a first timbre of the plurality of candidate timbres, determine the first timbre as the timbre of the voice assistant of the vehicle.
[0084] In some embodiments, the plurality of candidate timbres includes a timbre of a vocal sound provided by a user.
[0085] In some embodiments, the text generation module 310 includes a text generation sub-module.
[0086] The text generation sub-module is configured to generate standard text content corresponding to the trigger instruction based on the trigger instruction.
[0087] The text generation submodule is further configured to adjust the content of the standard text based on the tone of the vehicle-mounted voice assistant, and generate the text content.
[0088] In some embodiments, the text generation submodule is configured to:
[0089] obtain common expressions corresponding to the tone of the vehicle-mounted voice assistant, and add the common expressions into the standard text content to generate the text content.
[0090] Alternatively,
[0091] obtain a user address of a character or a person corresponding to the tone of the vehicle-mounted voice assistant, and add the user address into the standard text content to generate the text content.
[0092] Alternatively,
[0093] obtain a self-address of an IP character or a person corresponding to the tone of the vehicle-mounted voice assistant, and add the self-address into the standard text content to generate the text content.
[0094] In some embodiments, the voice playing module 330 is configured to:
[0095] In a case where there is no passenger in a front passenger seat of a vehicle where the vehicle-mounted terminal is located, the voice is played by using a first loudspeaker.
[0096] The first loudspeaker is a loudspeaker corresponding to the front passenger seat.
[0097] In summary, the technical solution provided by the embodiments of the present application generates text content based on a trigger instruction and a specific tone, so that the generated text content takes into account the tone characteristics of the vehicle-mounted voice assistant, so that the played voice content and tone are more suitable and more harmonious, thereby improving the flexibility of voice interaction function.
[0098] It should be noted that the device provided in the above embodiments is only used as an example to illustrate the division of the above functional modules in realizing its functions. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.
[0099] For reference Figure 4 which shows a structure block diagram of a vehicle-mounted terminal 400 provided by an embodiment of the present application. The vehicle-mounted terminal is used to implement the voice playing method provided in the above embodiments. Specifically:
[0100] Generally, the vehicle terminal 400 comprises a processor 401 and a memory 402.
[0101] The processor 401 can comprise one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 can be implemented in the form of at least one of a DSP (Digital Signal Processing), an FPGA (Field Programmable Gate Array), a PLA (Programmable Logic Array). The processor 401 can also comprise a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 401 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by a display screen. In some embodiments, the processor 401 can further comprise an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0102] The memory 402 can comprise one or more computer-readable storage media, which can be non-transitory. The memory 402 can further comprise a high-speed random access memory, and a non-volatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is configured to store a computer program, and is configured to be executed by one or more processors to implement the voice playing method described above.
[0103] In some embodiments, the vehicle terminal 400 can further optionally comprise a peripheral device interface 403 and at least one peripheral device. The processor 401, the memory 402 and the peripheral device interface 403 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 403 through a bus, a signal line or a circuit board. Specifically, the peripheral device comprises at least one of a radio frequency circuit 404, a display screen 405, an audio circuit 406 and a power supply 407.
[0104] Those skilled in the art can understand that the structure shown in the above description is not a limitation on the vehicle terminal 400, and the vehicle terminal 400 can comprise more or fewer components than those shown in the drawings, or combine certain components, or adopt a different arrangement of components. Figure 4 The structure shown in the above description is not a limitation on the vehicle terminal 400, and the vehicle terminal 400 can comprise more or fewer components than those shown in the drawings, or combine certain components, or adopt a different arrangement of components.
[0105] In an example embodiment, a computer readable storage medium is also provided, the storage medium storing a computer program, the computer program being executed by a processor to implement the voice playing method.
[0106] In an example embodiment, a computer program product is also provided, the computer program product being loaded and executed by a processor to implement the voice playing method.
[0107] It should be understood that "multiple" referred to herein means two or more. "And / or" describes the association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after are in an "or" relationship.
[0108] The above only describes example embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A voice play method characterized by, The method is executed by a vehicle terminal, and the method comprises: In a case where a trigger instruction for triggering a vehicle voice assistant in the vehicle terminal is received, generating, based on the trigger instruction, standard text content corresponding to the trigger instruction; Based on the timbre of the vehicle voice assistant, adjusting the content of the standard text content to generate text content; Based on the timbre of the vehicle voice assistant and the text content, generating a voice corresponding to the text content; Playing the voice; The method further comprises: Displaying the plurality of candidate timbres; In response to a selection operation on a first timbre of the plurality of candidate timbres, determining the first timbre as the timbre of the vehicle voice assistant. The plurality of candidate timbres comprises a timbre of a user-provided human voice. The playing of the voice comprises: In a case where there is no passenger in a front passenger seat of a vehicle in which the vehicle terminal is located, playing the voice by using a first loudspeaker; 2. The method of claim 1, wherein, The first loudspeaker is a loudspeaker corresponding to the front passenger seat. The device comprises: A text generation module configured to, in a case where a trigger instruction for triggering a vehicle voice assistant in a vehicle terminal is received, generate, based on the trigger instruction, standard text content corresponding to the trigger instruction; and adjust the content of the standard text content based on the timbre of the vehicle voice assistant to generate text content; The text generation module is further configured to: acquire common expressions corresponding to the timbre of the vehicle voice assistant; add the common expressions into the standard text content to generate the text content; or acquire a user's name called by a character corresponding to the timbre of the vehicle voice assistant; add the name into the standard text content to generate the text content; or acquire a self-reference of the character corresponding to the timbre of the vehicle voice assistant; and add the self-reference into the standard text content to generate the text content; 3. The method of claim 2, wherein, A voice generation module configured to generate, based on the timbre of the vehicle voice assistant and the text content, a voice corresponding to the text content; 4. The method of claim 1, wherein, A voice playing module configured to play the voice. The vehicle terminal comprises a processor and a memory, and the memory stores a computer program, which is loaded and executed by the processor to implement the voice playing method according to any one of claims 1 to 4. 5. A speech playback apparatus characterized by comprising: 6. A vehicle terminal, characterized by comprising: 7. A computer readable storage medium characterized by The computer program is stored in the computer readable storage medium and loaded and executed by the processor to realize the voice playing method in any one of claims 1 to 4.
8. A computer program product, characterised in that, The computer program product is loaded and executed by the processor to realize the voice playing method in any one of claims 1 to 4.
Citation Information
Patent Citations
Speech playing method and apparatus for article, and device, storage medium and program product
WO2022184055A1