Voice personalization methods, devices, electronic devices and vehicles
By recognizing the wake-up audio signal in the sound region and detecting the personalized voice image and voice set by the user, the problem of existing technologies being unable to meet user customization needs is solved, realizing personalized voice interaction and broadcasting, and improving the user experience.
Patent Information
- Application Number
- CN202310523213.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-05-10
AI Technical Summary
Existing technologies cannot meet users' personalized voice avatar customization needs, resulting in a poor user experience.
By recognizing the wake-up audio signal in the sound region, the system detects whether the user has made personalized settings, displays the personalized voice image and voice set by the user, and realizes voice interaction and broadcasting.
To meet the personalized customization needs of all users and improve the user experience.
Smart Images

Figure CN116504244B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice interaction technology, and in particular to a voice personalization method, device, electronic device, and vehicle. Background Technology
[0002] With the development of in-vehicle voice recognition, voice interaction is gradually shifting from abstract to anthropomorphic, and the voice characters in in-vehicle scenarios are also being given emotional designs.
[0003] Currently, existing technologies for displaying voice avatars mainly fall into two categories. One method utilizes a hardware platform, but since hardware avatars are inherently monotonous, they cannot be personalized. The other method uses software, which, while incorporating voice source localization for features like zone control, still cannot meet the personalized customization needs of all users.
[0004] Therefore, it can be seen that the existing technology for displaying voice images cannot meet the personalized customization needs of all users, resulting in a poor user experience. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a voice personalization method, device, electronic device, and vehicle to meet the personalized customization needs of all users and improve the user experience.
[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0007] The first aspect of this invention discloses a voice personalization method, the method comprising:
[0008] When a wake-up audio signal is detected in any audio region, it is determined whether the user in that audio region has made personalized settings for that audio region.
[0009] If it is detected that the user has made personalized settings for the audio region, the first personalized voice image and the first personalized voice set by the user are determined, the first personalized voice image is displayed, the first personalized voice image is used to interact with the user via voice, and the first personalized voice image performs voice broadcast according to the first personalized voice.
[0010] Optional, also includes:
[0011] If no personalized settings are detected from the user for the audio region, the default voice image and corresponding default broadcast voice of the audio region are obtained, the default voice image is displayed, and the user is interacted with by voice using the default voice image and the corresponding default broadcast voice.
[0012] Optional, also includes:
[0013] Acquire multiple speech information from the same speech region, where each speech information corresponds to a speech emission time;
[0014] By comparing the duration of each speech emission time, the speech emission times are arranged in ascending order.
[0015] According to the order of the speech emission times, the speech information corresponding to each speech emission time is identified sequentially;
[0016] When one or more of the wake-up audio signals are detected in any of the aforementioned voice information, the voice region is controlled to execute the voice personalization method as described in the first aspect of the present invention.
[0017] Optional, also includes:
[0018] Acquire speech information from the first and second voice regions that exist simultaneously outside the aforementioned voice region;
[0019] Compare the speech emission time corresponding to the speech information in the first voice region with the speech emission time corresponding to the speech information in the second voice region;
[0020] If the speech time corresponding to the speech information in the first voice region is less than the speech time corresponding to the speech information in the second voice region, the speech information in the first voice region is identified.
[0021] When the wake-up audio signal is detected in the voice information of the first voice region, the first voice region is controlled to execute the voice personalization method as described in the first aspect of the present invention.
[0022] If the speech time corresponding to the speech information in the first voice region is greater than the speech time corresponding to the speech information in the second voice region, the speech information in the second voice region is identified.
[0023] When the wake-up audio signal is detected in the voice information of the second voice region, the second voice region is controlled to execute the voice personalization method as described in the first aspect of the present invention.
[0024] If the speech time corresponding to the speech information in the first voice region is equal to the speech time corresponding to the speech information in the second voice region, then the speech information in the first voice region and the speech information in the second voice region are recognized simultaneously.
[0025] When the wake-up audio signal is detected in both the voice information in the first voice region and the voice information in the second voice region, the voice personalization method described in the first aspect of the present invention is executed in parallel in the first voice region and the second voice region.
[0026] When the wake-up audio signal is detected in the voice information in the first voice region or the voice information in the second voice region, the voice personalization method as described in the first aspect of the present invention is controlled to be executed in the first voice region or the second voice region.
[0027] Optionally, determining the user-defined first personalized voice avatar and first personalized voice, displaying the first personalized voice avatar, using the first personalized voice avatar to interact with the user via voice, and having the first personalized voice avatar perform voice broadcasts according to the first personalized voice, includes:
[0028] The user-defined first personalized voice avatar is determined from a plurality of personalized voice avatars preset in the voice region.
[0029] Obtain the image data and voice interaction data corresponding to the first personalized voice avatar;
[0030] Based on the image data and the voice interaction data, the first personalized voice image is displayed;
[0031] The user-defined first personalized voice is determined from a plurality of personalized voices preset in the sound region.
[0032] Obtain the voice packet corresponding to the first personalized voice;
[0033] Based on the voice package, the user interacts with the first personalized voice avatar through the broadcast voice set in the voice package.
[0034] Optionally, determining the user-defined first personalized voice avatar and first personalized voice, displaying the first personalized voice avatar, using the first personalized voice avatar to interact with the user via voice, and having the first personalized voice avatar perform voice broadcasts according to the first personalized voice, includes:
[0035] Obtain the personalized voice profile parameters set by the user;
[0036] The personalized voice image parameters are matched with the personalized voice image library preset in the voice region to obtain a personalized voice image that meets the preset similarity.
[0037] Obtain the image data and voice interaction data corresponding to the personalized voice image that meets the preset similarity;
[0038] Based on the image data and the voice interaction data, the personalized voice image that meets the preset similarity is displayed;
[0039] Obtain the personalized voice parameters corresponding to the personalized voice image that meets the preset similarity set by the user.
[0040] The personalized voice parameters are matched with a pre-defined personalized voice library for the voice region to obtain a personalized voice that meets a preset similarity.
[0041] Using the personalized voice image that meets the preset similarity, the user can interact with the user via voice through the corresponding personalized voice that meets the preset similarity.
[0042] Optionally, determining the user-defined first personalized voice avatar and first personalized voice, displaying the first personalized voice avatar, using the first personalized voice avatar to interact with the user via voice, and having the first personalized voice avatar perform voice broadcasts according to the first personalized voice, includes:
[0043] Obtain the voice pack corresponding to the personalized voice image that meets the preset similarity;
[0044] Using the personalized voice avatar that meets the preset similarity, the user can interact with the voice through the broadcast voice set in the voice package.
[0045] A second aspect of this invention discloses a voice personalization device, the device comprising:
[0046] The detection module is used to detect whether the user in any sound zone has made personalized settings for the sound zone when a wake-up audio signal is detected in any sound zone.
[0047] The first display and broadcast module is used to determine the first personalized voice image and the first personalized voice set by the user if it is detected that the user has made personalized settings for the audio region, display the first personalized voice image, use the first personalized voice image to interact with the user by voice, and have the first personalized voice image broadcast the message according to the first personalized voice.
[0048] A third aspect of the present invention discloses an electronic device for running a program, wherein the program executes a voice personalization method as described in any of the first aspects of the present invention.
[0049] A fourth aspect of the present invention discloses a vehicle, the vehicle including the electronic equipment of the third aspect of the present invention.
[0050] Based on the above embodiments of the present invention, a voice personalization method, device, electronic device, and vehicle are provided. The method includes: when a wake-up audio signal is detected in any voice region, detecting whether a user in that voice region has personalized settings for that voice region; if the user has personalized settings for the voice region, determining the first personalized voice avatar and the first personalized voice set by the user, displaying the first personalized voice avatar, interacting with the user using the first personalized voice avatar, and having the first personalized voice avatar perform voice broadcasts according to the first personalized voice. In this solution, when a wake-up audio signal is detected in a certain voice region, the first personalized voice avatar is directly displayed according to the user's settings for displaying the first personalized voice avatar in that voice region. Voice interaction is performed with the user based on the first personalized voice avatar, and voice broadcasts are performed according to the user's settings for broadcasting the first personalized voice in that voice region, thereby satisfying the personalized customization needs of all users, realizing personalized voice avatars and voice broadcasts, and improving the user experience. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0052] Figure 1 A schematic diagram of a vehicle architecture provided for an embodiment of the present invention;
[0053] Figure 2 A schematic diagram of the architecture of a voice personalization system provided in an embodiment of the present invention;
[0054] Figure 3 A flowchart illustrating a voice personalization method provided in an embodiment of the present invention;
[0055] Figure 4 This is a schematic diagram of the structure of a voice personalization device provided in an embodiment of the present invention. Detailed Implementation
[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0058] As the background technology shows, there are two existing ways to display voice avatars. One is to display voice avatars through a hardware platform. This method generally focuses on interaction between the driver and front passenger, and cannot take into account the rear and third rows. For example, in a wake-up scenario, after the user says the wake-up word, the voice avatar will face the driver or front passenger and broadcast the wake-up feedback. However, the rear and third rows cannot fully interact with the avatar. Moreover, although the hardware avatar can be defined with different actions according to the scenario, such as the avatar holding a guitar in a music scenario, the hardware avatar itself is relatively monotonous and cannot achieve personalization. Another approach involves using software to display voice-activated avatars. When the vehicle is equipped with rear-seat screens (multi-zone + multi-screen), the avatar displayed across all zones and screens is the same. From an interactive experience perspective, users in each seating position have specific needs. For example, children in the second and third rows require avatars more age-appropriate, and the voice might be a child's voice or a replica of a parent's voice. Furthermore, while there are media for displaying voice-activated avatars for the driver, front passenger, second-row, and third-row users, and while this method incorporates voice source localization for zoned control, it still cannot meet the personalized customization needs of all users. Therefore, the existing methods for displaying voice-activated avatars cannot satisfy the personalized customization needs of all users, resulting in a poor user experience.
[0059] Therefore, embodiments of the present invention provide a voice personalization method, device, electronic device, and vehicle. In this solution, when a wake-up audio signal is detected in a certain voice zone, the first personalized voice image is directly displayed according to the user's setting for displaying the first personalized voice image in that voice zone. Voice interaction is conducted with the user based on the first personalized voice image, and voice broadcast is performed according to the user's setting for broadcasting the first personalized sound in that voice zone, so as to meet the personalized customization needs of all users, realize the personalization of voice image and voice broadcast, and improve the user experience.
[0060] First, let's introduce the overall vehicle architecture. In practical applications, based on the seating arrangement, the vehicle is divided into different audio zones: the driver's audio zone corresponding to the driver's seat, the passenger audio zone corresponding to the front passenger seat, the second-row left audio zone corresponding to the second-row left seat, the second-row right audio zone corresponding to the second-row right seat, the third-row left audio zone corresponding to the third-row left seat, and the third-row right audio zone corresponding to the third-row right seat. Infotainment terminals are installed in the corresponding positions of different audio zones according to actual needs, and each infotainment terminal is connected to the vehicle's in-vehicle system.
[0061] In some embodiments, the vehicle also includes a central audio control area for controlling audio zones other than the central audio control area, which may be located between the driver's seat and the passenger seat.
[0062] like Figure 1 As shown, the vehicle architecture diagram includes the vehicle 100, the in-vehicle system 101, the driver's side audio zone 102, the passenger side audio zone 103, the second row left audio zone 104, the second row right audio zone 105, the third row left audio zone 106, the third row right audio zone 107, the infotainment terminal corresponding to the driver's side audio zone 108, the infotainment terminal corresponding to the passenger side audio zone 109, the infotainment terminal corresponding to the second row left audio zone 110, the infotainment terminal corresponding to the second row right audio zone 111, the infotainment terminal corresponding to the third row left audio zone 112, and the infotainment terminal corresponding to the third row right audio zone 113.
[0063] The following is combined with Figure 1 The schematic diagram of the vehicle shown illustrates the voice personalization system to which the voice personalization method / device of this application is applicable, such as... Figure 2 As shown, the voice personalization system includes an infotainment terminal 20 and a cloud platform 21.
[0064] The HUT (HeadUnit, Intelligent Cockpit Module) host serves as the infotainment terminal 20 in the vehicle system. The HUT host integrates a voice recognition module 200, a broadcast module 201, and a personalization module 202. The broadcast module 201 and the personalization module 202 are connected to the voice recognition module 200, respectively. The HUT host is also equipped with a display screen 203 and a sound module 204.
[0065] Data is transmitted between the HUT host and the display screen 203, which enables personalized settings and display of the voice image for each audio zone, as well as personalized settings of the broadcast tone (i.e., the broadcast sound).
[0066] The HUT host transmits audio data with the sound module 204, and the sound module 204 outputs the broadcast sound.
[0067] The infotainment terminal 20 communicates wirelessly with the cloud 21 via a wireless network.
[0068] In practical applications, the broadcast module 201 and the image personalization module 202 communicate wirelessly with the cloud 21 via a wireless network.
[0069] The speech recognition module 200 includes speech features such as wake-up, recognition, sound source localization, and semantic understanding. The speech recognition module 200 completes the overall speech interaction chain and interaction.
[0070] Among them, sound source localization is used to confirm the effective sound direction, locate the sound source position, and form a pickup beam with the help of a narrow beam; sound source localization can realize the functions of directional sound pickup, zone control, and personalized interaction through positioning.
[0071] Narrow beam: It forms a pickup beam at a certain angle in space. The signal inside the beam is preserved, while the signal outside the beam is suppressed. It can shield against outside noise, improve the recognition rate, and ensure call quality.
[0072] The image personalization module 202 stores the pre-set voice image of the vehicle system, that is, the image personalization module 202 stores the factory-preset voice image of the vehicle system.
[0073] The pre-installed voice avatars in the vehicle system are built into the system and cannot be deleted. The pre-installed voice avatars are categorized into male, female, and robot avatars.
[0074] The personalization module 202 also stores user-defined voice avatars. For user-defined voice avatars, users can set their own tag names to distinguish voice avatars generated by different users.
[0075] It should be noted that there is no upper limit to the number of voice images stored in the image personalization module 202, and this invention does not impose any limitation on this.
[0076] The personalized image module 202 allows users to select the voice image set for each voice zone. If the user does not set the voice image for each voice zone, the voice image for each voice zone will be in the default state, that is: the voice images of the passenger voice zone, the second row left voice zone, the second row right voice zone, the third row left voice zone, and the third row right voice zone will be consistent with the voice image of the driver voice zone.
[0077] In practical applications, the image personalization module 202 connects to the cloud 21 via data traffic to upload and generate the voice image to be generated. Specifically, the user uploads the voice image to be generated to the infotainment terminal 20 according to the set requirements. The infotainment terminal 20 then uploads the voice image to the cloud 21. The cloud 21 generates a new voice image and sends it to the image personalization module 202 after the new voice image is successfully generated. When the image personalization module 202 receives the new voice image sent by the cloud 21, the corresponding memory will be synchronized and updated.
[0078] The personalized image module 202 also supports downloading and deleting voice images stored in the cloud.
[0079] The voice recognition module 200 and the image personalization module 202 can achieve the following: when a voice requirement is detected in a certain voice area (driver's voice area, passenger's voice area, second row left voice area, second row right voice area, third row left voice area, third row right voice area), the voice image set for that voice area can be displayed and interacted with.
[0080] It should be noted that the need to use voice can be understood as the user waking up the voice (the process of the voice recognition module 200 being invoked) by using a wake word, and then performing certain control or interactive actions by inputting commands.
[0081] The broadcast module 201 stores the broadcast sounds preset by the vehicle system, that is, the broadcast module 201 stores the broadcast sounds preset by the vehicle system at the factory.
[0082] The pre-set announcement sounds of the vehicle system are built into the system and cannot be deleted. The pre-set announcement sounds of the vehicle system are divided into male voice, female voice, boy's voice and girl's voice.
[0083] The broadcast module 201 also stores user-defined broadcast sounds, where the user-defined broadcast sounds refer to the sounds that are synchronized to the HUT host through sound replication.
[0084] For user-defined broadcast voices, users can set their own tag names to distinguish the voices replicated by different users.
[0085] It should be noted that there is no upper limit to the number of broadcast sounds stored in the broadcast module 201, and this invention does not impose any limitation on this.
[0086] In practical applications, the broadcast module 201 connects to the cloud 21 via data traffic to achieve voice replication. Specifically, the user uploads a pre-recorded voice to the infotainment terminal 20, which then uploads the pre-recorded voice to the cloud 21. The cloud 21 trains and synthesizes the pre-recorded voice uploaded by the user, ultimately generating a voice package, which is then sent to the broadcast module 201. This process is called voice replication.
[0087] The speech recognition module 200 and the broadcasting module 201 can achieve the following: when a need for voice is detected in a certain voice zone (driver's voice zone, passenger's voice zone, second row left voice zone, second row right voice zone, third row left voice zone, third row right voice zone), the broadcasting will be carried out according to the broadcasting effect set for that voice zone.
[0088] In practical applications, the effect of the broadcast depends on the content being broadcast and its related settings. For example, there are differences in sound effects; some people prefer a mature and steady voice, while others prefer a youthful and lively voice. The broadcast content is also related to the instructions being executed, such as broadcasting that certain actions have been completed or that the broadcast voice does not currently support this function.
[0089] Understandably, the broadcast effect refers to the effect of the speaker. The broadcast can be performed according to user-defined settings. For example, the user can set different broadcast effects for each vocal range, or if the user does not set any settings, the broadcast will be performed according to the unified default broadcast effect of the vehicle system.
[0090] based on Figure 2 The process by which publicly available voice personalization systems achieve voice personalization is as follows:
[0091] When the voice recognition module 200 detects a wake-up audio signal in any audio region, the infotainment terminal 20 selects the target screen corresponding to the audio region from the multiple screens connected to the in-vehicle system for wake-up.
[0092] If the infotainment terminal 20 detects that a user in the audio zone has personalized settings for the audio zone through the target screen, it determines the user's first personalized voice avatar and first personalized voice, displays the first personalized voice avatar, uses the first personalized voice avatar to interact with the user via voice, and the first personalized voice avatar performs voice broadcasts according to the first personalized voice.
[0093] If no personalized settings for the audio zone are detected by the user through the target screen, the system obtains the default voice image and corresponding default voice prompts for that audio zone in the vehicle system, and uses the default voice image and corresponding default voice prompts to interact with the user via voice.
[0094] In practical applications, the sound module corresponding to the target screen plays the voice message delivered by the default voice avatar through the corresponding default voice prompt.
[0095] The voice broadcast by the default voice avatar through the corresponding default voice is pre-recorded by the user and uploaded to the infotainment terminal 20. The infotainment terminal 20 then uploads the pre-recorded voice to the cloud 21, whereby the cloud 21 trains and synthesizes the pre-recorded voice to generate a voice package, which is then sent to the infotainment terminal 20.
[0096] According to the embodiment of the present invention, when a wake-up audio signal is detected in a certain voice region, the system directly displays the first personalized voice image set by the user for that voice region, performs voice interaction with the user based on the first personalized voice image, and performs voice broadcast according to the first personalized voice set by the user for that voice region, so as to meet the personalized customization needs of all users, realize personalized voice image and voice broadcast, and improve user experience.
[0097] Based on the voice personalization system shown above, such as Figure 3 The diagram shown is a flowchart of a voice personalization method provided in an embodiment of the present invention, which is applied to a voice personalization device.
[0098] It should be noted that voice personalization devices can provide... Figure 2 The infotainment terminal in China.
[0099] This voice personalization method mainly includes the following steps:
[0100] Step S301: Determine whether a wake-up audio signal is detected in any audio region. If yes, proceed to step S302; otherwise, end the operation.
[0101] In step S301, the wake-up audio signal includes, but is not limited to, a wake-up word.
[0102] In the specific implementation of step S301, the vehicle system collects the audio signal of the user in any voice zone through the voice recognition module in the in-vehicle infotainment terminal, and recognizes the collected audio signal in real time. At this time, if the voice recognition module recognizes that the collected audio signal contains a wake-up word, it means that the voice recognition module has recognized a wake-up audio signal in any voice zone, and it also means that the user in the vehicle has a need to use voice. Then step S302 is executed.
[0103] If the voice recognition module does not detect a wake-up word in the collected audio signal during the real-time recognition of various sound zones in the vehicle, it means that the user in the vehicle does not have a need to use voice, and the operation will end directly.
[0104] Optionally, after determining that the voice recognition module has detected an audio signal in any sound zone, the target screen corresponding to the sound zone can be selected from the multiple screens connected to the in-vehicle system for wake-up.
[0105] Specifically, once the voice recognition module detects a wake-up audio signal in any audio region, meaning it determines that the user in the vehicle has a need to use voice, the system selects the target screen in the user's audio region from among the multiple screens connected to the in-vehicle system and wakes it up.
[0106] For example, when the voice recognition module recognizes that a user within the range of the second row left audio area emits a wake-up audio signal containing a wake-up word, it identifies the screen in front of the second row left audio area as the target screen and wakes up the target screen.
[0107] In practical applications, waking up the target screen can be done by changing the target screen from a sleep state to a wake-up state, or by changing the target screen's display interface from a standby interface to the woke-up interface. It should be noted that this is merely an illustrative description of target screen waking up; the specific waking-up method can be determined according to the specific application environment and user needs. This application does not impose any limitations on this method, and all such methods fall within the scope of protection of this application.
[0108] Optionally, the process of selecting the target screen corresponding to the audio zone from the multiple screens connected to the in-vehicle system for wake-up is performed, mainly including the following steps:
[0109] Step S11: Identify the wake-up audio signal and determine the target sound zone where the user who issued the wake-up audio signal is located.
[0110] In the specific implementation step S11, after the speech recognition module recognizes the presence of a wake-up audio signal in any sound zone, it identifies the wake-up audio signal, which involves text conversion and sound source localization of the wake-up audio signal, thereby determining the target sound zone where the user who issued the wake-up audio signal is located.
[0111] For example, when a user issues a wake-up word such as "Hello, please turn on the device", the target sound zone is determined by the sound zone corresponding to the user's current location.
[0112] Step S12: Wake up the target screen corresponding to the target audio region based on the target audio region.
[0113] Step S302: Detect whether the user in the audio region has made personalized settings for the audio region. If yes, proceed to step S303; otherwise, proceed to step S304.
[0114] In the specific implementation of step S302, after determining that a wake-up audio signal exists in any sound zone and wakes up the target screen corresponding to the target sound zone, the user in that sound zone can make personalized settings on the target screen according to their needs. At this time, if it is detected that the user sets personalization through the target screen, step S303 is executed; if it is not detected that the user sets personalization through the target screen, step S304 is executed.
[0115] Step S303: Determine the first personalized voice avatar and the first personalized voice set by the user, display the first personalized voice avatar, use the first personalized voice avatar to interact with the user via voice, and have the first personalized voice avatar broadcast according to the first personalized voice.
[0116] In step S303, the first personalized voice image is any one of the multiple personalized voice images of the voice assistant that are preset in the voice region.
[0117] The first personalized voice is any one of the multiple personalized voices of the voice assistant that are pre-set in the voice zone.
[0118] Among them, the personalized voice assistant is used to interact with users via voice.
[0119] In the specific implementation of step S303, when it is determined that a user in the detected audio region has made personalized settings for the audio region through the target screen, the user's first personalized voice image and first personalized voice are determined, the first personalized voice image is displayed on the target screen, the user interacts with the first personalized voice image through voice, and the audio module is controlled to play the voice broadcast by the first personalized voice image according to the first personalized voice based on the content of the voice interaction.
[0120] For example, when the user who wakes up the voice assistant is in the target audio zone of the second row left audio zone, the first personalized voice image is displayed on the target screen corresponding to the second row left audio zone.
[0121] When the user who wakes up the voice assistant is in the driver's voice zone, the first personalized voice avatar is displayed on the target screen corresponding to the driver's voice zone.
[0122] When the user who wakes up the voice assistant is in the target voice zone of the third row right voice zone, the first personalized voice image is displayed on the target screen corresponding to the third row right voice zone.
[0123] Preferably, once the screen for displaying the personalized voice image is determined, the voice assistant uses the sound module to broadcast the message according to the first personalized voice.
[0124] Optionally, after performing step S303 of displaying the first personalized voice avatar and using the first personalized voice avatar to interact with the user via voice, the method further includes:
[0125] Step S21: Obtain multiple speech information from the same vocal range.
[0126] In step S21, each voice message corresponds to a voice transmission time.
[0127] In the specific implementation step S21, voice information from multiple users in the same audio region is obtained, resulting in multiple voice information.
[0128] Step S22: Compare the duration of each speech sound to obtain the order of speech sound emission times.
[0129] In step S22, the order of speech emission times is sorted from smallest to largest.
[0130] Step S23: Identify the speech information corresponding to each speech time in the order of speech times.
[0131] Step S24: When one or more wake-up audio signals are detected in any voice information, the control area executes steps S302 to S304; otherwise, the operation ends directly.
[0132] Understandably, when multiple users emit voice information in the same or different voice zones, the order of voice emission times is obtained by comparing the voice emission times corresponding to each voice information. Then, the voice information corresponding to each voice emission time is identified in turn according to the order of voice emission times. If one or more wake-up audio signals are identified in one of the voice information, steps S302 to S304 are repeated. Otherwise, the operation ends directly.
[0133] Optionally, after performing step S303 of displaying the first personalized voice avatar and using the first personalized voice avatar to interact with the user via voice, the method further includes:
[0134] Step S31: Obtain speech information from the first and second vocal registers that exist simultaneously outside the vocal register.
[0135] In step S31, each voice message corresponds to a voice transmission time.
[0136] Step S32: Compare the speech time corresponding to the speech information in the first vocal range with the speech time corresponding to the speech information in the second vocal range.
[0137] If the speech time corresponding to the speech information in the first voice region is less than the speech time corresponding to the speech information in the second voice region, proceed with steps S33 to S34.
[0138] If the speech time corresponding to the speech information in the first voice region is greater than the speech time corresponding to the speech information in the second voice region, proceed with steps S35 to S36.
[0139] If the speech time corresponding to the speech information in the first vocal range is equal to the speech time corresponding to the speech information in the second vocal range, proceed with steps S37 to S39.
[0140] Step S33: Identify speech information in the first vocal register.
[0141] Step S34: Determine whether a wake-up audio signal is detected in the voice information of the first voice zone. If yes, control the first voice zone to execute steps S302 to S304. If no, end the operation directly.
[0142] Step S35: Identify speech information in the second vocal register.
[0143] Step S36: Determine whether a wake-up audio signal is detected in the voice information of the second voice zone. If yes, control the second voice zone to execute steps S302 to S304. If no, end the operation directly.
[0144] Step S37: Simultaneously identify speech information in the first and second vocal registers.
[0145] Step S38: Determine whether wake-up audio signals are detected in both the voice information in the first and second voice regions. If yes, control the first and second voice regions to execute steps S302 to S304 in parallel. If no, execute step S39.
[0146] Step S39: Determine whether a wake-up audio signal is detected in the voice information in the first or second voice region. If yes, control the first or second voice region to execute steps S302 to S304. If no, end the operation directly.
[0147] It is understandable that when users in the first and second audio zones outside the designated audio zones simultaneously emit voice information, there may be a time difference in the time when the users emit voice information. In this case, the voice emission time corresponding to the voice information in the first audio zone is compared with the voice emission time corresponding to the voice information in the second audio zone. If the voice emission time corresponding to the voice information in the first audio zone is less than the voice emission time corresponding to the voice information in the second audio zone, the voice information in the first audio zone is identified. If a wake-up audio signal is detected in the voice information in the first audio zone, the first audio zone is controlled to execute steps S302 to S304. Otherwise, the operation is terminated directly.
[0148] If the speech time corresponding to the speech information in the first voice region is greater than the speech time corresponding to the speech information in the second voice region, the speech information in the second voice region is identified. If a wake-up audio signal is detected in the speech information in the second voice region, the second voice region is controlled to execute steps S302 to S304; otherwise, the operation is terminated directly.
[0149] It is also understandable that when users in the first and second audio regions outside the audio region simultaneously emit voice information, there is no time difference in the time when the users emit voice information. In other words, the time when the voice information in the first audio region is emitted is the same as the time when the voice information in the second audio region is emitted. At this time, the voice information in the first and second audio regions is recognized simultaneously.
[0150] If a wake-up audio signal is detected in each voice message, then the first and second voice regions are processed simultaneously, that is, the first and second voice regions are controlled in parallel to execute steps S302 to S304.
[0151] If a wake-up audio signal is detected in the voice information of the first voice zone, then the first voice zone is controlled to execute steps S302 to S304.
[0152] If a wake-up audio signal is detected in the voice information of the second voice zone, then the second voice zone is controlled to execute steps S302 to S304.
[0153] If no wake-up audio signal is detected in any of the voice messages, the operation will end immediately.
[0154] Optionally, step S303, which involves determining the user-defined first personalized voice avatar and first personalized voice, displaying the first personalized voice avatar, using the first personalized voice avatar to interact with the user via voice, and having the first personalized voice avatar perform voice broadcasts according to the first personalized voice, may include:
[0155] Step S31: Select the first personalized voice image set by the user from among the multiple personalized voice images preset in the voice region.
[0156] Step S32: Obtain the image data and voice interaction data corresponding to the first personalized voice image.
[0157] Step S33: Based on the image data and voice interaction data, display the first personalized voice image.
[0158] Step S34: Determine the first personalized voice set by the user from among the multiple personalized voices preset in the sound zone.
[0159] Step S35: Obtain the voice package corresponding to the first personalized voice.
[0160] Step S36: Based on the voice package, use the first personalized voice avatar to interact with the user via voice through the broadcast voice set in the voice package.
[0161] Optionally, step S303, which involves determining the user-defined first personalized voice avatar and first personalized voice, displaying the first personalized voice avatar, using the first personalized voice avatar to interact with the user via voice, and having the first personalized voice avatar perform voice broadcasts according to the first personalized voice, may include:
[0162] Step S41: Obtain the personalized voice image parameters set by the user.
[0163] Step S42: Match the personalized voice image parameters with the personalized voice image library preset in the voice region to obtain a personalized voice image that meets the preset similarity.
[0164] Specifically, features are extracted from the personalized voice image parameters set by the user to obtain one or more image feature tags. These image feature tags are then matched with a pre-defined personalized voice image library for different voice regions to obtain a personalized voice image that meets the preset similarity.
[0165] It should be noted that each personalized voice avatar pre-defined in the voice region has its own corresponding image characteristics, such as a cute personalized voice avatar and a gentle personalized voice avatar. Then, the extracted image feature tags are matched with the image features of each personalized voice avatar to obtain the matching results.
[0166] It should be noted that the matching result achieves a preset similarity with the personalized voice image formed by the user's set personalized voice image parameters.
[0167] For example, the extracted image feature labels are "children" and "nursery rhymes". Based on "children" and "nursery rhymes", the image features "cute" and "gentle" of multiple personalized voice images that are preset in the voice range are matched. Finally, the matching result is determined to be "cute", which means that a cute personalized voice image is determined.
[0168] Step S43: Obtain the image data and voice interaction data corresponding to the personalized voice image that meets the preset similarity.
[0169] Step S44: Based on image data and voice interaction data, display personalized voice images that meet the preset similarity.
[0170] Step S45: Obtain the personalized voice parameters corresponding to the personalized voice image that meets the preset similarity set by the user.
[0171] Step S46: Match the personalized voice parameters with the personalized voice library preset in the sound range to obtain a personalized voice that meets the preset similarity.
[0172] Specifically, users set corresponding personalized voice parameters for the matched personalized voice image, extract features from the personalized voice parameters to obtain one or more timbre feature tags, and match the timbre feature tags with a pre-set personalized voice library of voice regions to obtain a personalized voice that meets the preset similarity.
[0173] It should be noted that each personalized voice preset in the sound region has its own corresponding timbre characteristics, and each personalized voice image is associated with and stored with its corresponding personalized voice.
[0174] For example, when a cute personalized voice avatar is matched, the extracted timbre feature tags are matched with the individual timbre features of each personalized voice to obtain the matching result.
[0175] It should be noted that the matching results achieve a preset similarity between the personalized voice and the user-defined personalized voice parameters.
[0176] For example, if the extracted timbre feature label is "lively", then the timbre feature "cheerful" of multiple personalized voices preset in the vocal range is matched, and the final matching result is "cheerful", which means that a cheerful personalized voice is determined.
[0177] Step S47: Using a personalized voice image that meets the preset similarity, interact with the user via voice through the corresponding personalized voice that meets the preset similarity.
[0178] Step S48: Obtain the voice pack corresponding to the personalized voice image that meets the preset similarity.
[0179] Step S49: Using a personalized voice avatar that meets the preset similarity, interact with the user via voice through the broadcast voice set in the voice package.
[0180] Step S304: Obtain the default voice image and corresponding default broadcast voice preset in the audio region, display the default voice image, and use the default voice image to conduct voice interaction with the user through the corresponding default broadcast voice.
[0181] In step S304, the default voice images preset in the audio zone of the vehicle system include, but are not limited to, male, female, and robot voice images.
[0182] The default broadcast voices preset in the audio zone of the in-vehicle system include, but are not limited to, male voices, female voices, boy voices, and girl voices.
[0183] In the specific implementation of step S304, if it is determined that the user has not made personalized settings for the audio zone through the target screen, the default voice image and corresponding default broadcast voice preset in the vehicle system for that audio zone are obtained, and the default voice image is displayed on the target screen. Using the default voice image, the user is interacted with via voice through the corresponding default broadcast voice.
[0184] In other words, if the user does not set the voice image for the target voice area, the voice image for the target voice area will be in the default state, that is: the voice images for the passenger voice area, the second row left voice area, the second row right voice area, the third row left voice area, and the third row right voice area will be consistent with the voice image for the driver voice area.
[0185] For example, when the user who wakes up the voice assistant is in the second row left voice zone, a male image is displayed on the target screen corresponding to the second row left voice zone, and a male image is also displayed on the target screen corresponding to other voice zones.
[0186] When the user who activates the voice assistant is in the driver's voice zone, a female image is displayed on the target screen corresponding to the driver's voice zone, and a female image is also displayed on the target screen corresponding to other voice zones.
[0187] When the user who has activated the voice assistant is in the target audio region of the third row right audio region, the robot image is displayed on the target screen corresponding to the third row right audio region, and the robot image is also displayed on the target screen corresponding to other audio regions.
[0188] For example, when the user who wakes up the voice assistant is in the target audio zone of the second row left audio zone, a male image is displayed on the target screen corresponding to the second row left audio zone. Using the male image, the voice assistant interacts with the user through a corresponding male voice. In actual application, the sound module is controlled to play the voice message read by the voice assistant in a male voice.
[0189] When the user who has activated the voice assistant is in the driver's voice zone, a male image is displayed on the target screen corresponding to the driver's voice zone. Using the male image, the voice assistant interacts with the user through a corresponding male voice. In actual application, the voice module is controlled by the voice assistant broadcasting the voice in a male voice.
[0190] When the user who has activated the voice assistant is in the target audio zone of the third row right audio zone, a male image is displayed on the target screen corresponding to the third row right audio zone. Using the male image, the voice assistant interacts with the user through a corresponding male voice. In actual application, the voice module is controlled by the voice assistant broadcasting the voice in the male voice.
[0191] According to the embodiment of the present invention, when a wake-up audio signal is detected in a certain voice region, the first personalized voice image is directly displayed according to the user's setting for displaying the first personalized voice image in that voice region. Voice interaction is conducted with the user based on the first personalized voice image, and voice broadcast is performed according to the user's setting for broadcasting the first personalized voice in that voice region, so as to meet the personalized customization needs of all users, realize the personalization of voice image and voice broadcast, and improve the user experience.
[0192] Preferred, based on Figure 3 The voice personalization method also includes the following steps:
[0193] Step S51: Obtain the voice image to be generated uploaded by the user through the target screen according to preset requirements.
[0194] In the specific implementation step S51, the user designs the voice image to be generated on a mobile terminal (such as a mobile phone), uploads the voice image to be generated to the infotainment terminal according to preset requirements via a wireless network, and displays the voice image uploaded by the user on the target screen of the infotainment terminal.
[0195] Step S52: Based on the voice image to be generated, initiate a new voice image generation request to the cloud.
[0196] In the specific implementation step S52, based on the voice image to be generated, the voice image to be generated is uploaded to the cloud according to preset requirements via a wireless network, and a new voice image generation request is initiated to the cloud.
[0197] Step S53: Receive the new voice image issued by the cloud based on the new voice image generation request, and store and update it.
[0198] In step S53, the new voice avatar is generated by the cloud based on the new voice avatar generation request.
[0199] In the specific implementation step S53, after receiving the voice image to be generated and the request to generate a new voice image, the cloud generates the voice image to be generated based on the request to generate a new voice image, obtains a new voice image, and sends the new voice image to the infotainment terminal. The infotainment terminal receives the new voice image sent by the cloud and stores the new voice image in the corresponding memory, and updates the voice image already stored in the memory.
[0200] Step S54: Obtain the tag names corresponding to the different new voice avatars set by the user through the target screen.
[0201] During the specific implementation of step S54, the user can view the newly generated voice avatar on the target screen and set the label name for the new voice avatar according to their needs to distinguish different new voice avatars.
[0202] For example, three new voice avatars have been generated: new voice avatar 1, new voice avatar 2, and new voice avatar 3. Based on the requirements, the label name of new voice avatar 1 is set as a child avatar, the label name of new voice avatar 2 is set as a teenager avatar, and the label name of new voice avatar 3 is set as a middle-aged avatar.
[0203] Step S55: For the new voice image, store the new voice image and the corresponding tag name.
[0204] For example, establish a relationship between the new voice image 1 and the child image, and store the new voice image 1 and the child image accordingly based on the established relationship.
[0205] For example, a relationship can be established between the new voice image 2 and the youth image, and the new voice image 2 and the youth image can be stored accordingly based on the established relationship.
[0206] Preferably, in some embodiments, the user sets corresponding tag names for the stored voice images via the target screen. Specifically, the new tag names corresponding to different stored voice images set by the user via the target screen are obtained, the original tag names corresponding to different stored voice images are determined, and the original tag names corresponding to different stored voice images are changed to new tag names.
[0207] Preferred, based on Figure 3 The voice personalization method also includes the following steps:
[0208] Step S61: Obtain the audio to be generated uploaded by the user through the target screen.
[0209] In the specific implementation step S61, the user pre-records the sound to be generated on the mobile terminal, and uploads the sound to be generated to the infotainment terminal according to preset requirements via wireless network on the mobile terminal. The target screen of the infotainment terminal displays the sound uploaded by the user.
[0210] Step S62: Based on the voice to be generated, initiate a new voice generation request to the cloud.
[0211] In the specific implementation of step S62, based on the sound to be generated, the sound to be generated is uploaded to the cloud according to preset requirements via a wireless network, and a new voice generation request is initiated to the cloud.
[0212] Step S63: Receive the new voice packet sent by the cloud based on the new voice generation request, and store and update it.
[0213] In step S63, the new speech package is generated by the cloud after training and synthesizing the voice to be generated based on the new speech generation request.
[0214] In the specific implementation step S63, after receiving the voice to be generated and the request to generate a new voice image, the cloud generates the voice to be generated based on the new voice generation request, obtains a new voice package, and sends the new voice package to the infotainment terminal. The infotainment terminal receives the new voice package sent by the cloud and stores the new voice package in the corresponding memory, and updates the voice package already stored in the memory.
[0215] Step S64: Obtain the tag names corresponding to the different new voice packs set by the user through the target screen.
[0216] During the specific implementation of step S64, the user can view the generated new voice pack on the target screen and set the tag name for the new voice pack according to their needs to distinguish different new voice packs.
[0217] For example, two new voice packs have been generated: New Voice Pack 1 and New Voice Pack 2. As needed, the tag name of New Voice Pack 1 is set to Cartoon Boy Voice, and the tag name of New Voice Pack 2 is set to Cartoon Girl Voice.
[0218] Step S65: For the new voice package, store the new voice package and the corresponding tag name.
[0219] For example, establish a relationship between the new voice pack 1 and the cartoon boy's voice, and store the new voice pack 1 and the cartoon boy's voice accordingly based on the established relationship.
[0220] For example, a relationship can be established between the new voice pack 2 and the cartoon girl's voice, and the new voice pack 2 and the cartoon girl's voice can be stored accordingly based on the established relationship.
[0221] Preferably, in some embodiments, the user sets corresponding tag names for the stored voice packs via the target screen. Specifically, the new tag names corresponding to different stored voice packs set by the user via the target screen are obtained, the original tag names corresponding to different stored voice packs are determined, and the original tag names corresponding to different stored voice packs are changed to the new tag names.
[0222] Preferred, based on Figure 3 The voice personalization method also includes the following steps:
[0223] Step S71: Set up the online download portal.
[0224] In the specific implementation step S71, an online download portal for downloading voice images from the cloud is set up in advance on the infotainment terminal.
[0225] Step S72: Initiate a voice image download request to the cloud based on the online download portal.
[0226] Step S73: Receive the voice image from the cloud in response to the voice image download request.
[0227] In the specific implementation step S73, the voice image feedback from the cloud based on the voice image download request is received, and the voice image feedback from the cloud is stored in the memory.
[0228] Step S74: Send a request to the cloud to delete the voice image.
[0229] Step S75: Receive the voice image deletion result from the cloud based on the voice image deletion request.
[0230] In the specific implementation step S75, after receiving the voice image deletion request, the cloud deletes the voice image that needs to be deleted as indicated in the voice image deletion request, obtains the voice image deletion result, and feeds back the voice image deletion result to the infotainment terminal. The infotainment terminal receives the voice image deletion result fed back by the cloud based on the voice image deletion request.
[0231] According to the voice personalization method provided by the embodiments of the present invention, when a wake-up audio signal is detected in a certain voice region, the first personalized voice image is directly displayed according to the user's setting for displaying the first personalized voice image in that voice region. Voice interaction is conducted with the user based on the first personalized voice image, and voice broadcast is performed according to the first personalized voice set by the user for broadcasting in that voice region, so as to meet the personalized customization needs of all users, realize the personalization of voice image and voice broadcast, improve the user experience, and enrich the personalized customization options by constructing new voice packages and new voice images, which can better meet the personalized customization needs of users.
[0232] Compared with the above embodiments of the present invention Figure 3 Corresponding to the voice personalization method shown, this embodiment of the invention also provides a voice personalization device, such as... Figure 4 As shown, the voice personalization device includes a detection module 401 and a first display and broadcast module 402.
[0233] The detection module 401 is used to detect whether the user in any sound zone has made personalized settings when a wake-up audio signal is detected in any sound zone.
[0234] The display and broadcast module 402 is used to determine the first personalized voice image and the first personalized voice set by the user if it is detected that the user has made personalized settings for the audio region, display the first personalized voice image, use the first personalized voice image to interact with the user by voice, and have the first personalized voice image broadcast the message according to the first personalized voice.
[0235] Optionally, based on the above Figure 4 The voice personalization device shown combines Figure 4 The voice personalization device also includes a default settings module.
[0236] The default setting module is used to obtain the default voice image and corresponding default broadcast voice for the pre-set voice area if no personalized settings are detected by the user. The default voice image is then displayed, and the user is interacted with via voice using the corresponding default broadcast voice.
[0237] Optionally, based on the above Figure 4 The voice personalization device shown combines Figure 4 The voice personalization device also includes a first processing module.
[0238] The first processing module is used to acquire multiple voice information from the same audio region, each voice information corresponding to a voice emission time; compare the magnitude of each voice emission time to obtain the order of the voice emission times, which is sorted in ascending order; according to the order of the voice emission times, sequentially identify the voice information corresponding to each voice emission time; when one or more wake-up audio signals are detected for any voice information, control the audio region to perform the above-described actions. Figure 3 Voice personalization methods.
[0239] Optionally, based on the above Figure 4 The voice personalization device shown combines Figure 4 The voice personalization device is further equipped with a second processing module, a third processing module, a fourth processing module and a fifth processing module.
[0240] The second processing module is used to acquire speech information that exists simultaneously in the first and second speech regions outside the existing speech regions; and to compare the speech emission time corresponding to the speech information in the first speech region with the speech emission time corresponding to the speech information in the second speech region.
[0241] The third processing module is used to identify the speech information in the first audio region if the speech emission time corresponding to the speech information in the first audio region is less than the speech emission time corresponding to the speech information in the second audio region; when a wake-up audio signal is detected in the speech information in the first audio region, it controls the first audio region to perform the above-mentioned actions. Figure 3 Voice personalization methods.
[0242] The fourth processing module is used to identify the speech information in the second voice region if the speech emission time corresponding to the speech information in the first voice region is greater than the speech emission time corresponding to the speech information in the second voice region; when a wake-up audio signal is detected in the speech information in the second voice region, it controls the second voice region to perform the above-described actions. Figure 3 Voice personalization methods.
[0243] The fifth processing module is used to simultaneously identify the speech information in the first and second sound regions if the speech emission time corresponding to the speech information in the first sound region is equal to the speech emission time corresponding to the speech information in the second sound region; when both the speech information in the first and second sound regions are identified to have wake-up audio signals, it controls the first and second sound regions to perform the above-described actions in parallel. Figure 3 The voice personalization method; when a wake-up audio signal is detected in the voice information of the first voice region or the voice information of the second voice region, the first voice region or the second voice region is controlled to perform the above-described actions. Figure 3 Voice personalization methods.
[0244] Optionally, based on the above Figure 4 The display and broadcast module 402 shown includes:
[0245] The first determining unit is used to determine the first personalized voice image set by the user from multiple personalized voice images preset in the voice region.
[0246] The first acquisition unit is used to acquire the image data and voice interaction data corresponding to the first personalized voice image.
[0247] The first display unit is used to showcase the first personalized voice avatar based on image data and voice interaction data.
[0248] The second determining unit is used to determine the first personalized sound set by the user among multiple personalized sounds preset in the sound zone.
[0249] The second acquisition unit is used to acquire the voice package corresponding to the first personalized voice.
[0250] The broadcasting unit is used to interact with the user via voice based on a voice package and using a first personalized voice avatar, through the broadcasting voice set in the voice package.
[0251] Optionally, based on the above Figure 4 The display and broadcast module 402 shown further includes:
[0252] The third acquisition unit is used to acquire the personalized voice image parameters set by the user.
[0253] The first matching unit is used to match the personalized voice image parameters with the personalized voice image library that is preset in the voice region to obtain a personalized voice image that meets the preset similarity.
[0254] The fourth acquisition unit is used to acquire image data and voice interaction data corresponding to personalized voice images that meet preset similarity.
[0255] The second display unit is used to display personalized voice images that meet preset similarity based on image data and voice interaction data.
[0256] The fifth acquisition unit is used to acquire personalized voice parameters corresponding to the personalized voice image that meets the preset similarity set by the user.
[0257] The second matching unit is used to match personalized voice parameters with a pre-set personalized voice library for the voice region to obtain personalized voices that meet the preset similarity.
[0258] The first voice interaction unit is used to interact with the user through a personalized voice image that meets a preset similarity score.
[0259] Optionally, based on the above Figure 4 The display and broadcast module 402 shown further includes:
[0260] The sixth acquisition unit is used to acquire the voice package corresponding to the personalized voice image that meets the preset similarity.
[0261] The second voice interaction unit is used to interact with the user via voice using a personalized voice image that meets a preset similarity score and the broadcast voice set in the voice package.
[0262] It should be noted that the specific principles and execution processes of each module in the voice personalization device disclosed in the above embodiments of the present invention are the same as those of the voice personalization method implemented in the above embodiments of the present invention. Please refer to the corresponding parts of the voice personalization method disclosed in the above embodiments of the present invention, and they will not be repeated here.
[0263] According to the embodiment of the present invention, when a wake-up audio signal is detected in a certain voice region, the first personalized voice image is directly displayed according to the user's setting for displaying the first personalized voice image in that voice region. The user interacts with the user based on the first personalized voice image and broadcasts the first personalized voice according to the user's setting for broadcasting the first personalized voice in that voice region, so as to meet the personalized customization needs of all users, realize the personalization of voice image and voice broadcast, and improve the user experience.
[0264] This application also discloses an electronic device for running a program; wherein, when the program runs, it executes the voice personalization method described in any of the above embodiments.
[0265] It should be noted that the relevant explanations regarding the voice personalization method can be found in the corresponding embodiments described above, and will not be repeated here.
[0266] This application also discloses a vehicle that includes the aforementioned electronic equipment.
[0267] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0268] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0269] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of voice personalization, characterized by, The method comprises: When identifying that any sound area exists a wake-up audio signal, detecting whether a user in the sound area personalizes the sound area; If detecting that the user personalizes the sound area, determining a first personalized voice image and a first personalized sound set by the user, displaying the first personalized voice image, using the first personalized voice image to interact with the user by voice, and using the first personalized voice image to broadcast by voice according to the first personalized sound; If not detecting that the user personalizes the sound area, obtaining a default voice image and a corresponding default broadcast sound pre-set in the sound area, displaying the default voice image, and using the default voice image to interact with the user by voice through the corresponding default broadcast sound.
2. The method of claim 1, wherein, Further comprising: Obtaining multiple voice information in the same sound area, each voice information corresponding to a voice emission time; Comparing the size of each voice emission time to obtain the arrangement order of the voice emission time, which is sequentially sorted from small to large; According to the arrangement order of the voice emission time, sequentially identifying the voice information corresponding to each voice emission time; When identifying that any voice information exists one or more wake-up audio signals, controlling the sound area to execute the voice personalization method of claim 1.
3. The method of claim 1, wherein, Further comprising: Obtaining voice information existing in a first sound area and a second sound area outside the sound area at the same time; Comparing the voice emission time corresponding to the voice information in the first sound area and the voice emission time corresponding to the voice information in the second sound area; If the voice emission time corresponding to the voice information in the first sound area is less than the voice emission time corresponding to the voice information in the second sound area, identifying the voice information in the first sound area; When identifying that the voice information in the first sound area exists the wake-up audio signal, controlling the first sound area to execute the voice personalization method of claim 1; If the voice emission time corresponding to the voice information in the first sound area is greater than the voice emission time corresponding to the voice information in the second sound area, identifying the voice information in the second sound area; When identifying that the voice information in the second sound area exists the wake-up audio signal, controlling the second sound area to execute the voice personalization method of claim 1; If the voice emission time corresponding to the voice information in the first sound area is equal to the voice emission time corresponding to the voice information in the second sound area, identifying the voice information in the first sound area and the voice information in the second sound area at the same time; When identifying that the voice information in the first sound area and the voice information in the second sound area both exist the wake-up audio signal, controlling the first sound area and the second sound area to execute the voice personalization method of claim 1 in parallel; When identifying that the voice information in the first sound area or the voice information in the second sound area exists the wake-up audio signal, controlling the first sound area or the second sound area to execute the voice personalization method of claim 1.
4. The method of claim 1, wherein, The determining of the first personalized voice image and the first personalized voice set by the user, the displaying of the first personalized voice image, the voice interaction between the first personalized voice image and the user, and the voice broadcast by the first personalized voice image according to the first personalized voice set, comprise: determining the first personalized voice image set by the user from the plurality of pre-set personalized voice images in the sound area; obtaining the image data and voice interaction data corresponding to the first personalized voice image; displaying the first personalized voice image based on the image data and the voice interaction data; determining the first personalized voice set by the user from the plurality of pre-set personalized voices in the sound area; obtaining the voice package corresponding to the first personalized voice; voice interaction between the first personalized voice image and the user through the broadcast voice set in the voice package based on the voice package.
5. The method of claim 1, wherein, The determining of the first personalized voice image and the first personalized voice set by the user, the displaying of the first personalized voice image, the voice interaction between the first personalized voice image and the user, and the voice broadcast by the first personalized voice image according to the first personalized voice set, comprise: obtaining the personalized voice image parameter set by the user; matching the personalized voice image parameter with the personalized voice image library pre-set in the sound area to obtain the personalized voice image satisfying the pre-set similarity; obtaining the image data and voice interaction data corresponding to the personalized voice image satisfying the pre-set similarity; displaying the personalized voice image satisfying the pre-set similarity based on the image data and the voice interaction data; obtaining the personalized voice parameter corresponding to the personalized voice image satisfying the pre-set similarity set by the user; matching the personalized voice parameter with the personalized voice library pre-set in the sound area to obtain the personalized voice satisfying the pre-set similarity; voice interaction between the personalized voice image satisfying the pre-set similarity and the user through the corresponding personalized voice satisfying the pre-set similarity.
6. The method of claim 5, wherein, The determining of the first personalized voice image and the first personalized voice set by the user, the displaying of the first personalized voice image, the voice interaction between the first personalized voice image and the user, and the voice broadcast by the first personalized voice image according to the first personalized voice set, comprise: obtaining the voice package corresponding to the personalized voice image satisfying the pre-set similarity; voice interaction between the personalized voice image satisfying the pre-set similarity and the user through the broadcast voice set in the voice package.
7. A voice personalization apparatus, characterized by comprising: The device comprises: a detection module for detecting whether the user in the sound area has personalized the sound area when recognizing that there is a wake-up audio signal in any sound area; The first display and broadcast module is configured to, if the user personalizes the sound area, determine a first personalized voice image and a first personalized sound set by the user, display the first personalized voice image, perform voice interaction with the user by using the first personalized voice image, and perform voice broadcast by the first personalized voice image in the first personalized sound. The default setting module is configured to, if the user does not personalize the sound area, obtain a default voice image and a corresponding default broadcast sound set pre-set for the sound area, display the default voice image, and perform voice interaction with the user by using the default voice image and the corresponding default broadcast sound.
8. An electronic device, comprising: The electronic device is configured to run a program, and the program is configured to execute the voice personalization method according to any one of claims 1-6.
9. A vehicle characterized by comprising: The vehicle comprises the electronic device according to claim 8.
Citation Information
Patent Citations
Customized prompt generation method, device and equipment
CN111145721A
Atmosphere lamp control method and device, electronic equipment and storage medium
CN113709954A