Voice packet switching method, device and electronic equipment

CN118609560BActive Publication Date: 2026-08-07CHINA FAW CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA FAW CO LTD
Filing Date
2024-06-06
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本发明实施例提供了一种语音包切换方法、装置和电子设备,以至少解决相关技术中缺乏个性化定制选项,对地理位置的感知不够智能的技术问题

Benefits of technology

[0023] According to one embodiment of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the voice packet switching method described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118609560B_ABST
    Figure CN118609560B_ABST
Patent Text Reader

Abstract

The application discloses a voice package switching method and device and electronic equipment, and relates to the technical field of vehicles. The method comprises the following steps: acquiring a first position of a vehicle, wherein the first position is in a first region; in response to determining that the vehicle drives out of the first region through the first position, acquiring a second position of the vehicle, wherein the second position is in a second region; determining a target voice package based on the second position, wherein the target voice package is generated based on a speaking feature corresponding to the second region; and switching an initial voice package of an in-vehicle infotainment system to the target voice package, wherein the initial voice package is generated based on a speaking feature corresponding to the first region. The application solves the technical problems that there is a lack of personalized customization options and the perception of geographical positions is not intelligent enough in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle technology, and more specifically, to a voice packet switching method, apparatus, and electronic device. Background Technology

[0002] With the development of artificial intelligence, personalized in-vehicle voice packages can be customized based on the vehicle's location, providing relevant information and services according to the vehicle's location. Personalized voice packages can be tailored to the driver's preferences and needs, providing a personalized voice interaction experience. For example, when a user is driving to their destination, a personalized in-vehicle voice package can provide real-time traffic information, road condition prompts, and navigation guidance based on the user's location. Furthermore, it can enhance localized dialect-based dialogue interaction based on the user's location.

[0003] Current in-vehicle infotainment systems lack personalized customization options, preventing users from tailoring the voice assistant's speech rate, tone, and other aspects to their preferences and needs. Furthermore, these systems lack intelligent geolocation awareness, failing to automatically switch to the appropriate voice package based on the vehicle's location, resulting in a lack of personalized and localized voice interaction experiences.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This invention provides a voice pack switching method, apparatus, and electronic device to at least address the technical problems of lacking personalized customization options and insufficient intelligence in the perception of geographical location in related technologies.

[0006] According to one embodiment of the present invention, a voice packet switching method is provided, comprising: obtaining a first location of a vehicle, wherein the first location is within a first region; in response to determining that the vehicle has left the first region through the first location, obtaining a second location of the vehicle, wherein the second location is within a second region; determining a target voice packet based on the second location, wherein the target voice packet is generated based on speech features corresponding to the second region; and switching an initial voice packet of the vehicle system to the target voice packet, wherein the initial voice packet is generated based on speech features corresponding to the first region.

[0007] Optionally, the method further includes: acquiring speech audio data, wherein the speech audio data is used to record the user's speech content; extracting features from the speech audio data to obtain speech features; and generating a target speech package set based on the speech features, wherein the target speech package set includes speech packages corresponding to different regions.

[0008] Optionally, the speech features include prosodic features and accent features. The speech features are obtained by extracting features from the speech audio data, including: extracting prosodic features from the speech audio data; converting the speech audio data into text to obtain the target text; and extracting accent features from the target text.

[0009] Optionally, generating a target speech packet set based on speech features includes: obtaining a speech feature set, wherein the speech feature set includes speech features corresponding to different regions; and generating a target speech packet set based on the speech features and the speech feature set.

[0010] Optionally, the method further includes: displaying a prompt message in the graphical user interface of the vehicle system, wherein the prompt message is used to indicate that the target voice package set has been generated.

[0011] Optionally, the method further includes at least one of the following: in response to receiving a confirmation instruction, importing the target voice packet set into a preset storage location; in response to receiving a modification instruction, updating the target voice packet set.

[0012] Optionally, the method further includes updating the target speech packet set based on the user's speech content.

[0013] According to one embodiment of the present invention, a voice packet switching device is also provided, comprising: a first acquisition module for acquiring a first location of a vehicle, wherein the first location is within a first region; a second acquisition module for acquiring a second location of the vehicle in response to determining that the vehicle has left the first region through the first location, wherein the second location is within the second region; a determination module for determining a target voice packet based on the second location, wherein the target voice packet is generated based on speech features corresponding to the second region; and a switching module for switching an initial voice packet of the vehicle system to the target voice packet, wherein the initial voice packet is generated based on speech features corresponding to the first region.

[0014] Optionally, the method further includes: a generation module for acquiring speech audio data, wherein the speech audio data is used to record the user's speech content; performing feature extraction on the speech audio data to obtain speech features; and generating a target speech package set based on the speech features, wherein the target speech package set includes speech packages corresponding to different regions.

[0015] Optionally, the generation module is also used to extract features from the speech audio data to obtain prosodic features; to convert the speech audio data into text to obtain target text; and to extract features from the target text to obtain accent features.

[0016] Optionally, the generation module is also used to obtain a speech feature set, wherein the speech feature set includes speech features corresponding to different regions; and to generate a target speech packet set based on the speech features and the speech feature set.

[0017] Optionally, the device further includes a display module for displaying prompt information in the graphical user interface of the vehicle system, wherein the prompt information indicates that the target voice package set has been generated.

[0018] Optionally, the device further includes: an instruction module, configured to import the target voice package set into a preset storage location in response to receiving a confirmation instruction; and to update the target voice package set in response to receiving a modification instruction.

[0019] Optionally, the device further includes an update module for updating the target speech packet set based on the user's speech content.

[0020] According to one embodiment of the present invention, a computer-readable storage medium is also provided, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the voice packet switching method described above when running on a computer or processor.

[0021] According to one embodiment of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the voice packet switching method in the embodiments of the present invention.

[0022] According to one embodiment of the present invention, a vehicle is also provided, which is used to perform the voice packet switching method described in any of the above claims.

[0023] According to one embodiment of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the voice packet switching method described above.

[0024] In this embodiment of the invention, a technical solution is achieved by obtaining a first location of the vehicle, wherein the first location is within a first region; in response to determining that the vehicle has left the first region based on the first location, obtaining a second location of the vehicle, wherein the second location is within a second region; determining a target voice package based on the second location, wherein the target voice package is generated based on the speech features corresponding to the second region; and switching the initial voice package of the vehicle system to the target voice package, wherein the initial voice package is generated based on the speech features corresponding to the first region. This achieves the technical effect of generating a unique voice package for the user based on their daily conversations with the vehicle system, and automatically switching to the corresponding regional voice package based on the vehicle's location. This provides personalized and intelligent voice interaction for the vehicle system, thereby solving the technical problems of lacking personalized customization options and insufficient intelligent perception of geographical location in related technologies. Attached Figure Description

[0025] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0026] Figure 1 This is a flowchart of a voice packet switching method according to an embodiment of the present invention;

[0027] Figure 2 This is a schematic diagram of a voice packet switching method according to an embodiment of the present invention;

[0028] Figure 3 This is a timing diagram of a voice packet switching method according to an embodiment of the present invention;

[0029] Figure 4 This is a structural block diagram of a voice packet switching device according to an embodiment of the present invention. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. In the description of these embodiments, unless otherwise stated, "a plurality of" means two or more. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] According to one embodiment of the present invention, an embodiment of a voice packet switching method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0033] This method embodiment can be executed in an electronic device, similar control device, or system that includes a memory and a processor. Taking an electronic device as an example, the electronic device may include one or more processors and a memory for storing data. Optionally, the electronic device may also include a communication device for communication functions and a display device. Those skilled in the art will understand that the above structural description is merely illustrative and does not limit the structure of the electronic device. For example, the electronic device may include more or fewer components than described above, or have a different configuration than described above.

[0034] A processor may include one or more processing units. For example, a processor may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP) chip, a microcontroller unit (MCU), a field-programmable gate array (FPGA), a neural network processing unit (NPU), a tensor processing unit (TPU), or an artificial intelligence (AI) processor. Different processing units may be independent components or integrated into one or more processors. In some instances, electronic devices may also include one or more processors.

[0035] The memory can be used to store computer programs, such as the computer program corresponding to the voice packet switching method in this embodiment of the invention. The processor implements the aforementioned voice packet switching method by running the computer program stored in the memory. The memory may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to electronic devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0036] Communication devices are used to receive or send data via a network. Specific examples of such networks may include wireless networks provided by the mobile terminal's communication provider. In one example, the communication device includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the communication device may be a radio frequency (RF) module used for wireless communication with the Internet.

[0037] The display device can be, for example, a touchscreen liquid crystal display (LCD) and a touch display (also referred to as a "touchscreen" or "touch screen"). This LCD allows the user to interact with the user interface of the mobile terminal. In some embodiments, the mobile terminal has a graphical user interface (GUI), which allows the user to interact with the GUI by touching and / or gesturing on a touch-sensitive surface. Optional human-computer interaction functions include: creating web pages, drawing, word processing, creating electronic documents, playing games, video conferencing, instant messaging, sending and receiving emails, a call interface, playing digital video, playing digital music, and / or web browsing, etc. Executable instructions for performing the above human-computer interaction functions are configured / stored in one or more processor-executable computer program products or readable storage media.

[0038] This embodiment provides a voice packet switching method running on an electronic device. Figure 1 This is a flowchart of a voice packet switching method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0039] Step S12: Obtain the first location of the vehicle, wherein the first location is within a first region;

[0040] Step S14: In response to determining that the vehicle has left the first region through the first location, obtain the second location of the vehicle, wherein the second location is within the second region;

[0041] Step S16: Determine the target speech packet based on the second location, wherein the target speech packet is generated based on the speech features corresponding to the second region;

[0042] Step S18: Switch the initial voice packet of the vehicle system to the target voice packet, wherein the initial voice packet is generated based on the speech features corresponding to the first region.

[0043] For example, the vehicle's initial position, i.e., the first position, is obtained, which is located within a pre-defined first region. When the vehicle leaves the first region, its current position, i.e., the second position, is obtained, which is located within a second region. That is, when the vehicle's location changes, a target voice package is determined based on the vehicle's current location. This target voice package represents a voice package generated using speech synthesis technology, incorporating the speech characteristics (i.e., local dialect) corresponding to the current location, as well as the user's personalized voice features. The initial voice package set by the vehicle's infotainment system when the vehicle was in its initial position is then switched to the target voice package.

[0044] Vehicle location can be determined using positioning technologies. Common positioning technologies include: Global Positioning System (GPS): This uses satellite signals to determine a vehicle's location and is widely used in navigation, mapping, aviation, and maritime fields. Regional positioning technologies: These include wireless Fidelity (WiFi) positioning, Bluetooth positioning, and Radio Frequency Identification (RFID) positioning, which can achieve positioning indoors or in dense buildings. Inertial navigation technology: This uses sensors such as accelerometers and gyroscopes to measure the acceleration and angular velocity of an object, thereby calculating the vehicle's position. Visual positioning technology: This uses devices such as cameras or laser scanners to capture images or 3D data of the vehicle, thereby determining its position. Ultrasonic positioning technology: This uses the transmission and reception of ultrasonic signals to measure the distance between the vehicle and sensors, thereby determining the object's position. Other positioning technologies include geomagnetic positioning, topographic mapping positioning, radio positioning, and various other technologies. The appropriate positioning technology can be selected based on specific needs.

[0045] Different regions can be divided according to latitude and longitude coordinates. For example, a region can be divided into areas with longitude intervals of 5° or latitude intervals of 10°. In particular, the same region corresponds to the same speech characteristics, that is, the same local dialect is spoken within the same region.

[0046] In another embodiment, different regions can be divided by administrative areas within a country or region. For example, if the current location of the vehicle is within the city of Xi'an, then the speech packet can be determined to be generated based on the local dialect of Xi'an.

[0047] Based on the above steps, the technical solution involves obtaining a first location of the vehicle, wherein the first location is within a first region; in response to determining that the vehicle has left the first region using the first location, obtaining a second location of the vehicle, wherein the second location is within a second region; determining a target voice package based on the second location, wherein the target voice package is generated based on the speech features corresponding to the second region; and switching the initial voice package of the vehicle system to the target voice package, wherein the initial voice package is generated based on the speech features corresponding to the first region. This achieves the technical effect of generating a unique voice package for the user based on their daily conversations with the vehicle system, and automatically switching to the corresponding regional voice package based on the vehicle's location. This provides personalized and intelligent voice interaction for the vehicle system, thereby solving the technical problems of lacking personalized customization options and insufficient intelligent perception of geographical location in related technologies.

[0048] Optionally, the method further includes step S15, which specifically includes performing the following steps:

[0049] Step S151: Obtain voice audio data, wherein the voice audio data is used to record the user's speech content;

[0050] Step S152: Extract features from the speech audio data to obtain speech features;

[0051] Step S153: Generate a target speech packet set based on speech features, wherein the target speech packet set includes speech packets corresponding to different regions.

[0052] For example, voice audio data is acquired through the vehicle's infotainment system. This voice audio data records the dialogue between the user and the infotainment system, or between users themselves. Feature extraction is then performed on the acquired voice audio data to obtain the voice features of the spoken content, and a personalized target voice package is generated accordingly. This target voice package set includes voice packages corresponding to the user's geographical location and incorporating the user's personalized voice characteristics; that is, a set of local dialect voices that integrate the user's personalized voice characteristics.

[0053] Optionally, in step S152, feature extraction of the speech audio data to obtain speech features includes performing the following steps:

[0054] Step S1521: Extract features from the speech audio data to obtain prosodic features;

[0055] Step S1522: Convert the speech audio data into text to obtain the target text;

[0056] Step S1523: Extract features from the target text to obtain accent features.

[0057] For example, speech features include prosodic features and personalized spoken language features.

[0058] The in-vehicle infotainment system extracts unique prosodic features from the acquired voice audio data, such as tone, speech rate, and intonation. These prosodic features are influenced by individual language habits, cultural background, and emotional state, resulting in unique characteristics in each person's voice. Some people may speak quickly and rhythmically, others slowly with noticeable pauses, and still others with a high-pitched and expressive voice. Personalized voice customization can be provided based on these individual prosodic features.

[0059] The methods for feature extraction from speech audio data depend on the specific application scenario and data format, and no specific restrictions are imposed here. For example, frequency domain analysis methods, such as short-time Fourier transform and autocorrelation function methods, or time domain analysis methods, such as pitch period detection and cross-correlation function methods, can be used to extract pitch features. Speech rate features can be extracted by calculating the short-time energy or short-time zero-crossing rate of the speech signal. Fundamental frequency extraction methods include fundamental frequency estimation algorithms based on autocorrelation functions and fundamental frequency estimation algorithms based on short-time Fourier transforms; and vocal tract parameter extraction methods, such as linear predictive analysis, can be used to extract pitch features.

[0060] Furthermore, speech audio data can be converted into text to obtain the target text. For example, a speech recognition engine can be used to convert speech data into text. Alternatively, a natural language processing model (such as a deep learning model, recurrent neural network, etc.) can be used to process the speech data and convert it into text. Or, open-source speech-to-text libraries can be used for speech conversion. Alternatively, a speech-to-text model can be built and trained independently according to specific needs to achieve more accurate conversion results. These methods allow for the selection of appropriate models and algorithms for speech-to-text conversion based on specific application scenarios and requirements.

[0061] After analysis and processing, feature extraction is performed on the target text to extract the user's personalized accent features. This allows the synthesized target speech package to include redundant words, interactivity, and modal particles, making the target speech package more in line with the user's speaking habits and sound more approachable and interesting. For example, the frequent use of conjunctions such as "then," "so," and "because" or modal particles such as "ah," "oh," "ba," and "ya" in the conversation can all be reflected in the target text and then extracted and integrated into the target speech package.

[0062] Optionally, in step S153, generating the target speech packet set based on speech features includes performing the following steps:

[0063] Step S1531: Obtain the speech feature set, wherein the speech feature set includes speech features corresponding to different regions;

[0064] Step S1532: Generate a target speech packet set based on speech features and speech feature set.

[0065] For example, the speech feature set is a set of dialects corresponding to different regions. That is, after obtaining the language packs of multiple dialects, the speech features extracted from the speech audio data are integrated into the speech packs of multiple dialects through speech synthesis technology. The resulting target speech pack set is the user-personalized speech pack set based on the geographical location.

[0066] Optionally, in step S15, the method further includes step S154: displaying prompt information in the graphical user interface of the vehicle system, wherein the prompt information is used to indicate that the target voice package set has been generated.

[0067] Once the target voice package set is generated, a notification message indicating completion will be displayed to the user through the vehicle's graphical user interface. The user can easily receive the notification message on the graphical user interface, allowing them to promptly check whether the generated target voice package set meets expectations, thus enabling continuous optimization of the vehicle's voice packages.

[0068] Optionally, in step S154, the method further includes at least one of the following:

[0069] Step S1541: In response to receiving the confirmation command, import the target voice packet set into the preset storage location;

[0070] Step S1542: In response to receiving the modification instruction, update the target speech packet set.

[0071] After the user sees the prompt message on the graphical interface of the vehicle system, they can view the generated target voice package set. When the user is satisfied with the generated target voice package set, they can issue a confirmation command by clicking the confirmation button on the graphical user interface. After receiving the confirmation command, the vehicle system will import the generated target voice package into the pre-set storage location so that when the vehicle location is changed in the future, the target voice package corresponding to the current vehicle location can be obtained in a timely manner.

[0072] When a user is not satisfied with the generated target voice package set, they can issue a modification command by clicking the modification button on the graphical user interface. Upon receiving the modification command, the vehicle system updates the target voice package set. This involves reacquiring the voice audio data, extracting the voice features from the voice audio data, and regenerating the target voice package set based on the voice features and the speech feature set.

[0073] Optionally, in step S15, the method further includes updating the target speech packet set based on the user's speech content.

[0074] For example, voice audio data is not static; users' speaking habits change with factors such as age and location, and the vehicle's driver may also change. Therefore, the in-vehicle infotainment system needs to continuously monitor the voice audio data in the vehicle and update the target voice package set in a timely manner. Once a user's speaking habits or accent change, the system can automatically adjust the target voice package set to ensure that voice interaction always maintains the best personalized experience.

[0075] Furthermore, in accordance with relevant requirements, the voice package switching method also includes user data privacy protection to ensure the security and privacy of user voice data, including measures such as data encryption and anonymization. For example, before acquiring the user's voice audio data, a pop-up window appears on the vehicle's graphical user interface, requesting the user's permission to collect audio. Simultaneously, user feedback on the voice interaction function is collected, including satisfaction with the voice package and the accuracy of voice package switching when the user's location changes, in order to continuously optimize the system's performance.

[0076] In a specific embodiment, Xiaoming is driving his car from Beijing to a tourist attraction in Zhangjiakou City, Hebei Province. When Xiaoming's vehicle enters Zhangjiakou City, the in-vehicle system senses the change in location and automatically switches to a Zhangjiakou dialect voice pack. It has already recorded and analyzed Xiaoming's daily conversations with the system, including his unique voice characteristics. Based on Xiaoming's voice features, the system automatically generates a personalized voice pack incorporating Zhangjiakou dialect, for example: "Hello Mr. Xiaoming! Welcome to Zhangjiakou City. Which attraction are you going to?" Xiaoming is surprised and happy to hear the system use his own voice combined with Zhangjiakou dialect for this voice prompt. This personalized and localized voice interaction experience meets user expectations and enhances user trust and satisfaction with the in-vehicle system.

[0077] Figure 2 This is a schematic diagram of a voice packet switching method according to an embodiment of the present invention, such as... Figure 2As shown, the vehicle's infotainment system first records daily conversations, acquiring voice audio data. After analyzing the voice audio data, it extracts voice features and then generates voice packages based on these features, which are then fed back to the user for settings (confirmation or modification). Finally, it obtains the vehicle's current location. When the vehicle's location changes and moves outside the preset area, it switches to the voice package corresponding to the current location; if the vehicle's location does not change or it has not moved outside the preset area, it retains the original voice package. In addition, the vehicle's infotainment system continuously monitors the voice audio data in the vehicle and updates the target voice package set in a timely manner.

[0078] Figure 3 This is a timing diagram of a voice packet switching method according to an embodiment of the present invention, such as... Figure 3 As shown, the vehicle and user begin their daily conversations. The vehicle records these conversations, storing the voice and audio data in the vehicle's infotainment system, and analyzes the content. Voice features are extracted from the infotainment system to generate voice packets. Once generated, a prompt message is sent to the user, who can then confirm or modify the generated voice packet and set it in the vehicle. The vehicle's location is determined using positioning technology. When the vehicle reaches a different location (i.e., a change in geographical location), the system switches to the voice packet corresponding to that location. Simultaneously, the voice packets are updated in real-time to continuously optimize system performance.

[0079] This feature integrates the user's personalized voice into the voice interaction and intelligently switches languages ​​based on the region, providing a more personalized and considerate voice interaction experience, further enhancing user satisfaction and overall user experience. By offering voice prompts and interactions that conform to local language habits, the system helps drivers obtain information more conveniently, improving driving safety and comfort. Especially in different regions or language environments, drivers can understand voice prompts more quickly, reducing distractions.

[0080] In this embodiment of the invention, by obtaining a first location of the vehicle, wherein the first location is within a first region; in response to determining that the vehicle has left the first region through the first location, obtaining a second location of the vehicle, wherein the second location is within a second region; determining a target voice package based on the second location, wherein the target voice package is generated based on the speech features corresponding to the second region; and switching the initial voice package of the vehicle system to the target voice package, wherein the initial voice package is generated based on the speech features corresponding to the first region, the technical solution achieves the technical effect of generating a unique voice package for the user based on the user's daily dialogue with the vehicle system, and automatically switching to the voice package of the corresponding region based on the vehicle's location. This provides personalized and intelligent voice interaction for the vehicle system, thereby solving the technical problems of lacking personalized customization options and insufficient intelligent perception of geographical location in related technologies.

[0081] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0082] This embodiment also provides a voice packet switching device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0083] Figure 4 This is a structural block diagram of a voice packet switching device according to an embodiment of the present invention, such as... Figure 4 As shown, a voice packet switching device 40 is used as an example. The device includes: a first acquisition module 42, used to acquire a first location of the vehicle, wherein the first location is within a first region; a second acquisition module 44, used to acquire a second location of the vehicle in response to determining that the vehicle has left the first region through the first location, wherein the second location is within the second region; a determination module 46, used to determine a target voice packet based on the second location, wherein the target voice packet is generated based on the speech features corresponding to the second region; and a switching module 48, used to switch the initial voice packet of the vehicle system to the target voice packet, wherein the initial voice packet is generated based on the speech features corresponding to the first region.

[0084] Optionally, the method further includes: a generation module for acquiring speech audio data, wherein the speech audio data is used to record the user's speech content; performing feature extraction on the speech audio data to obtain speech features; and generating a target speech package set based on the speech features, wherein the target speech package set includes speech packages corresponding to different regions.

[0085] Optionally, the generation module is also used to extract features from the speech audio data to obtain prosodic features; to convert the speech audio data into text to obtain target text; and to extract features from the target text to obtain accent features.

[0086] Optionally, the generation module is also used to obtain a speech feature set, wherein the speech feature set includes speech features corresponding to different regions; and to generate a target speech packet set based on the speech features and the speech feature set.

[0087] Optionally, the device further includes a display module for displaying prompt information in the graphical user interface of the vehicle system, wherein the prompt information indicates that the target voice package set has been generated.

[0088] Optionally, the device further includes: an instruction module, configured to import the target voice package set into a preset storage location in response to receiving a confirmation instruction; and to update the target voice package set in response to receiving a modification instruction.

[0089] Optionally, the device further includes an update module for updating the target speech packet set based on the user's speech content.

[0090] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0091] According to an embodiment of the present invention, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored computer program, wherein the device where the non-volatile storage medium is located executes a voice packet switching method by running the computer program.

[0092] Optionally, the device containing the non-volatile storage medium executes the following steps by running the computer program:

[0093] Step S12: Obtain the first location of the vehicle, wherein the first location is within a first region;

[0094] Step S14: In response to determining that the vehicle has left the first region through the first location, obtain the second location of the vehicle, wherein the second location is within the second region;

[0095] Step S16: Determine the target speech packet based on the second location, wherein the target speech packet is generated based on the speech features corresponding to the second region;

[0096] Step S18: Switch the initial voice packet of the vehicle system to the target voice packet, wherein the initial voice packet is generated based on the speech features corresponding to the first region.

[0097] Optionally, the device containing the non-volatile storage medium may also execute the following by running the computer program: acquiring voice audio data, wherein the voice audio data is used to record the user's speech; extracting features from the voice audio data to obtain voice features; and generating a target voice packet set based on the voice features, wherein the target voice packet set includes voice packets corresponding to different regions.

[0098] Optionally, the device containing the non-volatile storage medium may also perform the following actions by running the computer program: extracting features from the speech audio data to obtain prosodic features; converting the speech audio data into text to obtain target text; and extracting features from the target text to obtain accent features.

[0099] Optionally, the device containing the non-volatile storage medium may also execute the computer program to perform actions such as acquiring a speech feature set, wherein the speech feature set includes speech features corresponding to different regions; and generating a target speech packet set based on the speech features and the speech feature set.

[0100] Optionally, the device containing the non-volatile storage medium can also execute the computer program to display prompt information in the graphical user interface of the vehicle system, wherein the prompt information is used to indicate that the target voice package set has been generated.

[0101] Optionally, the device containing the non-volatile storage medium, by running the computer program, is also configured to perform at least one of the following: in response to receiving a confirmation instruction, importing the target voice packet set into a preset storage location; in response to receiving a modification instruction, updating the target voice packet set.

[0102] Optionally, the device containing the non-volatile storage medium can also execute the computer program to perform updates to the target speech packet set based on the user's speech content.

[0103] Embodiments of the present invention also provide a vehicle for performing the steps in any of the above method embodiments.

[0104] According to embodiments of the present invention, a computer program product is also provided, including a computer program, which is executed by a processor through the steps of any of the above method embodiments.

[0105] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0106] Step S12: Obtain the first location of the vehicle, wherein the first location is within a first region;

[0107] Step S14: In response to determining that the vehicle has left the first region through the first location, obtain the second location of the vehicle, wherein the second location is within the second region;

[0108] Step S16: Determine the target speech packet based on the second location, wherein the target speech packet is generated based on the speech features corresponding to the second region;

[0109] Step S18: Switch the initial voice packet of the vehicle system to the target voice packet, wherein the initial voice packet is generated based on the speech features corresponding to the first region.

[0110] Optionally, the computer program is configured to, when running on a computer or processor, also acquire speech audio data, wherein the speech audio data is used to record the user's speech content; extract features from the speech audio data to obtain speech features; and generate a target speech packet set based on the speech features, wherein the target speech packet set includes speech packets corresponding to different regions.

[0111] Optionally, the computer program is configured to, when running on a computer or processor, also perform feature extraction on speech audio data to obtain prosodic features; perform text conversion on speech audio data to obtain target text; and perform feature extraction on the target text to obtain accent features.

[0112] Optionally, the computer program is configured to, when running on a computer or processor, also acquire a speech feature set, wherein the speech feature set includes speech features corresponding to different regions; and generate a target speech packet set based on the speech features and the speech feature set.

[0113] Optionally, the computer program is configured to, when running on a computer or processor, also display prompt information in the graphical user interface of the vehicle system, wherein the prompt information indicates that the target voice package set has been generated.

[0114] Optionally, the computer program is configured to, when running on a computer or processor, also perform at least one of the following: importing the target voice packet set into a preset storage location in response to receiving a confirmation instruction; and updating the target voice packet set in response to receiving a modification instruction.

[0115] Optionally, the computer program is configured to also update the target speech package set based on the user's speech content when running on a computer or processor.

[0116] Optionally, in this embodiment, the computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0117] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0118] Optionally, in this embodiment, the processor in the above-described electronic device may be configured to run a computer program to perform the following steps:

[0119] Step S12: Obtain the first location of the vehicle, wherein the first location is within a first region;

[0120] Step S14: In response to determining that the vehicle has left the first region through the first location, obtain the second location of the vehicle, wherein the second location is within the second region;

[0121] Step S16: Determine the target speech packet based on the second location, wherein the target speech packet is generated based on the speech features corresponding to the second region;

[0122] Step S18: Switch the initial voice packet of the vehicle system to the target voice packet, wherein the initial voice packet is generated based on the speech features corresponding to the first region.

[0123] Optionally, the processor is configured to run a computer program that also performs the following actions: acquiring speech audio data, wherein the speech audio data is used to record the user's speech content; extracting features from the speech audio data to obtain speech features; and generating a target speech packet set based on the speech features, wherein the target speech packet set includes speech packets corresponding to different regions.

[0124] Optionally, the processor is configured to run a computer program that also performs the following actions: extracting features from the speech audio data to obtain prosodic features; converting the speech audio data into text to obtain target text; and extracting features from the target text to obtain accent features.

[0125] Optionally, the processor is configured to run a computer program that also performs the following: acquiring a speech feature set, wherein the speech feature set includes speech features corresponding to different regions; and generating a target speech packet set based on the speech features and the speech feature set.

[0126] Optionally, the processor is configured to run a computer program that also executes a prompt message displayed in the graphical user interface of the vehicle system, wherein the prompt message indicates that the target voice package set has been generated.

[0127] Optionally, the processor is configured to run a computer program that also performs at least one of the following: importing the target voice packet set into a preset storage location in response to receiving a confirmation instruction; and updating the target voice packet set in response to receiving a modification instruction.

[0128] Optionally, the processor is configured to run computer programs that also perform actions to update the target speech packet set based on the user's speech content.

[0129] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0130] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0131] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0132] In the several embodiments provided by this invention, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces; the indirect coupling or communication connection of units or modules can be electrical or other forms.

[0133] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0134] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0135] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0136] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A voice packet switching method, characterized in that, include: Obtain the first location of the vehicle, wherein the first location is within a first region; In response to determining that the vehicle has left the first area through the first location, a second location of the vehicle is obtained, wherein the second location is within the second area; The target speech packet is determined based on the second location, wherein the target speech packet is generated based on the speech features corresponding to the second region; The initial voice packet of the vehicle system is switched to the target voice packet, wherein the initial voice packet is generated based on the speech features corresponding to the first region; The method further includes: acquiring speech audio data, wherein the speech audio data is used to record the user's speech content; performing feature extraction on the speech audio data to obtain prosodic features, wherein the prosodic features include pitch features, speech rate features, and intonation features, wherein the pitch features are extracted based on frequency domain analysis methods or time domain analysis methods, wherein the frequency domain analysis methods include short-time Fourier transform and autocorrelation function methods, and the time domain analysis methods include pitch period detection and cross-correlation function methods; the speech rate features are extracted by calculating the short-time energy or short-time zero-crossing rate of the speech signal; and the intonation features are based on autocorrelation function... The fundamental frequency estimation algorithm, the fundamental frequency estimation algorithm based on short-time Fourier transform, or the linear predictive analysis method are used to extract features; the speech audio data is converted into text to obtain target text; features are extracted from the target text to obtain accent features, wherein the speech package synthesized using the accent features contains redundancy and modal particles; a speech feature set is obtained, wherein the speech feature set includes speech features corresponding to different regions; a target speech package set is generated based on the speech features and the speech feature set, wherein the speech features include prosodic features and accent features, and the target speech package set includes speech packages corresponding to different regions.

2. The method according to claim 1, characterized in that, The method further includes: A prompt message is displayed in the graphical user interface of the vehicle system, wherein the prompt message is used to indicate that the target voice package set has been generated.

3. The method according to claim 2, characterized in that, The method further includes at least one of the following: In response to receiving a confirmation command, the target voice packet set is imported into a preset storage location; In response to receiving a modification instruction, the target speech packet set is updated.

4. The method according to claim 1, characterized in that, The method further includes: The target speech packet set is updated based on the user's speech content.

5. A voice packet switching device, characterized in that, include: The first acquisition module is used to acquire the first location of the vehicle, wherein the first location is within a first region; The second acquisition module is configured to acquire a second location of the vehicle in response to determining that the vehicle has left the first region through the first location, wherein the second location is within the second region. The determining module is used to determine a target speech packet based on the second location, wherein the target speech packet is generated based on the speech features corresponding to the second region; A switching module is used to switch the initial voice package of the vehicle system to the target voice package, wherein the initial voice package is generated based on the speaking features corresponding to the first region; The device is further configured to acquire speech audio data, wherein the speech audio data is used to record the user's speech content; to extract features from the speech audio data to obtain prosodic features, wherein the prosodic features include pitch features, speech rate features, and intonation features, wherein the pitch features are extracted based on frequency domain analysis methods or time domain analysis methods, wherein the frequency domain analysis methods include short-time Fourier transform and autocorrelation function methods, and the time domain analysis methods include pitch period detection and cross-correlation function methods; the speech rate features are extracted by calculating the short-time energy or short-time zero-crossing rate of the speech signal, and the intonation features are based on autocorrelation function... The fundamental frequency estimation algorithm, the fundamental frequency estimation algorithm based on short-time Fourier transform, or the linear predictive analysis method are used to extract features; the speech audio data is converted into text to obtain target text; features are extracted from the target text to obtain accent features, wherein the speech package synthesized using the accent features contains redundancy and modal particles; a speech feature set is obtained, wherein the speech feature set includes speech features corresponding to different regions; a target speech package set is generated based on the speech features and the speech feature set, wherein the speech features include prosodic features and accent features, and the target speech package set includes speech packages corresponding to different regions.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the voice packet switching method as described in any one of claims 1 to 4 when run on a computer or processor.

7. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the voice packet switching method according to any one of claims 1 to 4.

8. A vehicle, characterized in that, The vehicle is used to perform the voice packet switching method as described in any one of claims 1 to 4.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the voice packet switching method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Voice playing method and device and electronic equipment

    CN111429882A

  • Voice broadcasting method and device, electronic equipment and storage medium

    CN111782175A