Control method and apparatus, and vehicle
By acquiring text content and cockpit location information, the sound-emitting device is controlled to emit stereo sound and sound effects at multiple sound image locations, solving the problem of single audio data generation in existing technologies and enhancing the user's auditory experience and sense of immersion.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- YINWANG INTELLIGENT TECHNOLOGIES CO LTD
- Filing Date
- 2025-01-27
- Publication Date
- 2026-07-30
AI Technical Summary
The existing voice broadcasting methods generate relatively simple audio data, resulting in a poor auditory experience for users in the cockpit.
By acquiring text content and the user's location information within the cockpit, the stereo and sound image paths are determined, and the sound-emitting device is controlled to emit corresponding audio data and sound effects at multiple sound image emission points, thereby enhancing the stereo effect of the audio content.
It enhances the user's auditory experience and immersion in the cabin, providing a richer auditory experience through the combination of stereo sound and sound effects.
Smart Images

Figure CN2025075514_30072026_PF_FP_ABST
Abstract
Description
Control methods, devices and vehicles Technical Field
[0001] This application relates to the field of electronic devices, and more specifically, to an audio data generation method, apparatus, and vehicle. Background Technology
[0002] Current voice narration (e.g., audiobooks) typically generates audio data using text-to-speech (TTS) technology. This method results in a relatively simple audiobook narration format, leading to a poor auditory experience for users. Summary of the Invention
[0003] This application provides a control method, device, and vehicle that helps enhance the user's auditory experience and sense of immersion within the cockpit.
[0004] In a first aspect, a control method is provided, the method comprising: acquiring first text content and acquiring user location information within the cockpit; determining a first stereo sound based on the first text content and the location information; and controlling a sound-generating device based on the first stereo sound.
[0005] Based on the above technical solution, stereo sound can be obtained by analyzing the text content and the user's position information within the cockpit. Thus, by controlling the sound output device in a stereo environment, the audio content heard by the user in the cockpit can have a stereo effect, helping to enhance the user's auditory experience and sense of immersion within the cockpit.
[0006] In some possible implementations, determining the first stereo sound based on the first text content and location information includes: determining the first stereo sound and the first sound image path based on the first text content and location information, wherein the first sound image path indicates the movement path of the sound image emitting position point of the first stereo sound; and controlling the sound-emitting device based on the first stereo sound includes: controlling the sound-emitting device based on the first stereo sound and the first sound image path.
[0007] In some possible implementations, the first stereo sound is audio emitted from multiple sound image emission points within the cockpit. Based on the first stereo sound, controlling the sound-emitting devices includes controlling multiple sound-emitting devices within the cockpit to emit first audio data corresponding to the first text content at the multiple sound image emission points.
[0008] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: determining a first sound effect based on the first text content; wherein, controlling the sound-generating device based on the first stereo sound includes: controlling the sound-generating device based on the first sound effect and the first stereo sound.
[0009] Based on the above technical solution, by analyzing the text content, sound effects that match the current text content can be generated. This allows users to hear first audio data with corresponding sound effects and stereo sound, helping to further enhance the user's auditory experience and sense of immersion.
[0010] In some possible implementations, the first stereo sound is audio emitted from multiple sound image emission points within the cockpit. Based on the first sound effect and the first stereo sound, the sound-emitting device is controlled, including: controlling multiple sound-emitting devices within the cockpit to emit the first audio data corresponding to the first text content at the multiple sound image emission points and controlling the first sound-emitting device within the cockpit to play the first sound effect.
[0011] In some possible implementations, the first sound effect includes any one of the following: ocean waves, falling leaves, war, fighting, pastoral, forest, desert, etc.
[0012] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: determining a first sound effect based on the first text content; wherein, determining a first stereo sound based on the first text content and location information includes: determining a first stereo sound based on the first text content, the first sound effect, and location information.
[0013] Based on the above technical solution, by analyzing the text content, sound effects that match the current text content can be generated. This allows users to hear sound effects with a stereo effect, or to hear sound effects with a stereo effect combined with primary audio data, further enhancing the user's auditory experience and sense of immersion.
[0014] In some possible implementations, determining the first stereo sound based on the first text content, the first sound effect, and location information includes: determining multiple sound image emission points based on the first text content and location information; and determining the first stereo sound based on the first sound effect and the multiple sound image emission points. In some possible implementations, the first stereo sound is audio emitted from multiple sound image emission points within the cockpit. Controlling the sound-emitting device based on the first stereo sound includes: controlling multiple sound-emitting devices within the cockpit to emit the first sound effect at the multiple sound image emission points and controlling the first sound-emitting device within the cockpit to play the first audio data corresponding to the first text content.
[0015] In some possible implementations, determining the first stereo sound based on the first text content, the first sound effect, and location information includes: determining multiple sound image emission point locations based on the first text content and location information; and determining the first stereo sound based on the first audio data corresponding to the first sound effect and the first text content, and the multiple sound image emission point locations.
[0016] In some possible implementations, the first stereo sound is audio emitted from multiple sound image emission points within the cockpit. Based on the first stereo sound, controlling the sound-emitting devices includes controlling multiple sound-emitting devices within the cockpit to emit a first sound effect and first audio data corresponding to the first text content at the multiple sound image emission points.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the first sound effect is determined based on the first text content, including: inputting the first text content into the first inference model to obtain the first sound effect.
[0018] Based on the above technical solution, by inputting the text content into the inference model, the corresponding sound effect can be obtained. In this way, obtaining the sound effect corresponding to the text content through model inference ensures that users can hear the corresponding sound effect while listening to the audio content of the first text content.
[0019] In some possible implementations, inputting the first text content into the first inference model to obtain the first sound effect includes: inputting the first text content into the first inference model to obtain the first sound effect, the start time of playing the first sound effect, and the end time of playing the first sound effect.
[0020] In some possible implementations, inputting the first text content into the first inference model to obtain the first sound effect includes: inputting the first text content into the first inference model to obtain the first sound effect, the start time of playing the first sound effect, and the playback duration of the first sound effect.
[0021] In some possible implementations, inputting the first text content into the first inference model to obtain the first sound effect includes: inputting the first text content into the first inference model to obtain the first sound effect and the target semantics. For example, when the first sound-emitting device plays the audio content corresponding to the target semantics, multiple sound-emitting devices are controlled to play the first sound effect at the multiple sound image sound-emitting position points.
[0022] In conjunction with the first aspect, in some implementations of the first aspect, determining the first sound effect based on the first text content includes: inputting the first text content into the first inference model to obtain the first sound effect type; wherein, the method further includes: determining the first sound effect from multiple sound effects stored in the sound effect library based on the first sound effect type.
[0023] Based on the above technical solution, the sound effect type corresponding to the text content is obtained through model reasoning, and the first sound effect can be determined from the sound effect library based on the sound effect type.
[0024] In some possible implementations, determining a first sound effect from multiple sound effects stored in a sound effect library based on the first sound effect type includes: determining a second sound effect from multiple sound effects stored in a sound effect library based on the first sound effect type, wherein the multiple sound effects include the second sound effect; when the playback duration of the second sound effect is greater than the playback duration of the first audio data corresponding to the first text content, compressing the playback duration of the second sound effect to obtain the first sound effect; or, extracting a portion of the second sound effect to obtain the first sound effect.
[0025] In conjunction with the first aspect, in some implementations of the first aspect, determining the first stereo sound based on the first text content and location information includes: inputting the first text content and location information into a second inference model to obtain the first stereo sound.
[0026] Based on the above technical solution, by inputting text content and location information into the inference model, stereo sound corresponding to the first text content can be obtained. In this way, obtaining stereo sound corresponding to the text content through model inference allows the audio content heard by users in the cockpit to have a stereo effect, helping to enhance the user's auditory experience and sense of immersion within the cockpit.
[0027] In some possible implementations, the first text content and location information are input into the second inference model to obtain the first stereo sound, including: inputting the first text content and location information into the second inference model to obtain the first stereo sound, the start time of playing the first stereo sound, and the end time of playing the first stereo sound.
[0028] In some possible implementations, the first text content and location information are input into the second inference model to obtain the first stereo sound, including: inputting the first text content and location information into the second inference model to obtain the first stereo sound, the start time of playing the first stereo sound, and the playback duration of the first stereo sound.
[0029] In some possible implementations, inputting the first text content and location information into a second inference model to obtain the first stereo sound includes: inputting the first text content and location information into the second inference model to obtain the first stereo sound and the target semantics. For example, when controlling the first sound-emitting device to play the audio content corresponding to the target semantics, multiple sound-emitting devices are controlled to emit the first stereo sound at the multiple sound image emission points.
[0030] In conjunction with the first aspect, in some implementations of the first aspect, determining the first stereo sound based on the first text content and location information includes: inputting the first text content and location information into a second inference model to obtain the first stereo sound and a first sound image path, wherein the first sound image path indicates the movement path of the sound image emitting position point of the first stereo sound.
[0031] In conjunction with the first aspect, in some implementations of the first aspect, determining the first stereo based on the first text content and location information includes: inputting the first text content and location information into a second inference model to obtain the first stereo type; and determining the first stereo from multiple stereos stored in a stereo library based on the first stereo type and location information.
[0032] Based on the above technical solution, by inputting the text content into the inference model, the stereo type corresponding to the first text content can be obtained. Therefore, the first stereo can be determined from multiple stereos stored in the sound effects library using the first stereo type and location information.
[0033] In some possible implementations, determining the first stereo from multiple stereos stored in a stereo library based on the first stereo type and location information includes: determining the second stereo from multiple stereos stored in a stereo library based on the first stereo type and location information, wherein the multiple stereos include the second stereo; when the playback duration of the second stereo is longer than the playback duration of the audio data corresponding to the first text content, compressing the playback duration of the second stereo to obtain the first stereo; or, extracting a portion of the second stereo to obtain the first stereo.
[0034] In conjunction with the first aspect, in some implementations of the first aspect, the first text content includes multiple sub-text contents, the first stereo includes multiple sound image path sequences, and the multiple sub-text contents correspond one-to-one with the multiple sound image path sequences.
[0035] In some possible implementations, the multiple sound image path sequences include a first sound image path sequence and a second sound image path sequence, and the sound image emission points corresponding to the first sound image path sequence and the second sound image path sequence are different.
[0036] In some possible implementations, determining the first stereo sound based on the first text content and location information includes: inputting the first text content into a second inference model to obtain multiple stereo sounds, wherein the multiple stereo sounds correspond one-to-one with the multiple sub-text contents.
[0037] In conjunction with the first aspect, in some implementations of the first aspect, obtaining the first text content includes: obtaining the first text content output by the content generation model; or, obtaining the first text content from an audio application; or, obtaining the first audio data received by a radio and determining the first text content based on the first audio data.
[0038] In some possible implementations, before obtaining the first text content output by the content generation model, the method further includes: obtaining a first input from the user, which instructs the generation of the first text content. For example, the first input is the user's voice input.
[0039] In conjunction with the first aspect, in some implementations of the first aspect, determining the first stereo sound based on the first text content and location information includes: determining the orientation of the target based on the first text content; and determining the first stereo sound based on the orientation and location information of the target.
[0040] Based on the above technical solution, the location of the target is determined by the first text content, and thus the first stereo sound can be determined based on the location of the target. In this way, the user can not only hear the audio content corresponding to the first text content, but also feel the audio location provided by the stereo sound, creating an immersive experience for the user.
[0041] In some possible implementations, determining the first stereo sound based on the first text content and location information includes: when there are multiple users in the cockpit, determining the center point of the area where the multiple users are located; and determining the first stereo sound based on the center point and the first text content.
[0042] In conjunction with the first aspect, in certain implementations of the first aspect, determining the first stereo sound based on the first text content and location information includes: determining playback parameters for a plurality of sound-generating devices based on the first text content and location information, wherein the playback parameters include at least one of time delay, volume, and phase. For example, the playback parameters of the plurality of sound-generating devices can enable the formation of the first stereo sound at the plurality of sound image emitting location points.
[0043] Secondly, this application provides a control device, which includes: an acquisition unit for acquiring first text content and acquiring the user's position information in the cockpit; a determination unit for determining a first stereo sound based on the first text content and the position information; and a control unit for controlling a sound-generating device based on the first stereo sound.
[0044] In conjunction with the second aspect, in some implementations of the second aspect, the determining unit is also used to determine the first sound effect based on the first text content; the control unit is used to control the sound-generating device based on the first sound effect and the first stereo sound.
[0045] In conjunction with the second aspect, in some implementations of the second aspect, a determining unit is used to: determine a first sound effect based on the first text content; and determine a first stereo sound based on the first text content, the first sound effect, and location information.
[0046] In conjunction with the second aspect, in some implementations of the second aspect, a unit is defined for: inputting the first text content into the first inference model to obtain the first sound effect.
[0047] In conjunction with the second aspect, in some implementations of the second aspect, a determining unit is used to: input the first text content into the first inference model to obtain the first sound effect type; and determine the first sound effect from multiple sound effects stored in the sound effect library based on the first sound effect type.
[0048] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes a sound effect processing unit and a determining unit, configured to determine a second sound effect from a plurality of sound effects stored in a sound effect library according to a first sound effect type, wherein the plurality of sound effects includes the second sound effect; the sound effect processing unit is configured to compress the playback duration of the second sound effect to obtain the first sound effect when the playback duration of the second sound effect is greater than the playback duration of the audio data corresponding to the first text content; or, to extract a portion of the second sound effect to obtain the first sound effect.
[0049] In conjunction with the second aspect, in some implementations of the second aspect, a unit is defined for: inputting the first text content and location information into the second inference model to obtain the first stereo sound.
[0050] In conjunction with the second aspect, in some implementations of the second aspect, a determining unit is used to: input the first text content and location information into the second inference model to obtain the first stereo sound and the first sound image path, wherein the first sound image path indicates the movement path of the sound image emitting position point of the first stereo sound.
[0051] In conjunction with the second aspect, in some implementations of the second aspect, a determining unit is used to: input the first text content and location information into the second inference model to obtain the first stereo type; and determine the first stereo from multiple stereos stored in the stereo library based on the first stereo type and location information.
[0052] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes a stereo processing unit and a determining unit, configured to determine a second stereo from a plurality of stereos stored in a stereo library based on the first stereo type and location information, wherein the plurality of stereos includes the second stereo; the stereo processing unit is configured to compress the playback duration of the second stereo to obtain the first stereo when the playback duration of the second stereo is greater than the playback duration of the audio data corresponding to the first text content; or, to extract a portion of the second stereo to obtain the first stereo.
[0053] In conjunction with the second aspect, in some implementations of the second aspect, the first text content includes multiple sub-text contents, the first stereo includes multiple sound image path sequences, and the multiple sub-text contents correspond one-to-one with the multiple sound image path sequences.
[0054] In conjunction with the second aspect, in some implementations of the second aspect, the acquisition unit is used to: acquire the first text content output by the content generation model; or, acquire the first text content from an audio application; or, acquire the first audio data received by a radio and determine the first text content based on the first audio data.
[0055] In conjunction with the second aspect, in some implementations of the second aspect, a determining unit is used to: determine the orientation of the target based on the first text content; and determine the first stereo based on the orientation and location information of the target.
[0056] In conjunction with the second aspect, in some implementations of the second aspect, a determining unit is used to: determine playback parameters of a plurality of sound-generating devices based on the first text content and location information, wherein the playback parameters include at least one of time delay, volume and phase.
[0057] Thirdly, this application provides a control device including a processor and a memory, wherein the memory is used to store instructions, and the processor executes the instructions stored in the memory to cause the device to perform any of the possible methods in the first aspect.
[0058] Fourthly, this application provides a control system including a computing platform and a sound-generating device, wherein the computing platform includes any of the possible devices in the second or third aspect.
[0059] Fifthly, this application provides a vehicle that includes any of the possible devices of the second or third aspect, or includes the system described in the fourth aspect.
[0060] In a sixth aspect, this application provides a computer program product comprising: computer program code, which, when executed on a computer, causes the computer to perform any of the possible methods described in the first aspect above.
[0061] It should be noted that the above-mentioned computer program code can be stored in whole or in part on the first storage medium, wherein the first storage medium can be packaged together with the processor or packaged separately from the processor. This application embodiment does not specifically limit this.
[0062] In a seventh aspect, this application provides a computer-readable medium storing program code that, when run on a computer, causes the computer to perform any of the possible methods described in the first aspect above.
[0063] Eighthly, this application provides a chip system including a processor for calling a computer program or computer instructions stored in a memory to cause the processor to perform any of the possible methods in the first aspect above.
[0064] In conjunction with the eighth aspect, in one possible implementation, the processor is coupled to the memory via an interface.
[0065] In conjunction with the eighth aspect, in one possible implementation, the chip system also includes a memory in which computer programs or computer instructions are stored.
[0066] Ninthly, this application provides a chip system including circuitry for performing any of the possible methods described in the first aspect above. Attached Figure Description
[0067] Figure 1 is a functional block diagram of the vehicle provided in an embodiment of this application.
[0068] Figure 2 is a schematic flowchart of the control method provided in an embodiment of this application.
[0069] Figure 3 is a schematic diagram of the vehicle cabin provided in the application embodiment.
[0070] Figure 4 is another schematic diagram of the vehicle cabin provided in the application embodiment.
[0071] Figure 5 is another schematic diagram of the vehicle cabin provided in the application embodiment.
[0072] Figure 6 is a schematic block diagram of the control device provided in an embodiment of this application. Detailed Implementation
[0073] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. "At least one" refers to one or more. For example, "at least one of A and B," similar to "A and / or B," describes the association relationship between related objects, indicating that three relationships can exist. For example, at least one of A and B can represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0074] The prefixes such as "first" and "second" used in this application embodiment are merely for distinguishing different descriptive objects and do not limit the position, order, priority, quantity, or content of the described objects. The use of ordinal numbers and other prefixes used to distinguish descriptive objects in this application embodiment does not constitute a limitation on the described objects. The description of the described objects is given in the claims or the context of the embodiments, and should not constitute unnecessary restrictions due to the use of such prefixes. Furthermore, in the description of this embodiment, unless otherwise stated, "multiple" means two or more.
[0075] Figure 1 is a functional block diagram of a vehicle 100 provided in an embodiment of this application. The vehicle 100 may include a sensing system 110, a computing platform 120, and a sound-emitting device 130. The sensing system 110 may include one or more sensors for sensing information about the environment surrounding the vehicle 100. For example, the sensing system 110 may include a positioning system, which may be a Global Positioning System (GPS), a BeiDou Navigation Satellite System, or another positioning system. As another example, the sensing system 110 may include one or more of an inertial measurement unit (IMU), an accelerometer, a lidar, millimeter-wave radar, ultrasonic radar, and a camera device. For example, the accelerometer may include a sensor for detecting acceleration signals from the air suspension system, or it may include a sensor for detecting ESC acceleration signals.
[0076] In this embodiment, the sensing system 110 may include a camera in the cockpit and / or a pressure sensor on the seat to obtain the user's location information in the cockpit.
[0077] Some or all of the functions of vehicle 100 can be controlled by computing platform 120. Computing platform 120 may include one or more processors, such as processors 121 to 12n (n being a positive integer). A processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU) (which can be understood as a type of microprocessor), or digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field-programmable gate array (FPGA). In reconfigurable hardware circuits, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, the processor can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), tensor processing unit (TPU), deep learning processing unit (DPU), etc. In addition, the computing platform 120 may also include a memory for storing instructions. Some or all of the processors 121 to 12n can call the instructions in the memory to implement the corresponding functions.
[0078] For example, the sound-generating device 130 may include multiple speakers in the cabin of the vehicle 100.
[0079] The above schematic diagram of vehicle 100 is merely illustrative. For example, vehicle 100 may also include more devices than those shown in Figure 1.
[0080] Figure 2 is a schematic flowchart of the control method 200 provided in an embodiment of this application. The method 200 can be executed by the vehicle 100; or by the computing platform 120; or by a processor, chip, or circuit in the computing platform 120; or by a system composed of the computing platform 120 and the sensing system 110. The method 200 includes:
[0081] S210, obtain the first text content and obtain the user's location information in the cockpit.
[0082] Optionally, obtaining the first text content includes: obtaining the first text content output by the content generation model; or obtaining the first text content from an audio application; or obtaining the first audio data received by a radio and determining the first text content based on the first audio data.
[0083] Optionally, before obtaining the first text content output by the content generation model, the method 200 further includes: obtaining a first input from the user, which instructs the generation of the first text content.
[0084] For example, the first input is the user's voice input. For instance, when the user's voice input "Xiao A, Xiao A, tell me a story" is detected, the story content can be obtained through a content generation model. The story content can be text content; where Xiao A is the name of the voice assistant in the cockpit.
[0085] For example, when a user's voice input "Xiao A, Xiao A, tell me a stand-up comedy" is detected, the stand-up comedy content can be obtained through a content generation model, and the stand-up comedy content can be text content.
[0086] Optionally, obtaining the user's location information within the cabin includes: obtaining the user's location information within the cabin based on images captured by cameras within the cabin; or obtaining the user's location information within the cabin based on data collected by pressure sensors on the seat.
[0087] S220, determine the first stereo sound based on the first text content and location information.
[0088] For example, the first stereo is a fixed-position (e.g., fixed rear, fixed front, etc.) or surround stereo.
[0089] Optionally, the first stereo sound is composed of sound source data emitted from multiple sound image emission points. The sound source data can be a first sound effect, or the sound source data can be first audio data corresponding to the first text content.
[0090] For example, the first sound effect includes any one of the following: ocean waves, falling leaves, war, fighting, pastoral scenery, forest, desert, or the sound of something breaking through the air.
[0091] Optionally, determining the first stereo sound based on the first text content and location information includes: determining multiple sound image emission point locations based on the first text content and location information; and determining the first stereo sound based on the first sound effect and the multiple sound image emission point locations.
[0092] For example, Figure 3 shows a schematic diagram of a vehicle cabin provided in an embodiment of this application. The vehicle 100 can determine that only the driver's area is currently occupied by a user based on images captured by a camera. Before playing the audio data 1 corresponding to text content 1 (e.g., Tom in the jungle hears Lily shouting behind him), the vehicle 100 can determine the stereo sound as fixed rear stereo with stereo sound image emitting points 1-3 based on the text content 1 and the user's position information within the cabin. For example, the stereo sound image emitting points 1-3 are located behind the driver's area. The vehicle can determine stereo 1 based on sound effect 1 (e.g., falling leaves sound effect) and the stereo sound image emitting points 1-3. For example, the vehicle can control multiple sound devices in the cabin to emit sound effect 1 at the stereo sound image emitting points 1-3 while controlling the sound device 1 in the cabin to play the audio data 1.
[0093] Optionally, the plurality of sound-generating devices may include sound-generating device 1, or the plurality of sound-generating devices may not include sound-generating device 1.
[0094] Optionally, determining the first stereo based on the first text content and location information includes: determining playback parameters of a plurality of sound-generating devices based on the first text content and location information, wherein the playback parameters include at least one of time delay, volume, and phase.
[0095] For example, as shown in Figure 3, the playback parameters of the multiple sound-generating devices can make stereo sound 1 be formed at sound image emission points 1-3.
[0096] For example, vehicle 100 can control multiple sound-emitting devices in the cabin to simultaneously emit the sound effect of falling leaves and the first audio data at multiple sound image emission points.
[0097] For example, Figure 4 shows another schematic diagram of a vehicle cabin provided in an embodiment of this application. The vehicle can determine the presence of a user in the driver's seat and passenger seat areas based on images captured by a camera. Before playing the audio data 2 corresponding to the text content 2 (e.g., Tom suddenly hears the sound of a violin playing around him), the vehicle can determine the sound image emission points 5-9 based on the text content 2 and the user's position information in the cabin. For example, the sound image emission points 5-9 surround the driver's seat and passenger seat areas. The vehicle can determine stereo 2 based on the sound effect 2 (e.g., a violin sound effect) and the sound image emission points 5-9. For example, the vehicle can control multiple sound devices in the cabin to emit sound effect 2 at the sound image emission points 5-9 while controlling the sound device 2 in the cabin to play the audio data 2.
[0098] Optionally, the plurality of sound-generating devices may include sound-generating device 2, or the plurality of sound-generating devices may not include sound-generating device 2.
[0099] Optionally, method 200 further includes: determining a first sound effect based on the first text content; wherein determining a first stereo sound based on the first text content and location information includes: determining a first stereo sound based on the first text content, the first sound effect, and location information.
[0100] Taking the scenario shown in Figure 3 as an example, the vehicle can determine that only the driver's area is currently occupied by a user based on images captured by the camera. Before playing the audio data 1 corresponding to the text content 1 (e.g., Tom in the jungle hears Lily shouting behind him), the vehicle can determine the stereo sound as a fixed rear stereo sound with stereo sound image emitting points 1-3 based on the text content 1 and the user's position information in the cabin. For example, the stereo sound image emitting points 1-3 are located behind the driver's area. The vehicle can determine stereo sound 1 based on sound effect 1 (e.g., falling leaves sound effect), text content 1, and stereo sound image emitting points 1-3. For example, the vehicle can control multiple sound-emitting devices in the cabin to emit sound effect 1 and audio data 1 at stereo sound image emitting points 1-3.
[0101] Taking the scenario shown in Figure 4 as an example, the vehicle can determine stereo 2 based on sound effect 2 (e.g., violin sound effect), text content 2, and sound image emission point 5-9. For example, the vehicle can control multiple sound-emitting devices in the cabin to emit sound effect 2 and the audio data 2 at the sound image emission point 5-9.
[0102] Optionally, the method 200 includes: determining multiple sound image emission location points based on the first text content and location information; and determining a first stereo sound based on the multiple sound image emission location points and the first audio data corresponding to the first text content.
[0103] For example, as shown in Figure 3, the vehicle can determine that only the driver's area is currently occupied by a user through images captured by a camera. After determining the sound image emission points 1-3 based on the first text content and the user's location information in the cabin, the vehicle can control multiple sound-emitting devices in the cabin to simultaneously emit audio data 1 corresponding to the first text content at the sound image emission points 1-3. Optionally, the vehicle can control the sound-emitting device 3 in the cabin to emit a first sound effect.
[0104] Optionally, the plurality of sound-generating devices may include sound-generating device 3, or the plurality of sound-generating devices may not include sound-generating device 3.
[0105] Optionally, the first stereo can be a progressive, progressive, left-to-right, or right-to-left stereo.
[0106] Optionally, determining the first stereo sound based on the first text content and location information includes: determining the first stereo sound and a first sound image path based on the first text content and location information, wherein the first sound image path indicates the movement path of multiple sound image emitting position points of the first stereo sound; controlling the sound-emitting device includes: controlling the sound-emitting device based on the first stereo sound and the movement path.
[0107] For example, consider the scenario shown in Figure 4. During the first time period while the sound-emitting device 2 in the cabin is playing audio data 2, the vehicle 100 can control multiple sound-emitting devices in the cabin to emit sound effect 2 at sound image emitting positions 5-9. Based on this movement path, the vehicle 100 can control the sound image emitting positions 5-9 to move closer to the center points of the driver's and passenger's areas, respectively. Therefore, during the second time period while the sound-emitting device 2 in the cabin is playing audio data 2, the vehicle 100 can control multiple sound-emitting devices in the cabin to emit sound effect 2 at the multiple sound image emitting positions after the position change. This allows the user to hear the audio data 2 in the cabin and also experience the sound effect moving towards their area, thus enhancing the user's auditory experience.
[0108] For example, the first time period and the second time period are two time periods in the process of audio data 2. These two time periods may be adjacent or not.
[0109] For example, Figure 5 shows another schematic diagram of a vehicle cabin provided in an embodiment of this application.
[0110] As shown in Figure 5, vehicle 100 can determine the presence of users in the driver's seat and passenger seat areas based on images captured by cameras. Before playing the audio data 3 corresponding to the text content 3 (e.g., Tom suddenly hears footsteps approaching), vehicle 100 can determine the sound image emitting point 10-12 and its movement path based on the text content 3 and the user's position information within the cabin. For example, the sound image emitting point 10-12 is located behind the center point of the driver's seat and passenger seat areas. Vehicle 100 can determine stereo sound 3 based on the sound effect 3 (e.g., footstep sound effect) and the sound image emitting point 10-12.
[0111] For example, vehicle 100 can be in T a The system controls multiple sound-emitting devices in the cockpit to simultaneously emit footstep sound effects at the sound image emission point 10-12, and controls the sound-emitting device 4 in the cockpit to play the audio data 3 corresponding to the text content 3.
[0112] For example, vehicle 100 can be in T b The system continuously controls multiple sound-emitting devices in the cockpit to simultaneously emit footstep sound effects at sound image emission points 13-15, and controls sound-emitting device 4 in the cockpit to play the audio data 3. Among them, sound image emission points 13-15 are closer to the center point of the driver's area and passenger's area than sound image emission points 10-12.
[0113] For example, vehicle 100 can be in T c The system continuously controls multiple sound-emitting devices in the cockpit to simultaneously emit footstep sound effects at sound image emission points 16-18 and controls sound-emitting device 4 in the cockpit to play the audio data 3. Among them, sound image emission points 16-18 are closer to the center point of the driver's area and passenger's area than sound image emission points 13-15.
[0114] As can be seen from Figure 6, sound image emitting point 10, sound image emitting point 13 and sound image emitting point 16 move along movement path 1, sound image emitting point 11, sound image emitting point 14 and sound image emitting point 17 move along movement path 2, and sound image emitting point 12, sound image emitting point 15 and sound image emitting point 18 move along movement path 3.
[0115] Optionally, controlling the sound-emitting device 4 in the cockpit to play the audio data 3 corresponding to the text content 3 includes: controlling the sound-emitting device 4 in the cockpit to play the audio data 3 within a first preset duration, wherein, during T a Time, T b Time and T c The time is within the first preset duration.
[0116] Optionally, the method 200 further includes: determining the moving speed of the sound image emitting position point on the moving path based on the first text content.
[0117] For example, if the text content 3 is "Tom suddenly heard footsteps approaching him", the vehicle 100 can determine the moving speed of the sound image emitting point on the moving path 1-3 as speed 1 based on the text content 3 and the location information; if the text content 4 is "Tom suddenly heard footsteps approaching him", the vehicle 100 can determine the moving speed of the sound image emitting point on the moving path 1-3 as speed 2 based on the text content 4 and the location information, wherein the moving speed 2 is greater than the moving speed 1.
[0118] Optionally, determining the first stereo sound based on the first text content and location information includes: inputting the first text content and location information into a second inference model to obtain the first stereo sound.
[0119] Optionally, determining the first stereo sound based on the first text content and location information includes: inputting the first text content and location information into a second inference model to obtain the first stereo sound and a first sound image path, wherein the first sound image path indicates the movement path of the sound image emission point of the first stereo sound.
[0120] For example, the first stereo sound can be a stereo sound emitted at sound image emitting position points 10-12 as shown in FIG. 5, and the first sound image path can include the movement path 1-3 as shown in FIG. 5.
[0121] Optionally, inputting the first text content and location information into the second inference model to obtain the first stereo sound includes: inputting the first text content and location information into the second inference model to obtain the first stereo sound, the start time of playing the first stereo sound, and the end time of playing the first stereo sound.
[0122] For example, Table 1 shows the input and output of the inference model.
[0123] Table 1
[0124] For example, consider the scenario shown in Figure 3. Vehicle 100 can determine that only the driver's area is currently occupied by a user through images captured by a camera. The vehicle can input the text content 1 (e.g., Tom in the jungle hears Lily shouting behind him) and the user's position information in the cabin into a second inference model to obtain the fixed rear stereo sound, the start time of playing the fixed rear stereo sound, and the end time of playing the fixed rear stereo sound. The sound image emission points of the fixed rear stereo sound include sound image emission point 1-3, and the sound source of the fixed rear stereo sound is the audio data 1 corresponding to the text content 1. For example, sound image emission point 1-3 is located behind the driver's area. Vehicle 100 can control multiple sound-emitting devices in the cabin to simultaneously emit audio data 1 at sound image emission point 1-3 at time T1, and control multiple sound-emitting devices in the cabin to stop emitting audio data 1 at sound image emission point 1-3 at time T2. In this example, the time period (T1, T2) is the time period for emitting audio data 1.
[0125] For example, consider the scenario shown in Figure 3. Vehicle 100 can input the text content 1 and the user's position information within the cabin into a second inference model to obtain the fixed rear stereo sound, the start time of playing the fixed rear stereo sound, and the end time of playing the fixed rear stereo sound. The sound image emitting point of the fixed rear stereo sound is sound image emitting point 1-3, and the sound source of the fixed rear stereo sound is sound effect 1. For example, sound image emitting point 1-3 is located behind the driver's area. The vehicle can control multiple sound-emitting devices in the cabin to simultaneously emit sound effect 1 at sound image emitting point 1-3 at time T1, and control multiple sound-emitting devices in the cabin to stop emitting sound effect 1 at sound image emitting point 1-3 at time T2. For example, the start time of playing the audio data 1 corresponding to the text content 1 is time T0, and the end time of playing the audio data 1 is time T3. In this example, the time period (T1, T2) is the time period during which stereo 1 is emitted, the time period (T0, T3) is the time period during which audio data 1 is emitted, and the time period (T1, T2) can be located within the time period (T0, T3).
[0126] Optionally, the T1 time can correspond to the playback time of the target semantic. For example, the target semantic can be "hear".
[0127] Optionally, inputting the first text content and location information into the second inference model to obtain the first stereo sound includes: inputting the first text content and location information into the second inference model to obtain the first stereo sound, the start time of playing the first stereo sound, and the playback duration of the first stereo sound.
[0128] For example, Table 2 shows the input and output of the inference model.
[0129] Table 2
[0130] For example, vehicle 100 can input the text content 1 and the user's position information in the cabin into the second inference model to obtain the fixed rear stereo sound, the start time of playing the fixed rear stereo sound, and the playback duration of the fixed rear stereo sound. The sound image emitting point of the fixed rear stereo sound is sound image emitting point 1-3, and the sound source of the fixed rear stereo sound is sound effect 1. For example, sound image emitting point 1-3 is located behind the driver's area. The vehicle can control multiple sound-emitting devices in the cabin to simultaneously emit sound effect 1 at sound image emitting point 1-3 at time T1, and control multiple sound-emitting devices in the cabin to stop emitting sound effect 1 at sound image emitting point 1-3 at time T1+ΔT. For example, the playback start time of the audio data 1 corresponding to the text content 1 is time T0, and the playback end time of the audio data 1 is time T3. The stereo 1 can start playing from time T1 and the playback duration can be △T. The time period (T1, T1+△T) can be located within the time period (T0, T3).
[0131] Optionally, inputting the first text content and location information into the second inference model to obtain the first stereo sound includes: inputting the first text content and location information into the second inference model to obtain the first stereo sound, the start time of playing the first stereo sound, the playback duration of the first stereo sound, and the number of loops.
[0132] For example, Table 3 shows the input and output of the inference model.
[0133] Table 3
[0134] For example, the playback start time of the audio data 1 corresponding to the text content 1 is time T0, and the playback end time of the audio data 1 is time T3. The stereo 1 can start playing from time T1 and loop N times. The playback duration of the stereo 1 can be N×△T. The time period (T1, T1+N×△T) can be located within the time period (T0, T3).
[0135] Optionally, inputting the first text content and location information into a second inference model to obtain the first stereo sound includes: inputting the first text content and location information into the second inference model to obtain the first stereo sound and the target semantics. For example, when the audio content corresponding to the target semantics is played by the sound-emitting device in the cockpit, multiple sound-emitting devices can be controlled to play the fixed rear stereo sound.
[0136] For example, Table 4 shows the input and output of the inference model.
[0137] Table 4
[0138] For example, the playback start time of the audio data 1 corresponding to the text content 1 is time T0. When the audio data 1 plays "dart", multiple sound-generating devices can be controlled to play fixed rear stereo sound at sound image emission points 1-3.
[0139] Optionally, determining the first stereo based on the first text content and location information includes: inputting the first text content and location information into a second inference model to obtain the first stereo type; and determining the first stereo from multiple stereos stored in a sound effects library based on the first stereo type and location information.
[0140] For example, the first stereo type includes stereo types such as indicating a fixed direction (e.g., fixed rear, fixed front, etc.), surround, approach, away, left to right, right to left, etc.
[0141] For example, taking method 200 as executed by vehicle 100, vehicle 100 may include a stereo library, which includes stereo sounds such as fixed orientation (e.g., fixed rear, fixed front, etc.), surround, approach, distance, left to right, and right to left.
[0142] For example, Table 5 shows the correspondence between stereo and stereo playback duration in the sound effects library.
[0143] Table 5
[0144] For example, vehicle 100 can input the first text content and the user's location information in the cabin into the second inference model to obtain the type of the first stereo sound, which indicates the type of fixed rear stereo sound. If the location information indicates that only the driver's area has a user, vehicle 100 can determine the fixed driver's area rear stereo sound from the correspondence shown in Table 5 based on the type and location information of the first stereo sound. Optionally, the sound image emission point of the fixed rear stereo sound in the stereo sound library can be sound image emission point 1-3, the playback duration of the fixed rear stereo sound can be ΔT1, and the sound source of the fixed rear stereo sound can be the first sound effect.
[0145] Optionally, determining a first stereo from multiple stereos stored in a stereo library based on the first stereo type and location information includes: determining a second stereo from multiple stereos stored in a stereo library based on the first stereo type and location information, wherein the multiple stereos include the second stereo; compressing the playback duration of the second stereo when the playback duration of the second stereo is longer than the playback duration of the audio data corresponding to the first text content to obtain the first stereo; or, extracting a portion of the second stereo to obtain the first stereo.
[0146] For example, taking the first stereo sound as surround sound, the playback duration of the surround sound in the driver's area stored in the stereo library is ΔT2, and the playback duration of the first audio data corresponding to the first text content is ΔT7. If ΔT2 is greater than ΔT7, the playback duration of the surround sound in the driver's area can be compressed, so that the duration of the surround sound in the driver's area is compressed to ΔT7. At this time, the first sound-emitting device can be controlled to play the first audio data, and multiple sound-emitting devices can be controlled to emit the compressed surround sound at the multiple sound image emission points.
[0147] For example, taking the first stereo sound as surround sound around the driver's area, the playback duration of the surround sound around the driver's area stored in the stereo sound library is △T2, and the playback duration of the first audio data corresponding to the first text content is △T7. If △T2 is greater than △T7, a portion of the surround sound around the driver's area can be extracted, and the playback duration of this portion of the stereo sound is △T7. This allows control of the first sound-emitting device to play the first audio data and control of multiple sound-emitting devices to emit the extracted surround sound at the multiple sound image emission points.
[0148] In this embodiment, by inputting text content into the inference model, the stereo type corresponding to the first text content can be obtained. Thus, the first stereo can be determined from multiple stereos stored in the sound effects library.
[0149] Optionally, the first text content includes multiple sub-text contents, and the first stereo includes multiple sound image path sequences, with each sub-text content corresponding to one of the multiple sound image path sequences.
[0150] For example, the method 200 includes: inputting the first text content into an inference model to obtain N stereo sounds, wherein the N stereo sounds correspond one-to-one with the N sub-text contents in the first text content.
[0151] For example, Table 6 shows the correspondence between N sub-text contents and N stereo sounds in the first text content.
[0152] Table 6
[0153] Optionally, determining the first stereo sound based on the first text content includes: inputting the first text content into a second inference model to obtain multiple stereo sounds, wherein the multiple stereo sounds correspond one-to-one with the multiple sub-text contents.
[0154] S230 controls the sound-generating device according to the first stereo sound.
[0155] Optionally, the first stereo sound is audio emitted from multiple sound image emission points within the cockpit. Based on the first stereo sound, controlling the sound-emitting devices includes controlling multiple sound-emitting devices within the cockpit to emit the first audio data corresponding to the first text content at the multiple sound image emission points.
[0156] For example, as shown in Figure 3, the vehicle 100 can control multiple sound-emitting devices in the cabin to simultaneously emit audio data 1 at the sound image emission points 1-3.
[0157] Optionally, the first stereo sound is audio emitted from multiple sound image emission points within the cockpit. Based on the first stereo sound, controlling the sound-emitting devices includes controlling multiple sound-emitting devices within the cockpit to emit a first sound effect and first audio data at the multiple sound image emission points.
[0158] For example, as shown in Figure 3, the vehicle 100 can control multiple sound-emitting devices in the cabin to simultaneously emit audio data 1 and sound effect 1 at the sound image emission points 1-3.
[0159] Optionally, method 200 further includes: determining a first sound effect based on the first text content; wherein controlling the sound-generating device based on the first stereo sound includes: controlling the sound-generating device based on the first sound effect and the first stereo sound.
[0160] Optionally, the first stereo sound is audio played at multiple sound image emitting points within the cockpit. Based on the first sound effect and the first stereo sound, controlling the sound-emitting device includes: controlling multiple sound-emitting devices within the cockpit to play the first sound effect at the multiple sound image emitting points and controlling the first sound-emitting device within the cockpit to play the first audio data corresponding to the first text content.
[0161] For example, as shown in Figure 3, the vehicle 100 can control multiple sound-emitting devices in the cabin to emit a falling leaf sound effect at the sound image emission point 1-3 and control the sound-emitting device 1 in the cabin to play the audio data 1 corresponding to the text content 1.
[0162] For example, as shown in Figure 4, the vehicle 100 can control multiple sound-emitting devices in the cabin to emit violin sound effects at the sound image emission points 5-9 and control the sound-emitting device 2 in the cabin to play the audio data 2 corresponding to the text content 2.
[0163] Optionally, determining the first sound effect based on the first text content includes: inputting the first text content into the first inference model to obtain the first sound effect.
[0164] The first reasoning model and the second reasoning model above can be the same reasoning model, or they can be different reasoning models.
[0165] In this embodiment, by inputting the text content into the inference model, the corresponding sound effect can be obtained. Thus, by obtaining the sound effect corresponding to the text content through model inference, it can be ensured that the user can hear the corresponding sound effect while listening to the audio content of the first text content.
[0166] Optionally, inputting the first text content into the first inference model to obtain the first sound effect includes: inputting the first text content into the first inference model to obtain the first sound effect, the start time of playing the first sound effect, and the end time of playing the first sound effect.
[0167] For example, Table 7 shows the input and output of the inference model.
[0168] Table 7
[0169] For example, the playback start time of the audio data 1 corresponding to the text content 1 "Suddenly a sound occurred, and darts were thrown simultaneously from all directions" is time T0, and the playback end time of the audio data 1 is time T3. The time period (T4, T5) can be located within the time period (T0, T3). For example, the vehicle 100 can control multiple sound-emitting devices to emit audio data 1 at sound image emission points 1-3 within the time period (T0, T3), and control a sound-emitting device in the cabin to play the falling leaf sound effect within the time period (T4, T5).
[0170] Optionally, inputting the first text content into the first inference model to obtain the first sound effect includes: inputting the first text content into the first inference model to obtain the first sound effect, the start time of playing the first sound effect, and the playback duration of the first sound effect.
[0171] For example, Table 8 shows the input and output of the inference model.
[0172] Table 8
[0173] Optionally, inputting the first text content into the first inference model to obtain the first sound effect includes: inputting the first text content into the first inference model to obtain the first sound effect, the start time of playing the first sound effect, the playback duration of the first sound effect, and the number of loops.
[0174] For example, Table 9 shows the input and output of the inference model.
[0175] Table 9
[0176] For example, the playback start time of the audio data 1 corresponding to the text content 1 is time T0, and the playback end time of the audio data 1 is time T3. The falling leaf sound effect can start playing from time T4 and loop M times. The playback duration of the falling leaf sound effect can be M×△T.
[0177] Optionally, inputting the first text content into the first inference model to obtain the first sound effect includes: inputting the first text content into the first inference model to obtain the first sound effect and the target semantics.
[0178] For example, when a sound-emitting device plays audio content corresponding to the target semantics, vehicle 100 can control a sound-emitting device to play the falling leaf sound effect. For example, time T4 can correspond to the playback time of the target semantics. For example, the target semantics can be "jungle".
[0179] Optionally, determining the first sound effect based on the first text content includes: inputting the first text content into a first inference model to obtain a first sound effect type; wherein, the method further includes: determining the first sound effect from multiple sound effects stored in a sound effect library based on the first sound effect type.
[0180] Based on the above technical solution, the sound effect type corresponding to the text content is obtained through model reasoning, and the first sound effect can be determined from the sound effect library based on the sound effect type.
[0181] Optionally, determining a first sound effect from multiple sound effects stored in a sound effect library according to the first sound effect type includes: determining a second sound effect from multiple sound effects stored in a sound effect library according to the first sound effect type, wherein the multiple sound effects include the second sound effect; when the playback duration of the second sound effect is greater than the playback duration of the first audio data corresponding to the first text content, compressing the playback duration of the second sound effect to obtain the first sound effect; or, extracting a portion of the second sound effect to obtain the first sound effect.
[0182] For example, taking method 200 executed by vehicle 100 as an example, vehicle 100 may include a sound effects library, which includes sound effects such as ocean waves, falling leaves, war, fighting, pastoral scenes, forests, deserts, and the sound of air being cut.
[0183] For example, Table 10 shows the correspondence between stereo and stereo playback duration in the sound effects library.
[0184] Table 10
[0185] For example, taking the sound effect determined by vehicle 100 from the sound effect library based on the first text content as a "sound of air breaking" as an example, the playback duration of the "sound of air breaking" stored in the sound effect library is △T15, and the playback duration of the first audio data corresponding to the first text content is △T7. If △T15 is greater than △T7, the playback duration of the "sound of air breaking" can be compressed, so that the playback duration of the "sound of air breaking" is compressed to △T7. At this time, multiple sound-emitting devices can be controlled to emit the first audio data and the compressed "sound of air breaking" at multiple sound-emitting positions.
[0186] For example, taking the sound effect determined by vehicle 100 from the sound effect library based on the first text content as a "sound breaking through the air," the playback duration of the surround sound stored in the sound effect library is △T15, and the playback duration of the first audio data corresponding to the first text content is △T7. If △T15 is greater than △T7, a portion of the first sound effect can be extracted, and the playback duration of this portion of the sound effect is less than or equal to △T7. This allows multiple sound-emitting devices to emit the first audio data and the extracted "sound breaking through the air" at multiple sound-emitting locations.
[0187] Optionally, determining the first stereo sound based on the first text content includes: determining the location of the target based on the first text content; and determining the first stereo sound based on the target's location and position information.
[0188] For example, taking the first text content as text content 1, vehicle 100 can determine that the target character Lily is located behind the protagonist Tom based on text content 1. Vehicle 100 can take the driver's area as the location of the protagonist Tom, and thus determine the sound image emission point 1-3 based on the position of the target character Lily relative to the protagonist Tom. Corresponding audio can then be emitted at sound image emission point 1-3, allowing the user in the driver's area to be immersed in the story narrated by text content 1, enhancing the user's auditory experience and sense of immersion.
[0189] In this embodiment, the location of the target is determined by the first text content, thereby enabling the determination of the first stereo sound based on the target's location. This allows the user to not only hear the audio content corresponding to the first text content but also experience the audio location provided by the stereo sound, creating an immersive experience.
[0190] Figure 6 shows a schematic block diagram of a control device 600 provided in an embodiment of this application. The device 600 includes: an acquisition unit 610, configured to acquire first text content and acquire the user's position information in the cockpit; a determination unit 620, configured to determine a first stereo sound based on the first text content and the position information; and a control unit 630, configured to control a sound-generating device based on the first stereo sound.
[0191] Optionally, the determining unit 620 is further configured to determine a first sound effect based on the first text content; the control unit 630 is configured to control the sound-generating device based on the first sound effect and the first stereo sound.
[0192] Optionally, the determining unit 620 is configured to: determine a first sound effect based on the first text content; and determine a first stereo sound based on the first text content, the first sound effect, and location information.
[0193] Optionally, the determining unit 620 is used to: input the first text content into the first inference model to obtain the first sound effect.
[0194] Optionally, the determining unit 620 is used to: input the first text content into the first inference model to obtain the first sound effect type; and determine the first sound effect from multiple sound effects stored in the sound effect library according to the first sound effect type.
[0195] Optionally, the device 600 further includes a sound effect processing unit and a determining unit 620, which is used to determine a second sound effect from a plurality of sound effects stored in a sound effect library according to a first sound effect type, wherein the plurality of sound effects includes the second sound effect; the sound effect processing unit is used to compress the playback duration of the second sound effect to obtain the first sound effect when the playback duration of the second sound effect is greater than the playback duration of the audio data corresponding to the first text content; or, to extract a portion of the second sound effect to obtain the first sound effect.
[0196] Optionally, the determining unit 620 is used to: input the first text content and location information into the second inference model to obtain the first stereo sound.
[0197] Optionally, the determining unit 620 is used to: input the first text content and location information into the second inference model to obtain the first stereo sound and the first sound image path, wherein the first sound image path indicates the movement path of the sound image emitting position point of the first stereo sound.
[0198] Optionally, the determining unit 620 is used to: input the first text content into the second inference model to obtain the first stereo type; and determine the first stereo from multiple stereos stored in the stereo library based on the first stereo type and location information.
[0199] Optionally, the device 600 further includes a stereo processing unit and a determining unit 620, configured to determine a second stereo from a plurality of stereos stored in a stereo library based on the first stereo type and location information, wherein the plurality of stereos includes the second stereo; the stereo processing unit is configured to compress the playback duration of the second stereo to obtain the first stereo when the playback duration of the second stereo is greater than the playback duration of the audio data corresponding to the first text content; or, to extract a portion of the second stereo to obtain the first stereo.
[0200] Optionally, the first text content includes multiple sub-text contents, and the first stereo includes multiple sound image path sequences, with each sub-text content corresponding to one of the multiple sound image path sequences.
[0201] Optionally, the acquisition unit 610 is configured to: acquire the first text content output by the content generation model; or, acquire the first text content from an audio application; or, acquire the first audio data received by the radio and determine the first text content based on the first audio data.
[0202] Optionally, the determining unit 620 is used to: determine the orientation of the target based on the first text content; and determine the first stereo based on the orientation and location information of the target.
[0203] Optionally, the determining unit 620 is configured to: determine playback parameters of a plurality of sound-generating devices based on the first text content, the playback parameters including at least one of time delay, volume and phase.
[0204] It should be understood that the division of units in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units in the device can be implemented by a processor calling software; for example, the device includes a processor connected to memory, which stores instructions. The processor calls the instructions stored in memory to implement any of the above methods or to implement the functions of each unit in the device. The processor can be, for example, a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented as hardware circuits. The functions of some or all units can be implemented through the design of the hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all units are implemented through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby implementing the functions of some or all units. All units of the above devices can be implemented entirely through processor calling software, or entirely through hardware circuits, or partially through processor calling software with the remaining parts implemented through hardware circuits.
[0205] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.
[0206] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0207] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together as a System-on-a-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and AI processor, CPU and GPU, etc.
[0208] This application also provides a control device, which includes a processing unit and a storage unit. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the device to perform the methods or steps described in the above embodiments.
[0209] Alternatively, if the control device is located in a vehicle, the processing unit may be the processor 121-12n shown in FIG1.
[0210] This application also provides a control system, which may include a computing platform and a display screen, and the computing platform may include the control device 600 described above.
[0211] This application also provides a vehicle that may include the control device 600 or control system described above.
[0212] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.
[0213] This application also provides a computer-readable medium storing program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.
[0214] This application also provides a chip, which includes a circuit for performing the methods described in the above embodiments.
[0215] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, power-on erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.
[0216] It should be understood that in the embodiments of this application, the memory may include read-only memory and random access memory, and provides instructions and data to the processor.
[0217] It should also be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0218] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0219] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0220] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0221] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0222] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0223] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0224] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be covered. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A control method, characterized in that, include: Obtain the first text content and the user's location information within the cockpit; Based on the first text content and the location information, the first stereo sound is determined; The sound-generating device is controlled according to the first stereo sound.
2. The method according to claim 1, characterized in that, The method further includes: Based on the content of the first text, determine the first sound effect; The step of controlling the sound-generating device according to the first stereo sound includes: The sound-generating device is controlled based on the first sound effect and the first stereo sound.
3. The method according to claim 1, characterized in that, The method further includes: Based on the content of the first text, determine the first sound effect; The step of determining the first stereo sound based on the first text content and the location information includes: The first stereo sound is determined based on the first text content, the first sound effect, and the location information.
4. The method according to claim 2 or 3, characterized in that, The step of determining the first sound effect based on the first text content includes: The first text content is input into the first inference model to obtain the first sound effect.
5. The method according to claim 2 or 3, characterized in that, The step of determining the first sound effect based on the first text content includes: Input the first text content into the first inference model to obtain the first sound effect type; The method further includes: Based on the first sound effect type, the first sound effect is determined from multiple sound effects stored in the sound effect library.
6. The method according to claim 5, characterized in that, The step of determining the first sound effect from multiple sound effects stored in the sound effect library according to the first sound effect type includes: Based on the first sound effect type, a second sound effect is determined from a plurality of sound effects stored in the sound effect library, wherein the plurality of sound effects includes the second sound effect; When the playback duration of the second sound effect is longer than the playback duration of the audio data corresponding to the first text content, the playback duration of the second sound effect is compressed to obtain the first sound effect; or, a portion of the second sound effect is extracted to obtain the first sound effect.
7. The method according to any one of claims 1 to 6, characterized in that, Determining the first stereo sound based on the first text content and the location information includes: The first text content and the location information are input into the second inference model to obtain the first stereo sound.
8. The method according to claim 7, characterized in that, The step of inputting the first text content and the location information into the second inference model to obtain the first stereo sound includes: The first text content and the location information are input into the second inference model to obtain the first stereo sound and the first sound image path. The first sound image path indicates the movement path of the sound image emitting point of the first stereo sound.
9. The method according to any one of claims 1 to 6, characterized in that, Determining the first stereo sound based on the first text content and the location information includes: Input the first text content into the second inference model to obtain the first stereo type; Based on the first stereo type and the location information, the first stereo is determined from multiple stereos stored in the stereo library.
10. The method according to claim 9, characterized in that, The step of determining the first stereo sound from multiple stereo sounds stored in the stereo sound library based on the first stereo sound type and the location information includes: Based on the first stereo type and the location information, the second stereo is determined from a plurality of stereos stored in the stereo library, wherein the plurality of stereos includes the second stereo. When the playback duration of the second stereo is longer than the playback duration of the audio data corresponding to the first text content, the playback duration of the second stereo is compressed to obtain the first stereo; or, a portion of the second stereo is extracted to obtain the first stereo.
11. The method according to any one of claims 1 to 10, characterized in that, The first text content includes multiple sub-text contents, and the first stereo includes multiple sound image path sequences, with each of the multiple sub-text contents corresponding to one of the multiple sound image path sequences.
12. The method according to any one of claims 1 to 11, characterized in that, The process of obtaining the first text content includes: Obtain the first text content output by the content generation model; or, Obtain the first text content from the audio application; or, The first audio data received by the radio is acquired, and the first text content is determined based on the first audio data.
13. The method according to any one of claims 1 to 12, characterized in that, Determining the first stereo sound based on the first text content and the location information includes: Based on the content of the first text, determine the location of the target; The first stereo sound is determined based on the orientation of the target and the location information.
14. The method according to any one of claims 1 to 13, characterized in that, Determining the first stereo sound based on the first text content and the location information includes: Based on the first text content and the location information, playback parameters for multiple sound-generating devices are determined, and the playback parameters include at least one of time delay, volume, and phase.
15. A control device, characterized in that, include: The acquisition unit is used to acquire the first text content and the user's location information inside the cockpit; The determining unit is configured to determine the first stereo sound based on the first text content and the location information; The control unit is used to control the sound-generating device according to the first stereo sound.
16. The apparatus according to claim 15, characterized in that, The determining unit is further configured to determine a first sound effect based on the first text content; The control unit is configured to control the sound-generating device based on the first sound effect and the first stereo sound.
17. The apparatus according to claim 15, characterized in that, The determining unit is used for: Based on the content of the first text, determine the first sound effect; The first stereo sound is determined based on the first text content, the first sound effect, and the location information.
18. The apparatus according to claim 16 or 17, characterized in that, The determining unit is used for: The first text content is input into the first inference model to obtain the first sound effect.
19. The apparatus according to claim 16 or 17, characterized in that, The determining unit is used for: Input the first text content into the first inference model to obtain the first sound effect type; Based on the first sound effect type, the first sound effect is determined from multiple sound effects stored in the sound effect library.
20. The apparatus according to claim 19, characterized in that, The device also includes a sound processing unit. The determining unit is configured to determine a second sound effect from a plurality of sound effects stored in the sound effect library according to the first sound effect type, wherein the plurality of sound effects includes the second sound effect; The sound effect processing unit is used to compress the playback duration of the second sound effect to obtain the first sound effect when the playback duration of the second sound effect is greater than the playback duration of the audio data corresponding to the first text content; or, to extract a portion of the second sound effect to obtain the first sound effect.
21. The apparatus according to any one of claims 15 to 20, characterized in that, The determining unit is used for: The first text content and the location information are input into the second inference model to obtain the first stereo sound.
22. The apparatus according to claim 21, characterized in that, The determining unit is used for: The first text content and the location information are input into the second inference model to obtain the first stereo sound and the first sound image path. The first sound image path indicates the movement path of the sound image emitting point of the first stereo sound.
23. The apparatus according to any one of claims 15 to 20, characterized in that, The determining unit is used for: The first text content and the location information are input into the second inference model to obtain the first stereo type; Based on the first stereo type and the location information, the first stereo is determined from multiple stereos stored in the stereo library.
24. The apparatus according to claim 23, characterized in that, The device also includes a stereo processing unit. The determining unit is configured to determine the second stereo sound from a plurality of stereo sounds stored in the stereo sound library based on the first stereo sound type and the location information, wherein the plurality of stereo sounds includes the second stereo sound. The stereo processing unit is used to compress the playback duration of the second stereo sound to obtain the first stereo sound when the playback duration of the second stereo sound is greater than the playback duration of the audio data corresponding to the first text content; or, to extract a portion of the second stereo sound to obtain the first stereo sound.
25. The apparatus according to any one of claims 15 to 24, characterized in that, The first text content includes multiple sub-text contents, and the first stereo includes multiple sound image path sequences, with each of the multiple sub-text contents corresponding to one of the multiple sound image path sequences.
26. The apparatus according to any one of claims 15 to 25, characterized in that, The acquisition unit is used for: Obtain the first text content output by the content generation model; or, Obtain the first text content from the audio application; or, The first audio data received by the radio is acquired, and the first text content is determined based on the first audio data.
27. The apparatus according to any one of claims 15 to 26, characterized in that, The determining unit is used for: Based on the content of the first text, determine the location of the target; The first stereo sound is determined based on the orientation of the target and the location information.
28. The apparatus according to any one of claims 15 to 27, characterized in that, The determining unit is used for: Based on the first text content and the location information, playback parameters for multiple sound-generating devices are determined, and the playback parameters include at least one of time delay, volume, and phase.
29. A control device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program stored in the memory to cause the apparatus to perform the method as described in any one of claims 1 to 14.
30. A control system, characterized in that, It includes a sound-generating device and a computing platform, wherein the computing platform includes the device as described in any one of claims 15 to 29.
31. A vehicle, characterized in that, Includes the apparatus as described in any one of claims 15 to 29, or the system as described in claim 30.
32. A computer-readable storage medium, characterized in that, It stores instructions that, when executed by a processor, cause the processor to implement the method as described in any one of claims 1 to 14.
33. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 14.
34. A chip, characterized in that, The chip includes circuitry for performing the method as described in any one of claims 1 to 14.