Vehicle-mounted audio generation method and device, electronic equipment and vehicle
By extracting and fusion of the features of the environmental images during the vehicle's driving process, coherent ambient sound is generated, which solves the problem of strong repetitiveness and easy fatigue driving in the existing car noise reduction methods, and improves the driving experience and noise reduction effect.
Patent Information
- Application Number
- CN202311585073.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing automotive noise reduction methods are more repetitive and rigid, which can easily cause users to drive fatigue.
By acquiring the environmental images collected by the image acquisition device during driving, the environmental image features of the current and previous moments are extracted, and the features are fused, and audio synthesis or classification is performed based on the fusion features is performed to generate coherent atmosphere sounds and play them.
The generated ambient sound is more coherent, and corresponding on-board audio can be generated based on the surrounding environment and scenery of the vehicle, which improves the repetitive problems in the existing noise reduction methods and improves the driving experience and noise reduction effect.
Smart Images

Figure CN120048244A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicles, and particularly to a method, device, electronic device, and vehicle for generating in-vehicle audio. Background Art
[0002] Currently, in the automotive industry, vehicle noise reduction mainly includes passive noise reduction by filling sound-absorbing materials and active noise reduction by playing sounds opposite to the noise. The active noise reduction method is more common, mainly using the "masking effect" between different sounds to reduce the impact of noise on users. However, since most current active noise reduction methods use long-term and repetitive ambient sounds, such as long-term bird calls or repetitive rustling of leaves. This repetitive ambient sound can make users sleepy and cause fatigue driving. Summary of the Invention
[0003] In view of this, this application provides a method, device, electronic device, and vehicle for generating in-vehicle audio, which can improve the problem of being repetitive and rigid in the commonly used noise reduction methods of current automobiles, and easily causing fatigue driving for users.
[0004] In a first aspect, this application provides a method for generating in-vehicle audio, including:
[0005] Obtaining environmental images collected by an image acquisition device at preset time intervals during driving;
[0006] Performing feature extraction on the environmental image at the current moment and the environmental image at the previous moment respectively, and superimposing the extracted features to obtain the combined feature corresponding to the environmental image at the current moment and the environmental image at the previous moment;
[0007] Performing preprocessing based on the combined feature to obtain the ambient sound at the current moment; where the preprocessing includes any one of audio synthesis and audio classification;
[0008] Playing the ambient sound based on the ambient sound.
[0009] Optionally, the preprocessing includes audio synthesis;
[0010] The performing preprocessing based on the combined feature to obtain the ambient sound at the current moment includes: using the combined feature as timbre information for audio synthesis to obtain the ambient sound at the current moment.
[0011] Optionally, the preprocessing includes audio classification;
[0012] Performing preprocessing based on the fusion features to obtain the ambient sound at the current moment includes: classifying the fusion features to obtain the audio category corresponding to the fusion features; querying in a pre-configured database based on the audio category to obtain the theoretical audio corresponding to the audio category; and using the theoretical audio as the ambient sound at the current moment.
[0013] Optionally, playing the ambient sound based on the ambient sound includes: when the ambient sound at the current moment is different from the ambient sound at the previous moment, performing a connection operation on the ambient sound at the current moment and the ambient sound at the previous moment, and playing the ambient sound at the current moment.
[0014] Optionally, the connection operation includes one or more of fade-in and fade-out, and mixing the tail of the ambient sound at the previous moment with the head of the ambient sound at the current moment.
[0015] Optionally, playing the ambient sound based on the ambient sound includes: when the ambient sound at the current moment is the same as the ambient sound at the previous moment, continuing to play the ambient sound at the previous moment.
[0016] In a second aspect, the present application provides a vehicle-mounted audio generation device, including:
[0017] An acquisition unit configured to acquire environmental images collected by an image acquisition device at preset time intervals during driving;
[0018] An extraction unit configured to respectively extract features from the environmental image at the current moment and the environmental image at the previous moment, and superimpose the extracted features to obtain the fusion features corresponding to the environmental image at the current moment and the environmental image at the previous moment;
[0019] A processing unit configured to perform preprocessing based on the fusion features to obtain the ambient sound at the current moment; where the preprocessing includes any one of audio synthesis and audio classification;
[0020] A playback unit configured to play the ambient sound based on the ambient sound.
[0021] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the vehicle-mounted audio generation method described in the first aspect is implemented.
[0022] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the vehicle-mounted audio generation method described in the first aspect is implemented.
[0023] In a fifth aspect, the present application provides a vehicle, including the in-vehicle audio generation device mentioned in the second aspect above or the electronic device mentioned in the fourth aspect.
[0024] By means of the above technical solution, for an in-vehicle audio generation method, device, electronic device and vehicle provided by the present application, first, environmental images collected by an image acquisition device at preset time intervals during driving are obtained. Then, feature extraction is respectively performed on the environmental image at the current moment and the environmental image at the previous moment, and the extracted features are superimposed to obtain the fusion features corresponding to the environmental image at the current moment and the environmental image at the previous moment. Next, preprocessing is performed based on the fusion features to obtain the ambient sound at the current moment. Wherein, the preprocessing includes any one of audio synthesis and audio classification. Finally, playback is performed based on the ambient sound. Compared with the related art, by performing feature extraction on the environmental images collected at the current moment and the previous moment and fusing the features of the two, the generated ambient sound is more coherent, and corresponding in-vehicle audio can be generated according to the surrounding environmental scenery of the vehicle, thereby effectively improving the problem of being relatively repetitive and rigid in the commonly used noise reduction methods of current automobiles, which is likely to cause users to drive fatigued, and improving the driving experience.
[0025] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Description of the Drawings
[0026] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application and used together with the description to explain the principles of the present application.
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 Shows a schematic flow chart of an in-vehicle audio generation method provided by an embodiment of the present application;
[0029] Figure 2 Shows a schematic structural diagram of an in-vehicle audio generation device provided by an embodiment of the present application. Detailed Embodiments
[0030] To more clearly understand the above objects, features, and advantages of the present application, the solutions of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other. In addition, in order to more fully understand the characteristics and technical content of the embodiments of the present disclosure, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The attached drawings are only for reference and explanation purposes and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to give a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner to simplify the drawings.
[0031] The in-vehicle audio generation method provided in this embodiment is applied to an in-vehicle audio generation device, which can be installed in an electronic control unit (ECU) in the vehicle or in an in-vehicle audio generator for noise reduction that communicates with the vehicle, as long as noise reduction can be performed according to the in-vehicle audio generation method during vehicle driving.
[0032] To improve the problem of being relatively repetitive and rigid in the commonly used noise reduction methods for current automobiles, which is likely to cause users to drive fatigued and to enhance the user driving experience, this embodiment proposes an in-vehicle audio generation method. As Figure 1 shown, the method includes:
[0033] S101, obtain the environmental images collected by the image acquisition device at preset time intervals during driving.
[0034] The image acquisition device can be a front camera already set on the vehicle, or any device such as a panoramic camera or a fish-eye camera that can capture the surrounding images during driving, and there is no limitation on its position and device. For example, the preset time interval can be set to 5 seconds, that is, the surrounding environmental images are captured every 5 seconds through the vehicle's camera (multiple cameras can be set at different positions of the vehicle to collect the current surrounding environmental images of the vehicle). Among them, the preset time interval is used to delimit the collected images in the time dimension and is also convenient for subsequent selection of corresponding ambient sounds according to "the current moment" and "the previous moment". Setting the preset time interval to 5 seconds, 10 seconds, or 15 seconds can all be preferred values according to the actual use experience.
[0035] S102, respectively perform feature extraction on the environmental image at the current moment and the environmental image at the previous moment, and superimpose the extracted features to obtain the combined features corresponding to the environmental image at the current moment and the environmental image at the previous moment.
[0036] The extracted features are used to represent the scene features of the environmental image, and then the features are superimposed to obtain the fusion features corresponding to the environmental image at the current moment and the environmental image at the previous moment. The purpose of fusion is to keep the generated sound style continuous and avoid the jump of the sound style.
[0037] S103, perform preprocessing based on the fusion features to obtain the ambient sound at the current moment.
[0038] Specifically, the preprocessing includes any one of audio synthesis and audio classification. For audio synthesis, specifically, through the audio generation model Riffusion, taking the fused features as the input of the model, the ambient sound at the current moment can be obtained. Audio classification can be carried out through some high-precision image recognition models (such as CNN models for classification), querying the theoretical audio corresponding to the audio category in the pre-configured database. The audio categories include forest, city, seaside, countryside, etc. The theoretical audio corresponding to the audio category is also pre-stored in the database. For example, if the sea is recognized from the picture, the theoretical audio corresponding to the seaside category can be obtained as the ambient sound at the current moment. It is worth noting that these audio categories can be preset or customized by the user.
[0039] S104, play the ambient sound based on the ambient sound.
[0040] During the vehicle driving process, the ambient sound is always generated through the above steps and connected in chronological order, so as to obtain the in-vehicle audio during the vehicle driving process. That is, the in-vehicle audio is generated while driving according to the surrounding environment. Specifically, the playback of the final in-vehicle audio can be selected considering various situations. For example, when the ambient sound at the current moment is different from the ambient sound at the previous moment, perform a connection operation on the ambient sound at the current moment and the ambient sound at the previous moment, and play the ambient sound at the current moment, and fade-in / fade-out, head / tail mixing and other methods can also be added. If the ambient sound at the current moment is the same as the ambient sound at the previous moment, the ambient sound at the previous moment can also be continued to be played.
[0041] In this embodiment, first, environmental images collected by an image acquisition device at preset time intervals during driving are obtained. Then, feature extraction is respectively performed on the environmental image at the current moment and the environmental image at the previous moment, and the extracted features are superimposed to obtain the fusion features corresponding to the environmental image at the current moment and the environmental image at the previous moment. Next, preprocessing is performed based on the fusion features to obtain the ambient sound at the current moment. Among them, the preprocessing includes any one of audio synthesis and audio classification. Finally, playback is performed based on the ambient sound. Compared with the related art, in this embodiment, by performing feature extraction on the environmental images collected at the current moment and the previous moment and fusing the features of the two, the generated ambient sound is made more coherent, and corresponding in-vehicle audio can be generated according to the surrounding environment of the vehicle, thereby effectively improving the problem of being relatively repetitive and rigid in the currently commonly used noise reduction methods for automobiles, which is likely to cause the user to drive fatigued, and improving the driving experience.
[0042] It should also be noted that the method proposed in this embodiment, in addition to being able to solve the problem that the current noise reduction method is relatively repetitive and rigid and has caused the user to drive fatigued, can also improve the noise reduction effect compared with the noise reduction methods in the related art. Currently, the common noise reduction methods usually only play bird calls and leaf sounds sufficiently to cover up the mechanical noise of the vehicle itself. However, the "masking effect" of this method is poor. The "masking effect" means that in a quiet environment, even if the sound pressure level of a sound is very low, it can be heard, indicating that the auditory threshold of the human ear for this sound can be very low. However, when listening to a sound, if there is another sound (called the masking sound) at the same time, it will affect the hearing effect of the sound being listened to by the human ear, and at this time, the auditory threshold for the sound being listened to will increase. This phenomenon where the auditory sensitivity of the human ear to other sounds is reduced due to the existence of a certain sound. In this embodiment, by making the generated sound style always continuous and generated according to the surrounding environment, it can improve the attraction to the user, have a better masking effect on mechanical noise, and improve the noise reduction effect. In addition, this embodiment can also improve the noise reduction effect by enhancing the ambient sound.
[0043] Optionally, the preprocessing includes audio synthesis;
[0044] Performing preprocessing based on the fusion features to obtain the ambient sound at the current moment includes: using the fusion features as timbre information for audio synthesis to obtain the ambient sound at the current moment.
[0045] In this embodiment, taking the Riffusion model as an example, the fused features are used as the input of the Riffusion model. Fundamentally, the timbre information is that the fused image features may correspond to a certain timbre. For example, the timbre corresponding to the tunnel may be thicker and heavier. Furthermore, the Riffusion model can perform audio synthesis to obtain the ambient sound at the current moment. Other image-to-music models (such as Stable Diffusion, etc.) can also be used.
[0046] Optionally, the preprocessing includes audio classification;
[0047] Performing preprocessing based on the fused features to obtain the ambient sound at the current moment includes: performing audio classification on the fused features to obtain the audio category corresponding to the fused features; querying in a preconfigured database based on the audio category to obtain the theoretical audio corresponding to the audio category; and using the theoretical audio as the ambient sound at the current moment.
[0048] In this embodiment, audio classification can be performed through some high-precision image recognition models (such as CNN models for classification). Querying the theoretical audio corresponding to the audio category in a preconfigured database, the audio categories include forest, city, seaside, countryside, etc. Multiple ambient sounds corresponding to feature tags can be pre-entered in the database. For example, the ambient sounds of highways, cities, etc. are set in advance, so that when the user enters the city from the highway, the in-vehicle audio can be flexibly and naturally connected and changed, greatly improving the user's driving experience. The duration of each ambient sound can also be selected and adjusted according to a preset time interval, which can be specifically set according to the user's needs. For example, if the sea is recognized from the picture, the theoretical audio corresponding to the seaside category can be obtained as the ambient sound at the current moment. It should be noted that these audio categories can be preset or custom-set by the user.
[0049] Optionally, playing the ambient sound based on the ambient sound includes: when the ambient sound at the current moment is different from the ambient sound at the previous moment, performing a connection operation on the ambient sound at the current moment and the ambient sound at the previous moment, and playing the ambient sound at the current moment.
[0050] In this embodiment, when the ambient sound at the current moment is different from the ambient sound at the previous moment, the ambient sound at the current moment and the ambient sound at the previous moment can be connected to play the in-vehicle audio.
[0051] Optionally, the connection operation includes one or more of fade-in / fade-out, mixing the tail of the ambient sound at the previous moment with the head of the ambient sound at the current moment.
[0052] In this embodiment, operations such as fade-in / fade-out and head-tail mixing can also be added during the connection process to avoid the in-vehicle audio from being abrupt.
[0053] Optionally, based on the ambient sound, play the ambient sound, including: when the ambient sound at the current moment is the same as the ambient sound at the previous moment, continue to play the ambient sound at the previous moment.
[0054] In this embodiment, when the ambient sound at the current moment is the same as the ambient sound at the previous moment, continue to play the ambient sound at the previous moment, so that the in-vehicle audio is smoother. The main emphasis is that throughout the driving process, the ambient sound at the current moment is always connected to the end of the ambient sound at the previous moment, so as to generate a complete, flexible and natural in-vehicle audio.
[0055] Further, as Figure 1 a specific implementation of the method shown, this embodiment provides an in-vehicle audio generation device, as Figure 2 shown, the device includes: an acquisition unit 201, an extraction unit 202, a processing unit 203 and a playback unit 204.
[0056] The acquisition unit 201 is configured to acquire environmental images collected by an image acquisition device at preset time intervals during driving;
[0057] The extraction unit 202 is configured to respectively perform feature extraction on the environmental image at the current moment and the environmental image at the previous moment, and superimpose the extracted features to obtain the combined features corresponding to the environmental image at the current moment and the environmental image at the previous moment;
[0058] The processing unit 203 is configured to perform preprocessing based on the combined features to obtain the ambient sound at the current moment; wherein, the preprocessing includes any one of audio synthesis and audio classification;
[0059] The playback unit 204 is configured to play the ambient sound based on the ambient sound.
[0060] In a specific application scenario, the processing unit 203 is specifically configured to perform audio synthesis using the combined features as timbre information to obtain the ambient sound at the current moment.
[0061] In a specific application scenario, the processing unit 203 is further specifically configured to perform audio classification on the combined features to obtain the audio category corresponding to the combined features; query in a pre-configured database based on the audio category to obtain the theoretical audio corresponding to the audio category; and use the theoretical audio as the ambient sound at the current moment based on the theoretical audio.
[0062] In a specific application scenario, the playback unit 204 is further specifically configured to perform a connection operation on the ambient sound at the current moment and the ambient sound at the previous moment when the ambient sound at the current moment is different from the ambient sound at the previous moment, and play the ambient sound at the current moment.
[0063] In a specific application scenario, the playback unit 204 is further specifically configured to continue playing the ambient sound at the previous moment when the ambient sound at the current moment is the same as the ambient sound at the previous moment.
[0064] It should be noted that for other corresponding descriptions of each functional unit involved in the in-vehicle audio generation device provided in this embodiment, reference can be made to Figure 1 the corresponding description in, which will not be elaborated here.
[0065] Based on the method as shown in Figure 1 above, correspondingly, this embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method as shown in Figure 1 above is implemented.
[0066] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in various implementation scenarios of this application.
[0067] Based on the method as shown in Figure 1 above, and Figure 2 the virtual device embodiment as shown in, for the purpose of achieving the above object, this embodiment of the application also provides an electronic device, which can be configured on the computer side, or the vehicle side, etc. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as shown in Figure 1 above.
[0068] Based on the above electronic device, this embodiment of the application also provides a vehicle, which specifically may include: the device as shown in Figure 2 above or the above electronic device. The vehicle can specifically be a new energy vehicle or a traditional vehicle, etc.
[0069] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen, an input unit such as a keyboard, etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0070] Those skilled in the art can understand that the structure of the above-mentioned physical device provided in this embodiment does not constitute a limitation on the physical device, and it may include more or fewer components, or combine some components, or have different component arrangements.
[0071] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication between other hardware and software in the information processing physical device.
[0072] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. Applying the solution of this embodiment, first obtain the environmental images collected by the image acquisition device at preset time intervals during driving. Then, respectively extract the features of the environmental image at the current moment and the environmental image at the previous moment, and superimpose the extracted features to obtain the combined features corresponding to the environmental image at the current moment and the environmental image at the previous moment. Then, based on the combined features, perform preprocessing to obtain the ambient sound at the current moment. Among them, the preprocessing includes any one of audio synthesis and audio classification. Finally, based on the ambient sound, play it. Compared with the related art, in this embodiment, by extracting the features of the environmental images collected at the current moment and the previous moment and fusing the features of the two, the generated ambient sound is more coherent, and the in-vehicle audio can be generated according to the surrounding environmental scenery of the vehicle, thus effectively improving the problem of being relatively repetitive and rigid in the commonly used noise reduction methods of current automobiles, which is likely to cause the user to drive tired, and improving the driving experience.
[0073] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0074] The above description is only a specific implementation manner of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments described herein, but rather will conform to the broadest scope consistent with the principles and novel features claimed herein.
[0075] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments only represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Some parts and features of some embodiments may be included in or replace parts and features of other embodiments. As used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise", its variants "comprises" and / or "comprising", etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groupings thereof. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device comprising the element. In this text, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, devices, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.
[0076] Those skilled in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner can depend on the specific application and design constraints of the technical solution. The skilled person can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. The skilled person can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0077] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components can be or can not be physically separated, and the components shown as units can be or can not be physical units, that is, they can be located in one place or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for generating in-vehicle audio, characterized in that, it includes: Obtain the environmental images collected by the image acquisition device at preset time intervals during driving; Extract features from the environmental image at the current moment and the environmental image at the previous moment respectively, and superimpose the extracted features to obtain the combined features corresponding to the environmental image at the current moment and the environmental image at the previous moment; Perform preprocessing based on the combined features to obtain the ambient sound at the current moment; wherein, the preprocessing includes any one of audio synthesis and audio classification; Play the ambient sound based on the ambient sound.
2. The method for generating in-vehicle audio according to claim 1, characterized in that, the preprocessing includes audio synthesis; The performing preprocessing based on the combined features to obtain the ambient sound at the current moment includes: Use the combined features as timbre information for audio synthesis to obtain the ambient sound at the current moment.
3. The method for generating in-vehicle audio according to claim 1, characterized in that, the preprocessing includes audio classification; The performing preprocessing based on the combined features to obtain the ambient sound at the current moment includes: Perform audio classification on the combined features to obtain the audio category corresponding to the combined features; Query in a pre-configured database based on the audio category to obtain the theoretical audio corresponding to the audio category; Use the theoretical audio as the ambient sound at the current moment based on the theoretical audio.
4. The method for generating in-vehicle audio according to claim 3, characterized in that, the playing the ambient sound based on the ambient sound includes: In the case where the ambient sound at the current moment is different from the ambient sound at the previous moment, perform a connection operation on the ambient sound at the current moment and the ambient sound at the previous moment, and play the ambient sound at the current moment.
5. The method for generating in-vehicle audio according to claim 4, characterized in that, the connection operation includes one or more of fade-in and fade-out, and mixing the tail of the ambient sound at the previous moment with the head of the ambient sound at the current moment.
6. The method for generating in-vehicle audio according to claim 3, characterized in that, the playing the ambient sound based on the ambient sound includes: In the case where the ambient sound at the current moment is the same as the ambient sound at the previous moment, continue to play the ambient sound at the previous moment.
7. An in-vehicle audio generation device, characterized in that, it includes: An acquisition unit configured to obtain the environmental images collected by the image acquisition device at preset time intervals during driving; An extraction unit configured to extract features from the environmental image at the current moment and the environmental image at the previous moment respectively, and superimpose the extracted features to obtain the combined features corresponding to the environmental image at the current moment and the environmental image at the previous moment; A processing unit configured to perform preprocessing based on the combined features to obtain the ambient sound at the current moment; wherein, the preprocessing includes any one of audio synthesis and audio classification; A playback unit configured to play the ambient sound based on the ambient sound.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 6.
9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein, when the processor executes the computer program, it implements the method according to any one of claims 1 to 6.
10. A vehicle, wherein, comprising: the device according to claim 7, or the electronic device according to claim 9.