Audio generation method and device
By generating sound effects scripts and performing spatial audio rendering, the problem of insufficient diversity of the in-car soundscape is solved, a personalized soundscape experience is achieved, and the user's listening experience is enhanced.
Patent Information
- Application Number
- CN202410309981.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-16
AI Technical Summary
In the in-car environment, existing technologies find it difficult to provide a diverse soundscape that meets user needs, affecting the user's listening experience.
By obtaining user instructions, generating a sound effect script, and performing spatial audio rendering based on the script, a soundscape that meets user expectations is generated, supporting user personalized customization of the soundscape.
It improves the user's listening experience, provides a personalized soundscape, and enhances the entertainment effect in the car.
Smart Images

Figure CN120659005A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to in-vehicle audio technology, and in particular to an audio generation method and device. Background Art
[0002] As technology advances and the number of cars on the road continues to grow, so too has the demand for in-car entertainment, creating a rich environment for the development of in-car audio technology. On the one hand, the enclosed interior of a car creates a naturally excellent mobile sound field, fertile ground for experiencing the ambiance of spatial audio. On the other hand, the high-pressure pace of society has led many users to choose to listen to audio in their cars, with relaxation and increased enjoyment being the two main drivers of audio listening.
[0003] Therefore, providing diverse and demand-oriented soundscapes in the car has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The present application provides an audio generation method and device to obtain a soundscape that meets user expectations and enhance the user's listening experience.
[0005] In a first aspect, the present application provides an audio generation method, comprising: obtaining a first script in response to a user instruction, wherein the first script is a sound effect script; obtaining a target sound effect according to the first script; and performing spatial audio rendering on the target sound effect to obtain a target audio.
[0006] In the embodiment of the present application, by generating a first script for sound effects and generating target sound effects based on the guidance of the script, a soundscape that meets the user's expectations can be obtained, thereby improving the user's listening experience.
[0007] Optionally, the user instruction can be a voice instruction input through the microphone. For example, the user says "Hello Xiaoyi, help me record this journey", "Hello Xiaoyi, play a sound effect of a forest environment", etc. The device has semantic recognition of voice, so when the preset trigger voice (for example, "Hello Xiaoyi") is recognized, the subsequent voice can be recognized to extract keywords (for example, record this journey, sound effects of a forest environment) to perform corresponding operations.
[0008] Optionally, the user command may be a control command input through the central control screen. For example, the central control provides an interactive interface for the soundscape. The user can click a corresponding button or icon on the central control screen to open the interactive interface, and then select a soundscape scene, input requirements, etc. on the interactive interface. It should be understood that the interactive interface for the soundscape can provide a variety of ways for the user to input commands, including but not limited to menu selection, control triggering, text input, etc., and the layout and content of the interactive interface can also be designed in a variety of ways, which are not specifically limited in the embodiments of the present application.
[0009] It should be noted that in addition to the two methods of obtaining user commands mentioned above, other methods can also be used to obtain user commands, for example, through a car-machine application installed on a mobile phone and a communication link between the mobile phone and the car-machine, to control the soundscape function on the car-machine. Therefore, the embodiment of the present application does not specifically limit the method of obtaining user commands.
[0010] The first script can be a sound effect script, and then the audio signal can be edited based on the first script to obtain the expected sound effect. In the embodiment of the present application, the first script can be obtained by the following two methods:
[0011] 1. Generate a first scenario based on multimodal information collected by multiple sensors.
[0012] Multimodal information includes at least one of the following: audio information collected by a microphone; video information collected by a camera; or control information obtained by the central control screen, where the control information includes at least one of text, images, video, or audio. Sensors can be various sensors installed on a device (such as a vehicle or a mobile phone), including but not limited to cameras, microphones, touch screens, radars, etc.
[0013] In an embodiment of the present application, semantic recognition can be performed on audio, video, text, etc. to obtain keywords in the aforementioned carriers, and corresponding sound effect objects can be obtained based on the keywords, thereby generating a first script.
[0014] Exemplarily, when the multimodal information is text, the keywords contained in the text can provide a sound effect object, based on which the attribute information of the sound effect object such as its corresponding virtual duration, virtual position, virtual volume, etc. is predicted, and these attribute information can constitute the first script.
[0015] Alternatively, when the multimodal information is audio, semantic recognition is performed on the audio, and the obtained keywords can provide a sound effect object. Based on the sound effect object, its corresponding attribute information such as real duration, real position, real volume, and real timbre is predicted. These attribute information can constitute the first script.
[0016] Alternatively, when the multimodal information is a video / image, semantic recognition is performed on the video / image, and the obtained keywords can provide a sound effect object. Based on the sound effect object, its corresponding real duration, real position and other attribute information of the sound effect object are predicted. These attribute information can constitute the first script.
[0017] In an embodiment of the present application, the first script may include descriptive information of at least one sound effect object (for example, the identification, text or description of a cat, dog, bird, person, etc.). In addition, the first script may also include attribute information of at least one sound effect object (including at least one of volume, duration, timbre, and may also include position). This information can guide the generation of the target sound effect.
[0018] It should be noted that the first script can be described in a variety of ways, and the content included is related to the application scenario and the information obtained. The embodiment of the present application does not specifically limit the content and description method of the first script.
[0019] 2. Get the pre-set first script.
[0020] In the embodiment of the present application, in order to improve efficiency, first scripts for various scenes can be pre-set, for example, a first script for a rest scene, a first script for a forest atmosphere scene, a first script for an opera house scene, etc. Attribute information is set for the sound effect objects that may appear in the scene, and then the first script is generated based on this attribute information and stored in the memory.
[0021] An interactive interface for scene selection is provided to the user on the central control screen. When the user selects a scene on the interactive interface, the first script corresponding to the scene can be called from the memory.
[0022] In a possible implementation, the device may further obtain a second script, where the second script is a video script generated based on visual information, where the visual information includes at least one of video, image, or text.
[0023] The device can acquire visual information through a camera (for example, sensors around the vehicle can capture video, a mobile phone camera can capture video, a dedicated video camera can capture video, etc.) or through a user's touch operation on the central control screen (for example, a user operates on the central control screen to select a local video, a user operates on the central control screen to download a video from the Internet, etc.). It should be understood that while the above example uses the acquisition of video as an example, the above method can also be used to acquire images or text.
[0024] In an embodiment of the present application, semantic recognition can be performed on videos, images, texts, etc. to obtain keywords in the aforementioned carriers, and corresponding sound effect objects can be obtained based on the keywords, thereby generating a second script.
[0025] Exemplarily, when the multimodal information is text, the keywords contained in the text can provide a sound effect object, based on which the attribute information of the sound effect object such as its corresponding virtual duration and virtual position is predicted, and these attribute information can constitute the second script.
[0026] Alternatively, when the multimodal information is a video / image, semantic recognition is performed on the video / image, and the obtained keywords can provide a sound effect object. Based on the sound effect object, its corresponding real duration, real position and other attribute information of the sound effect object are predicted. These attribute information can constitute the second script.
[0027] In an embodiment of the present application, the second script may include descriptive information of at least one sound effect object (for example, an identifier, text, or description of a cat, dog, bird, or person, etc.). In addition, the second script may also include attribute information of at least one sound effect object (including at least one of duration and position), which can guide the generation of the target video. The second script may also refer to the description method of Table 1, without specific limitation. Compared with the first script, the attribute information of the sound effect object in the second script does not include volume and timbre.
[0028] In an embodiment of the present application, the first script can be used as a script for sound effects to guide the generation of target sound effects, and the second script can be used as a script for videos to guide the generation of target videos. In addition, since the second script also contains the attribute information of the sound effect object, the second script can be referenced when generating the target sound effect to generate a target sound effect that matches the sound effect object in the second script. For example, the duration and volume of the target sound effect match the duration and volume of the sound effect object in the second script, respectively, the timbre of the target sound effect matches the sound effect object in the second script, and the position of the target sound effect matches the position where the sound effect object appears in the second script, so that the sound effect and video are highly consistent, enhancing the user's audio-visual experience. Optionally, the second script can serve as a supplement to the first script. When a sound effect object that is not in the first script appears in the second script, the sound of the sound effect object can be added to the target sound effect to improve the soundscape effect of the target sound effect.
[0029] In a possible implementation, there is a first script but no second script. A respective sub-sound effect can be generated for at least one sound effect object appearing in the first script to constitute a target sound effect.
[0030] For example, a first script generated based on text describes N sound effect objects, where sound effect object A is a dog, appears from the 2nd to the 5th second, for a total of 3 seconds, and is located on the left side. The AIGC model can convert the dog into a barking sound, ultimately generating the barking sound. This barking sound can constitute sub-sound effect A in the target sound effect, corresponding to the dog; the aforementioned appearance time indicates the temporal position of the barking sound within the target sound effect; this position can be used in subsequent rendering processing to give the user the impression that the three-second barking sound is occurring on the left side.
[0031] In one possible implementation, there are a first script and a second script. A union of at least one sound effect object appearing in the first script and at least one sound effect object appearing in the second script may be obtained; and sub-sound effects may be generated for the multiple sound effect objects in the union to form a target sound effect.
[0032] For example, in addition to the first script generated based on text, the second script describes M sound effect objects, and L sound effect objects are not included in the first script. Among them, sound effect object B is thunder, which appears from the 6th to the 8th second, for a total of 2 seconds, and is located above. The AIGC model can realize the conversion from thunder to thunder sound, and finally generate the thunder sound. The thunder sound can constitute the sub-sound effect B in the target sound effect, corresponding to thunder; the aforementioned appearance time indicates the time position of the thunder sound in the target sound effect; the aforementioned position can be used for subsequent rendering processing, so that the user hears the two-second thunder sound as if it is happening above. At this time, the user will hear the barking of a dog on the left from the 2nd to the 5th second, and the thunder sound from above from the 6th to the 8th second, accompanied by a thunder picture.
[0033] The device can perform spatial audio rendering on the target sound effect to produce the target audio. For example, in the first script, there is a sound effect object of a dog, and its corresponding position is on the left. After spatial audio rendering, the dog's barking sound will be perceived as being on the left.
[0034] In one possible implementation, after acquiring the target audio, the target audio can be played to the user via a speaker. If the user finds any modifications needed during the listening process, they can enter a modification instruction. Accordingly, the device receives the user's modification instruction; modifies the first script in response to the modification instruction to obtain a third script; acquires the modified target sound effect based on the third script; and acquires the modified target audio based on the modified target sound effect.
[0035] The difference from the first script generated above is that this embodiment can add factors to modify the instructions (such as keywords in the modification instructions) on the basis of the first script, thereby obtaining a third script. For example, the first script is generated based on the instruction of "sound effects of the forest environment", and the user enters the modification instruction of "adding children's playful sounds". At this time, the third script adds the sound effect information of children's playful sounds on the basis of the first script. After obtaining the third script, the device again generates a modified target sound effect based on the third script. Compared with the target sound effect before modification, the modified target sound effect adds the sound effect of children's playful sounds, and then obtains the modified target audio based on the modified target sound effect and plays it for the user.
[0036] The above description of the process of a user modifying a created soundscape is intended to illustrate that the embodiment of the present application supports modification of the soundscape, realizes the personalized customization function of the soundscape, and makes the soundscape function more in line with the user's audio-visual needs.
[0037] In one possible implementation, after obtaining the second script, the device can edit the visual information according to the second script to obtain an edited video; and fuse the edited video and the target sound effect to obtain the target video, wherein the timestamp of the sound effect in the target video is synchronized with the timestamp of the sound effect object and conforms to the sound characteristics of the sound effect object.
[0038] In an embodiment of the present application, in addition to generating the target sound effect, the captured video can be edited, and then the edited video and the target sound effect can be fused to obtain the target video. The picture and sound effect of the target video are synchronously matched, that is, the final video is generated at one time, including the picture and the sound effect matching the sound effect object in the picture, the background music matching the picture content, etc., thereby improving the efficiency of video editing.
[0039] In a second aspect, the present application provides an audio generation device, comprising: an acquisition module for acquiring a first script in response to a user instruction, wherein the first script is a sound effect script; a processing module for acquiring a target sound effect based on the first script; and a rendering module for performing spatial audio rendering on the target sound effect to obtain a target audio.
[0040] In a possible implementation, the acquisition module is specifically configured to generate the first script based on multimodal information, where the multimodal information is collected by a variety of sensors; or to acquire a preset first script.
[0041] In one possible implementation, the multimodal information includes at least one of the following information:
[0042] Audio information collected through a microphone; or
[0043] Video information collected by a camera; or
[0044] The control information obtained through the central control screen includes at least one of text, image, video or audio.
[0045] In a possible implementation, the first script includes at least one sound effect object, where the sound effect object is a physical object that emits a sound effect.
[0046] In a possible implementation, the first script further includes attribute information of each of the at least one sound effect objects, where the attribute information includes at least one of volume, duration, or timbre.
[0047] In one possible implementation, the first script also includes the respective positions of the at least one sound effect object; the rendering module is specifically configured to perform spatial audio rendering on the target sound effect according to the respective positions of the at least one sound effect object to obtain the target audio.
[0048] In one possible implementation, the acquisition module is further used to acquire a second script, which is a video script generated based on visual information, and the visual information includes at least one of video, image or text; the processing module is further used to acquire the target sound effect based on the first script and the second script.
[0049] In a possible implementation, the visual information is collected through a camera or a central control screen.
[0050] In a possible implementation, the second script includes at least one sound effect object, where the sound effect object is a physical object that emits a sound effect in the visual information.
[0051] In a possible implementation, the second script further includes attribute information of each of the at least one sound effect objects, where the attribute information includes duration and / or position.
[0052] In a possible implementation, the processing module is specifically configured to generate respective sub-sound effects for at least one sound effect object appearing in the first script to constitute the target sound effect.
[0053] In one possible implementation, the processing module is specifically used to obtain the union of at least one sound effect object appearing in the first script and at least one sound effect object appearing in the second script; and generate respective sub-sound effects for the multiple sound effect objects in the union to constitute the target sound effect.
[0054] In a possible implementation, the user instruction includes at least one of the following instructions:
[0055] Voice commands entered through a microphone; or,
[0056] Control commands entered through the central control screen.
[0057] In a third aspect, the present application provides a device comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement a method as described in any one of the above-mentioned first aspects.
[0058] In a fourth aspect, the present application provides a computer-readable storage medium, characterized in that it includes a computer program, and when the computer program is executed on a computer, it enables the computer to perform any one of the methods described in the first aspect.
[0059] In a fifth aspect, the present application provides a computer program product, characterized in that the computer program product includes computer program code, and when the computer program code is run on a computer, the computer executes any one of the methods in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is an exemplary functional block diagram of the vehicle 100 according to an embodiment of the present application;
[0061] Figure 2 This is an exemplary functional block diagram of the vehicle-mounted assisted driving system according to an embodiment of the present application;
[0062] Figure 3 This is a schematic structural diagram of a mobile phone 300 according to an embodiment of the present application;
[0063] Figure 4 This is a block diagram of the software structure used by the device in the embodiment of the present application;
[0064] Figure 5 Flowchart of process 500 of the audio generation method provided in an embodiment of the present application;
[0065] Figure 6 A schematic diagram of a personalized rendering system according to an embodiment of the present application;
[0066] Figure 7 A schematic diagram of an interactive interface for a soundscape according to an embodiment of the present application;
[0067] Figure 8 This is a structural diagram of a personalized rendering system according to an embodiment of the present application;
[0068] Figure 9 This is a structural diagram of a personalized rendering system according to an embodiment of the present application;
[0069] Figure 10 Schematic diagram of the structure of the audio generating device 1000 of this application. DETAILED DESCRIPTION
[0070] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0071] The terms "first," "second," and the like in the description, embodiments, claims, and drawings of this application are used solely for descriptive purposes and are not to be construed as indicating or implying relative importance or order. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0072] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0073] The following are terms used in the embodiments of this application:
[0074] 1. Sound effects: refers to the sound effects used in visual art media such as movies, television, music production or games to enhance the emotional effect.
[0075] 2. Sound effect object: refers to the physical object that produces sound effects in visual information such as videos, images and descriptive text.
[0076] 3. Sound effect synthesis: refers to the automatic generation of corresponding sound effects through multiple modal input conditions.
[0077] 4. Spatial Audio Rendering: A technology that locates and processes audio signals in three-dimensional space. It can simulate the position, size, shape, and movement of the sound source in space, giving listeners a more realistic audio experience.
[0078] 5. Soundscape, also known as soundscape, is a field of auditory ecology research that focuses on how people perceive and understand the sound environment and how sound influences their subjective experience. Unlike traditional noise control, soundscape emphasizes the perceived value of sound, encompassing not only physical quantities but also positive and harmonious sounds, viewing the sound environment as a resource. Soundscape aims to enhance people's comfort and satisfaction in urban environments through planning and design, enabling them to experience a high-quality sound ecology.
[0079] Before describing the technical solutions of the embodiments of the present application, the devices (eg, a vehicle and a mobile phone) of the embodiments of the present application will be described with reference to the accompanying drawings.
[0080] Figure 1 This is an exemplary functional block diagram of the vehicle 100 according to an embodiment of the present application. Figure 1 As shown, components coupled to or included in the vehicle 100 may include a propulsion system 110, a sensor system 120, a control system 130, peripheral devices 140, a power source 150, a computing device 160, and a driver interface 170. The components of the vehicle 100 may be configured to operate in a manner interconnected with each other and / or with other components coupled to the various systems. For example, the power source 150 may provide power to all components of the vehicle 100. The computing device 160 may be configured to receive data from the propulsion system 110, the sensor system 120, the control system 130, and the peripheral devices 140 and to control them. The computing device 160 may also be configured to generate a display of images on the driver interface 170 and to receive input from the driver interface 170.
[0081] It should be noted that in other examples, the vehicle 100 may include more, fewer, or different systems, and each system may include more, fewer, or different components. In addition, the systems and components shown may be combined or divided in any manner, and this application does not specifically limit this.
[0082] Computing device 160 may include a processor 161, a transceiver 162, and a memory 163. Computing device 160 may be a controller or part of a controller of vehicle 100. Memory 163 may store instructions 1631 executed by processor 161 to execute various functional applications and data processing of vehicle 100. It may also store data generated during use of vehicle 100 (e.g., map data 1632), an operating system (e.g., an embedded operating system such as Android, Apple Mobile Platform (iOS), Microsoft Windows, or a UNIX-like operating system (Linux)), and applications required for at least one function. Processor 161 included in computing device 160 may include one or more general-purpose processors and / or one or more specialized processors (e.g., an image processor, a digital signal processor, etc.). To the extent processor 161 includes more than one processor, such processors may operate individually or in combination. Computing device 160 may implement functions for controlling vehicle 100 based on input received through driver interface 170. Transceiver 162 facilitates communication between computing device 160 and various systems. The memory 163, in turn, may include one or more volatile storage components and / or one or more non-volatile storage components, such as optical, magnetic, and / or organic storage devices, and the memory 163 may be fully or partially integrated with the processor 161. The memory 163 may contain instructions 1631 (e.g., program logic) executable by the processor 161 to perform various vehicle functions, including any of the functions or methods described herein.
[0083] The propulsion system 110 can provide powered movement for the vehicle 100. Figure 1 As shown, the propulsion system 110 may include an engine / motor 114, an energy source 113, a transmission 112, and wheels / tires 111. In addition, the propulsion system 110 may additionally or alternatively include Figure 1 This application does not specifically limit other components than those shown.
[0084] The sensor system 120 may include several sensors for sensing information about the environment in which the vehicle 100 is located. Figure 1As shown, the sensors of the sensor system 120 include a global positioning system (GPS) 126, an inertial measurement unit (IMU) 125, a lidar sensor 124, a camera sensor 123, a millimeter wave radar sensor 122, and an actuator 121 for modifying the position and / or orientation of the sensors. The GPS 126 can be any sensor used to estimate the geographic location of the vehicle 100. To this end, the GPS 126 can include a transceiver that estimates the position of the vehicle 100 relative to the earth based on satellite positioning data. In an example, the computing device 160 can be used to use the GPS 126 in conjunction with map data 1632 to estimate the road traveled by the vehicle 100. The IMU 125 can be used to sense changes in the position and orientation of the vehicle 100 based on inertial acceleration and any combination thereof. In some examples, the combination of sensors in the IMU 125 can include, for example, an accelerometer and a gyroscope. In addition, other combinations of sensors in the IMU 125 are also possible. The lidar sensor 124 can be considered an object detection system that uses light sensing to detect objects in the environment in which the vehicle 100 is located. The lidar sensor 124 is typically an optical remote sensing technology that can measure the distance to a target or other properties of the target by illuminating the target with light. By way of example, the lidar sensor 124 can include a laser source and / or laser scanner configured to emit laser pulses, and a detector for receiving reflections of the laser pulses. For example, the lidar sensor 124 can include a laser rangefinder reflected by a rotating mirror and scan the laser in one or two dimensions around a digitized scene, thereby collecting distance measurements at specified angular intervals. In an example, the lidar sensor 124 can include components such as a light (e.g., laser) source, a scanner and optical system, light detectors and receiver electronics, and a positioning and navigation system. The lidar sensor 124 determines the distance to an object by scanning the laser light reflected from the object, and can form a 3D image of the environment with up to centimeter-level accuracy. The camera sensor 123 can include any camera (e.g., a pinhole camera, a fisheye camera, a still camera, a video camera, etc.) for capturing images of the environment in which the vehicle 100 is located. To this end, the camera sensor 123 can be configured to detect visible light, or can be configured to detect light from other parts of the spectrum (such as infrared light or ultraviolet light). Other types of camera sensors 123 are also possible. The camera sensor 123 can be a two-dimensional detector, or can have a three-dimensional spatial range detection function. In some examples, the camera sensor 123 can be, for example, a distance detector that is configured to generate a two-dimensional image indicating the distance from the camera sensor 123 to several points in the environment. To this end, the camera sensor 123 can use one or more distance detection technologies.For example, the camera sensor 123 can be configured to use structured light technology, in which the vehicle 100 uses a predetermined light pattern, such as a grid or checkerboard pattern, to illuminate objects in the environment, and uses the camera sensor 123 to detect the reflection of the predetermined light pattern from the object. Based on the distortion in the reflected light pattern, the vehicle 100 can be configured to detect the distance of a point on the object. The predetermined light pattern may include infrared light or light of other wavelengths. The millimeter-wave radar sensor (Millimeter-Wave Radar) 122 generally refers to an object detection sensor with a wavelength of 1 to 10 mm and a frequency range of approximately 10 GHz to 200 GHz. The measurement value of the millimeter-wave radar sensor 122 has depth information and can provide the distance to the target; secondly, because the millimeter-wave radar sensor 122 has a significant Doppler effect and is very sensitive to speed, the speed of the target can be directly obtained, and the speed of the target can be extracted by detecting its Doppler frequency shift. Currently, the two mainstream automotive millimeter-wave radar application frequency bands are 24GHz and 77GHz. The former has a wavelength of about 1.25cm and is mainly used for short-range perception, such as the surrounding environment of the vehicle body, blind spots, parking assistance, lane change assistance, etc.; the latter has a wavelength of about 4mm and is used for medium and long-range measurement, such as automatic following, adaptive cruise control (ACC), and automatic emergency braking (AEB).
[0085] The sensor system 120 may also include additional sensors, including, for example, sensors that monitor internal systems of the vehicle 100 (e.g., an O2 monitor, a fuel gauge, an oil temperature, etc.). The sensor system 120 may also include other sensors. This application does not make specific limitations on this.
[0086] The control system 130 may be configured to control the operation of the vehicle 100 and its components. To this end, the control system 130 may include a steering unit 136, a throttle 135, a brake unit 134, a sensor fusion algorithm 133, a computer vision system 132, and a navigation / route control system 131. The control system 130 may additionally or alternatively include Figure 1 This application does not specifically limit other components than those shown.
[0087] The peripheral devices 140 may be configured to allow the vehicle 100 to interact with external sensors, other vehicles, and / or the driver. To this end, the peripheral devices 140 may include, for example, a lighting system 145, a wireless communication system 144, a touch screen 143, a microphone 142, and / or a speaker 141. The peripheral devices 140 may additionally or alternatively include a Figure 1 This application does not specifically limit other components than those shown.
[0088] Power source 150 can be configured to provide power to some or all components of vehicle 100. To this end, power source 150 can include, for example, a rechargeable lithium-ion or lead-acid battery. In some examples, one or more battery packs can be configured to provide power. Other power source materials and configurations are also possible. In some examples, power source 150 and energy source 113 can be implemented together, as in some all-electric vehicles.
[0089] The components of the vehicle 100 can be configured to work in an interconnected manner with other components within and / or outside of their respective systems. To this end, the components and systems of the vehicle 100 can be communicatively linked together through a system bus, a network, and / or other connection mechanisms.
[0090] Figure 2 This is an exemplary functional block diagram of the vehicle-mounted assisted driving system according to an embodiment of the present application, such as Figure 2 As shown, the components coupled to or included in the vehicle-mounted assisted driving system may include a computing unit, a sensor, a central control screen, a lighting system, and an audio system, wherein the computing unit corresponds to Figure 1 In the control system 130 of the embodiment shown, the sensors correspond to Figure 1 The sensor system 120 in the embodiment shown mainly includes a camera sensor 123 (including a pinhole camera, a fisheye camera), a millimeter wave radar sensor 122, and a lidar sensor 124. Figure 1 The touch screen 143 in the embodiment shown provides the driver with an interface for human-computer interaction. Figure 1 The lighting system 145 in the embodiment shown corresponds to the sound system Figure 1 Loudspeaker 141 in the embodiment shown.
[0091] The vehicle-mounted assisted driving system in the embodiment of the present application may also be referred to as a vehicle computer, a vehicle central control system (abbreviated as central control), etc., without any specific limitation.
[0092] Figure 3 This is a schematic diagram of the structure of the mobile phone 300 according to an embodiment of the present application. It should be understood that: Figure 3 The illustrated cell phone 300 is merely an example, and the cell phone 300 may have more or fewer components than shown, may combine two or more components, or may have a different configuration of components. Figure 3 The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits. It should be understood that the embodiments of the present application are described using a mobile phone as an example, but the structure is also applicable to other terminal devices, such as a tablet computer.
[0093] The mobile phone 300 may include: a processor 310, an external memory interface 320, an internal memory 321, a universal serial bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 342, an antenna 1, an antenna 2, a mobile communication module 350, a wireless communication module 360, an audio module 370, a speaker 370A, a receiver 370B, a microphone 370C, an earphone interface 370D, a sensor module 380, a button 390, a motor 391, an indicator 392, a camera 393, a display screen 394, and a subscriber identification module (SIM) card interface 395, etc. The sensor module 380 may include a pressure sensor 380A, a gyroscope sensor 380B, an air pressure sensor 380C, a magnetic sensor 380D, an acceleration sensor 380E, a distance sensor 380F, a proximity light sensor 380G, a fingerprint sensor 380H, a temperature sensor 380J, a touch sensor 380K, an ambient light sensor 380L, a bone conduction sensor 380M, etc.
[0094] The processor 310 may include one or more processing units. For example, the processor 310 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0095] The controller may be the nerve center and command center of the mobile phone 300. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0096] Processor 310 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 310 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 310. If processor 310 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 310 latency, and thus improves system efficiency.
[0097] In some embodiments, the processor 310 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0098] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 310 may include multiple I2C busses. The processor 310 may be coupled to the touch sensor 380K, the charger, the flash, the camera 393, and the like via different I2C bus interfaces. For example, the processor 310 may be coupled to the touch sensor 380K via the I2C interface, enabling communication between the processor 310 and the touch sensor 380K via the I2C bus interface, thereby enabling the touch function of the mobile phone 300.
[0099] The I2S interface can be used for audio communication. In some embodiments, the processor 310 can include multiple I2S buses. The processor 310 can be coupled to the audio module 370 via the I2S bus to enable communication between the processor 310 and the audio module 370. In some embodiments, the audio module 370 can transmit audio signals to the wireless communication module 360 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.
[0100] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 370 and the wireless communication module 360 can be coupled via a PCM bus interface. In some embodiments, the audio module 370 can also transmit audio signals to the wireless communication module 360 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0101] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 310 and the wireless communication module 360. For example, the processor 310 communicates with the Bluetooth module in the wireless communication module 360 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 370 can transmit audio signals to the wireless communication module 360 via the UART interface, enabling the function of playing music through Bluetooth headphones.
[0102] The MIPI interface can be used to connect the processor 310 to peripheral devices such as the display 394 and the camera 393. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 310 and the camera 393 communicate via the CSI interface to enable the camera function of the mobile phone 300. The processor 310 and the display 394 communicate via the DSI interface to enable the display function of the mobile phone 300.
[0103] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 310 to the camera 393, the display 394, the wireless communication module 360, the audio module 370, the sensor module 380, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0104] USB interface 330 is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. USB interface 330 can be used to connect a charger to charge mobile phone 300 and to transfer data between mobile phone 300 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect to other user devices, such as AR devices.
[0105] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present application is merely an illustrative illustration and does not constitute a structural limitation on the mobile phone 300. In other embodiments of the present application, the mobile phone 300 may also adopt a different interface connection method from the above embodiment, or a combination of multiple interface connection methods.
[0106] The charging management module 340 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 340 can receive charging input from the wired charger via the USB interface 330. In some wireless charging embodiments, the charging management module 340 can receive wireless charging input via the wireless charging coil of the mobile phone 300. While charging the battery 342, the charging management module 340 can also provide power to the user device via the power management module 341.
[0107] The power management module 341 is used to connect the battery 342, the charging management module 340, and the processor 310. The power management module 341 receives input from the battery 342 and / or the charging management module 340 and provides power to the processor 310, the internal memory 321, the external memory, the display 394, the camera 393, and the wireless communication module 360. The power management module 341 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 341 can also be set in the processor 310. In other embodiments, the power management module 341 and the charging management module 340 can also be set in the same device.
[0108] The wireless communication function of the mobile phone 300 can be implemented through the antenna 1, the antenna 2, the mobile communication module 350, the wireless communication module 360, the modem processor and the baseband processor.
[0109] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 300 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0110] The mobile communication module 350 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the mobile phone 300. The mobile communication module 350 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 350 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 350 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 350 can be set in the processor 310. In some embodiments, at least some of the functional modules of the mobile communication module 350 can be set in the same device as at least some of the modules of the processor 310.
[0111] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 370A, the receiver 370B, etc.) or displays an image or video through the display screen 394. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 310 and be provided in the same device as the mobile communication module 350 or other functional modules.
[0112] The wireless communication module 360 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the mobile phone 300. The wireless communication module 360 can be one or more devices that integrate at least one communication processing module. The wireless communication module 360 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 310. The wireless communication module 360 can also receive the signal to be sent from the processor 310, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0113] In some embodiments, antenna 1 of mobile phone 300 is coupled to mobile communication module 350, and antenna 2 is coupled to wireless communication module 360, so that mobile phone 300 can communicate with a network and other devices via wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0114] Mobile phone 300 implements display functionality through a GPU, display screen 394, and an application processor. The GPU is a microprocessor for image processing that connects display screen 394 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 310 may include one or more GPUs that execute program instructions to generate or modify display information.
[0115] Display screen 394 is used to display images, videos, etc. Display screen 394 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, mobile phone 300 may include one or N display screens 394, where N is a positive integer greater than one.
[0116] The mobile phone 300 can realize the shooting function through the ISP, camera 393, video codec, GPU, display 394 and application processor.
[0117] The ISP processes data fed back by camera 393. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 393.
[0118] The camera 393 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the mobile phone 300 may include 1 or N cameras 393, where N is a positive integer greater than 1.
[0119] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the mobile phone 300 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0120] Video codecs are used to compress or decompress digital video. Mobile phone 300 may support one or more video codecs. This allows mobile phone 300 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0121] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications such as image recognition, face recognition, speech recognition, and text comprehension in the mobile phone 300.
[0122] The external memory interface 320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 300. The external memory card communicates with the processor 310 via the external memory interface 320 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0123] The internal memory 321 can be used to store computer executable program codes, which include instructions. The processor 310 executes various functional applications and data processing of the mobile phone 300 by running the instructions stored in the internal memory 321. The internal memory 321 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the mobile phone 300 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 321 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0124] The mobile phone 300 can implement audio functions such as music playback and recording through the audio module 370, speaker 370A, receiver 370B, microphone 370C, headphone jack 370D, and application processor.
[0125] The audio module 370 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 370 can also be used to encode and decode audio signals. In some embodiments, the audio module 370 can be provided in the processor 310, or some functional modules of the audio module 370 can be provided in the processor 310.
[0126] The speaker 370A, also called a "horn," is used to convert audio electrical signals into sound signals. The mobile phone 300 can listen to music or make hands-free calls through the speaker 370A.
[0127] The receiver 370B, also called the "earpiece", is used to convert audio electrical signals into sound signals. When the mobile phone 300 receives a call or a voice message, the voice can be heard by placing the receiver 370B close to the human ear.
[0128] Microphone 370C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 370C to input the sound signal into the microphone 370C. The mobile phone 300 can be provided with at least one microphone 370C. In other embodiments, the mobile phone 300 can be provided with two microphones 370C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the mobile phone 300 can also be provided with three, four or more microphones 370C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0129] The headphone jack 370D is used to connect a wired headphone. The headphone jack 370D can be a USB interface 330 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0130] Pressure sensor 380A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 380A can be located on display screen 394. There are many types of pressure sensors 380A, such as resistive, inductive, and capacitive. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When force is applied to pressure sensor 380A, the capacitance between the electrodes changes. Mobile phone 300 determines the intensity of the pressure based on this change in capacitance. When a touch operation is applied to display screen 394, mobile phone 300 detects the touch intensity based on pressure sensor 380A. Mobile phone 300 can also calculate the touch location based on the detection signal from pressure sensor 380A. In some embodiments, touch operations applied to the same touch location but with different touch intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, a command to view short messages is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to a short message application icon, a command to create a new short message is executed.
[0131] The gyroscope sensor 380B can be used to determine the motion posture of the mobile phone 300. In some embodiments, the angular velocity of the mobile phone 300 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 380B. The gyroscope sensor 380B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 380B detects the angle of the mobile phone 300 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the mobile phone 300 through reverse movement to achieve anti-shake. The gyroscope sensor 380B can also be used for navigation and somatosensory game scenes.
[0132] The air pressure sensor 380C is used to measure air pressure. In some embodiments, the mobile phone 300 calculates the altitude using the air pressure value measured by the air pressure sensor 380C to assist in positioning and navigation.
[0133] The magnetic sensor 380D includes a Hall effect sensor. The mobile phone 300 can use the magnetic sensor 380D to detect the opening and closing of the flip case. In some embodiments, if the mobile phone 300 is a flip phone, the mobile phone 300 can detect the opening and closing of the flip cover using the magnetic sensor 380D. Based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.
[0134] Accelerometer 380E detects the magnitude of acceleration of the phone 300 in all directions (generally three axes). When the phone 300 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the user's device posture, enabling applications such as switching between landscape and portrait modes and pedometers.
[0135] Distance sensor 380F is used to measure distance. Mobile phone 300 can measure distance using infrared or laser. In some embodiments, when shooting a scene, mobile phone 300 can use distance sensor 380F to measure distance to achieve fast focus.
[0136] The proximity light sensor 380G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The mobile phone 300 emits infrared light outward through the light emitting diode. The mobile phone 300 uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the mobile phone 300. When insufficient reflected light is detected, the mobile phone 300 can determine that there is no object near the mobile phone 300. The mobile phone 300 can use the proximity light sensor 380G to detect when the user holds the mobile phone 300 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 380G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.
[0137] Ambient light sensor 380L senses ambient light levels. Mobile phone 300 can adaptively adjust the brightness of display screen 394 based on the perceived ambient light level. Ambient light sensor 380L can also automatically adjust white balance when taking photos. Ambient light sensor 380L can also work with proximity sensor 380G to detect whether mobile phone 300 is in a pocket, preventing accidental touches.
[0138] The fingerprint sensor 380H is used to collect fingerprints. The mobile phone 300 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.
[0139] Temperature sensor 380J is used to detect temperature. In some embodiments, the mobile phone 300 uses the temperature detected by temperature sensor 380J to implement a temperature management strategy. For example, when the temperature reported by temperature sensor 380J exceeds a threshold, the mobile phone 300 reduces the performance of a processor located near temperature sensor 380J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature falls below another threshold, the mobile phone 300 heats the battery 342 to prevent the low temperature from causing an abnormal shutdown of the mobile phone 300. In other embodiments, when the temperature falls below yet another threshold, the mobile phone 300 boosts the output voltage of the battery 342 to prevent an abnormal shutdown caused by the low temperature.
[0140] The touch sensor 380K is also called a "touch panel." The touch sensor 380K can be disposed on the display screen 394. The touch sensor 380K and the display screen 394 form a touch screen, also called a "touch screen." The touch sensor 380K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 394. In other embodiments, the touch sensor 380K can also be disposed on the surface of the mobile phone 300, in a location different from that of the display screen 394.
[0141] The bone conduction sensor 380M can acquire vibration signals. In some embodiments, the bone conduction sensor 380M can acquire vibration signals from the vibrating bones of the human body's vocal cords. The bone conduction sensor 380M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 380M can also be set in headphones to form bone conduction headphones. The audio module 370 can parse out voice signals based on the vibration signals of the vibrating bones of the vocal cords acquired by the bone conduction sensor 380M to implement voice functions. The application processor can parse heart rate information based on the blood pressure signals acquired by the bone conduction sensor 380M to implement heart rate detection functions.
[0142] Keys 390 include a power button, a volume button, and the like. Keys 390 may be mechanical keys or touch-sensitive keys. Mobile phone 300 may receive key inputs and generate key signal inputs related to user settings and function control of mobile phone 300.
[0143] Motor 391 can generate vibration prompts. Motor 391 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 394, motor 391 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0144] Indicator 392 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0145] The SIM card interface 395 is used to connect a SIM card. The SIM card can be connected to and disconnected from the mobile phone 300 by inserting it into or removing it from the SIM card interface 395. The mobile phone 300 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 395 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 395 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 395 can also be compatible with different types of SIM cards. The SIM card interface 395 can also be compatible with external memory cards. The mobile phone 300 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the mobile phone 300 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the mobile phone 300 and cannot be separated from the mobile phone 300.
[0146] It is understood that the structures illustrated in the embodiments of the present application do not constitute specific limitations on the terminal device. In other embodiments of the present application, the terminal device may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0147] It should be noted that the above exemplary description of the equipment to which the technical solutions of the embodiments of the present application are applicable includes vehicles (especially the central control in the vehicle) and mobile phones, but this does not constitute a limitation on the equipment to which the technical solutions of the embodiments of the present application are applicable, and the embodiments of the present application do not make specific limitations on this.
[0148] Figure 4 This is a block diagram of the software structure used by the device in the embodiment of the present application.
[0149] The device's layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other via software interfaces. In some embodiments, the Android system is divided into four layers: from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0150] The application layer can include a series of application packages.
[0151] like Figure 4 As shown, the application package may include applications such as calls, maps, navigation, WLAN, Bluetooth, music, video, and assisted driving.
[0152] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0153] like Figure 4 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0154] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0155] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, maps, audio, calls made and received, etc.
[0156] The view system includes visual controls, such as those that display text and images. The view system is used to build applications. A user interface can consist of one or more views. For example, a user interface for a notification icon might include a view that displays text and a view that displays an image.
[0157] The phone manager is used to provide communication functions for the device, such as managing call status (including answering, hanging up, etc.).
[0158] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0159] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically, without driver interaction. For example, the Notification Manager can be used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the system's top status bar as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, audible alerts, and flashing indicators.
[0160] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0161] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0162] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0163] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0164] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0165] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0166] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0167] A 2D graphics engine is a drawing engine for 2D drawings.
[0168] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0169] It is understandable that Figure 4 The components included in the system framework layer, system library and runtime layer shown do not constitute a specific limitation on the device. In other embodiments of the present application, the device may include more or fewer components than shown, or combine some components, split some components, or arrange the components differently.
[0170] The present application provides a soundscape function that can create a soundscape that meets the needs of users in various application scenarios, allowing users to achieve an immersive experience. The various application scenarios may include:
[0171] (1) During driving, the vehicle's sensors collect video and audio to create a video log (vlog), such as a self-driving tour vlog. In this application scenario, the sensors installed on the vehicle can record high-quality video materials, but cannot record high-quality sound materials (for example, the sound source is too far away or the sound is too noisy to achieve good pickup of high-quality sound materials, or it is limited by the microphone's capabilities, etc.). This requires the user to retrieve matching sound effect materials from the sound effect database and merge the video material and sound effect material to create a video log (vlog) that matches the picture and sound. However, the retrieval method is inefficient and the content length, timbre, and sound position of the retrieved audio are difficult to align with the video content, which cannot meet the audio-visual needs of a high degree of fit between the soundscape and the video picture.
[0172] (2) In the driving / parking state, in a confined space, play sounds to create a soundscape that matches the user's instructions. For example, play the sound effects of a forest environment to create a soundscape that matches the forest environment in which the vehicle is located, allowing the user to achieve an immersive experience. In this application scenario, based on the user's instructions, matching sound materials are retrieved from the system's pre-made ambient sounds for playback. However, the system's pre-made ambient sounds are limited in variety and cannot match user-defined scenes.
[0173] (3) When users create short videos, they need to dub the edited videos. In this application scenario, the video material and the sound effect material are separate, which requires users to retrieve matching sound effect materials from the sound effect database and merge the video material and sound effect material to create a short video with matching images and sounds. However, the same problem exists as when making a vlog for a vehicle: the retrieval method is inefficient and it is difficult to align the content length, timbre, and sound position of the retrieved audio with the video content, which cannot meet the audio-visual needs of a high degree of fit between the soundscape and the video image.
[0174] (4) Generate a dubbing based on the text content. In this application scenario, the text content is first semantically recognized to find the sound effect objects that can produce sound. Then, matching sound materials are retrieved from the system's pre-made sound library to create a soundscape that matches the sound effect objects, allowing the text to produce sound and achieve a more vivid expression. However, the system's pre-made sound library is limited in variety and cannot achieve a good match with the diverse sound effect objects in the text.
[0175] In addition to the several application scenarios and the technical problems described above, there are also many soundscape demand scenarios, and they also have similar problems. In order to solve these technical problems, the embodiments of the present application propose an audio generation method and device.
[0176] Figure 5Flowchart of process 500 of the audio generation method provided in an embodiment of the present application. Process 500 can be performed by the device described above, for example, vehicle 100 (especially the vehicle computer in the vehicle) or mobile phone 300. Process 500 is described as a series of steps or operations. It should be understood that process 500 can be performed in various orders and / or occur simultaneously, not limited to Figure 5 The execution order shown is as follows. Process 500 may include:
[0177] Step 501: In response to a user instruction, obtain a first script.
[0178] Optionally, the user instruction can be a voice instruction input through the microphone. For example, the user says "Hello Xiaoyi, help me record this journey", "Hello Xiaoyi, play a sound effect of a forest environment", etc. The device has semantic recognition of voice, so when the preset trigger voice (for example, "Hello Xiaoyi") is recognized, the subsequent voice can be recognized to extract keywords (for example, record this journey, sound effects of a forest environment) to perform corresponding operations.
[0179] Optionally, the user command may be a control command input through the central control screen. For example, the central control provides an interactive interface for the soundscape. The user can click a corresponding button or icon on the central control screen to open the interactive interface, and then select a soundscape scene, input requirements, etc. on the interactive interface. It should be understood that the interactive interface for the soundscape can provide a variety of ways for the user to input commands, including but not limited to menu selection, control triggering, text input, etc., and the layout and content of the interactive interface can also be designed in a variety of ways, which are not specifically limited in the embodiments of the present application.
[0180] It should be noted that in addition to the two methods of obtaining user commands mentioned above, other methods can also be used to obtain user commands, for example, through a car-machine application installed on a mobile phone and a communication link between the mobile phone and the car-machine, to control the soundscape function on the car-machine. Therefore, the embodiment of the present application does not specifically limit the method of obtaining user commands.
[0181] The first script can be a sound effect script, and then the audio signal can be edited based on the first script to obtain the expected sound effect. In the embodiment of the present application, the first script can be obtained by the following two methods:
[0182] 1. Generate a first scenario based on multimodal information collected by multiple sensors.
[0183] Multimodal information includes at least one of the following: audio information collected by a microphone; video information collected by a camera; or control information obtained by the central control screen, where the control information includes at least one of text, images, video, or audio. Sensors can be various sensors installed on a device (such as a vehicle or a mobile phone), including but not limited to cameras, microphones, touch screens, radars, etc.
[0184] In an embodiment of the present application, semantic recognition can be performed on audio, video, text, etc. to obtain keywords in the aforementioned carriers, and corresponding sound effect objects can be obtained based on the keywords, thereby generating a first script.
[0185] Exemplarily, when the multimodal information is text, the keywords contained in the text can provide a sound effect object, based on which the attribute information of the sound effect object such as its corresponding virtual duration, virtual position, virtual volume, etc. is predicted, and these attribute information can constitute the first script.
[0186] Alternatively, when the multimodal information is audio, semantic recognition is performed on the audio, and the obtained keywords can provide a sound effect object. Based on the sound effect object, its corresponding attribute information such as real duration, real position, real volume, and real timbre is predicted. These attribute information can constitute the first script.
[0187] Alternatively, when the multimodal information is a video / image, semantic recognition is performed on the video / image, and the obtained keywords can provide a sound effect object. Based on the sound effect object, its corresponding real duration, real position and other attribute information of the sound effect object are predicted. These attribute information can constitute the first script.
[0188] In an embodiment of the present application, the first script may include descriptive information of at least one sound effect object (for example, the identification, text or description of a cat, dog, bird, person, etc.). In addition, the first script may also include attribute information of at least one sound effect object (including at least one of volume, duration, timbre, and may also include position). This information can guide the generation of the target sound effect.
[0189] Table 1
[0190] Audio Type Logo Location volume Duration Sound Effect Object Description Sub-sound effects 0 background -40 20 xxx Sub-sound effects 1 prospect -10 10 yyy …… …… …… …… …… …… Sub-sound effects N prospect -15 15 zzz
[0191] For example, referring to Table 1, the first script describes a sound effect script, including N+1 sound effect objects, namely sub-sound effects identified as 0-N+1, and each sub-sound effect describes its attribute information, including position, volume, duration and sound effect object description.
[0192] It should be noted that the first script can be described in a variety of ways, and the content included is related to the application scenario and the information obtained. The embodiment of the present application does not specifically limit the content and description method of the first script.
[0193] 2. Get the pre-set first script.
[0194] In the embodiment of the present application, in order to improve efficiency, first scripts for various scenes can be pre-set, for example, a first script for a rest scene, a first script for a forest atmosphere scene, a first script for an opera house scene, etc. Attribute information is set for the sound effect objects that may appear in the scene, and then the first script is generated based on this attribute information and stored in the memory.
[0195] An interactive interface for scene selection is provided to the user on the central control screen. When the user selects a scene on the interactive interface, the first script corresponding to the scene can be called from the memory.
[0196] In a possible implementation, the device may further obtain a second script, where the second script is a video script generated based on visual information, where the visual information includes at least one of video, image, or text.
[0197] The device can acquire visual information through a camera (for example, sensors around the vehicle can capture video, a mobile phone camera can capture video, a dedicated video camera can capture video, etc.) or through a user's touch operation on the central control screen (for example, a user operates on the central control screen to select a local video, a user operates on the central control screen to download a video from the Internet, etc.). It should be understood that while the above example uses the acquisition of video as an example, the above method can also be used to acquire images or text.
[0198] In an embodiment of the present application, semantic recognition can be performed on videos, images, texts, etc. to obtain keywords in the aforementioned carriers, and corresponding sound effect objects can be obtained based on the keywords, thereby generating a second script.
[0199] Exemplarily, when the multimodal information is text, the keywords contained in the text can provide a sound effect object, based on which the attribute information of the sound effect object such as its corresponding virtual duration and virtual position is predicted, and these attribute information can constitute the second script.
[0200] Alternatively, when the multimodal information is a video / image, semantic recognition is performed on the video / image, and the obtained keywords can provide a sound effect object. Based on the sound effect object, its corresponding real duration, real position and other attribute information of the sound effect object are predicted. These attribute information can constitute the second script.
[0201] In an embodiment of the present application, the second script may include descriptive information of at least one sound effect object (for example, an identifier, text, or description of a cat, dog, bird, or person, etc.). In addition, the second script may also include attribute information of at least one sound effect object (including at least one of duration and position), which can guide the generation of the target video. The second script may also refer to the description method of Table 1, without specific limitation. Compared with the first script, the attribute information of the sound effect object in the second script does not include volume and timbre.
[0202] In an embodiment of the present application, the first script can be used as a script for sound effects to guide the generation of target sound effects, and the second script can be used as a script for videos to guide the generation of target videos. In addition, since the second script also contains the attribute information of the sound effect object, the second script can be referenced when generating the target sound effect to generate a target sound effect that matches the sound effect object in the second script. For example, the duration and volume of the target sound effect match the duration and volume of the sound effect object in the second script, respectively, the timbre of the target sound effect matches the sound effect object in the second script, and the position of the target sound effect matches the position where the sound effect object appears in the second script, so that the sound effect and video are highly consistent, enhancing the user's audio-visual experience. Optionally, the second script can serve as a supplement to the first script. When a sound effect object that is not in the first script appears in the second script, the sound of the sound effect object can be added to the target sound effect to improve the soundscape effect of the target sound effect.
[0203] For example, the user inputs "play the sound effects of the forest environment" by voice. At this time, the corresponding sound effects (for example, the sound of wind, rain, birds, branches, etc.) are obtained according to the sound effect objects that may exist in the forest environment (for example, wind, rain, birds, branches, etc.). At this time, only the first script can be generated. The first script can include a combination of the aforementioned multiple sound effects, including the sound effect timestamp, spatial position, etc.
[0204] For another example, a user inputs a text by voice, and based on text recognition, it is found that the text contains a sound effect object. Sound content that matches the sound effect object can be synchronously added. At this time, a first script and a second script can be generated. The first script can include the timestamp and timbre information of the sound content that is added when the sound effect object appears, and the second script can include the sound effect object in the text and its timestamp, etc.
[0205] For another example, when dubbing a video, in addition to adding one or more background music that fits the picture, when a sound effect object (for example, fireworks, puppies, people, etc.) appears in the picture, sound content that matches the sound effect object (for example, the sound of fireworks bursting, the barking of puppies, the sound of dialogue that distinguishes the timbre of characters, etc.) can be synchronously added. At this time, a first script and a second script can be generated. The first script may include information about the aforementioned background music, the timestamp and timbre information of the sound content that is matched with the appearance of the sound effect object, etc. The second script may include the sound effect object in the picture, its timestamp, spatial position, etc.
[0206] In a possible implementation, the first script may meet the preset conditions, including the following situations:
[0207] (1) At least one keyword included in the user instruction. The preset condition satisfied by the first script can come from the above-mentioned user instruction. That is, when the user enters the user instruction, the user can include the requirement in the instruction, for example, voice input "record this journey" and text input "play a sound effect of a forest environment". After the device performs semantic analysis on the user instruction and recognizes the keyword in the instruction, the keyword can be used as the condition satisfied by the first script (preset condition).
[0208] (2) At least one timbre information included in the sound collected by the microphone. The timbre information included in the sound collected by the microphone can also be used as a condition (preset condition) satisfied by the first script. For example, in a conversation between a man and a woman, in order to make the soundscape more realistic, the difference between male and female voices can also be reflected in the target sound effect. In this case, extracting timbre information from the sound collected by the microphone can assist in the generation of the first script.
[0209] (3) At least one sound effect object described in the second script. The sound effect object in the second script is a physical object that emits sound. The sound effect corresponding to it can be added to the first script. For example, if the second script describes the explosion of fireworks, the sound effect of the fireworks explosion can be included in the first script. In this way, the target sound effect includes the sound effect of the fireworks explosion synchronized with the fireworks explosion in the picture, achieving a dual experience of visual and soundscape. Therefore, the sound effect object can also serve as a condition (preset condition) satisfied by the first script.
[0210] It should be noted that, in addition to the above three preset conditions, the first script may also meet other preset conditions, and the embodiment of the present application does not make specific limitations on this.
[0211] Step 502: Obtain target sound effects according to the first script.
[0212] In a possible implementation, there is a first script but no second script. A respective sub-sound effect can be generated for at least one sound effect object appearing in the first script to constitute a target sound effect.
[0213] For example, a first script generated based on text describes N sound effect objects, where sound effect object A is a dog, appears from the 2nd to the 5th second, for a total of 3 seconds, and is located on the left side. The AIGC model can convert the dog into a barking sound, ultimately generating the barking sound. This barking sound can constitute sub-sound effect A in the target sound effect, corresponding to the dog; the aforementioned appearance time indicates the temporal position of the barking sound within the target sound effect; this position can be used in subsequent rendering processing to give the user the impression that the three-second barking sound is occurring on the left side.
[0214] In one possible implementation, there are a first script and a second script. A union of at least one sound effect object appearing in the first script and at least one sound effect object appearing in the second script may be obtained; and sub-sound effects may be generated for the multiple sound effect objects in the union to form a target sound effect.
[0215] For example, in addition to the first script generated based on text, the second script describes M sound effect objects, and L sound effect objects are not included in the first script. Among them, sound effect object B is thunder, which appears from the 6th to the 8th second, for a total of 2 seconds, and is located above. The AIGC model can realize the conversion from thunder to thunder sound, and finally generate the thunder sound. The thunder sound can constitute the sub-sound effect B in the target sound effect, corresponding to thunder; the aforementioned appearance time indicates the time position of the thunder sound in the target sound effect; the aforementioned position can be used for subsequent rendering processing, so that the user hears the two-second thunder sound as if it is happening above. At this time, the user will hear the barking of a dog on the left from the 2nd to the 5th second, and the thunder sound from above from the 6th to the 8th second, accompanied by a thunder picture.
[0216] Step 503: Perform spatial audio rendering on the target sound effect to obtain target audio.
[0217] The device can perform spatial audio rendering on the target sound effect to produce the target audio. For example, in the first script, there is a sound effect object of a dog, and its corresponding position is on the left. After spatial audio rendering, the dog's barking sound will be perceived as being on the left.
[0218] In the embodiment of the present application, by generating a first script for sound effects and generating target sound effects based on the guidance of the script, a soundscape that meets the user's expectations can be obtained, thereby improving the user's listening experience.
[0219] In one possible implementation, after acquiring the target audio, the target audio can be played to the user via a speaker. If the user finds any modifications needed during the listening process, they can enter a modification instruction. Accordingly, the device receives the user's modification instruction; modifies the first script in response to the modification instruction to obtain a third script; acquires the modified target sound effect based on the third script; and acquires the modified target audio based on the modified target sound effect.
[0220] The difference from the first script generated above is that this embodiment can add factors to modify the instructions (such as keywords in the modification instructions) on the basis of the first script, thereby obtaining a third script. For example, the first script is generated based on the instruction of "sound effects of the forest environment", and the user enters the modification instruction of "adding children's playful sounds". At this time, the third script adds the sound effect information of children's playful sounds on the basis of the first script. After obtaining the third script, the device again generates a modified target sound effect based on the third script. Compared with the target sound effect before modification, the modified target sound effect adds the sound effect of children's playful sounds, and then obtains the modified target audio based on the modified target sound effect and plays it for the user.
[0221] The above description of the process of a user modifying a created soundscape is intended to illustrate that the embodiment of the present application supports modification of the soundscape, realizes the personalized customization function of the soundscape, and makes the soundscape function more in line with the user's audio-visual needs.
[0222] In one possible implementation, after obtaining the second script, the device can edit the visual information according to the second script to obtain an edited video; and fuse the edited video and the target sound effect to obtain the target video, wherein the timestamp of the sound effect in the target video is synchronized with the timestamp of the sound effect object and conforms to the sound characteristics of the sound effect object.
[0223] In an embodiment of the present application, in addition to generating the target sound effect, the captured video can be edited, and then the edited video and the target sound effect can be fused to obtain the target video. The picture and sound effect of the target video are synchronously matched, that is, the final video is generated at one time, including the picture and the sound effect matching the sound effect object in the picture, the background music matching the picture content, etc., thereby improving the efficiency of video editing.
[0224] The technical solution of this application is described below through several specific embodiments.
[0225] Figure 6 This is a schematic diagram of a personalized rendering system according to an embodiment of the present application. Figure 6 As shown, taking a vehicle as an example, the personalized rendering system consists of four modules: input module, multimodal module, sound synthesis module and spatial rendering module.
[0226] The input module includes multimodal sensors such as a panoramic camera, microphone, and Global Positioning System (GPS) installed on the outside of the vehicle body, as well as an interactive interface provided by a touch screen. The input module can collect multimodal information, which includes at least one of the following: audio information collected by a microphone; or video information collected by a camera; or control information obtained by the central control screen, where the control information includes at least one of text, images, video, or audio. The sensor can be any type of sensor installed on a device (such as around the vehicle or on a mobile phone), including but not limited to a camera, microphone, touch screen, radar, etc.
[0227] The multimodal module is one of the main modules of the personalized rendering system, which is equivalent to the "brain". It analyzes and interprets the above multimodal information and outputs the sound effect script (corresponding to the above first script). The content of the sound effect script can be referred to Table 1. The sound effect script describes multiple sub-sound effects, which are identified as 0-N+1, the layout of multiple sub-sound effects (including background, foreground, etc.), volume, duration, and a description of the sound effect object corresponding to the sub-sound effect.
[0228] The sound effect synthesis module can translate the configurations of the multiple sub-sound effects in Table 1 into sound effect objects with corresponding attributes, and obtain mono audio of the multiple sound effect objects.
[0229] The spatial rendering module can mix the mono audio of multiple sound effect objects into multi-channel audio, namely the target audio, through audio rendering technology.
[0230] Based on the above process, the multi-directional speakers in the car decode and play the target audio, creating a soundscape in the car.
[0231] In addition, the user can input user instructions, which can act on the multimodal information to affect the generation of the sound effect script, or act on the sound effect script to modify it. The specific process can be referred to the above description and will not be repeated here.
[0232] Figure 7 This is a schematic diagram of the interactive interface of the soundscape of the embodiment of the present application, such as Figure 7 As shown, the interactive interface includes controls for two soundscape modes, namely, controls corresponding to journey Vlog and free space, and users can select the corresponding soundscape mode by clicking the controls.
[0233] Figure 8 This is a structural diagram of the personalized rendering system of an embodiment of the present application, such as Figure 8 As shown, this embodiment is a personalized rendering system of the soundscape inside the vehicle when the vehicle is in motion.
[0234] Multimodal information I includes information collected by sensors during vehicle operation, such as cameras, GPS, microphones, and other sensors. Multimodal information I can be processed locally or uploaded to the cloud.
[0235] 1. After starting the vehicle, the user enters the soundscape interface and clicks the "Travel Vlog" control. Based on this operation, the camera, microphone and other sensors on the vehicle body start working and retain the collected multimodal information I until the soundscape rendering function ends;
[0236] 2. During the driving process, the vehicle can obtain user instructions through voice interaction and edit the video content without touching the central control screen. User instructions include I v1 ,…,I vn , for example, "Hello Xiaoyi, please help me record the journey I just took";
[0237] 3. When the user stops the vehicle or ends the soundscape rendering function, the multimodal module extracts the sound effect object and its spatial location information R from the video information collected by the camera. v ; The multimodal module extracts the timbre information of the original sound effect from the audio information collected by the microphone a ;
[0238] 4. Multimodal module according to R v , and user instructions I v1 ,…,I vn (User instructions are optional), generate image script S v (corresponding to the second scenario above); according to R a , and user instructions I v1 ,…,I vn (User instructions are optional), generate sound effect script S a (corresponding to the first script above);
[0239] 5. Video editing model X based on image script S v The video clips captured by the camera are processed to obtain the vlog video V;
[0240] 6. Audio synthesis model Y based on sound effect script S a and Image Script S v Generate target sound effects, which include sound effects A1, ..., A1 of multiple sound effect objects. n ;
[0241] 7. The spatial rendering module transforms the sound effects A1, ..., A n According to the sound script S a Mix the audio to get the vlog audio A;
[0242] 8. Combine the video V and audio A to obtain vlog.
[0243] In this embodiment, a vehicle in motion collects signals through sensors and inputs these signals and user commands into an interactive, personalized rendering system, which automatically edits a journey vlog with realistic sound effects. This process restores and processes the original audio, allowing for interactive editing of sound effects, spatial positions, and other aspects of the vlog audio during or after the process. The final video is generated in one go, including the frame, sound effects matching the sound effects, and background music matching the content, thereby improving video editing efficiency.
[0244] Figure 9 This is a structural diagram of the personalized rendering system of an embodiment of the present application, such as Figure 9 As shown, this embodiment is a personalized rendering system of the soundscape inside the vehicle when the vehicle is stopped.
[0245] The system collects audio information I through the microphone a And user instructions I v (For example, "Hello Xiaoyi, help me generate a beach environment") Generate target sound effects.
[0246] 1. The user enters the soundscape interface in a stationary car and clicks the "Free Space" control. Based on this operation, the microphone and other sensors on the car body start working and transmit the collected audio information to the a Retained until the soundscape rendering function is completed;
[0247] 2. Without touching the central control screen, the car computer can obtain user instructions through voice interaction and edit the audio content. User instructions include I v1 ,…,I vn , for example, "Hello Xiaoyi, help me generate a beach environment";
[0248] 3. The multimodal module collects audio information from the microphone I a Extract sound effect object R i ;
[0249] 4. Multimodal module according to R i , and user instructions I v1 ,…,I vn (User instructions are optional), generate sound effect script S a (corresponding to the first script above);
[0250] 5. Audio generation model Y based on sound effect script S a Generate target sound effects, which include sound effects A1, ..., A1 of multiple sound effect objects. n ;
[0251] 6. The spatial rendering module transforms the sound effects A1, ..., A n According to the sound script S a Target audio A obtained by mixing;
[0252] 7. The user reserves the target audio A. And according to the feedback of playing the target audio A, input the modification instruction I a 'To modify the sound effect script S a ;
[0253] 8. The audio generation model Y generates a modified sound effect based on the modified sound effect script and returns to step 7 to implement multiple modifications to the audio.
[0254] In this embodiment, a parked vehicle collects audio signals through a microphone and inputs the audio signals and user commands into an interactive personalized rendering system. The system can generate user-defined sound effects and allow users to modify the generated sound effects, realizing the personalized customization function of the soundscape, making the soundscape function more in line with the user's audio-visual needs.
[0255] Figure 10 This is a structural diagram of the audio generating device 1000 of the present application, as shown in FIG. Figure 10 As shown, the audio generation device 1000 of this embodiment can be applied to the above devices, such as vehicles or mobile phones. The audio generation device 1000 may include: an acquisition module 1001, a processing module 1002 and a rendering module 1003.
[0256] The acquisition module 1001 is used to obtain a first script in response to a user instruction, where the first script is a sound effect script; the processing module 1002 is used to obtain a target sound effect according to the first script; and the rendering module 1003 is used to perform spatial audio rendering on the target sound effect to obtain a target audio.
[0257] In a possible implementation, the acquisition module 1001 is specifically configured to generate the first script based on multimodal information, where the multimodal information is collected by a variety of sensors; or to acquire a preset first script.
[0258] In one possible implementation, the multimodal information includes at least one of the following information:
[0259] Audio information collected through a microphone; or
[0260] Video information collected by a camera; or
[0261] The control information obtained through the central control screen includes at least one of text, image, video or audio.
[0262] In a possible implementation, the first script includes at least one sound effect object, where the sound effect object is a physical object that emits a sound effect.
[0263] In a possible implementation, the first script further includes attribute information of each of the at least one sound effect objects, where the attribute information includes at least one of volume, duration, or timbre.
[0264] In one possible implementation, the first script also includes the respective positions of the at least one sound effect object; the rendering module is specifically configured to perform spatial audio rendering on the target sound effect according to the respective positions of the at least one sound effect object to obtain the target audio.
[0265] In one possible implementation, the acquisition module 1001 is also used to acquire a second script, which is a video script generated based on visual information, and the visual information includes at least one of video, image or text; the processing module is also used to acquire the target sound effect based on the first script and the second script.
[0266] In a possible implementation, the visual information is collected through a camera or a central control screen.
[0267] In a possible implementation, the second script includes at least one sound effect object, where the sound effect object is a physical object that emits a sound effect in the visual information.
[0268] In a possible implementation, the second script further includes attribute information of each of the at least one sound effect objects, where the attribute information includes duration and / or position.
[0269] In a possible implementation, the processing module 1002 is specifically configured to generate respective sub-sound effects for at least one sound effect object appearing in the first script to constitute the target sound effect.
[0270] In one possible implementation, the processing module 1002 is specifically used to obtain the union of at least one sound effect object appearing in the first script and at least one sound effect object appearing in the second script; and generate respective sub-sound effects for the multiple sound effect objects in the union to constitute the target sound effect.
[0271] In a possible implementation, the user instruction includes at least one of the following instructions:
[0272] Voice commands entered through a microphone; or,
[0273] Control commands entered through the central control screen.
[0274] The device of this embodiment can be used to perform Figure 5 The technical solution of the method embodiment shown has similar implementation principles and technical effects, which will not be repeated here.
[0275] During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of the hardware in the processor or an instruction in the form of software. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware coding processor, or being executed by a combination of hardware and software modules in the coding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0276] The memory mentioned in the above embodiments may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0277] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0278] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0279] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0280] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0281] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0282] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0283] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An audio generation method, characterized in that: include: In response to a user instruction, obtaining a first script, wherein the first script is a sound effect script; Obtaining target sound effects according to the first script; Perform spatial audio rendering on the target sound effect to obtain target audio.
2. The method according to claim 1, characterized in that The obtaining of the first script includes: generating the first scenario based on multimodal information collected by multiple sensors; or Get the pre-set First Script.
3. The method according to claim 2, characterized in that The multimodal information includes at least one of the following information: Audio information collected through a microphone; or Video information collected by a camera; or The control information obtained through the central control screen includes at least one of text, image, video or audio.
4. The method according to any one of claims 1 to 3, characterized in that The first script includes at least one sound effect object, which is a physical object that emits sound effects.
5. The method according to claim 4, characterized in that The first script also includes attribute information of each of the at least one sound effect objects, and the attribute information includes at least one of volume, duration or timbre.
6. The method according to claim 4 or 5, characterized in that The first script further includes a position of each of the at least one sound effect objects; and performing spatial audio rendering on the target sound effect to obtain target audio includes: Perform spatial audio rendering on the target sound effect according to the respective positions of the at least one sound effect object to obtain the target audio.
7. The method according to any one of claims 1 to 6, characterized in that Also includes: Obtaining a second script, where the second script is a video script generated based on visual information, where the visual information includes at least one of a video, an image, or text; The obtaining of the target sound effect according to the first script includes: The target sound effect is obtained according to the first script and the second script.
8. The method according to claim 7, characterized in that The visual information is collected through a camera or a central control screen.
9. The method according to claim 7 or 8, characterized in that The second script includes at least one sound effect object, where the sound effect object is a physical object that emits a sound effect in the visual information.
10. The method according to claim 9, characterized in that The second script further includes attribute information of each of the at least one sound effect object, wherein the attribute information includes duration and / or position.
11. The method according to any one of claims 1 to 6, characterized in that The step of obtaining a target sound effect according to the first script includes: A respective sub-sound effect is generated for at least one sound effect object appearing in the first script to constitute the target sound effect.
12. The method according to any one of claims 7 to 10, characterized in that The obtaining of the target sound effect according to the first script and the second script includes: Obtaining a union of at least one sound effect object appearing in the first script and at least one sound effect object appearing in the second script; Respective sub-sound effects are generated for the plurality of sound effect objects in the union to constitute the target sound effect.
13. The method according to any one of claims 1 to 12, characterized in that The user instruction includes at least one of the following instructions: Voice commands entered through a microphone; or, Control commands entered through the central control screen.
14. An audio generating device, characterized in that: include: An acquisition module, configured to acquire a first script in response to a user instruction, wherein the first script is a sound effect script; a processing module, configured to obtain a target sound effect according to the first script; A rendering module is used to perform spatial audio rendering on the target sound effect to obtain target audio.
15. The device according to claim 14, characterized in that The acquisition module is specifically configured to generate the first script based on multimodal information, where the multimodal information is collected by a variety of sensors; or to acquire a preset first script.
16. The device according to claim 15, characterized in that The multimodal information includes at least one of the following information: Audio information collected through a microphone; or Video information collected by a camera; or The control information obtained through the central control screen includes at least one of text, image, video or audio.
17. The device according to any one of claims 14 to 16, characterized in that The first script includes at least one sound effect object, which is a physical object that emits sound effects.
18. The device according to claim 17, characterized in that The first script also includes attribute information of each of the at least one sound effect objects, and the attribute information includes at least one of volume, duration or timbre.
19. The device according to claim 17 or 18, characterized in that The first script also includes the respective positions of the at least one sound effect object; the rendering module is specifically used to perform spatial audio rendering on the target sound effect according to the respective positions of the at least one sound effect object to obtain the target audio.
20. The device according to any one of claims 14 to 19, characterized in that The acquisition module is further configured to acquire a second script, where the second script is a video script generated based on visual information, where the visual information includes at least one of a video, an image, or a text; The processing module is further used to obtain the target sound effect according to the first script and the second script.
21. The device according to claim 20, characterized in that The visual information is collected through a camera or a central control screen.
22. The device according to claim 20 or 21, characterized in that The second script includes at least one sound effect object, where the sound effect object is a physical object that emits a sound effect in the visual information.
23. The device according to claim 22, characterized in that The second script further includes attribute information of each of the at least one sound effect object, wherein the attribute information includes duration and / or position.
24. The device according to any one of claims 14 to 19, characterized in that The processing module is specifically configured to generate respective sub-sound effects for at least one sound effect object appearing in the first script to constitute the target sound effect.
25. The device according to any one of claims 20 to 23, characterized in that The processing module is specifically used to obtain the union of at least one sound effect object appearing in the first script and at least one sound effect object appearing in the second script; and generate respective sub-sound effects for the multiple sound effect objects in the union to constitute the target sound effect.
26. The device according to any one of claims 14 to 25, characterized in that The user instruction includes at least one of the following instructions: Voice commands entered through a microphone; or, Control commands entered through the central control screen.
27. A device, characterized in that include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 13.
28. A computer-readable storage medium, characterized in that The invention comprises a computer program which, when executed on a computer, causes the computer to perform the method according to any one of claims 1 to 13.
29. A computer program product, characterized in that The computer program product comprises a computer program code, and when the computer program code is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 13.