Audio control method and system, storage medium, electronic equipment and vehicle

By recognizing user commands and adjustments to the audio source location, the system addresses the issues of insufficient spatial awareness and immersion in in-vehicle audio systems, achieving personalized audio playback effects and enhancing the user's immersive listening experience.

CN121742791APending Publication Date: 2026-03-27BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing in-vehicle audio systems fall short in terms of spatiality, immersion, and personalization, making it difficult to meet users' higher demands.

Method used

By recognizing the user's commands, the system controls the audio playback effects of the sound sources in the fused audio, separates the audio data of multiple sound sources from the fused audio, plays the sound data according to the position of the sound sources on the stage and in the space, and adjusts the audio effects using a spatial sound field model and speakers.

Benefits of technology

It enhances the spatial sense and immersion of audio playback, realizes a personalized listening experience, and improves interactivity and entertainment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121742791A_ABST
    Figure CN121742791A_ABST
Patent Text Reader

Abstract

The invention discloses an audio control method and system, a storage medium, electronic equipment and a vehicle. The audio control method comprises the step of controlling an audio playing effect of at least one sound source in fused audio according to a command action of a user. According to the invention, the audio data of the plurality of sound sources can be separated from the fused audio, and the audio data of the plurality of sound sources can be played, so that the playing of the fused audio is realized. Through sound source separation, the position of each sound source and the audio playing effect can be flexibly adjusted, so that a user can feel sound from different directions and different sound sources, the sense of space and immersion of audio playing are enhanced, and the immersive listening experience is realized. Besides, according to the embodiment of the invention, the command action of the user is identified, the audio playing effect of at least one sound source is controlled according to the command action, and the command process of the user is fused during audio playing, so that the user can feel that the user becomes a music commander, the interactivity and entertainment are enhanced, and the personalized sound listening experience is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology, and more particularly to an audio control method, system, storage medium, electronic device, and vehicle. Background Technology

[0002] With the rapid development of smart cockpit technology, the level of vehicle intelligence is constantly improving, and in-vehicle entertainment systems are becoming increasingly sophisticated. In the field of in-vehicle music, users' demand for immersive music experiences is growing. However, most current in-vehicle audio systems are still limited to playing single audio data, and they are significantly lacking in spatial sense, immersion, and personalization, making it difficult to meet users' higher demands. Therefore, how to improve audio playback quality is a technical problem that the industry is currently dedicated to researching. Summary of the Invention

[0003] This application provides an audio control method, system, storage medium, electronic device, and vehicle that enhances the spatial sense and immersion of audio playback and achieves a personalized listening experience, thereby at least partially solving the aforementioned technical problems.

[0004] To achieve the above objectives, according to a first aspect of this application, an audio control method is provided, the method comprising: controlling the audio playback effect of at least one sound source in a fused audio according to a user's command action.

[0005] Optionally, the audio playback effect includes at least one of the following: audio playback speed, audio playback volume, audio surround playback, audio playback pitch, audio playback intensity, and audio playback status.

[0006] Optionally, the command action includes at least one of the following: hand waving action, hand clenching action, clenching and releasing action, and hand rotation action.

[0007] Optionally, the method further includes: displaying a command interface; wherein the command interface is used to display at least one of the following: the command area corresponding to each of the sound sources, and the command actions.

[0008] Optionally, the method further includes: separating audio data of multiple sound sources from the fused audio; and playing the audio data of each sound source according to the first position of each sound source on the stage.

[0009] Optionally, the method further includes: displaying a stage interface; wherein the stage interface is used to display the first position of each of the sound sources on the stage; and updating the first position of the sound source on the stage in response to a touch operation on the sound source.

[0010] Optionally, playing the audio data of each sound source according to its first position on the stage includes: determining the relative position between each sound source and a first listening position based on the first position on the stage; wherein the first listening position corresponds to a first space with the stage; determining the relative position between each sound source and a second listening position based on the relative position between the sound sources and the first listening position; determining the second position of each sound source in a spatial sound field model based on the relative position between the sound sources and the second listening position; wherein the second listening position corresponds to a second space with the spatial sound field model; and playing the audio data of each sound source according to its second position.

[0011] Optionally, the method further includes: displaying a first selection interface for displaying at least one listening position in the first space; and / or displaying a second selection interface for displaying at least one listening position in the second space.

[0012] Optionally, playing the audio data of each of the audio sources according to the second position of each of the audio sources includes: determining the speaker corresponding to each of the audio sources according to the second position of each of the audio sources; and controlling the speaker corresponding to each of the audio sources to play the audio data of each of the audio sources.

[0013] Optionally, separating the audio data of multiple sound sources from the fused audio includes: inputting the fused audio into a sound source separation model so that the sound source separation model outputs the audio data of multiple sound sources.

[0014] According to a second aspect of this application, a computer-readable storage medium is provided having a computer program or instructions stored thereon, which, when executed by a processor, implement the audio control method as described above.

[0015] According to a third aspect of this application, a computer program product is provided, comprising a computer program or instructions that, when executed by a processor, implement the audio control method as described above.

[0016] According to a fourth aspect of this application, an electronic device is provided, comprising: a memory having a computer program or instructions stored thereon; and a processor for executing the computer program or instructions in the memory to implement the audio control method as described above.

[0017] According to a fifth aspect of this application, an audio control system is provided, the audio control system comprising: a controller for: controlling the audio playback effect of at least one sound source in a fused audio according to a user's command.

[0018] Optionally, the audio control system further includes: a data acquisition device connected to the controller; wherein the data acquisition device is used to: send user images to the controller; and the controller is further used to: identify the command actions from the user images.

[0019] Optionally, the audio control system further includes: a speaker connected to the controller; wherein the controller is further configured to: control the speaker corresponding to each of the sound sources to play the audio data of each of the sound sources.

[0020] Optionally, the audio control system further includes: a display screen connected to the controller; wherein the display screen is used to display at least one of a command interface, a stage interface, a first selection interface, and a second selection interface; wherein the command interface is used to display at least one of the following: the command area corresponding to each of the sound sources and the command action; the stage interface is used to display the first position of each of the sound sources on the stage; the first selection interface is used to display at least one listening position in a first space; and the second selection interface is used to display at least one listening position in a second space.

[0021] According to a sixth aspect of this application, a vehicle is provided, the vehicle including the audio control system as described above, or including the electronic equipment as described above.

[0022] In summary, the technical solution provided in this application can separate audio data from multiple sound sources in fused audio and play the audio data from multiple sound sources to achieve fused audio playback. Sound source separation facilitates flexible adjustment of the position and audio playback effect of each sound source, allowing users to experience sounds from different directions and sources, enhancing the spatial sense and immersion of audio playback, and achieving an immersive listening experience. Furthermore, this application's embodiments recognize the user's command gestures and control the audio playback effect of at least one sound source based on these gestures, integrating the user's command process into the audio playback, allowing the user to feel like a "music conductor," enhancing interactivity and entertainment, and achieving a personalized listening experience.

[0023] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of an audio control method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a selection interface provided in an embodiment of this application; Figure 3 This is a flowchart of another audio control method provided in the embodiments of this application; Figure 4 This is a schematic diagram of the sound source distribution in a concert hall provided in an embodiment of this application; Figure 5 This is a schematic diagram of a stage interface provided in an embodiment of this application; Figure 6 This is a schematic diagram of a position mapping provided in an embodiment of this application; Figure 7 This is a schematic diagram illustrating the restoration of a listening experience provided in an embodiment of this application; Figure 8 This is a schematic diagram of a command interface provided in an embodiment of this application; Figure 9 This is a schematic diagram of another command interface provided in an embodiment of this application; Figure 10 This is a schematic diagram of another command interface provided in an embodiment of this application; Figure 11 This is a schematic diagram of another command interface provided in an embodiment of this application; Figure 12 This is a schematic diagram of another command interface provided in an embodiment of this application; Figure 13 This is a schematic diagram of an audio surround playback provided in an embodiment of this application; Figure 14 This is a schematic diagram of another audio surround playback provided in an embodiment of this application; Figure 15 This is a flowchart of another audio control method provided in the embodiments of this application; Figure 16 This is a schematic diagram illustrating the construction of a sound source separation dataset provided in an embodiment of this application; Figure 17 This is a flowchart of another audio control method provided in the embodiments of this application; Figure 18 This is a flowchart of another audio control method provided in the embodiments of this application; Figure 19 This is a schematic diagram of an audio control system provided in an embodiment of this application; Figure 20 This is a schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0027] According to a first aspect of this application, embodiments of this application provide an audio control method.

[0028] Please see Figure 1 , Figure 1 This is a flowchart of an audio control method provided in an embodiment of this application. Figure 1 As shown, the audio control method may include the following step S100.

[0029] Step S100: Control the audio playback effect of at least one sound source in the fused audio according to the user's command.

[0030] Fusion audio refers to audio data that combines or integrates different elements, styles, sources, etc. This application does not limit the specific type of fusion audio; for example, fusion audio may include orchestral audio, symphonic audio, pop music audio, etc. Typically, fusion audio corresponds to multiple sound sources. A sound source refers to the origin or source of sound, which can be a physical object in nature (such as animals, wind, flowing water, rain, people, etc.) or a man-made device or system (such as musical instruments, software synthesizers, industrial machines, etc.).

[0031] For example, taking the fused audio including orchestral audio as an example, the sound sources corresponding to the orchestral audio can include, but are not limited to, at least one of the following: string instruments (such as violin, viola, cello, double bass, etc.), woodwind instruments (such as flute, oboe, clarinet, bassoon, etc.), brass instruments (such as trumpet, trombone, French horn, tuba, etc.), percussion instruments (such as timpani, triangle, cymbals, etc.), keyboard instruments (such as piano, etc.), and others (such as vocals, etc.). Of course, in practical applications, the number and categories of sound sources can be flexibly set according to needs. For example, the sound sources corresponding to orchestral audio can be divided into four categories: string instruments, woodwind instruments, brass instruments, and others. That is, the aforementioned keyboard instruments and percussion instruments can all be classified as "other".

[0032] This application embodiment can separate audio data from multiple sound sources from fused audio and play the audio data from multiple sound sources to achieve fused audio playback. Sound source separation facilitates flexible adjustment of the position and audio playback effect of each sound source, allowing users to experience sounds from different directions and sound sources, enhancing the spatial sense and immersion of audio playback, and achieving an immersive listening experience. Furthermore, this application embodiment recognizes the user's command gestures and controls the audio playback effect of at least one sound source based on these gestures, integrating the user's command process into the audio playback, allowing the user to feel like a "music conductor," enhancing interactivity and entertainment, and achieving a personalized listening experience.

[0033] In some embodiments, the above-described audio control method may further include: separating audio data from multiple sound sources from the fused audio; and playing the audio data of each sound source according to its first position on the stage. This embodiment obtains audio data from multiple sound sources through sound source separation, and then plays the audio data of each sound source according to its first position on the stage, simulating the actual performance effect of each sound source in the stage space, creating a sense of space and immersion, and effectively achieving an immersive listening experience.

[0034] The stage is used to indicate the area where the sound source is located. The stage can be indoors or outdoors; correspondingly, the space where the stage is located (referred to as the first space in this embodiment) can be a relatively enclosed indoor space or a relatively open outdoor space, and this embodiment does not limit this. For example, the first space can be any of the following: concert hall, opera house, drama theater, auditorium, open-air venue, gymnasium, square, studio, etc.

[0035] In some embodiments, the audio control method described above may further include: displaying a space selection interface, which displays at least one space. A user can select a first space from the at least one space in the space selection interface to subsequently simulate the actual performance effect of the sound source in the first space. The user can select the first space through touch operations such as clicking or pressing, or through gestures, voice, etc., and this embodiment does not limit the specific methods used.

[0036] This application embodiment plays the audio data of each audio source according to its first position on the stage in the first space. The first position of each audio source on the stage can be preset, default, or user-defined; this application embodiment does not limit this. To facilitate users in setting the first position of each audio source, in some embodiments, the above audio control method may further include: displaying a stage interface for displaying the first position of each audio source on the stage; and updating the first position of the audio source on the stage in response to a touch operation on the audio source.

[0037] The stage interface initially displays the default first position of each sound source on the stage, which can be shown through icons or buttons. Users can then trigger touch operations on each sound source by manipulating its corresponding icon or button, such as swiping or clicking, to update the sound source's first position on the stage. Furthermore, the stage interface can also display the space in which the stage is located, i.e., the primary space, such as a concert hall; and at least one listening position within that primary space.

[0038] It should be noted that in this embodiment, the number of sound sources on the stage can be equal to, greater than, or less than, the number of sound sources separated from the blended audio. That is, the user can add or delete the separated sound sources. For example, if the audio data of string instruments, woodwind instruments, brass instruments, and other four sound sources are separated from the blended audio, the user can copy the audio data of one or more of these sound sources in the stage interface, or delete the audio data of one or more of these sound sources in the stage interface.

[0039] Since the space where the stage is located (the first space) and the actual playback space of the fused audio (referred to as the second space in this embodiment) may not be the same space, it is necessary to map the first position of each sound source to the second space to obtain the second position of each sound source, and then play the audio data of each sound source according to the second position of each sound source to ensure accurate reproduction of the actual performance effect of each sound source in the first space.

[0040] Based on this, in some embodiments, the above-mentioned playing audio data of each sound source according to its first position on the stage may include: determining the second position of each sound source in the spatial sound field model according to its first position on the stage; and playing audio data of each sound source according to its second position. The spatial sound field model is the sound field model of the second space, used to indicate the propagation and distribution characteristics of sound within the second space. This application does not limit the specific type of the second space; for example, the second space can be any of the following: a vehicle, cinema, theater, home entertainment room, recording studio, shopping mall, airplane, cruise ship, etc.

[0041] To ensure that the user's listening experience in the second space is consistent with that in the first space, the relative positions of each sound source and the user's second listening position in the second space, as well as the relative positions of each sound source and the user's first listening position in the first space, should be consistent. Based on this, in some embodiments, determining the second position of each sound source in the spatial sound field model based on its first position on the stage can include: determining the relative position between each sound source and its first listening position based on its first position on the stage, where both the first listening position and the stage correspond to the first space; determining the relative position between each sound source and its second listening position based on its relative position to the first listening position; and determining the second position of each sound source in the spatial sound field model based on its relative position to the second listening position, where both the second listening position and the spatial sound field model correspond to the second space.

[0042] The relative positions (including angles and / or distances) between each sound source and the second listening position in the second space are consistent with the relative positions (including angles and / or distances) between each sound source and the first listening position in the first space. This ensures that the distribution of each sound source in the first and second spaces is consistent, thereby realizing the reproduction of the listening experience of the first listening position in the first space on the second listening position in the second space.

[0043] The first and / or second listening positions can be preset, default, or user-defined; this embodiment does not limit this. To facilitate user setting of the first listening position, in some embodiments, the audio control method may further include: displaying a first selection interface, which displays at least one listening position in a first space; to facilitate user setting of the second listening position, in some embodiments, the audio control method may further include: displaying a second selection interface, which displays at least one listening position in a second space.

[0044] The first and second selection interfaces can be displayed simultaneously or sequentially; this application embodiment does not limit this. For example, the first and second selection interfaces can be displayed simultaneously without overlapping. The user can select a first listening position from at least one listening position in the first selection interface and a second listening position from at least one listening position in the second selection interface. Alternatively, the first selection interface can be displayed first, and after the user selects the first listening position from at least one listening position in the first selection interface, the second selection interface can be displayed, and then the user can select the second listening position from at least one listening position in the second selection interface. The user can select the first and / or second listening positions through touch operations such as clicking or pressing, or through gestures, voice, etc.; this application embodiment does not limit this.

[0045] This application embodiment can play audio data from various sound sources using speakers. To create a sense of space and immersion, the second space may include multiple speakers, each capable of playing audio data from one or more sound sources. In some embodiments, the above-mentioned playing audio data from each sound source according to its second position may include: determining the speaker corresponding to each sound source according to its second position; and controlling the speaker corresponding to each sound source to play the audio data from each sound source.

[0046] Specifically, based on the second position of each sound source in the second space and the position of each speaker in the second space, the distance between each sound source and each speaker can be determined. For each sound source, one or more speakers closest to that sound source can be selected to play its audio data. Furthermore, to further enhance the sense of space, for each speaker, the audio data from sound sources closer to that speaker can be played at a higher volume; and for each sound source, the audio data from speakers closer to that sound source can be played at a higher volume.

[0047] Below, we will use several examples to introduce and explain the above-mentioned spatial selection, listening position selection, position mapping, and audio playback. In this example, we will use a concert hall as the first space, a vehicle as the second space, and orchestral audio as the blended audio.

[0048] The vehicle can display a space selection interface on its screen. This interface can include icons or buttons corresponding to at least one concert hall, such as the Concertgebouw (Netherlands), Boston Symphony Hall (USA), the National Centre for the Performing Arts (China), the Berlin Philharmonic Hall (Germany), and the Paris Philharmonic Hall (France). Users can click on the icon or button corresponding to a concert hall to trigger a selection operation for that concert hall. The vehicle can then display a first selection interface for that concert hall and a second selection interface for the vehicle itself. The first selection interface includes at least one listening position within the concert hall, and the second selection interface includes at least one listening position within the vehicle.

[0049] Please see Figure 2 , Figure 2 This is a schematic diagram of a selection interface provided in an embodiment of this application. For example... Figure 2As shown, the selection interface may include a first selection interface 210 and a second selection interface 220. The first selection interface 210 includes at least one listening position 211 in the concert hall selected by the user. Each listening position 211 corresponds to a seat in the concert hall (each seat faces the stage). The user can select one listening position 211 as the first listening position through touch operations such as clicking or pressing. The second selection interface 220 includes at least one listening position 221 in the vehicle. Each listening position 221 corresponds to a seat in the vehicle, such as the driver's seat, front passenger seat, left rear, right rear, or middle seat (the middle seat can be used when more than one user in the vehicle needs to experience the spatial feel of the orchestral audio). The user can select one listening position 221 as the second listening position through touch operations such as clicking or pressing. Of course, the first and second listening positions can also be preset or default. Furthermore, during the playback of the orchestral audio, the selection interface can still be triggered at any time, allowing the user to change the first and / or second listening positions at any time.

[0050] Please see Figure 3 , Figure 3 This is a flowchart of another audio control method provided in an embodiment of this application. Based on Figure 3 The audio control method shown allows users to experience the listening experience from different seats in real time. For example... Figure 3 As shown, the audio control method may include the following steps S310 to S350.

[0051] Step S310: Construct geometric architectural models of each concert hall using architectural acoustics software or algorithms; Step S320: Calculate the impulse response of all seats in each concert hall; Step S330: Configure the shape outline of each concert hall and the impulse response of all seats in the vehicle; Step S340: Associate each seat in the first selection interface corresponding to each concert hall with the corresponding impulse response; Step S350: When the user selects the first listening position in the first selection interface, the audio data played by each speaker in the vehicle is convolved with the impulse response corresponding to the first listening position.

[0052] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating the distribution of sound sources in a concert hall, as provided in an embodiment of this application. Figure 4 As shown, according to the sound production methods of various sound sources in an orchestra, different sound sources can be distributed in different positions on the concert hall stage. For high fidelity reproduction... Figure 4 The distribution of sound sources on the concert hall stage is shown in the embodiment of this application, which separates the audio data of multiple sound sources from the orchestral audio.

[0053] Users can select any orchestral audio track and upload it to the vehicle via a music selection interface displayed on the vehicle's screen and / or on the screen of a mobile device such as a smartphone, through methods such as software, Bluetooth, or a USB flash drive. The vehicle can then separate audio data from multiple sound sources within the orchestral audio and display a stage interface that includes visual icons associated with each sound source.

[0054] Please see Figure 5 , Figure 5 This is a schematic diagram of a stage interface provided in an embodiment of this application. For example... Figure 5 As shown, the stage interface can include visual icons corresponding to various sound sources, with each icon indicating its initial position on the stage. The initial position of each sound source can be preset or default. For example, the initial position of each sound source is preset or default when the stage interface is initially displayed. Users can then move the visual icons corresponding to the sound sources using touch operations such as clicking and swiping to customize their initial positions. Furthermore, for sound sources separated from orchestral audio, users can copy any sound source to add to the stage, such as... Figure 5 As shown, the user has copied the string instrument. The stage interface shows two string instruments, and the audio data of these two string instruments is the same.

[0055] After determining the primary position of each sound source on the concert hall stage, the vehicle can control each speaker to play the corresponding audio data based on spatial audio algorithms. During the playback of orchestral audio, the user can still change the primary position of each sound source on the stage at any time, and the sound source corresponding to each speaker may change accordingly.

[0056] Please see Figure 6 , Figure 6 This is a schematic diagram of a position mapping provided in an embodiment of this application. In the above spatial audio algorithm, the vehicle can calculate the relative position (including angle and distance) between each sound source and the first listening position in the concert hall. Based on this relative position, the sound transfer function of the sound source relative to the head can be calculated. Furthermore, a spatial sound field model of the vehicle can be constructed based on the actual distribution of the vehicle's speakers. This spatial sound field model includes sound sources, speakers, and a second listening position. The spatial sound field model can construct a coordinate system with the second listening position as the origin. Figure 6 As shown, the relative positions between each sound source and the second listening position in the spatial sound field model are consistent with the relative positions between each sound source and the first listening position in the concert hall. Figure 6In the diagram, a, b, c, and d represent distances. In the spatial sound field model of a vehicle, the appropriate speaker can be selected for each sound source to play its audio data based on the distance between each sound source and each speaker.

[0057] Please see Figure 7 , Figure 7 This is a schematic diagram illustrating the restoration of a listening experience provided in an embodiment of this application. In the spatial sound field model of the vehicle, the sound source corresponding to the audio data played by each speaker is obtained by comprehensively considering the sound transfer function between the sound source and the first listening position in a concert hall, the impulse response of the first listening position, and the relative position of the sound source and the speaker. The purpose is to restore the listening experience of the first listening position in a concert hall from the second listening position in the vehicle. The sound transfer function is used to restore the sense of location of the listening position; the impulse response is used to restore the listening effect of the concert hall, achieving an immersive experience; the relative position is used to select the speaker and control the playback volume of the speaker, and this relative position includes angle and distance, such as... Figure 7 D1, D2, and D3 in the diagram represent distances. The closer the sound source is to the speaker, the louder the volume of the audio data played by the speaker.

[0058] During the playback of the merged audio, that is, during the playback of audio data from each sound source, the user can adjust the audio playback effect of at least one sound source through command gestures. In some embodiments, the above step S100 may include: responding to the user's command gesture, controlling the audio playback effect of each sound source in the merged audio according to the action parameters of the command gesture. By controlling the audio playback effect according to the action parameters of the command gesture, the audio playback effect can be adapted to the user's command intention, fully meeting the user's personalized needs.

[0059] In this embodiment, the user's command actions include, but are not limited to, at least one of the following: hand command actions, head command actions, foot command actions, and facial expression command actions. Head command actions may include, but are not limited to, head shaking, nodding left and right, and nodding up and down; facial expression command actions may include, but are not limited to, smiling, being serious, and laughing; and foot command actions may include, but are not limited to, lifting a foot and tiptoeing. Typically, command actions may include hand command actions. Taking hand command actions as an example, in some embodiments, the aforementioned command actions may include, but are not limited to, at least one of the following: hand waving actions, hand clenching / unclenching actions, and hand rotation actions. Hand waving actions may be two-handed waving actions or one-handed waving actions, and may be left-right waving actions or up-down waving actions; hand clenching / unclenching actions may be two-handed clenching / unclenching actions or one-handed clenching / unclenching actions; and hand rotation actions may be wrist rotation actions or fingertip rotation actions. By recognizing command actions from multiple dimensions, we can further improve and enrich the user's command methods, thereby enhancing the personalization of audio playback.

[0060] The motion parameters of a command gesture can be associated with the type of gesture. Taking hand gestures as an example, if a hand gesture includes waving, the motion parameters could include the waving speed and waving area; if a hand gesture includes hand rotation, the motion parameters could include the rotation area and rotation direction; if a hand gesture includes clenching, the motion parameters could include the clenching area. For further explanation of hand gestures and their parameters, please refer to the following examples, which will not be elaborated upon here.

[0061] To facilitate users' real-time monitoring of their command actions and the sound sources controlled by those actions, in some embodiments, the audio control method may further include: displaying a command interface, which displays at least one of the following: the command area corresponding to each sound source and the command action. The command interface may display the user's own real-time video feed, or the real-time feed of a virtual character corresponding to the user, to indicate the command actions performed by the user in real time. The command areas corresponding to each sound source in the command interface may be automatically divided based on the first position of each sound source on the stage, or they may be preset or default. The number of command areas in the command interface is the same as the number of sound sources on the stage in the first space.

[0062] In some embodiments, the aforementioned audio playback effects may include, but are not limited to, at least one of the following: audio playback speed, audio playback volume, audio surround playback, audio playback pitch, audio playback intensity, and audio playback status. The audio playback status may include stopping playback and resuming playback. By controlling multiple dimensions of audio playback effects, the audio playback effects of different sound sources can be further improved and enriched, further enhancing the spatial sense and immersiveness of audio playback.

[0063] Below, several examples will be used to illustrate the aforementioned conducting gestures, gesture parameters, and audio playback effects. In this example, the first space is a concert hall, the second space is a vehicle, the blended audio is orchestral audio, and the conducting gesture is a hand gesture.

[0064] In this embodiment, the user can choose to activate the conductor mode. After entering the conductor mode, the user's hand gestures can be detected and recognized, and the audio playback effects of the corresponding sound sources in the orchestral audio can be controlled according to the user's hand gestures, such as controlling the audio playback rhythm and volume of the sound sources, so that the user can experience an immersive and interactive experience of being a "music conductor" in the car.

[0065] Please see Figures 8 to 12 , Figures 8 to 12 These are schematic diagrams of a command interface provided in an embodiment of this application. The vehicle can display the command interface on a screen, which may include command areas 010 corresponding to each sound source and the user's own real-time video feed 020. The vehicle can automatically divide the command interface into command areas 010 corresponding to each sound source based on the first position of each sound source on the stage, with the number of command areas 010 being the same as the number of sound sources on the stage. Figure 8 As shown, the command interface includes two sound sources on the stage, each corresponding to a command area 010; as... Figure 9 As shown, the command interface includes three sound sources on the stage, each corresponding to a command area 010; as... Figure 10 As shown, the command interface includes four sound sources on the stage, each corresponding to a command area 010; as... Figure 11 As shown, the command interface includes five sound sources on the stage, each corresponding to a command area 010; as... Figure 12 As shown, the command interface includes six sound sources on the stage, each corresponding to a command area 010. Users can... Figures 8 to 12 The command interface shown clearly displays the hand gestures being executed and the audio sources being controlled. In this embodiment, the audio playback speed, volume, and status of the corresponding audio source can be controlled, as well as the surround sound playback of the corresponding audio source.

[0066] For example, hand gestures include hand waving actions, and the audio playback effect controlled by the hand waving actions includes audio playback speed. Therefore, in some embodiments, step S100 may include: controlling the audio playback speed of the target sound source in the blended audio in response to the user's hand waving action. To adapt the audio playback speed control to the user's command intention, in some embodiments, controlling the audio playback speed of the target sound source in the blended audio in response to the user's hand waving action may include: controlling the audio playback speed of the target sound source in the blended audio according to the waving parameters of the hand waving action in response to the user's hand waving action. The waving parameters may include, but are not limited to, waving speed and waving area. Therefore, in some embodiments, controlling the audio playback speed of the target sound source in the blended audio according to the waving parameters of the hand waving action may include: using the waving speed of the hand waving action as the audio playback speed of the target sound source in the blended audio.

[0067] The target sound source can be any sound source in the blended audio, or a specific sound source within the blended audio, such as the sound source corresponding to the conducting area corresponding to the waving area of ​​a hand gesture. The hand gesture can include left and right waving movements, such as waving both hands left and right or waving one hand left and right, provided the hand is not clenched into a fist. The faster the hand gesture, the faster the audio playback speed of all sound sources in the blended audio (such as orchestral audio).

[0068] To ensure that the audio playback speed matches the waving speed of the hand, when a user's two hands or one hand are detected waving left and right through visual recognition algorithms, an index α (where α ranges from 0 to 100, with a larger value indicating a faster waving speed) can be calculated in real time to indicate the waving speed. Then, the audio playback speed is determined based on α, and a corresponding control signal is generated. This control signal is then sent to the in-vehicle infotainment system (IVI) via the vehicle's CAN bus, automotive Ethernet, direct hardware interfaces (such as GPIO or I2C), or open interfaces of the in-vehicle system (such as the Android Automotive API), enabling real-time adjustment of the playback speed of all audio sources.

[0069] For example, hand gestures include hand waving actions, and the audio playback effect controlled by the hand waving actions includes audio playback volume. Therefore, in some embodiments, step S100 may include: controlling the audio playback volume of the target sound source in the blended audio in response to the user's hand waving action. To adapt the audio playback volume control to the user's command intention, in some embodiments, controlling the audio playback volume of the target sound source in the blended audio in response to the user's hand waving action may include: controlling the audio playback volume of the target sound source in the blended audio according to the waving parameters of the hand waving action in response to the user's hand waving action. The waving parameters may include, but are not limited to, waving speed, waving position, and waving area. Therefore, in some embodiments, controlling the audio playback volume of the target sound source in the blended audio according to the waving parameters of the hand waving action may include: determining the audio playback volume of the target sound source in the blended audio according to the waving position of the hand waving action.

[0070] The target audio source can be any audio source in the blended audio or a specific audio source within the blended audio. For example, a hand waving motion can be a two-handed waving motion, and the target audio source includes all audio sources in the blended audio; alternatively, a hand waving motion can be a single-handed waving motion, and the target audio source includes at least one audio source in the blended audio, such as the audio source corresponding to the command area of ​​the waving motion or the audio source corresponding to the command area of ​​the other hand. The hand waving motion can include left and right waving motions, such as left and right waving motions with both hands or left and right waving motions with one hand. The closer the waving position of the hand is to the upper boundary of the command interface, the louder the audio playback volume of the target audio source; the closer the waving position of the hand is to the lower boundary of the command interface, the quieter the audio playback volume of the target audio source.

[0071] Taking hand waving actions, including single-hand waving actions, as an example, to ensure that the audio playback volume matches the waving position of the hand, when a visual recognition algorithm detects that the user's hand A (e.g., not in a fist) is located in the command area of ​​a sound source (e.g., a stringed instrument) and the other hand B (e.g., in a fist) is waving left and right, an index β (e.g., β ranges from 0 to 100, and the closer the waving position of hand B is to the upper boundary of the command interface, the larger the value of β) can be calculated in real time. Then, based on β, the audio playback volume of the sound source (e.g., a stringed instrument) is determined, and a corresponding control signal is generated. This control signal is then sent to the in-vehicle infotainment system via the vehicle's CAN bus / Automotive Ethernet / direct hardware interface (e.g., GPIO or I2C) / open interface of the in-vehicle system (e.g., Android Automotive API), thereby achieving real-time adjustment of the audio playback volume of the sound source (e.g., a stringed instrument).

[0072] Taking hand waving actions, including waving with both hands, as an example, to ensure that the audio playback volume matches the waving position of the hand, when the visual recognition algorithm detects that the user's hands are waving left and right on the same horizontal line, an index γ (where γ ranges from 0 to 100, and the closer the waving position of the hands is to the upper boundary of the command interface, the larger the value of γ) can be calculated in real time. Then, based on γ, the audio playback volume of all sound sources is determined, and corresponding control signals are generated. These control signals are then sent to the in-vehicle infotainment system via the vehicle's CAN bus, AutomotiveEthernet, direct hardware interfaces (such as GPIO or I2C), or open interfaces of the in-vehicle system (such as Android AutomotiveAPI), to achieve real-time adjustment of the audio playback volume of all sound sources.

[0073] For example, hand gestures include clenching fists with both hands and then releasing them. Audio playback effects controlled by clenching fists include pausing playback, and audio playback effects controlled by releasing them include resuming playback. Therefore, in some embodiments, step S100 may include: pausing playback of each audio source in the blended audio in response to the user's clenching fists; and resuming playback of each audio source in the blended audio in response to the subsequent releasing of the fists.

[0074] To ensure that the audio playback status matches the hand gestures, when the user's hand position is detected by visual recognition algorithms, a control signal δ for the audio playback status can be output in real time (δ is 0 when both hands are clenched into fists; δ is 1 when at least one hand is not clenched into a fist). Then, the audio playback speed control signal is sent to the in-vehicle infotainment system through the vehicle's CAN bus / Automotive Ethernet / direct hardware interface (such as GPIO or I2C) / open interface of the in-vehicle system (such as Android Automotive API). When δ is 0, all audio sources are paused; when δ is 1, all audio sources are resumed, thus achieving real-time control of the audio playback status of all audio sources.

[0075] For example, hand gestures include hand rotation, and the audio playback effect controlled by the hand rotation includes surround sound playback. Therefore, in some embodiments, step S100 may include: responding to the user's hand rotation, controlling the target audio source in the blended audio to perform surround sound playback. To adapt the control of surround sound playback to the user's command intention, in some embodiments, responding to the user's hand rotation, controlling the target audio source in the blended audio to perform surround sound playback may include: responding to the user's hand rotation, controlling the target audio source in the blended audio to perform surround sound playback according to the rotation parameters of the hand rotation. The rotation parameters may include, but are not limited to, rotation area, rotation direction, etc. Therefore, in some embodiments, controlling the target audio source in the blended audio to perform surround sound playback according to the rotation parameters of the hand rotation may include: determining the command area corresponding to the hand rotation based on the rotation area of ​​the hand rotation; determining the target audio source in the blended audio corresponding to the command area; determining the surround direction of the target audio source according to the rotation direction of the hand rotation; and controlling the target audio source to perform surround sound playback according to the surround direction.

[0076] The hand rotation action can be a fingertip rotation action, and the rotation direction of the fingertip rotation action can be counterclockwise or clockwise; audio surround playback can refer to the sound source rotating counterclockwise or clockwise with the second listening position in the vehicle as the center, thereby achieving immersive surround sound for the sound source. The surround direction of the target sound source can be the same as the rotation direction of the hand rotation action.

[0077] Please see Figure 13 and Figure 14 , Figure 13 and Figure 14 These are schematic diagrams illustrating an audio surround playback method provided in an embodiment of this application. Figure 13 As shown, when a user rotates a finger with one hand, and the area of ​​rotation of that finger is located in the control area of ​​a sound source (such as a stringed instrument), then that sound source (such as a stringed instrument) is controlled to play audio in surround mode. The surround direction of the sound source (such as a stringed instrument) is the same as the direction of rotation of the finger, such as... Figure 13 As shown, all directions are counter-clockwise. It should be understood that in this embodiment, the hand rotation action can act as a switch for audio surround playback. The user only needs to perform a hand rotation action once for a specific sound source to control that sound source to continue playing audio surround music. That is, even if the user does not subsequently perform a hand rotation action in the command area corresponding to that sound source, the sound source can still continue playing audio surround music. Based on this, the user can sequentially perform hand rotation actions in the command areas corresponding to each sound source to trigger each sound source to play audio surround music sequentially, such as... Figure 14As shown, all sound sources in an orchestral audio can be played in surround sound within the same time period, but the surround sound playback of these sound sources can be triggered sequentially.

[0078] It should be noted that the above command scenarios can be used in combination or independently; furthermore, the above command scenarios and specific command rules are only examples and do not constitute a limitation on the embodiments of this application. Other command actions and / or other audio playback effects in actual applications should also fall within the protection scope of the embodiments of this application.

[0079] Please see Figure 15 , Figure 15 This is a flowchart of another audio control method provided in an embodiment of this application. For example... Figure 15 As shown, the audio control method may include the following steps S1510 to S1550.

[0080] Step S1510: Acquire a video frame containing the user's hands.

[0081] The system can capture video footage of the user using the vehicle's built-in camera, or it can add low-light / infrared cameras to ensure stable operation under different lighting conditions.

[0082] Step S1520: Detect and track key hand joints of the user based on a lightweight visual recognition model, and record the coordinate data of specific joints in real time.

[0083] It can detect and track user-specific hand joints (such as fingertips, wrists, etc.) based on self-developed or open-source visual recognition models (such as BlazePose and MoveNet), and record the coordinates of the wrist and each finger.

[0084] Step S1530: Generate control signals for audio playback based on the coordinate data of specific joint points.

[0085] The audio playback speed (Va) can be controlled based on the wrist's movement speed (Vh); the audio playback volume (Vol) can be controlled based on the wrist's coordinates. For example, the wrist's movement speed (Vh) can be calculated based on the wrist's displacement in the left-right direction and the time taken, a speed threshold can be set, and the speed can be mapped to the audio playback speed (Va). Using the bottom of the video screen as a reference point, the distance (D) of the wrist from the reference point can be mapped to the audio playback volume (Vol), and a maximum volume can be set to avoid clipping. In addition, a minimum displacement threshold can be set to determine a valid wrist movement.

[0086] Step S1540: Send a control signal for audio playback effects to the in-vehicle infotainment system.

[0087] Audio playback control signals can be transmitted to the in-vehicle infotainment system via communication methods such as CAN bus, Android AIDL / Binder interface, Socket, D-Bus, and Bluetooth. For example, CAN bus communication can be used for traditional in-vehicle infotainment systems or industrial control scenarios requiring high reliability; local communication can be achieved through the Android AIDL / Binder interface for Android-based in-vehicle infotainment systems; D-Bus communication can be used for Linux-based in-vehicle infotainment systems; and Socket communication can be used for networked or module-controlled in-vehicle infotainment systems.

[0088] Step S1550: The in-vehicle infotainment system controls the audio playback effect of at least one audio source in real time according to the control signal.

[0089] For audio playback speed control signals, the audio playback speed can be controlled in real time through the SoundTouch audio processing library (the audio playback pitch can remain unchanged); for audio playback volume control signals, the audio playback volume can be controlled in real time through AudioManager (for Android systems) or ALSA / PulseAudio (for Linux systems).

[0090] As can be seen from the above embodiments, the embodiments of this application require separating audio data from multiple audio sources from fused audio. To facilitate fast and accurate audio source separation, an audio source separation model can be used. Based on this, in some embodiments, the above-mentioned separation of audio data from multiple audio sources from fused audio may include: inputting the fused audio into the audio source separation model, so that the audio source separation model outputs audio data from multiple audio sources. The audio source separation model can automate and intelligently achieve audio source separation, improving the efficiency of audio source separation.

[0091] The audio source separation model can be an artificial intelligence model such as deep learning. This application does not limit the specific type of audio source separation model; in practical applications, it can be flexibly set according to requirements and the type of the second space. For example, if the second space includes a vehicle, the audio source separation model can be set to an end-to-end architecture more suitable for in-vehicle systems, such as Conv-TasNet, demucs, Wave-U-Net, etc.

[0092] To improve the accuracy and robustness of the audio source separation model, it can be trained first, and then the trained model can be used to process the fused audio to output audio data from multiple sources in real time. Based on this, in some embodiments, the audio control method may further include: training the audio source separation model based on the audio source separation dataset. Each training sample in the audio source separation dataset includes one sample fused audio and sample audio data from multiple audio sources. To ensure the accuracy of audio source separation, the number of audio sources corresponding to each training sample can be equal to the number of audio sources that need to be separated from the fused audio.

[0093] In this embodiment, the sound source separation dataset can be generated based on historical data, manual annotation, automated annotation, etc. In some embodiments, the audio control method described above may further include: constructing a sound source separation dataset based on original recording files of multiple sound-producing entities. Here, a sound-producing entity refers to the smallest unit that emits sound, such as a musical instrument, animal, wind, rain, or flowing water. By combining the original recording files of multiple sound-producing entities, mixing and other processing can be performed to construct a large and rich set of training samples, improving the efficiency of constructing the sound source separation dataset.

[0094] Since each training sample includes a sample fused audio and sample audio data from multiple sound sources, and the sound sources include a class of sound-producing bodies, in some embodiments, the above-mentioned construction of a sound source separation dataset based on the original recording files of multiple sound-producing bodies may include: grouping the original recording files of multiple sound-producing bodies; mixing the original recording files of each group to obtain sample audio data of each sound source; and mixing the sample audio data of multiple sound sources to obtain sample fused audio. Specifically, when grouping the original recording files of multiple sound-producing bodies, the original recording files of sound-producing bodies with the same pronunciation method can be grouped together based on the pronunciation method and frequency of the sound-producing bodies.

[0095] To facilitate the rapid construction of training samples and the rapid training of models, in some embodiments, the audio control method described above may further include: preprocessing the original recording files of multiple sound sources; and constructing a sound source separation dataset based on the preprocessed original recording files. The data preprocessing may include, but is not limited to, at least one of the following: format conversion and data length conversion. Through data preprocessing such as format conversion and / or data length conversion, the original recording files can be standardized for subsequent batch processing.

[0096] Furthermore, to avoid digital clipping distortion, in some embodiments, the above-mentioned audio control method may further include: performing a first linear scaling process on sample audio data from each audio source; performing a mixing process on the sample audio data after the first linear scaling process to obtain sample fused audio; and performing a second linear scaling process on the sample fused audio. The peak level thresholds corresponding to the first and second linear scaling processes may be the same or different, and this application embodiment does not limit this. For example, the first linear scaling process is used to linearly scale the reference group peak level of the sample audio data to a first peak level threshold, such as -3dBFS; the second linear scaling process is used to linearly scale the maximum level of the sample fused audio to a second peak level threshold, such as -1dBFS. Both the first and second peak level thresholds can be flexibly set according to actual needs, and can also be set according to test conditions during actual application, etc., and this application embodiment does not limit this. dBFS is a unit of volume, with a maximum value of 0 and other values ​​being negative. A value of 0 indicates a state where clipping will not occur.

[0097] The following examples illustrate the construction of the aforementioned audio source separation dataset and the training of the audio source separation model. In this example, we use a concert hall as the first space, a vehicle as the second space, and orchestral audio as the fused audio.

[0098] In practical applications, algorithms for automatically generating sound source separation datasets can be written using common languages ​​such as Python and C++. These algorithms can then efficiently convert raw orchestral recordings into sound source separation datasets needed for model training. The content of the training samples in the sound source separation dataset depends on the number of sound sources to be separated. For example, if the training samples in the sound source separation dataset contain four sound sources, then the sound source separation model trained on this dataset can be used to separate those four sound sources.

[0099] Please see Figure 16 and Figure 17 , Figure 16 This is a schematic diagram illustrating the construction of a sound source separation dataset provided in an embodiment of this application. Figure 17 This is a flowchart of another audio control method provided in an embodiment of this application. For example... Figure 16 As shown, sound-producing instruments can be divided into several groups according to their manner of sound production and frequency of use, with each group corresponding to a sound source: string instruments, woodwind instruments, brass instruments, and others (although the harp should be classified as a string instrument according to its manner of sound production, it is divided into other groups because its timbre differs significantly from other string instruments). For example... Figure 17 As shown, the audio control method may include the following steps S1710 to S1750.

[0100] Step S1710: Perform data preprocessing on the original audio file.

[0101] The original recording files can be in formats such as WAV, MP3, FLAC, and AIFF. Data preprocessing can be performed on the original recording files; for example, by adding silence at the end of the audio, the sample count of all original recording files can be standardized to the maximum sample count (ensuring that the length of each original recording file is consistent).

[0102] Step S1720: Group the original audio files.

[0103] The sound sources can be categorized into string instruments, woodwind instruments, brass instruments, and others (decreasing in frequency from left to right) based on their pronunciation and frequency of use. Then, the original recording files can be divided into multiple groups by constructing a feature lexicon (keywords of the original data names of different types of sound sources).

[0104] Step S1730: Mix the original recording files of each group to obtain sample audio data of each audio source.

[0105] When mixing each group of original recording files, a thread pool can be used to perform the mixing operation concurrently, directly superimposing all the audio tracks in the group according to the original amplitude ratio to generate sample audio data of the sound source.

[0106] Step S1740: Perform a first linear scaling process on the sample audio data of each sound source.

[0107] It can calculate the peak level of all sound sources, select the sound source with the highest peak level as the reference group, calculate the gain adjustment coefficient based on the first peak level threshold (such as -3dBFS), and normalize the peak level of the reference group to the first peak level threshold (such as -3dBFS) while keeping the relative proportion of the levels between each sound source unchanged.

[0108] Step S1750: Mix the sample audio data from multiple audio sources to obtain sample fused audio, and perform a second linear scaling process on the sample fused audio.

[0109] Mix the sample audio data from all audio sources to obtain the total audio signal (e.g., Figure 16 The example shown is mixture.wav, which is the sample-fused audio. The maximum level of the sample-fused audio is then limited to a second peak level threshold (e.g., -1dBFS) through a second linear scaling process to avoid digital clipping distortion.

[0110] The above method of creating audio source separation datasets can reduce dataset creation time by 90% and improve the SDR (Source-to-Distortion Ratio) score during training without loss of quality.

[0111] Based on the established audio source separation dataset and a real-time neural network, deep learning training is conducted to obtain a trained audio source separation model. The audio source separation model can adopt an end-to-end architecture more suitable for in-vehicle systems, such as Conv-TasNet, demucs, or Wave-U-Net.

[0112] The trained audio source separation model is then configured on the vehicle's infotainment system. For example, the model can be converted to an inference framework format suitable for the vehicle (such as ONNX, TensorRT, OpenVINO, TensorFlow Lite, PyTorch Mobile, etc.) and fitted with compatible computing hardware (such as NVIDIA Jetson series, ONNX Runtime, etc.). The trained audio source separation model is then converted to TensorRT format and deployed on the vehicle's infotainment system using NVIDIA Jetson. Users can then utilize NVIDIA Jetson's computing power to perform audio source separation on uploaded fused audio, thus meeting the requirements for real-time computation and output.

[0113] It should be noted that the audio source separation, visual recognition, and control signal calculation processes in the embodiments of this application can all be completed on a high-performance vehicle platform, such as a vehicle-grade computing platform that can utilize GPU acceleration, such as an NVIDIA Jetson, Qualcomm Snapdragon Ride / QCS series, or Renesas R-Car.

[0114] The audio control method provided in this application embodiment will be described below with an example. In this example, the first space is a concert hall, the second space is a vehicle, the blended audio is orchestral audio, and the conducting action is a hand conducting action.

[0115] Please see Figure 18 , Figure 18 This is a flowchart of another audio control method provided in an embodiment of this application. For example... Figure 18 As shown, the audio control method may include the following steps S1810 to S1850.

[0116] Step S1810: Process the original recording file based on the automatic generation algorithm to construct the sound source separation dataset; Step S1820: Conduct deep learning training based on the sound source separation dataset to obtain the sound source separation model, and deploy the sound source separation model to the vehicle system; Step S1830: Obtain the user's selection operation for the first listening position in the concert hall and the selection operation for the second listening position in the vehicle; Step S1840: Separate the audio data of multiple sound sources from the orchestral audio using the sound source separation model, and play the audio data of each sound source according to its first position, first listening position, and second listening position on the concert hall stage. Step S1850: In response to the user's hand gesture, control the audio playback effect of at least one sound source according to the action parameters of the hand gesture.

[0117] In summary, the beneficial effects of the embodiments of this application include at least the following: Firstly, the embodiments of this application enable a more immersive orchestral music listening experience in the vehicle: the in-vehicle concert hall mode in related technologies is based on the speaker units of the vehicle's surround sound channel, convolving the in-vehicle sound field signal with the impulse response of a concert hall to achieve a similar listening experience to a concert hall, but it cannot achieve the spatial sense of orchestral music in a real concert hall, and cannot allow users to hear the sounds of instruments from different directions; while the embodiments of this application, based on sound source separation technology and spatial audio algorithms, enable users to feel the sounds of orchestral instruments (such as string instruments, woodwind instruments, brass instruments, percussion instruments, and others) from different directions in the vehicle, achieving an immersive concert hall listening experience; Secondly, the embodiments of this application can achieve a more personalized orchestral music listening experience: in related technologies, the orchestral music in the car is either a stereo playback or a preset virtual surround sound, and users cannot feel a personalized listening experience; while the embodiments of this application, based on orchestral sound source separation technology and spatial audio algorithm, not only allow users to customize the position of various orchestral sound sources (moving sound sources on the car interface), but also allow users to customize the position of the user (themselves) in the concert hall or the path of walking (moving receiver). At the same time, the car speaker system can output immersive orchestral music in real time according to the user's custom settings, thereby realizing a highly customized in-car orchestral music listening experience; Third, the embodiments of this application can achieve a more interactive and entertaining orchestral music listening experience: There is no orchestral music mode involving in-vehicle human-computer interaction technology in the related technologies. The embodiments of this application are based on sound source separation technology and visual recognition algorithm. In the "music conductor" scenario, the in-vehicle camera is used to identify the user's left and right hand movements to control the playback rhythm and loudness of multiple orchestral sound sources, so that the user can feel the entertaining interactive experience of becoming a "music conductor".

[0118] According to a second aspect of this application, embodiments of this application also provide a computer-readable storage medium storing a computer program or instructions that, when executed by a processor, implement the above-described audio control method and have all the beneficial effects of the above-described audio control method, which will not be elaborated further here.

[0119] According to a third aspect of this application, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the above-described audio control method and have all the beneficial effects of the above-described audio control method, which will not be elaborated further here.

[0120] According to a fourth aspect of this application, embodiments of this application also provide an electronic device, including: a memory and a processor, wherein the memory stores a computer program or instructions; the processor is configured to execute the computer program or instructions in the memory to implement the steps of the above-described audio control method. This electronic device possesses all the beneficial effects of the above-described audio control method, which will not be elaborated upon further herein.

[0121] Computer-readable storage media can be, for example, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof, without particular limitation herein. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0122] In some embodiments of this application, a computer-readable storage medium may be any tangible medium that contains or stores a program that may be used or combined with an instruction execution system, apparatus, or device.

[0123] The aforementioned computer-readable storage medium may be included in the aforementioned electronic device or may exist independently without being assembled into the electronic device.

[0124] Computer program code for performing operations of some embodiments of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a Local Area Network (LAN) or a Wide Area Network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function.

[0126] It should also be noted that in some alternative implementations, the functions marked in the box may occur in a different order than those marked in the attached figures.

[0127] For example, two consecutively represented blocks can actually be executed in substantially parallel order, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, as well as combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or using a combination of dedicated hardware and computer instructions.

[0128] The units described in some embodiments of this application can be implemented in software or in hardware. The described units can also be located in a processor.

[0129] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0130] According to a fifth aspect of this application, embodiments of this application provide an audio control system.

[0131] Please see Figure 19 , Figure 19 This is a schematic diagram of an audio control system provided in an embodiment of this application. This audio control system can be used to execute any of the audio control methods described in the above embodiments. Figure 19 As shown, the audio control system may include a controller 1910.

[0132] When the second space includes a vehicle, the controller 1910 can be implemented as a domain controller, body controller, intelligent driving controller, vehicle controller, in-vehicle infotainment system (vehicle system), etc. The controller 1910 can be used to execute any of the audio control methods described in the above embodiments. For example, the controller 1910 can be used to control the audio playback effect of at least one sound source in the fused audio according to the user's command.

[0133] In some embodiments, such as Figure 19 As shown, the audio control system may further include a data acquisition device 1920 connected to the controller 1910. The data acquisition device 1920 may be connected to the controller 1910 via hardwired connections or a network, etc., as not limited in this embodiment. The data acquisition device 1920 is used to send user images to the controller 1910, and the controller 1910 is also used to recognize command actions from the user images. The data acquisition device 1920 may include sensors such as a camera.

[0134] In some embodiments, such as Figure 19As shown, the audio control system may further include a speaker 1930 connected to the controller 1910. The speaker 1930 may be connected to the controller 1910 via hardwired connections or a network, etc., and this embodiment does not limit this connection. The controller 1910 is also used to: control the speaker 1930 corresponding to each audio source to play the audio data of each audio source.

[0135] In some embodiments, such as Figure 19 As shown, the audio control system may further include a display screen 1940 connected to the controller 1910. The display screen 1940 may be connected to the controller 1910 via hardwired connections or a network, etc., and this embodiment does not limit this connection. The display screen 1940 may be a conventional display screen, a floating display screen, or a head-up display (HUD), etc. The display screen is used to display any of the interfaces described in the above embodiments, for example, to display at least one of a command interface, a stage interface, a space selection interface, a first selection interface, and a second selection interface. For a detailed description of each interface, please refer to the above embodiments; further details are omitted here.

[0136] It should be understood that the aforementioned controller 1910, speaker 1930, and / or display screen 1940, etc., can be components of an in-vehicle infotainment system. When the computing power of the in-vehicle infotainment system is sufficient, the audio control system in this embodiment can be implemented as an in-vehicle infotainment system; or, when the computing power of the in-vehicle infotainment system is insufficient to support audio source separation and other processing, the audio control system in this embodiment can include other controllers 1910, etc., with higher computing power, in addition to the in-vehicle infotainment system.

[0137] For further details regarding the steps performed by each component in the audio control system and their beneficial effects, please refer to the embodiments of the audio control method described above; these details will not be elaborated upon here.

[0138] According to the sixth aspect of this application, such as Figure 20 As shown, this application also provides a vehicle 10, which includes the aforementioned electronic equipment or the aforementioned audio control system. This vehicle possesses all the beneficial effects of the aforementioned electronic equipment and audio control system, etc., which will not be elaborated upon here.

[0139] The vehicle may be a gasoline-powered vehicle, a plug-in hybrid electric vehicle, or a new energy vehicle, etc., and this application does not make any specific restrictions.

[0140] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0141] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0142] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0143] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although the descriptions of each embodiment in this application have different focuses, and parts not described in detail in a certain embodiment can be referred to the relevant descriptions of other embodiments, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. An audio control method, characterized in that, The method includes: Control the audio playback effect of at least one sound source in the blended audio according to the user's command.

2. The method according to claim 1, characterized in that, The audio playback effects include at least one of the following: audio playback speed, audio playback volume, audio surround playback, audio playback pitch, audio playback intensity, and audio playback status.

3. The method according to claim 1, characterized in that, The command actions include at least one of the following: hand waving, hand clenching, fist clenching and releasing, and hand rotation.

4. The method according to claim 1, characterized in that, The method further includes: Display the command interface; The command interface is used to display at least one of the following: the command area corresponding to each of the sound sources, and the command actions.

5. The method according to claim 1, characterized in that, The method further includes: Separate audio data from multiple sound sources from the fused audio; Based on the first position of each of the aforementioned sound sources on the stage, the audio data of each of the aforementioned sound sources is played.

6. The method according to claim 5, characterized in that, The method further includes: Display stage interface; wherein, the stage interface is used to display the first position of each of the sound sources on the stage; In response to a touch operation on the sound source, update the first position of the sound source on the stage.

7. The method according to claim 5, characterized in that, The step of playing the audio data of each of the sound sources according to their first positions on the stage includes: Based on the first position of each sound source on the stage, the relative position between each sound source and the first listening position is determined; wherein, the first listening position corresponds to the first space of the stage; The relative positions of each sound source and the second sound position are determined according to the relative positions of each sound source and the first listening position; Based on the relative positions of each sound source and the second listening position, the second position of each sound source in the spatial sound field model is determined; wherein, the second listening position corresponds to the second space of the spatial sound field model. The audio data of each of the audio sources is played according to the second position of each of the audio sources.

8. The method according to claim 7, characterized in that, The method further includes: Display a first selection interface, which is used to display at least one listening position in the first space; And / or, Display a second selection interface, which is used to display at least one listening position in the second space.

9. The method according to claim 7, characterized in that, Playing the audio data of each of the sound sources according to their second positions includes: According to the second position of each of the sound sources, determine the speaker corresponding to each of the sound sources; Control the speaker corresponding to each of the sound sources to play the audio data of each sound source.

10. The method according to claim 5, characterized in that, The audio data from which multiple sound sources are separated from the fused audio includes: The fused audio is input into the audio source separation model so that the audio source separation model outputs audio data from multiple audio sources.

11. An audio control system, characterized in that, The audio control system includes: The controller is used to control the audio playback effect of at least one audio source in the blended audio according to the user's command.

12. The system according to claim 11, characterized in that, The audio control system further includes at least one of the following: A data acquisition device is connected to the controller; wherein the data acquisition device is used to: send user images to the controller; the controller is further used to: identify the command actions from the user images; A speaker is connected to the controller; wherein the controller is further configured to: control the speaker corresponding to each of the sound sources to play the audio data of each of the sound sources; A display screen is connected to the controller; wherein the display screen is used to display at least one of a command interface, a stage interface, a first selection interface, and a second selection interface; wherein the command interface is used to display at least one of the following: the command area corresponding to each of the sound sources and the command action; the stage interface is used to display the first position of each of the sound sources on the stage; the first selection interface is used to display at least one listening position in a first space; and the second selection interface is used to display at least one listening position in a second space.

13. A computer-readable storage medium, characterized in that, It stores a computer program or instructions thereon, which, when executed by a processor, implement the audio control method as described in any one of claims 1 to 10.

14. An electronic device, characterized in that, include: A memory on which computer programs or instructions are stored; A processor for executing the computer program or instructions in the memory to implement the audio control method as described in any one of claims 1 to 10.

15. A vehicle, characterized in that, The vehicle includes an audio control system as described in claim 11 or 12, or an electronic device as described in claim 14.