An audio processing method, apparatus and system

By playing audio in separate zones based on the positions of the singer and the listener in a karaoke setting, the problem of interference from the original singer's voice is solved, thus improving the experience for both the singer and the listener.

CN120526741BActive Publication Date: 2026-07-21YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YINWANG INTELLIGENT TECHNOLOGIES CO LTD
Filing Date
2025-03-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In karaoke settings, when singers listen to the original vocals, the listener's experience is disrupted, affecting their enjoyment.

Method used

By playing the original vocals and accompaniment audio at the singer's position and the singer's vocals and accompaniment audio at the listener's position, the system uses a microphone and camera to detect the singer's position, automatically divides the sound zones, and controls the audio playback to ensure that the singer can hear the original vocals while the listener is not disturbed.

Benefits of technology

It allows singers to clearly hear the original vocals while singing karaoke, enhancing their karaoke experience, while also improving the listener's experience by ensuring they are not disturbed by the original vocals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526741B_ABST
    Figure CN120526741B_ABST
Patent Text Reader

Abstract

Disclosed are an audio processing method, device and system, applied to the fields of audio processing and vehicle technology. The method comprises the following steps: obtaining singer voice audio sung by a singer for a target song, and obtaining original singer voice audio and accompaniment audio of the target song; controlling the playing of first audio in a first audio area and the playing of second audio in a second audio area, the first audio comprising the original singer voice audio, and the second audio comprising the singer voice audio and the accompaniment audio, the first audio area being a space area where the singer is located, and the second audio area comprising a space area where the singer and a listener are located. In this way, in a karaoke singing scene, the singer can hear the original singer voice, but the listener is not disturbed by the original singer voice, thus ensuring a good karaoke singing experience of the singer and improving the experience of the listener.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of audio processing and vehicle technology, and in particular to an audio processing method, apparatus and system. Background Technology

[0002] Karaoke is a popular form of entertainment. Karaoke venues include, but are not limited to, karaoke rooms, car seats, home theaters, and living rooms.

[0003] Currently, in karaoke settings, some singers habitually use the original vocals when singing, allowing them to calibrate their pitch and rhythm, thus enhancing the karaoke experience. However, for listeners, this results in a superposition of the original vocals and the singer's voice, negatively impacting the listener's experience.

[0004] Therefore, in karaoke scenarios, improving the listener's experience while ensuring the singer's karaoke experience is an urgent technical problem that needs to be solved. Summary of the Invention

[0005] This application discloses an audio processing method, apparatus, and system that enables singers to hear the original vocals in karaoke scenarios, while listeners are not disturbed by the original vocals, thus ensuring a better karaoke experience for singers and enhancing the listener's experience at the same time.

[0006] Firstly, this application provides an audio processing method, which includes: firstly, acquiring the original vocal audio and accompaniment audio of a target song, wherein the target song is a song sung by a singer; and secondly, acquiring the singer's vocal audio; and then controlling the playback of a first audio in a first audio region and a second audio in a second audio region. The first audio region is the spatial region where the singer is located, and the first audio includes the original vocal audio; the second audio region includes the spatial region where both the singer and the listener are located, and the second audio includes the singer's vocal audio and the accompaniment audio.

[0007] For example, since the first audio includes the original vocal audio, the first audio can be the audio obtained after applying sound effects processing to the original vocal audio. Since the second audio includes the singer's vocal audio and the accompaniment audio, the second audio can be the audio obtained after applying sound effects processing and / or mixing processing to the singer's vocal audio and the accompaniment audio.

[0008] The second vocal range includes the first vocal range, and the second audio played in the second vocal range includes the singer's vocal audio and the accompaniment audio. Therefore, at the singer's location, the original vocal audio, the accompaniment audio, and the singer's own vocal audio can be heard. Thus, the first audio can also be the audio obtained by performing sound effects processing and / or mixing processing on the original vocal audio, the accompaniment audio, and the singer's vocal audio. Alternatively, the first audio can be the audio obtained by performing sound effects processing and / or mixing processing on the target song (which includes the original vocal audio and the accompaniment audio) and the singer's vocal audio.

[0009] In the above method, the first vocal range is related to the singer's location, and the second vocal range is related to the locations of both the singer and the listener. The system controls the playback of a first audio file containing the original vocals within the first vocal range, and a second audio file containing both the singer's vocals and the accompaniment within the second vocal range. This ensures that the singer's location allows them to hear the original vocals, the singer's vocals, and the accompaniment, while the listener's location allows them to hear only the singer's vocals and the accompaniment. This achieves a karaoke experience where the singer can hear the original vocals without the listener being disturbed by them, thus enhancing both the singer's and listener's experience.

[0010] In one possible implementation of the first aspect, the sound pressure level of the original vocal audio at the singer's location is higher than the sound pressure level of the original vocal audio at the listener's location.

[0011] For example, the sound pressure level of the original vocal audio at the listener's location is relatively low, for example, it may be a value approaching zero, or even zero, so that the listener is less affected by the original vocal audio or even unaffected.

[0012] By implementing the above method, the singer can hear the original vocal audio, which allows the singer to clearly hear the lyrics during the singing process. It also helps the singer to better grasp the details of the performance (such as timbre, pitch, breath control, and articulation), while the listener can hardly perceive the original vocal audio, greatly reducing the interference of the original vocal audio on the listener.

[0013] In one possible implementation of the first aspect, the original vocal audio includes a first original vocal track and a second original vocal track; the singers include a first singer and a second singer; when the first singer sings, the singer's vocal audio is the first vocal audio; when the second singer sings, the singer's vocal audio is the second vocal audio; the first vocal audio matches the first original vocal track, and the second vocal audio matches the second original vocal track; the first vocal range includes a third vocal range and a fourth vocal range, wherein the third vocal range is the spatial region where the first singer is located, and the fourth vocal range is the spatial region where the second singer is located; the aforementioned control of playing the first audio in the first vocal range includes:

[0014] When the first singer sings, control the playback of the first audio track in the third vocal register; the first audio track includes the original vocal track; and / or

[0015] When the second singer sings, the first audio is played in the fourth vocal register. The first audio includes the second original vocal track.

[0016] For example, the first original vocal track audio and the second original vocal track audio are both portions of the original vocal audio.

[0017] Implementing the above method allows for the playback of different original vocal tracks from the locations of different singers. Thus, in a duet between the first and second singers, if the first singer performs the first original vocal track and the second singer performs the second, then when the first singer sings, only the first singer is in the third vocal range, while other participants are in the second vocal range. Therefore, only the first singer hears the first original vocal track, and other participants are not disturbed by it. Similarly, when the second singer sings, only the second singer is in the fourth vocal range, while other participants are in the second vocal range. Therefore, only the second singer hears the second original vocal track, and other participants are not disturbed by it, thus improving the listening experience for both singers and listeners.

[0018] In one possible implementation of the first aspect, the audio processing method further includes the following during the singer's performance:

[0019] Detect the singer's location using microphone arrays and / or cameras;

[0020] Determine the first vocal register based on the singer's location.

[0021] Implementing the above methods enables automatic detection of the singer's location. Furthermore, detecting the singer's location using a microphone array effectively prevents the recording and leakage of passengers' appearance, movements, and behaviors, while also requiring low hardware costs. Detecting the singer's location using a camera allows for accurate positioning of individuals in two-dimensional or three-dimensional space. Combining microphone arrays and cameras in the detection of the singer's location improves the accuracy and reliability of the detection, achieving high-precision audio-visual collaborative positioning.

[0022] In one possible implementation of the first aspect, before the karaoke session begins, the audio processing method further includes: receiving first setting information input by the user, the first setting information being used to indicate the singer's location; and determining a first vocal range based on the first setting information.

[0023] This implementation allows users to set the singer's location before the karaoke session begins. This ensures that the singer can hear the original vocals when the karaoke session starts, but the listener is not disturbed by the original vocals, thus enhancing user engagement.

[0024] In one possible implementation of the first aspect, receiving first setting information input by the user includes: displaying a first state control of multiple locations; receiving a first operation by the user on the first state control of a first target location to obtain the first setting information; wherein the first target location belongs to multiple locations, and the first operation causes the first state control of the first target location to indicate that the first target location is selected as the singer's location.

[0025] For example, in response to the first operation, the first state control at the first target location is displayed as selected, such as highlighted.

[0026] By implementing the above method, a user-friendly interactive interface with multiple position status controls is presented to the user, making it convenient for the user to set the singer's position. The status display of the position control shows whether the position has been selected by the user as the singer's position, which increases user-friendliness.

[0027] In one possible implementation of the first aspect, where the microphone used by the singer is a handheld microphone, the audio processing method further includes: identifying the position of the handheld microphone via a camera; and determining a first vocal range based on the position of the handheld microphone.

[0028] Implementing the above method, automatic positioning of the handheld microphone is equivalent to automatic positioning of the singer. When the singer is using a handheld microphone, if there are many people participating in karaoke, there may be few singers. This method of positioning the singer by locating the handheld microphone used by the singer is simpler and more efficient than the method of first detecting the positions of people and then determining the singer's position from the detected positions of people.

[0029] In one possible implementation of the first aspect, the audio processing method further includes: during the karaoke process, detecting the singer's position using a microphone array and / or a camera to obtain the singer's position detection result; updating the first vocal range based on the singer's position detection result; and controlling the playback of the first audio in the updated first vocal range.

[0030] By implementing the above method, which involves either setting the singer's location by the user or detecting the position of the handheld microphone, the first vocal register can be determined before the karaoke session begins. Therefore, when the karaoke session starts, the first audio file can be played in the first vocal register, providing a better karaoke experience for the singer. As the karaoke session progresses, the singer's current position can be detected in real time via the microphone array and / or camera. When the singer's position changes, the first vocal register can be updated promptly, further enhancing the singer's experience throughout the entire karaoke session.

[0031] In one possible implementation of the first aspect, the audio processing method further includes: receiving second setting information input by a user, the second setting information being used to indicate the location of a target person who is not participating in karaoke; and determining a second sound zone based on the second setting information, the second sound zone not including the spatial area where the target person is located.

[0032] The above implementation method allows users to set the location of people who do not participate in karaoke. For example, if the driver wants to concentrate on driving and does not want to participate in karaoke, the driver's position can be set to be unaffected by the original singer's voice, accompaniment, and the singer's voice. This not only improves the karaoke experience for both singers and listeners, but also takes into account the feelings of those who do not want to participate in karaoke, thus enriching the application scenarios of karaoke.

[0033] In one possible implementation of the first aspect, receiving second setting information input by the user includes: displaying a second status control for multiple locations; receiving a second operation by the user on the second status control of a second target location to obtain the second setting information; wherein the second target location belongs to multiple locations, and the second operation causes the second status control of the second target location to indicate that the target person at the second target location does not participate in karaoke.

[0034] For example, in response to the second operation, the second state control at the second target location is displayed as selected, such as highlighted.

[0035] Implementing the above method provides users with an intuitive, interactive interface featuring status controls for multiple locations. This allows users to easily set the locations of those not participating in the karaoke session, and the status controls indicate whether a location has been selected as the location for those not participating. Considering that in most cases, the number of people participating in the karaoke session (including both singers and listeners) is greater than the number of those not participating, providing an interface for setting the locations of those not participating reduces the number of clicks required, enabling faster setup and improving user-friendliness.

[0036] In one possible implementation of the first aspect, detecting the singer's location via a microphone array and / or camera includes: obtaining M audio tracks for M locations based on microphone signals collected by the microphone array, where M is a positive integer; determining the probability that singing exists at the i-th location based on the audio track at the i-th location, where i is a positive integer less than or equal to M; and determining the singer's location from the M locations based on the obtained M probabilities.

[0037] By implementing the above method, the singer's location is detected through a microphone array. Since the microphone array only collects sound information to determine the singer's location, it effectively avoids the recording and leakage of the passenger's appearance, movements, behavior, and other private information, thus protecting the passenger's privacy and security. Compared to image-based singer location detection, the microphone's sound signal acquisition is not dependent on lighting conditions or obstructions, making it more environmentally adaptable, with lower hardware costs and more flexible deployment.

[0038] In one possible implementation of the first aspect, determining the probability that singing exists at position i based on the audio track at position i includes: determining the probability that singing exists at position i based on the audio track at position i and the original vocal audio.

[0039] By implementing the above method, the probability of singing at position i is not only related to the audio track at position i, but also to the original singer's audio. For example, it is affected by the similarity between the audio track at position i and the original singer's audio, which can improve the accuracy of estimating the probability of singing at each position.

[0040] In one possible implementation of the first aspect, when the original vocal audio includes multiple original vocal tracks, the audio processing method further includes: determining the similarity between the speech track at the i-th position and each of the multiple original vocal tracks; and using the original vocal track with the highest similarity to the speech track at the i-th position as the original vocal track matched at the i-th position.

[0041] By implementing the above method, the most similar original vocal track audio is matched for each position. In a multi-person chorus scenario, the singer at the singer's position can hear the original vocal track of their own performance without being affected by other original vocal tracks in the original vocal audio, thus improving the singer's karaoke experience.

[0042] In one possible implementation of the first aspect, before obtaining the audio tracks at the M locations, the audio processing method further includes: identifying the aforementioned M locations of the person using a camera.

[0043] Compared to the method of analyzing the microphone signals collected by the microphone array to obtain the audio tracks of all locations and then analyzing them to determine the singer's location, the above method first identifies the location of the person through the camera, then analyzes the microphone signals collected by the microphone array to obtain the audio track of the person's location, and finally analyzes to determine the singer's location. This method can improve the efficiency of singer location detection, especially when there are many locations but few people.

[0044] In one possible implementation of the first aspect, detecting the singer's location via a microphone array and / or camera includes: identifying lip movement features of individuals at M locations using the camera, where M is a positive integer; obtaining the probability that singing exists at the i-th location based on the lip movement features of the individual at the i-th location and the song features of the original vocal audio, where i is a positive integer less than or equal to M; and determining the singer's location from the M locations based on the obtained M probabilities.

[0045] By implementing the above method, the location of the singer can be identified through a camera. The camera can directly capture visual images of the person's lip movements. Through image analysis technology, the specific location of the person in two-dimensional or three-dimensional space can be accurately detected. Furthermore, by extracting lip movement features and comparing them with the song features of the original vocal audio, the singer's location can be determined from multiple locations of multiple people based on the similarity between the lip movement features and the song features, thus enriching the methods for detecting the singer's location.

[0046] In one possible implementation of the first aspect, when the original vocal audio includes multiple original vocal tracks, the probability of singing at the i-th position is obtained based on the lip movement characteristics of the person at the i-th position and the song characteristics of each original vocal track. This includes: determining the similarity between the lip movement characteristics of the person at the i-th position and the song characteristics of each original vocal track; and determining the probability of singing at the i-th position based on the maximum similarity. For example, the probability of singing at the i-th position is the maximum similarity between the lip movement characteristics of the person at the i-th position and the song characteristics of each original vocal track.

[0047] By implementing the above method, the higher the similarity between the lip movement characteristics of a person at a certain location and the song characteristics of each original vocal track, the greater the probability that there is singing at that location.

[0048] In one possible implementation of the first aspect, the audio processing method further includes: taking the original vocal track audio with the highest similarity among multiple original vocal tracks as the original vocal track audio matched at the i-th position.

[0049] By implementing the above method, the most similar original vocal track audio is matched for each position. In a multi-person chorus scenario, the singer at the singer's position can hear the original vocal track of their own performance without being affected by other original vocal tracks in the original vocal audio, thus improving the singer's karaoke experience.

[0050] In one possible implementation of the first aspect, the singer's location is determined from the M locations based on the obtained M probabilities, including: taking the location corresponding to the probability greater than a probability threshold among the M probabilities as the singer's location.

[0051] By implementing the above method, the higher the probability corresponding to the location, the higher the probability that a singer exists at that location. Thus, by comparing with the probability threshold, the current location of the singer can be accurately determined.

[0052] As an example, the probabilities greater than the probability threshold among the M probabilities include the first probability and the second probability. The singer's position includes the first position corresponding to the first probability and the second position corresponding to the second probability. The first vocal range includes the fifth vocal range and the sixth vocal range. The fifth vocal range is associated with the first position, and the sixth vocal range is associated with the second position. The sound pressure level of the original singer's voice in the fifth vocal range is the same as the sound pressure level of the original singer's voice in the sixth vocal range.

[0053] Here, "same sound pressure level" can be understood as two sound pressure levels being equal, or it can be understood as the difference between two sound pressure levels being within the allowable error range.

[0054] By implementing the above method, for the positions corresponding to the probability judgment conditions, the original vocal audio uses a uniform sound pressure level, so that the detected singer's position can well perceive the original vocal audio.

[0055] As another example, the probabilities greater than the probability threshold among the M probabilities include the first probability and the second probability. The singer's position includes the first position corresponding to the first probability and the second position corresponding to the second probability. Among them, the first vocal range includes the fifth vocal range and the sixth vocal range. The fifth vocal range is associated with the first position, and the sixth vocal range is associated with the second position. The first probability is greater than the second probability. The sound pressure level of the original singer's voice in the fifth vocal range is higher than the sound pressure level of the original singer's voice in the sixth vocal range.

[0056] The above implementation method is applicable to scenarios where the singer's position changes slightly. For example, when a singer sings and moves their body to the music, the singer's position may gradually shift from the right back row to the left back row. At a certain moment during this movement, a probability greater than the probability threshold is detected, including the probability corresponding to the right back row position and the probability corresponding to the left back row position. The probability corresponding to the right back row position is greater than the probability corresponding to the left back row position. By controlling the sound pressure level of the original vocal audio in the right back row position area to be greater than the sound pressure level of the original vocal audio in the left back row position area, a smooth transition of the sound pressure level of the original vocal audio is achieved, which does not sound abrupt to the singer.

[0057] Secondly, this application provides an audio processing apparatus, comprising: an acquisition unit for acquiring the original vocal audio and accompaniment audio of a target song, wherein the target song is a song sung by a singer; and an acquisition unit for acquiring the singer's vocal audio; and a processing unit for controlling the playback of a first audio in a first audio region and the playback of a second audio in a second audio region. The first audio region is the spatial region where the singer is located, and the first audio includes the original vocal audio; the second audio region includes the spatial region where the singer and the listener are located, and the second audio includes the singer's vocal audio and the accompaniment audio.

[0058] In one possible implementation of the second aspect, the sound pressure level of the original vocal audio at the singer's location is higher than the sound pressure level of the original vocal audio at the listener's location.

[0059] In one possible implementation of the second aspect, the original vocal audio includes a first original vocal track and a second original vocal track; the singers include a first singer and a second singer; when the first singer sings, the singer's vocal audio is the first vocal audio; when the second singer sings, the singer's vocal audio is the second vocal audio; the first vocal audio matches the first original vocal track, and the second vocal audio matches the second original vocal track; the first vocal range includes a third vocal range and a fourth vocal range, wherein the third vocal range is the spatial region where the first singer is located, and the fourth vocal range is the spatial region where the second singer is located; the processing unit is specifically used to: when the first singer sings, control the playback of the first audio in the third vocal range, the first audio including the first original vocal track; and / or, when the second singer sings, control the playback of the first audio in the fourth vocal range, the first audio including the second original vocal track.

[0060] In one possible implementation of the second aspect, during the singer's performance, the processing unit is also used to: detect the singer's location via a microphone array and / or a camera; and determine the first vocal register based on the singer's location.

[0061] In one possible implementation of the second aspect, the acquisition unit is further configured to receive first setting information input by the user, the first setting information being used to indicate the singer's location; the processing unit is further configured to determine the first vocal range based on the first setting information.

[0062] In one possible implementation of the second aspect, the device further includes a display unit for displaying first status controls for multiple locations; the acquisition unit is specifically used to receive a first operation by the user on the first status control of the first target location to obtain first setting information; wherein the first target location belongs to multiple locations, and the first operation causes the first status control of the first target location to indicate that the first target location is selected as the singer's location.

[0063] In one possible implementation of the second aspect, where the microphone used by the singer is a handheld microphone, the processing unit is further configured to: identify the position of the handheld microphone via a camera; and determine the first vocal range based on the position of the handheld microphone.

[0064] In one possible implementation of the second aspect, the processing unit is further configured to: detect the singer's position through a microphone array and / or a camera during the karaoke process, and obtain the singer's position detection result; update the first vocal range based on the singer's position detection result; and control the playback of the first audio in the updated first vocal range.

[0065] In one possible implementation of the second aspect, the acquisition unit is further configured to: receive second setting information input by the user, the second setting information being used to indicate the location of the target person who is not participating in karaoke; the processing unit is further configured to: determine a second vocal range based on the second setting information, the second vocal range not including the spatial area where the target person is located.

[0066] In one possible implementation of the second aspect, the display unit is further configured to display second status controls for multiple locations; the acquisition unit is further configured to receive a second operation by the user on the second status control of the second target location to obtain second setting information; wherein the second target location belongs to multiple locations, and the second operation causes the second status control of the second target location to indicate that the target person at the second target location does not participate in karaoke.

[0067] In one possible implementation of the second aspect, the processing unit is specifically used to: obtain speech tracks at M locations based on the microphone signals collected by the microphone group, where M is a positive integer; determine the probability that singing exists at the i-th location based on the speech track at the i-th location, where i is a positive integer less than or equal to M; and determine the location of the singer from the M locations based on the obtained M probabilities.

[0068] In one possible implementation of the second aspect, the processing unit is specifically used to: determine the probability that singing exists at the i-th position based on the audio track and the original vocal audio at the i-th position.

[0069] In one possible implementation of the second aspect, when the original vocal audio includes multiple original vocal tracks, the processing unit is further configured to: determine the similarity between the speech track at the i-th position and each of the multiple original vocal tracks; and use the original vocal track with the highest similarity to the speech track at the i-th position as the original vocal track matched at the i-th position.

[0070] In one possible implementation of the second aspect, the processing unit is further configured to: identify the aforementioned M locations of the person via a camera.

[0071] In one possible implementation of the second aspect, the processing unit is specifically used to: identify the lip movement features of people at M locations using a camera, where M is a positive integer; obtain the probability that there is singing at the i-th location based on the lip movement features of the person at the i-th location and the song features of the original vocal audio, where i is a positive integer less than or equal to M; and determine the location of the singer from the M locations based on the obtained M probabilities.

[0072] In one possible implementation of the second aspect, when the original vocal audio includes multiple original vocal tracks, the processing unit is specifically used to: determine the similarity between the lip movement features of the person at the i-th position and the song features of each original vocal track; and determine the probability that singing exists at the i-th position based on the maximum similarity.

[0073] In one possible implementation of the second aspect, the processing unit is further configured to: use the original vocal track audio with the highest similarity among multiple original vocal track audios as the original vocal track audio matched at the i-th position.

[0074] In one possible implementation of the second aspect, the processing unit is specifically used to: take the position corresponding to the probability greater than the probability threshold among the M probabilities as the position of the singer.

[0075] As an example, the probabilities greater than the probability threshold among the M probabilities include the first probability and the second probability. The singer's position includes the first position corresponding to the first probability and the second position corresponding to the second probability. The first vocal range includes the fifth vocal range and the sixth vocal range. The fifth vocal range is associated with the first position, and the sixth vocal range is associated with the second position. The sound pressure level of the original singer's voice in the fifth vocal range is the same as the sound pressure level of the original singer's voice in the sixth vocal range.

[0076] As another example, the probabilities greater than the probability threshold among the M probabilities include the first probability and the second probability. The singer's position includes the first position corresponding to the first probability and the second position corresponding to the second probability. Among them, the first vocal range includes the fifth vocal range and the sixth vocal range. The fifth vocal range is associated with the first position, and the sixth vocal range is associated with the second position. The first probability is greater than the second probability. The sound pressure level of the original singer's voice in the fifth vocal range is higher than the sound pressure level of the original singer's voice in the sixth vocal range.

[0077] Thirdly, this application provides an apparatus for audio processing, the apparatus including a processor and a memory, wherein the memory is used to store program instructions; the processor calls the program instructions in the memory, causing the apparatus to execute the method in the first aspect or any possible implementation of the first aspect.

[0078] Fourthly, this application provides an audio processing system, which includes a control device, a microphone group, and a speaker group, wherein the control device is connected to the microphone group and the speaker group respectively, and the control device is used to execute the method in the first aspect or any possible implementation of the first aspect.

[0079] Optionally, the audio processing system also includes a camera, which is used to transmit captured image data containing people to the control device, and the image data is used to determine the location of the singer.

[0080] Optionally, the microphone set includes a vehicle-mounted microphone and / or a handheld microphone.

[0081] Fifthly, this application provides a vehicle that includes the apparatus as described in the second aspect or any possible implementation of the second aspect, or includes the apparatus described in the third aspect, or includes the audio processing system as described in the fourth aspect or any possible implementation of the fourth aspect.

[0082] In a sixth aspect, this application provides a computer-readable storage medium including computer instructions that, when executed by a processor, implement the method in the first aspect or any possible implementation thereof.

[0083] In a seventh aspect, this application provides a computer program product that, when executed by a processor, implements the methods described in the first aspect or any possible embodiment of the first aspect. The computer program product may, for example, be a software installation package. When the methods provided by any possible design of the first aspect are required, the computer program product can be downloaded and executed on a processor to implement the methods described in the first aspect or any possible embodiment of the first aspect.

[0084] The technical effects of the second to seventh aspects mentioned above can be referred to the description of the first aspect above, and will not be repeated here. Attached Figure Description

[0085] Figure 1A This is a schematic diagram of the architecture of an audio processing system provided in an embodiment of this application;

[0086] Figure 1B This is a schematic diagram of the architecture of another audio processing system provided in the embodiments of this application;

[0087] Figure 1C This is a schematic diagram of the architecture of another audio processing system provided in the embodiments of this application;

[0088] Figure 2A This is a schematic diagram of an in-vehicle karaoke scenario provided in an embodiment of this application;

[0089] Figure 2B This is a schematic diagram of a smart home karaoke scenario provided in an embodiment of this application;

[0090] Figure 3 This is a flowchart of a karaoke processing method provided in an embodiment of this application;

[0091] Figure 4 This is a schematic diagram of a display interface for setting the singer's position in a vehicle-mounted scenario, provided in an embodiment of this application.

[0092] Figure 5A This is a schematic diagram illustrating how to determine the location of a singer, provided in an embodiment of this application.

[0093] Figure 5B This is another schematic diagram provided in the embodiments of this application for determining the location of a singer;

[0094] Figure 5C This is a schematic diagram illustrating a process of detecting the singer's location using a microphone array, as provided in an embodiment of this application.

[0095] Figure 5D This is a schematic diagram illustrating a process of detecting the singer's location using a camera, as provided in an embodiment of this application.

[0096] Figure 6A This is a schematic diagram of a display interface for setting the location of people who are not participating in karaoke in a vehicle scenario, provided by an embodiment of this application.

[0097] Figure 6B These are schematic diagrams of some second vocal registers provided in embodiments of this application;

[0098] Figure 6C This is a schematic diagram of a display interface for setting the location of people participating in karaoke in a vehicle-mounted scenario, provided in an embodiment of this application.

[0099] Figure 7A This is a schematic diagram illustrating a method for original vocal track matching based on a microphone array, as provided in an embodiment of this application.

[0100] Figure 7B This is a schematic diagram illustrating original vocal track matching based on a camera, provided in an embodiment of this application.

[0101] Figure 7C This is a schematic diagram of an in-vehicle karaoke scenario provided in an embodiment of this application;

[0102] Figure 8A This is a schematic diagram of an in-vehicle karaoke scenario provided in an embodiment of this application;

[0103] Figure 8B This is a schematic diagram of an in-vehicle karaoke scenario provided in an embodiment of this application;

[0104] Figure 9 This is a schematic diagram of a vocal register division provided in an embodiment of this application;

[0105] Figure 10 This is a schematic diagram of another type of vocal register division provided in an embodiment of this application;

[0106] Figure 11 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application;

[0107] Figure 12 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation

[0108] In this scheme, prefixes such as "first" and "second" are used solely to distinguish different descriptive objects and do not impose any restrictions on the position, order, priority, quantity, or content of the described objects. For example, if the described object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the modified "fields" are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the described object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority of the "levels." Furthermore, the number of described objects is not limited by prefixes; it can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device," then "first device" and "second device" can be the same device, devices of the same type, or devices of different types. Similarly, if the object being described is "information," then "first information" and "second information" can be information with the same content or information with different content. In summary, the use of prefixes to distinguish the objects being described in the embodiments of this application does not constitute a limitation on the objects being described. The description of the objects being described is based on the claims or the context of the embodiments, and should not constitute an unnecessary limitation due to the use of such prefixes.

[0109] To facilitate understanding, the relevant terms that may be involved in the embodiments of this application will be introduced below.

[0110] (1) Vocal register

[0111] A sound zone refers to a spatial area where audio is played. For example, a specific spatial area (such as the cabin of a vehicle or airplane, a karaoke room, a home theater, or a living room) can be logically divided into areas with different audio characteristics or functions using specific technologies, and each area can be called a sound zone.

[0112] As an example, the vehicle cabin can be divided into audio zones based on occupants, resulting in different zones such as the driver's area, front passenger area, rear left seat area, and rear right seat area, to provide personalized audio services for occupants in different positions. In this case, one seat in the cabin corresponds to one audio zone. In some solutions, the vehicle cabin can also be divided according to functional needs, such as navigation audio zones, entertainment audio zones, and communication audio zones, to meet different functional requirements, such as ensuring the driver can clearly hear navigation instructions, passengers can enjoy entertainment content, and conduct private communications.

[0113] Here, the method of dividing vocal registers is not limited. For example, in a karaoke setting, vocal registers can be divided based on the location of the singer and the listener. As the singer moves, the vocal registers can also change accordingly. In this scheme, different vocal registers can partially overlap or not overlap at all, and one vocal register can be contained within another.

[0114] (2) Singer and listener

[0115] Singer refers to the actual person performing in a karaoke setting; there can be one or more singers. Listener refers to the audience in a karaoke setting.

[0116] It is understandable that when multiple singers participate in a karaoke session, each singer also acts as a listener while the other singers are singing, meaning that the singers also have the identity of listeners.

[0117] The above terms may be used in the following embodiments.

[0118] This solution provides an audio processing system that can determine different sound zones based on the location of the singer and the listener. By controlling the playback of audio (i.e., the singer's vocals, the original vocals, and the accompaniment) in different sound zones through speakers, the system enables the singer to hear the original vocals in a karaoke setting, while the listener is not disturbed by the original vocals. This ensures a good karaoke experience for the singer while also enhancing the listener's experience.

[0119] The following describes the components of an audio processing system.

[0120] See Figure 1A , Figure 1A This is a schematic diagram of the architecture of an audio processing system provided in an embodiment of this application. Figure 1A As shown, the audio processing system 10 includes a control device, a microphone group, and a speaker group, wherein the control device communicates with the microphone group in a wired and / or wireless manner, and the control device communicates with the speaker group in a wired and / or wireless manner.

[0121] Here, the control device can determine different vocal registers based on the location of the singer and the listener, and control the playback of audio (i.e., singer's vocals, original vocals, and accompaniment) in different registers through a speaker array. In some solutions, the control device can also identify / detect the singer's location through a microphone array.

[0122] A control device is a device with computing capabilities, such as a system-on-chip (SOC), an artificial intelligence (AI) accelerator chip, or a component within a chip, such as an integrated circuit or a processor.

[0123] For example, when applied to a vehicle, the control device can be a domain controller within the vehicle or a component within a domain controller. Components can be, for example, chips, control units, integrated circuits, etc. For instance, the domain controller can be a hardware-software integrated platform for providing in-vehicle multimedia services (such as at least one of head-up display, instrument panel display, and entertainment system), such as a cockpit domain controller (CDC, also referred to as the vehicle infotainment system). In some solutions, the domain controller can also be a controller that integrates at least one of a vehicle domain controller (VDC) and a mobile data center (MDC) with the CDC. Here, the VDC is an example of a hardware-software integrated platform supporting body control and chassis control, and the MDC is an example of a hardware-software integrated platform supporting intelligent driving, i.e., an in-vehicle computing platform. As an example, by integrating the VDC, MDC, and CDC, a domain controller capable of providing body control functions, autonomous driving control functions, and cockpit control functions can be obtained. In this case, the domain controller can also be referred to as a central computing unit.

[0124] The microphone array is used to capture the singer's vocal audio. This vocal audio comes from the microphone the singer uses while singing; the microphone can be a handheld microphone or a vehicle-mounted microphone. When used in a vehicle, the microphone array is used to capture in-vehicle voice signals, which may include the singer's vocal audio, user voice commands, or other audio inputs.

[0125] For example, the microphone array includes handheld microphones and / or in-vehicle microphones. For instance, the deployment locations of in-vehicle microphones include, but are not limited to, the overhead area of ​​the vehicle, the center console area, the steering wheel area, the inside of the doors, and near the seats.

[0126] The speaker assembly includes at least one speaker, each speaker being used to play audio. For example, when applied to a vehicle, the speaker assembly includes external speakers such as body speakers, headrest speakers, and shoulder pillow speakers.

[0127] For example, combined Figure 1A The control device acquires the original vocal audio, the accompaniment audio, and the singer's vocal audio from the microphone array. By sending control signals to the speaker array, the control device enables the playback of the original vocal audio in the first register and the playback of the singer's vocal audio and the accompaniment audio in the second register. The first register is the spatial region where the singer is located, and the second register includes the spatial region where both the singer and the listener are located. Figure 1AIn this diagram, the first register is represented by a light-colored circle, and the second register by a dark-colored ellipse. It can be seen that the second register encompasses the first register. Here, the representation of registers is merely illustrative; in some schemes, registers may also be represented by other planar or three-dimensional shapes.

[0128] As an example, the control signal mentioned above includes the original vocal audio, or the control signal includes the singer's vocal audio and the accompaniment audio. In some solutions, where the speaker does not have a mixing function, the control signal includes a mixture of the original vocal audio, the singer's vocal audio, and the accompaniment audio, or the control signal includes a mixture of the singer's vocal audio and the accompaniment audio.

[0129] In some possible embodiments, the audio processing system 10 described above also includes a camera. The camera is used to capture image data containing people.

[0130] See Figure 1B , Figure 1B This is a schematic diagram of the architecture of another audio processing system provided in the embodiments of this application. Compared to Figure 1A The audio processing system 10 shown, Figure 1B The audio processing system 10 shown also includes a camera, wherein the control device communicates with the camera in a wired and / or wireless manner.

[0131] For example, in Figure 1B In this system, the control device can identify / detect the singer's location via a camera, or a combination of a camera and a microphone. For example, the control device can receive image data captured by the camera as input, or it can receive image data captured by the camera and vocal signals captured by the microphone as input, and process this input data to determine the singer's location.

[0132] Furthermore, when there are multiple singers and multiple original vocal tracks are involved in the original vocal audio, the control device also has an original vocal matching function, which can match the corresponding original vocal track to the singer's vocal audio. Please refer to the following implementation method. Figure 1C The narrative.

[0133] See Figure 1C , Figure 1C This is an application diagram of an audio processing system provided in an embodiment of this application. Figure 1C In the diagram, the first vocal range is represented by a light-colored circle, and the second vocal range by a dark-colored rectangle. It can be seen that the first vocal range includes register 1 and register 2, where register 1 is the spatial area occupied by singer 1, and register 2 is the spatial area occupied by singer 2. There are four people in the second vocal range.

[0134] For example, in Figure 1C In this process, the original vocal audio acquired by the control device includes original vocal audio 1 and original vocal audio 2, meaning that there are two original vocalists involved in the original vocal audio. The control device detects the location of singer 1 and singer 2 using at least one of the camera and microphone arrays. Based on the location of singer 1, it determines the aforementioned vocal range 1, and based on the location of singer 2, it determines the aforementioned vocal range 2. Depending on the singer performing, the control device controls the playback of the audio in different vocal ranges differently; please refer to the descriptions of Case 1 and Case 2 below.

[0135] Scenario 1: When singer 1 performs, the aforementioned singer's vocal audio is singer's vocal audio 1. Assuming the control device recognizes that singer's vocal audio 1 matches the original singer's vocal audio 1, the control device plays the original singer's vocal audio 1 in zone 1 through the speaker array, and plays singer's vocal audio 1 and accompaniment audio in zone 2. In this karaoke scenario, singer 1 is the singer, and the other three are listeners.

[0136] Scenario 2: When singer 2 performs, the aforementioned singer's vocal audio is singer's vocal audio 2. Assuming the control device recognizes that singer's vocal audio 2 matches the original singer's vocal audio 2, the control device plays the original singer's vocal audio 2 in zone 2 via the speaker array, and plays singer's vocal audio 2 and accompaniment audio in zone 2. In this karaoke scenario, singer 2 is the singer, and the other three are listeners.

[0137] Figure 1A , Figure 1B or Figure 1C The audio processing system shown can be applied to a variety of application scenarios, such as: mobile internet (MI), self-driving, transportation safety, internet of things (IoT), smart city, or smart home.

[0138] Figure 1A , Figure 1B or Figure 1CThe audio processing system shown can be applied to various network types, such as one or more of the following: SparkLink, Long Term Evolution (LTE) networks, 5th generation mobile communication technology (5G), wireless local area networks (e.g., Wi-Fi), Bluetooth (BT), Zigbee, or vehicular short-range wireless communication networks, etc.

[0139] Figure 1A , Figure 1B or Figure 1C This is merely an illustrative architecture diagram. (The above...) Figure 1A , Figure 1B or Figure 1C The karaoke processing method provided in the following embodiments is applicable. In some solutions, the control device described above can also be divided into more modules according to its function. For example, for... Figure 1A The control device includes an audio signal processing module and a sound zone playback control module.

[0140] For example, the audio signal processing module includes a microphone signal processing module and a sound effects module. The microphone signal processing module is used to filter out noise from the acquired audio signal; the sound effects module can process and enhance sound effects by adding special effects, adjusting equalization, and controlling dynamic range, and supports multi-channel and surround sound processing to create an immersive listening scene, and can also complete the synthesis and mixing of multiple audio sources. In some solutions, the microphone signal processing module can also be integrated into the microphone, and the sound effects module can be integrated into the speaker.

[0141] The frequency range playback control module is used to control the speakers to play audio in different frequency ranges. Further, the frequency range playback control module includes an independent frequency range playback control module and a full-range playback control module. The independent frequency range playback control module controls the playback of audio in the first frequency range, and the full-range playback control module controls the playback of audio in the second frequency range. Figure 1B The control device also includes a voice processing module and a position detection module. The voice processing module is used to process the audio signals collected by the microphone group to obtain multiple audio tracks, and the position detection module is used to determine the singer's location.

[0142] The above Figures 1A-1C The audio processing system shown can be applied to various scenarios, such as in-vehicle scenarios, smart homes, and smart communities. For ease of understanding, the in-vehicle karaoke scenario and the smart home karaoke scenario are described below.

[0143] See Figure 2A , Figure 2A This is a schematic diagram of a car karaoke scenario provided in an embodiment of this application. Figure 2A In this system, the aforementioned control device can be the vehicle's infotainment system or integrated into it. The vehicle 200 is equipped with a camera 211 and a microphone 221. For example, a user (e.g., a driver or passenger) activates the karaoke function by operating the screen. When the passenger sings, the infotainment system acquires the original vocal audio and accompaniment audio of the target song being sung. The microphone 221 is used to capture the passenger's vocal audio and can also receive voice commands from passengers inside the vehicle. The camera 211 is used to capture image data of passengers inside the vehicle. Thus, the control device can determine the location of the passenger (i.e., the singer), thereby determining the vocal range, and finally control the playback of audio in different vocal ranges through the speakers, enabling the singer to hear the original vocals in the in-vehicle karaoke scenario without disturbing the listener.

[0144] For example, the screen includes one or more of the following: a physical screen (such as a central control screen), a projection system, a smart entity, or a button panel. The projection system includes, for example, a light field screen, a head-up display (HUD), or other projection systems.

[0145] Understandable. Figure 2A This is just one example. In some solutions, the vehicle's infotainment system can also have a connection device, which can be a user's mobile phone, tablet, PDA, or other user terminal. The user terminal is used to provide the target song that the user wants to sing to the infotainment system.

[0146] See Figure 2B , Figure 2B This is a schematic diagram of a smart home karaoke scenario provided in an embodiment of this application. See also... Figure 2B The home system includes a television 210 and a microphone 221. The control device is integrated into the television 210, and the microphone 221 is used to capture the singer's vocal video. The television 210 supports karaoke and is also equipped with a camera 211 and a speaker 231. The camera 211 is used to capture image data, which is used to analyze the singer's location. The control device controls the speaker 231 to play audio (e.g., the singer's vocal audio, the original singer's vocal audio, and the accompaniment audio), enabling the singer to hear the original vocals without disturbing the listener.

[0147] exist Figure 2A Passengers can also sing karaoke using handheld microphones. (The above...) Figure 2A , Figure 2B The scenario shown is merely an example; in some solutions, it can also be... Figure 2A or Figure 2BThe scenario shown illustrates a multi-person karaoke session. Besides the aforementioned scenario, this solution can also be applied to other karaoke settings, such as karaoke rooms, airplane cabins, and conference rooms.

[0148] See Figure 3 , Figure 3 This is a flowchart of a karaoke processing method provided in an embodiment of this application. This method can be applied to the aforementioned control device. Figure 3 The method shown includes, but is not limited to, the following steps S301-S304.

[0149] S301: The control device acquires the original vocal audio and accompaniment audio of the target song.

[0150] In this context, the target song is the song sung by the singer. For example, the vehicle's infotainment system detects a user's voice command including phrases like "I want to sing," or the user clicks on a karaoke app installed on the system. The system then launches the karaoke app, and the user selects the target song through it. Here, the user can be the singer, driver, front passenger, or other occupants of the vehicle. In some solutions, the target song can also be selected and confirmed by the user through the karaoke app on their mobile device.

[0151] In one implementation, when the control device is deployed independently of the vehicle's infotainment system, the control device acquires the original vocal audio and accompaniment audio of the target song, including: the control device acquiring the original vocal audio and accompaniment audio of the target song from the vehicle's infotainment system. If the control device is integrated into the vehicle's infotainment system, the control device acquires the original vocal audio and accompaniment audio of the target song locally.

[0152] For example, the original vocal audio and the accompaniment audio can be obtained by the vehicle system through vocal separation processing of the target song, or the vehicle system can directly obtain the original vocal audio and the accompaniment audio of the target song from the local karaoke application, or the vehicle system can receive the original vocal audio and the accompaniment audio sent by the user terminal.

[0153] In some solutions, when the control device is integrated into the vehicle's infotainment system, it acquires the original vocal and instrumental audio from the target song locally; alternatively, it receives the original vocal and instrumental audio from the target song sent by the user terminal. In another solution, the control device may also perform vocal separation processing on the acquired target song to obtain the original vocal and instrumental audio.

[0154] In some possible embodiments, when the target song involves a collaboration of multiple singers, the original vocal audio includes multiple original vocal tracks, each corresponding to one singer in the target song. For example, if the target song is sung by singer A and singer B, the original vocal audio includes the original vocal tracks of singer A and singer B, where singer A's original vocal track only includes singer A's original vocals, and singer B's original vocal track only includes singer B's original vocals.

[0155] S302: The control device acquires the singer's vocal audio.

[0156] Here, the singer's vocal audio refers to the a cappella performance of the target song, which highlights the singer's vocal qualities, pitch, rhythm, and emotional expression. For example, the singer's vocal audio includes the singer's vocals, breath sounds, and subtle sounds from the mouth and throat related to the performance of the target song.

[0157] The singer's vocal audio can come from a handheld microphone used by the singer or from a car microphone. A car microphone, also known as a cabin microphone, is used when a singer sings using a car microphone instead of a handheld microphone; this is also known as microphoneless karaoke.

[0158] In one implementation, a singer uses a handheld microphone to sing, and a control device acquires the singer's vocal audio. This includes the control device acquiring the singer's vocal audio from the handheld microphone used by the singer. Handheld microphones can be wired or wireless. For wired handheld microphones, the singer's vocal audio is transmitted to the control device via an audio cable; for wireless handheld microphones, the singer's vocal audio is transmitted to the control device wirelessly via Bluetooth or radio frequency technology. When there are multiple singers, they can share a single handheld microphone, or each singer can have their own handheld microphone.

[0159] In another implementation, the control device acquires the singer's vocal audio by obtaining the singer's vocal audio from the microphone signals collected by the cockpit microphone. In this case, the vehicle microphones do not correspond one-to-one with the seats in the vehicle. For example, if there are four seats in the vehicle but five vehicle microphones are deployed, the microphone signals collected by the five vehicle microphones can be processed to obtain the audio track at the singer's location, and this audio track carries the singer's vocal audio. In some solutions, the vehicle microphones may also correspond one-to-one with the seats in the vehicle. In this case, the control device can acquire the singer's vocal audio from the vehicle microphone corresponding to the singer's seat.

[0160] Optionally, S303 can also be executed.

[0161] S303: The control device obtains the singer's location and determines the first vocal range based on the singer's location.

[0162] For example, the number of singer positions is related to the number of singers at any given moment. For instance, if only one singer is singing at any given moment, the number of singer positions is one; if multiple singers are singing simultaneously, the number of singer positions is multiple.

[0163] For example, the singer's location can be represented by coordinate values ​​in any coordinate system (e.g., two-dimensional or three-dimensional coordinates). For instance, these coordinate values ​​could be the position of the singer's head, the center point of the singer's seat, a vertex of the headrest, or the position of the handheld microphone used by the singer. When applied to a vehicle cabin, the coordinate system can be a local coordinate system, and the singer's location can be represented as a three-dimensional coordinate system consisting of the X, Y, and Z coordinates. For example, the local coordinate system could have its origin at any point on the vehicle's seat or in the space above the seat, with the X direction pointing to the left side of the vehicle, the Y direction pointing to the front, and the Z direction pointing to the roof. In one implementation, the coordinate system could also be a vehicle body coordinate system, a World Geodetic System 1984 (WGS84), etc.

[0164] In some solutions, the singer's location can be represented by location identification information, which indicates the coordinates of a seat or the spatial range of a seat. For example, when applied to a vehicle cabin, each seat in the cabin is a location with its own identification information. The memory stores the correspondence between the seat identification information and the seat's coordinates (or spatial range). Given the identification information of the singer's seat, the coordinates or spatial range of that seat can be retrieved based on the seat's representation information, thus revealing the singer's location.

[0165] The first vocal range mentioned above is the spatial area where the singer is located. In other words, the first vocal range defines a range of movement for the singer at their current location, and the first vocal range is associated with the singer's location.

[0166] In one implementation, if the obtained location of the singer is a target spatial region, then the first vocal range is the target spatial region; if the obtained location of the singer is a target position coordinate, then the first vocal range is determined according to the location of the singer, including: determining the first vocal range according to the target position coordinate and a preset rule, wherein the preset rule is used to indicate the method of generating the first vocal range based on the target position coordinate.

[0167] For example, if the target position coordinates are the location of the singer's head, the preset rule could be to construct a spherical spatial region centered on the target position coordinates and with a first length as the radius; or, to construct a cubic spatial region centered on the target position coordinates and with a first length as the side length. As another example, if the target position coordinates are a vertex of the headrest of the singer's chair, the preset rule could be to construct a cuboid spatial region with the target position coordinates as the vertex, a first length as the length, a second length as the width, and a third length as the height. Here, the spatial region represented by the first vocal register can also be a cylinder, an ellipsoid, or other irregular shapes.

[0168] In another implementation, when applied to a vehicle cabin, each seat corresponds to a spatial region, and the spatial regions corresponding to different seats do not overlap; that is, each seat corresponds to an independent vocal range. In this case, determining the first vocal range based on the singer's location involves comparing the singer's location with the spatial region corresponding to each seat in the cabin. If the singer's location falls within the first spatial region corresponding to the first seat, then the first vocal range is determined as the first spatial region. Implementing this method only requires finding the corresponding spatial region based on the singer's location to obtain the first vocal range, rather than calculating it based on the singer's location, thus saving computational resources for the control device.

[0169] In some schemes, when there are multiple singers, the number of locations for each singer is also multiple. Taking multiple singers including a first singer and a second singer as an example, the first vocal range includes vocal range 1 and vocal range 2. Vocal range 1 is associated with the location of the first singer, and vocal range 2 is associated with the location of the second singer. For example, the association of vocal range 1 with the location of the first singer can be: determining vocal range 1 based on the location of the first singer. The method for determining vocal range 1 is similar to the aforementioned method for determining the first vocal range, and will not be repeated here.

[0170] For example, there is no overlap between register 1 and register 2, so that register 1 and register 2 do not interfere with each other.

[0171] In this solution, there are multiple ways to obtain the singer's location, such as user-defined location settings, location based on the position of the singer's handheld microphone, or detection of the singer's location using vehicle sensors (e.g., microphone arrays and / or cameras). The implementation methods are described below; please refer to the descriptions of implementation methods 1-5.

[0172] Implementation method 1: User sets the singer's position.

[0173] As an example, the control device obtains the singer's location by: receiving first setting information input by a user, the first setting information indicating the singer's location; and determining the singer's location based on the first setting information. When applied to a vehicle cabin, Figure 5A The process of implementation method 1 is shown.

[0174] This implementation allows users to set the singer's location before the karaoke session begins. This ensures that the singer can hear the original vocals when the karaoke session starts, but the listener is not disturbed by the original vocals. This not only ensures a good karaoke experience for the singer but also enhances the listener's experience.

[0175] For example, receiving first setting information input by a user includes receiving first setting information input by the user via at least one of a touchscreen, button, keyboard, and voice input. The user may be a singer, driver, front passenger, or other occupant of the vehicle.

[0176] Taking voice input as an example, voice input can be a user inputting a voice command such as "set the passenger seat as the singer's position" (i.e., the first setting information). The voice command contains keywords such as "seat" and "singer" to set the singer's position.

[0177] As an example, when a user inputs via a touchscreen, the system receives first setting information input by the user, including: displaying first status controls for multiple locations; receiving a first operation by the user on the first status control for a first target location to obtain the first setting information; wherein the first target location belongs to these multiple locations, and the first operation causes the first status control of the first target location to indicate that the first target location is selected as the singer's position.

[0178] The following is based on Figure 4 This indicates that the user inputs the first setting information mentioned above via a touchscreen. Here, the touchscreen is an example of the screen described above.

[0179] See Figure 4 , Figure 4 This is a schematic diagram of a display interface for setting the singer's position in a vehicle-mounted scenario, provided in an embodiment of this application. Figure 4 The displayed interface is a human-computer interaction interface. Figure 4 The interface shown displays status controls for seven positions within the vehicle's cabin: driver's seat, front passenger seat, front left seat, front right seat, rear left seat, rear center seat, and rear right seat.

[0180] for Figure 4The status control at any of the shown locations changes color from light to dark when clicked by the user. The dark color indicates that the user has selected that location as the singer's position. Clicking the location again will change the status control back from dark to light, indicating that the location is no longer selected. Alternatively, the user can click the clear button to cancel all current settings for the singer's position. It's understood that the user can... Figure 4 At least one of the positions shown is set to the location of the singer. After the singer's position is set, the user clicks the save button to store the current settings as the first setting information.

[0181] For example, suppose the user clicks Figure 4 In response to this operation, the status control for the "passenger seat" position in the interface shown changes from a light color to a dark color. When the user clicks the save button, the first setting information is generated, which indicates that the singer's position is the passenger seat area.

[0182] For example, Figure 4 The interface shown can automatically pop up when the user launches the karaoke app. In some solutions, it can also be displayed in response to the user clicking to set the singer's position in the karaoke app. Figure 4 The interface shown.

[0183] The above Figure 4 This is merely an example of a display interface provided for users to input initial settings information, and is not intended to limit the display interface to only one type of information. Figure 4 The format shown can also be other display interfaces that allow users to set the singer's position. In some solutions, Figure 4 The interface shown can also display more or less information than currently displayed. For example, Figure 4 The number of locations (or seats) shown on the interface can be more or less, or... Figure 4 The interface can also display prompts such as "The original vocals will be enabled at the singer's location".

[0184] The above implementation allows users to set the singer's location in advance, providing a good karaoke experience for both the singer and the listener when the karaoke session begins, and also enhancing user engagement.

[0185] Implementation method 2: Use the position of the identified handheld microphone as the singer's location.

[0186] As an example, when the singer is using a handheld microphone, the control device obtains the singer's location by: the control device identifying the position of the handheld microphone through a camera; and using the position of the handheld microphone as the singer's location. In this case, "determining the first vocal range based on the singer's location" is equivalent to "determining the first vocal range based on the position of the handheld microphone."

[0187] Here, the location of the handheld microphone refers to its position within a specific spatial area (such as the vehicle cabin).

[0188] For example, the control device identifies the position of the handheld microphone through a camera, including: the control device acquiring image data containing the handheld microphone captured by the camera; the control device extracting and identifying the features of the handheld microphone from the image data using a target detection algorithm, and locating the identified handheld microphone by combining spatial coordinate calculations to obtain the position of the handheld microphone.

[0189] Here, the object detection algorithm can be, for example, Faster R-CNN, YOLO, etc. For example, the features of the handheld microphone include, but are not limited to, one or more of the following: shape features, color features, brand logo features, etc. It can be understood that the position of the handheld microphone obtained through camera recognition is also called the visual localization result of the handheld microphone by the camera.

[0190] For example, the above spatial coordinate calculation process can be as follows: The pixel coordinates of the feature points of the handheld microphone are converted into actual three-dimensional spatial coordinates using the camera's calibration parameters. Based on the three-dimensional spatial coordinates corresponding to each feature point of the handheld microphone, the position of the handheld microphone is determined. In some solutions, when there are multiple cameras, different perspectives of the same feature point captured by different cameras can be used to obtain the three-dimensional spatial coordinates of each feature point based on the principles of triangulation and the camera's calibration parameters. Finally, the coordinates calculated by different cameras are transformed into the same coordinate system, and the coordinate data is fused to obtain the position of the handheld microphone. This improves the accuracy of handheld microphone position detection.

[0191] Here, the camera's calibration parameters include intrinsic parameters (such as focal length, principal point coordinates, etc.) and extrinsic parameters (such as rotation matrix, translation vector, etc.). Feature points of the handheld microphone include, for example, the endpoints of the handheld microphone and specific points on the edges of the handheld microphone.

[0192] It is understandable that when applied to vehicle cabins, Figure 5B The process of implementing method 2 is shown.

[0193] In the above implementation, locating the handheld microphone is equivalent to locating the singer. When the singer uses a handheld microphone, if there are many people participating in karaoke, there may be few singers. Locating the singer by locating the handheld microphone used by the singer is simpler and more efficient than first detecting the positions of the people and then determining the singer's location from the detected positions.

[0194] Method 3: Detect the singer's location using a microphone array.

[0195] As an example, the control device detects the singer's location using a microphone array, including: obtaining M audio tracks at M locations based on microphone signals collected by the microphone array, where M is a positive integer; determining the probability that singing exists at the i-th location based on the audio track at the i-th location, where i is a positive integer less than or equal to M; and determining the singer's location from the M locations based on the obtained M probabilities.

[0196] When applied to a vehicle cabin, the aforementioned M locations can be all locations within the vehicle cabin, or only some locations (e.g., locations occupied by occupants). Each location represents a seat. In the case where the M locations are only some locations within the vehicle cabin, the location of an occupant (or person) can be identified using a camera. Alternatively, the presence of an occupant at each location can be determined based on parameter changes from sensors (e.g., pressure sensors, capacitance sensors, etc.) installed in each location within the vehicle cabin.

[0197] When applied to a vehicle cabin, a microphone array primarily refers to an in-vehicle microphone. The microphone signals captured by the microphone array include the microphone array's collection of human voices within the vehicle cabin. For example, the microphone signals captured by the microphone array may include one or more of the following: human speech, singer's voice, ambient noise, etc.

[0198] In one implementation, determining the probability of singing at position i based on the audio track at position i includes: analyzing the audio track at position i to obtain the probability of a speech signal at position i and the speech energy of the audio track at position i; and obtaining the probability of singing at position i based on the probability of a speech signal at position i and the speech energy.

[0199] Here, the audio track at position i represents the audio data sequence recorded at position i, and the audio track at position i records the amplitude information of the sound signal (in the case of a sound signal) at position i that changes over time.

[0200] Speech energy refers to the actual physical energy carried in a speech signal, which is usually related to the amplitude of the sound. Speech energy reflects the power or intensity of the sound. The calculation of speech energy has wide applications in the field of speech, and will not be elaborated on here.

[0201] For example, the probability of a speech signal existing at position i is related to the speech energy of the speech track at position i. It can be understood that speech signals typically have a certain energy distribution. If a speech signal exists at a certain position, the speech energy of the speech track at that position is relatively high; if a speech signal does not exist at a certain position, the speech energy of the speech track at that position is low or close to the background noise level. As an example, determining the probability of a speech signal existing at position i based on the speech energy of the speech track at position i includes: dividing the speech track at position i into multiple short frames and calculating the speech energy of each short frame; using the proportion of short frames with speech energy greater than an energy threshold to the total number of frames as the probability of a speech signal existing at position i.

[0202] For example, the presence and probability of a speech signal at a given location can be detected based on the zero-crossing characteristics of the speech signal. Speech signals exhibit zero-crossing characteristics in the time domain, which is the number of times the signal crosses between positive and negative values. For instance, the speech track at the i-th location can be divided into multiple short frames, and the zero-crossing rate (i.e., the number of times the signal crosses zero within the frame) of each short frame can be calculated. The proportion of short frames with zero-crossing rates within a zero-crossing threshold range is taken as the probability of a speech signal at the i-th location. In some schemes, the probability of a speech signal at the i-th location can also be estimated using hidden Markov models, deep learning algorithms, or other models; alternatively, the probability of singing at the i-th location can be directly estimated using a model.

[0203] In another implementation, the probability of singing at position i is determined based on the audio track at position i. This includes determining the probability of singing at position i based on the audio track at position i and the original vocal audio. In other words, the probability of singing at position i is related not only to the audio track at position i but also to the original vocal audio. For example, it is influenced by the similarity between the audio track at position i and the original vocal audio, which can improve the accuracy of estimating the probability of singing at each position.

[0204] Furthermore, based on the audio track at position i and the original vocal audio, the probability of singing at position i is determined, including: analyzing the audio track at position i to obtain the probability of a speech signal at position i and the speech energy of the audio track at position i; calculating the similarity between the audio track at position i and the original vocal audio; and obtaining the probability of singing at position i based on the similarity, the probability of a speech signal, and the speech energy corresponding to position i.

[0205] For example, similarity calculation based on acoustic features can be used to obtain the similarity corresponding to the i-th position. Acoustic features include one or more of the following: rhythm features, pitch features, Mel-frequency cepstral coefficients (MFCC), fundamental frequency, and spectrum. For instance, the first acoustic feature of the speech track at the i-th position and the second acoustic feature of the original vocal audio are extracted. The similarity between the first and second acoustic features is calculated, and this similarity is used as the similarity between the speech track at the i-th position and the original vocal audio, i.e., the similarity corresponding to the i-th position. It can be understood that the first acoustic feature corresponds to the second acoustic feature; for example, if the first acoustic feature is a rhythm feature, then the second acoustic feature is also a rhythm feature.

[0206] When the acoustic features are rhythmic features, the strength of the rhythm is related to energy changes, with frames of higher energy often corresponding to the downbeats. For example, after aligning the audio track at position i with the original vocal audio in time, it is divided into multiple short frames. Taking the extraction of rhythmic features from the audio track at position i as an example, the speech energy of each frame of the audio track at position i is calculated, and autocorrelation function analysis is performed on multiple frames of the audio track at position i to obtain the autocorrelation analysis results. The autocorrelation analysis results are used to indicate the periodic components of the audio track at position i. Finally, based on the speech energy of each frame of the audio track at position i and the autocorrelation analysis results, the rhythmic features of the audio track at position i are obtained. The rhythmic features include the position of the beat and the rhythmic pattern (e.g., 4 / 4 time, 3 / 4 time, etc.). The rhythmic features of the original vocal audio can be obtained in the same way, thus allowing the calculation of the similarity between the rhythmic features of the audio track at position i and the rhythmic features of the original vocal audio.

[0207] When the acoustic feature is pitch, pitch is related to the fundamental frequency. The fundamental frequency refers to the lowest frequency component in a periodic signal, which can also be understood as the number of vibrations the signal completes in one cycle. The fundamental frequency determines the pitch of the sound. For example, after aligning the speech track at position i with the original vocal audio in time, it is divided into multiple short frames. Taking the pitch feature extraction of the speech track at position i as an example, the autocorrelation function is calculated for each frame of the speech track at position i to obtain the fundamental frequency of each frame. The fundamental frequencies of multiple frames of the speech track at position i constitute the first fundamental frequency sequence of the speech track at position i. The second fundamental frequency sequence of the original vocal audio can be obtained in the same way. The first and second fundamental frequency sequences are then processed using dynamic time warping (DTW) or the Pearson correlation coefficient to obtain the similarity between the first and second fundamental frequency sequences.

[0208] It can be understood that when only one acoustic feature is extracted, such as a rhythm feature, the similarity between the rhythm feature of the speech track at position i and the rhythm feature of the original vocal audio is used as the similarity corresponding to position i. When multiple acoustic features are extracted, such as rhythm and pitch features, the similarity corresponding to position i can be obtained based on the similarity between the rhythm feature of the speech track at position i and the rhythm feature of the original vocal audio (referred to as the first similarity), and the similarity between the first fundamental frequency sequence and the second fundamental frequency sequence (referred to as the second similarity). As an example, the similarity corresponding to position i is the result of a weighted sum of the first and second similarities.

[0209] For example, content-based similarity calculation can be used to obtain the similarity corresponding to the i-th position. Exemplarily, the content of the audio track at the i-th position is converted into first text using speech recognition technology, and the content of the original vocal audio is converted into second text using the same technology. The similarity between the first and second texts is then calculated, and this similarity is used as the similarity between the audio track at the i-th position and the original vocal audio, i.e., the similarity corresponding to the i-th position. Methods for calculating the similarity between the first and second texts can include, for example, the edit distance method, bidirectional encoder representations from transformers (BERT) models, etc.

[0210] For example, the probability that there is singing at position i can be a weighted sum of the similarity corresponding to position i, the probability that there is a speech signal at position i, and the signal energy of the speech track at position i.

[0211] In some solutions, an artificial intelligence (AI) model can be used to estimate the probability of singing at position i. This AI model could be, for example, a deep neural network (DNN), a convolutional neural network (CNN), or a Siamese neural network. In this case, determining the probability of singing at position i based on the audio track and the original vocal audio at position i involves inputting both into the AI ​​model to obtain the probability of singing at position i. The AI ​​model then outputs the similarity between the audio track and the original vocal audio at position i.

[0212] Taking the Siamese neural network model as an example, a Siamese neural network consists of two sub-networks sharing weights. Each sub-network can be a CNN, RNN, or similar structure. One sub-network takes the audio track at position i as input, while the other takes the original vocal audio as input. Each sub-network extracts features from its respective input, obtaining its own feature vector. The similarity between the audio track at position i and the original vocal audio is then determined by calculating the distance between these two feature vectors (e.g., cosine distance, Euclidean distance). The training data required for training the Siamese neural network model includes multiple triples. Each triple includes an anchor audio, a positive sample audio (similar to the anchor audio), and a negative sample audio (dissimilar to the anchor audio). Within each triple, the positive sample audio is semantically or feature-wise similar to the anchor audio, while the negative sample audio is semantically and feature-wise dissimilar to the anchor audio.

[0213] For example, anchor audio is represented by A, positive sample audio by P, and negative sample audio by N. In actual training, to improve training efficiency and stability, data is usually input in batches. Assuming each batch size is B, then each time B triples {(A1, P1, N1), (A2, P2, N2), ..., (A... B ,P B N B If the training process of the Siamese neural network model is as described in steps 1-3 below, then this is merely an example of the training process for an artificial intelligence model.

[0214] Step 1: For the anchor audio points A1, A2, ... A in these B triplet sets... B An audio batch of size B is formed and simultaneously input into a subnetwork of the Siamese neural network model; correspondingly, positive sample audio P1, P2, ..., P... B These audio batches are then combined into another audio batch and fed into another sub-network of the Siamese neural network model. Within the network, these audio batches pass through various network layers in parallel for feature extraction and other operations, resulting in corresponding feature vector batches f(A1), f(A2), ..., f(A... B ) and f(P1), f(P2), ..., f(P B ).

[0215] Step 2: Similarly, the negative sample audio N1, N2, ... N will be processed in a similar manner. B Composition of batch and anchor audio A1, A2, ... A B The inputs are respectively fed into these two subnetworks of the Siamese neural network model to obtain feature vector batches f(N1), f(N2), ..., f(N). B ) and f(A1), f(A2), ..., f(A B ).

[0216] Step 3: Based on the batch feature vectors obtained above, calculate the distance and loss for each triple within the batch. Then, update the parameters of the Siamese neural network model using the backpropagation algorithm, enabling the Siamese neural network model to optimize in the direction of reducing the triple loss. The goal of the triple loss is to minimize the distance between the feature vectors of anchor audio A and positive sample audio P in the feature space learned by the model, while maximizing the distance between the feature vectors of anchor audio A and negative sample audio N, and these two distances must meet a certain gap requirement.

[0217] After training the Siamese neural network model through steps 1-3, the model not only learns the common features among similar audio files but also the feature differences between different audio files. Furthermore, through the constraint of triplet loss, the model can learn more discriminative feature representations, better understand the intrinsic structure and semantic information of audio data, and help the model accurately determine the similarity when faced with unseen audio data, thus improving the model's generalization ability and robustness.

[0218] When applied to vehicle cabins Figure 5CThe above implementation method is clearly presented: For example, the control device is equipped with an in-vehicle voice processing module, a voice-to-original-singer similarity calculation module, and a voice activity and voice energy detection module. The in-vehicle voice processing module analyzes the microphone signals collected by the cabin microphone to obtain voice tracks at M locations. The in-vehicle voice processing module sequentially inputs the voice track at each location into the voice-to-original-singer similarity calculation module and the voice activity and voice energy detection module, respectively. Taking the input of the voice track at the i-th location as an example, the voice activity and voice energy detection module outputs the probability of a voice signal existing at the i-th location and the voice energy of the voice track at the i-th location based on the voice track at the i-th location. The voice-to-original-singer similarity calculation module outputs the similarity between the voice track at the i-th location and the original voice audio based on the original vocal audio and the voice track at the i-th location. Finally, the similarity, the probability of a voice signal existing, and the voice energy corresponding to the i-th location are weighted and summed to obtain the probability of singing at the i-th location. Here, Figure 5C This is merely an example of using a microphone array to determine the probability of singing at the i-th position. In some schemes, the similarity corresponding to the i-th position can be directly used as the probability of singing at the i-th position, or the weighted sum of the probability of speech signal at the i-th position and speech energy can be used as the probability of singing at the i-th position.

[0219] By using the above method, we can obtain the probability that singing exists at each of the M positions, that is, obtain M probabilities.

[0220] In one implementation, the singer's location is determined from M positions based on the obtained M probabilities. This includes selecting the position corresponding to the probability greater than a probability threshold from the M probabilities as the singer's location. Here, the probability threshold can be preset by the developers based on experience or it can be a factory default setting.

[0221] The above-described method detects the singer's location using a microphone array. Since the microphone array only collects sound information to determine the singer's location, it effectively prevents the recording and leakage of the passenger's appearance, movements, behavior, and other private information, thus protecting the passenger's privacy and security. Compared to image-based singer location detection, microphone sound signal acquisition is not dependent on lighting conditions or obstructions, making it more environmentally adaptable, with lower hardware costs and more flexible deployment.

[0222] Method 4: Detect the singer's location using a camera.

[0223] As an example, the control device detects the singer's location using a camera, including: the control device identifies the lip movement characteristics of people at M locations using the camera, where M is a positive integer; based on the lip movement characteristics of the person at the i-th location and the song characteristics of the original vocal audio, it obtains the probability that singing exists at the i-th location, where i is a positive integer less than or equal to M; based on the obtained M probabilities, it determines the singer's location from these M locations.

[0224] For example, lip movement features include at least one of lip movement features and lip shape features.

[0225] For example, lip movement characteristics include at least one of the following: lip opening / closing degree, lip movement rate, and lip movement trajectory. Lip opening / closing degree refers to the extent to which the lips open and close. Different phonemes require different lip opening / closing degrees during speech; for example, the lips open more when pronouncing "ah," while they are more rounded and less open when pronouncing "oo." Lip movement rate refers to the speed at which the lips complete various movements (e.g., opening, closing, pursing, stretching). Lip movements are faster when speaking quickly, and relatively slower when speaking slowly. Lip movement trajectory refers to the path and direction of the lips during movement. For example, when pronouncing retroflex consonants, the lips may have an inward curling movement trajectory; when pronouncing labiodental consonants, the lips will have specific contact and movement trajectories with the teeth.

[0226] For example, lip shape characteristics include the lip contour, lip width, lip height, and lip cleft angle, which refers to the angle formed between the upper and lower lips. The degree and manner of change in lip width and lip height also vary when speaking or making lip movements. For instance, when pronouncing certain vowels, lip height may increase significantly, while lip width may decrease accordingly. The lip cleft angle changes during different lip movements.

[0227] For example, song features include at least one of melody features, rhythmic features, harmonic features, and vowel formant features. In some schemes, song features may also include timbre features, which characterize the vocal characteristics of the singer (original singer), including timbre quality, brightness, thickness, etc.

[0228] For example, melodic features include pitch variations, pitch sequences, and interval relationships. A pitch sequence refers to the sequence of pitches of each note in a song; for instance, in the C major scale, the pitches of "do, re, mi" increase sequentially, forming a specific pitch variation pattern. Interval relationships refer to the pitch distance between adjacent notes (e.g., a major second, a minor third, etc.). Different interval combinations give melodies different directions and emotional colors; for example, large intervals usually bring a bright and open feeling, while small intervals may convey a soft and restrained emotion.

[0229] For example, rhythmic features include at least one of the following: syllable duration, rhythmic pauses, meter type (i.e., the pattern of strong and weak beats), and rhythmic pattern (i.e., the variation in rhythmic density). Syllable duration refers to the length of a syllable; rhythmic pauses refer to the location and duration of pauses in a song; meter type is the time signature used in a song, such as common 4 / 4, 3 / 4, and 6 / 8 time signatures; rhythmic pattern indicates the combination and arrangement of note lengths, such as syncopation and dotted rhythms. Different meter types and rhythmic patterns create different rhythms and dynamics.

[0230] For example, harmonic features can be the combination relationship when multiple notes are sounded at the same time, the construction and progression of chords, and chord progression refers to the order in which different chords are connected in a song, such as the chord progression "CG-Am-Em".

[0231] For example, vowel formant features refer to the frequency distribution and intensity of formants during vowel pronunciation in a song. By analyzing the vowel formant features in a song, a connection can be established with lip shape. When the lips exhibit a typical shape corresponding to a certain vowel, the corresponding vowel formant features should appear in the song.

[0232] It's understandable that there's a correspondence between lip movement features and song features. For example, the degree of lip opening and closing in lip movement features is related to melodic features (including pitch and volume). When singing high notes, the lips may open wider to create a broader resonance space; while when singing low notes, the degree of lip opening and closing is relatively smaller. Simultaneously, the degree of lip opening and closing also increases with volume; for example, when singing strong passages, the lips will open and close more forcefully to output more breath and energy, corresponding to accented notes or climaxes in the song. The rate of lip movement in lip movement features is related to rhythmic features; the faster the rhythm of the song, the faster the lip movement, to match rapid note and rhythmic changes. Lip shape features are related to melodic features, vowel formant features, and harmonic features.

[0233] For example, the probability of singing at the i-th position is obtained based on the lip movement features of the person at the i-th position and the song features of the original vocal audio, including: calculating the similarity between the lip movement features of the person at the i-th position and the song features of the original vocal audio, and using the similarity between the lip movement features of the person at the i-th position and the song features of the original vocal audio as the probability of singing at the i-th position.

[0234] The extraction of lip movement features can be achieved using computer vision methods (such as convolutional neural networks (CNNs) or network models specifically optimized for lip movement feature extraction) to process image data containing lip movements captured by a camera. The extraction of song features from the original vocal audio can be achieved using audio processing techniques and deep learning models, such as deep neural networks (DNNs), long short-term memory networks (LSTMs), and gated recurrent units (GRUs), to analyze the original vocal audio and extract song features such as pitch, timbre, rhythm, melody, and vowel formant features.

[0235] When calculating the similarity between the lip movement features of a person at position i and the song features of the original vocal audio, it is necessary to first perform temporal alignment on the lip movement features and the song features. This is because, ideally, a person's lip movements and the sound they produce are closely synchronized; that is, a specific lip movement corresponds to a specific audio content. Temporal alignment ensures that the extracted lip movement features and song features match each other in time. For example, when the lips make the shape to pronounce a certain vowel, the corresponding vowel pronunciation in the audio can be found. This allows for accurate analysis of the relationship between the two, providing a reliable basis for calculating similarity. Furthermore, feature alignment is also required for the lip movement features and song features. Since lip movement features and song features may have different dimensions and feature representations, feature alignment maps them to the same or compatible feature space, making them comparable in dimensionality. After performing time alignment and feature alignment, distance-based metrics (such as Euclidean distance, cosine similarity, etc.), machine learning-based classification methods (such as support vector machines, random forests, etc.), or deep learning model fusion methods (such as multimodal neural networks) can be used to calculate the similarity between the lip movement features of the person at the i-th position and the song features of the original singer's voice audio.

[0236] In some solutions, the three functions mentioned above—lip movement feature extraction, song feature extraction, and similarity calculation between lip movement features and song features—can be integrated into a single artificial intelligence model. This means that an AI model can be used to estimate the probability of singing at the i-th location. In this case, the process of obtaining the probability of singing at the i-th location can also be as follows: input the image data of the person's lip movements at the i-th location captured by the camera and the original singer's audio into the AI ​​model to obtain the probability of singing at the i-th location. The AI ​​model then outputs the similarity between the lip movement features of the person at the i-th location and the song features of the original singer's audio.

[0237] See Figure 5DAs an example, an artificial intelligence model is deployed on the control device. This model includes a lip movement feature extraction module, a song feature extraction module, and a similarity calculation module. The image data of the lip movements of a person at the i-th position, captured by a camera, and the original singer's audio are input into the artificial intelligence model to obtain the probability that singing exists at the i-th position. This includes: inputting the image data of the lip movements of a person at the i-th position, captured by a camera, into the lip movement feature extraction module; the lip movement feature extraction module outputs the lip movement features of the person at the i-th position to the similarity calculation module; inputting the original singer's audio into the song feature extraction module; the song feature extraction module outputs the song features of the original singer's audio to the similarity calculation module; and the similarity calculation module outputs the probability that singing exists at the i-th position based on the lip movement features of the person at the i-th position and the song features of the original singer's audio, where the probability that singing exists at the i-th position is the similarity between the lip movement features of the person at the i-th position and the song features of the original singer's audio. For example, the lip movement feature extraction module can use a convolutional neural network, the song feature extraction module can use a deep neural network or a variant of a recurrent neural network, and the similarity calculation module can use a fully connected neural network or a fusion network based on an attention mechanism.

[0238] Figure 5D This is merely one example of using a camera to determine the probability of singing at the i-th location. Figure 5D The various functional modules shown can also be deployed separately on the control device, rather than integrated into an artificial intelligence model.

[0239] By using the above method, we can obtain the probability that singing exists at each of the M positions, that is, obtain M probabilities.

[0240] In one implementation, the singer's location is determined from M positions based on the obtained M probabilities. This includes selecting the position corresponding to the probability greater than a probability threshold from the M probabilities as the singer's location. Here, the probability threshold can be preset by the developers based on experience or it can be a factory default setting.

[0241] The above-mentioned method identifies the singer's location through a camera. The camera can directly capture visual images of the person's lip movements. Through image analysis technology, it can accurately detect the person's specific location in two-dimensional or three-dimensional space. It can also compare the extracted lip movement features with the song features of the original vocal audio. By comparing the similarity between the lip movement features and the song features, the singer's location can be determined from multiple locations of multiple people, thus enriching the methods for detecting the singer's location.

[0242] Method 5: Detect the singer's location using a microphone array and a camera.

[0243] As an example, implementation methods 3 and 4 can be combined, i.e., execution Figure 5C The probability of singing at position i obtained from the process shown is denoted as the probability of the first branch corresponding to position i. The execution... Figure 5D The probability of singing at the i-th position obtained from the illustrated process is denoted as the second branch probability corresponding to the i-th position. The weighted sum of the first and second branch probabilities corresponding to the i-th position is taken as the overall probability of singing at the i-th position. Thus, for M positions where a person is located, M probabilities can be obtained (in this case, each probability is a comprehensive probability). Therefore, based on the obtained M probabilities, the singer's position can be determined from the M positions. The method for determining the singer's position is described in the preceding content and will not be repeated here.

[0244] Implementation method 5 combines the two methods to comprehensively determine the probability of singing at each location, making full use of the vehicle's sensors (i.e., microphones and cameras), which helps improve the accuracy and reliability of singer location detection and achieves high-precision sound-visual collaborative positioning.

[0245] Based on the descriptions of implementation methods 1-5 above, it can be seen that implementation methods 1 and 2 can be applied before the karaoke session begins, meaning the singer's location can be determined before the session starts. Therefore, as soon as the target song begins playing, the playback strategy described in S304 can be executed to provide a good karaoke experience for both the singer and the listener. Any of implementation methods 2-5 can be applied after the karaoke session begins. The corresponding application scenario could be: the user selects the target song to play, and during the playback of the target song (e.g., after a few seconds), implementation methods 2-5 determine the singer's location, thus implementing the playback strategy based on S304 to provide a good karaoke experience for both the singer and the listener. Furthermore, compared to implementation method 1, implementation methods 2 to 5 can automatically identify the singer's location, achieving adaptive adjustment of the first vocal register.

[0246] Furthermore, regarding the above implementation methods 2-5, they can be executed multiple times in a karaoke scenario. For example, the singer's location can be detected in real time. When the singer's location changes (including an increase in the number of singers at different locations, or a change in location without an increase in the number of singers at different locations), the first vocal range will also change accordingly. In actual karaoke scenarios, there may be situations where different singers take turns singing or multiple singers sing simultaneously during a multi-person duet. Taking singer 1 and singer 2 as an example, when it is detected that only singer 1 is singing, the first vocal range is determined based on singer 1's location; when it is detected that only singer 2 is singing, the first vocal range is determined based on singer 2's location; when it is detected that both singer 1 and singer 2 are singing simultaneously, the first vocal range is determined based on both singer 1's location and singer 2's location.

[0247] S304: The control device controls the playback of a first audio in a first audio zone and a second audio in a second audio zone. The first audio includes the original vocal audio and the second audio includes the singer's vocal audio and the accompaniment audio.

[0248] The first vocal range is the space where the singer is located, and the second vocal range includes the space where both the singer and the listener are located. Here, both the singer and the listener are participants in the karaoke session.

[0249] As an example, the second audio zone can be a full-coverage zone, in which case the second audio zone is set by default to include the entire spatial area of ​​the audio playback space. For example, when applied to a vehicle cabin, the full-coverage zone includes the entire spatial area within the vehicle cabin; when applied to a karaoke room, the full-coverage zone includes the entire spatial area within the karaoke room; when applied to a family living room, the full-coverage zone includes the entire spatial area of ​​the living room.

[0250] As another example, the second register can also be a partially covered register. In this case, the second register is set as a portion of the audio playback space. The determination of the second register is described below.

[0251] In one implementation, the control device receives second setting information input by the user, which indicates the location of a target user who is not participating in karaoke; the control device determines a second sound zone based on the second setting information, wherein the second sound zone does not include the spatial area where the target user is located. For example, the second sound zone is the spatial area outside the spatial area where the target user is located within the full-coverage sound zone.

[0252] For example, receiving second setting information input by the user includes receiving second setting information input by the user via at least one of touch screen, button, keyboard, and voice input. Users can refer to the foregoing descriptions of the corresponding content; further details will not be repeated here.

[0253] As an example, when a user inputs via a touchscreen, the system receives second setting information input by the user, including: displaying second status controls for multiple locations; receiving a second operation by the user on the second status control of a second target location to obtain second setting information; wherein the second target location belongs to these multiple locations, and the second operation causes the second status control of the second target location to indicate that the target person at the second target location does not participate in karaoke.

[0254] The following is based on Figure 6A This indicates that the user inputs the second setting information mentioned above via the touchscreen.

[0255] See Figure 6A , Figure 6A This is a schematic diagram of a display interface for setting the location of people who are not participating in karaoke in a vehicle-mounted scenario, as provided in an embodiment of this application. Figure 6A The displayed interface is a human-computer interaction interface. Figure 6A The interface shown displays status controls for four locations within the vehicle's cabin: driver's seat, front passenger seat, rear left seat, and rear right seat. This is to distinguish them from the aforementioned... Figure 4 The described state control, Figure 4 The described state control is called the first state control. Figure 6A The described state control is called the second state control.

[0256] for Figure 6A The status control at any location in the settings changes color from light to dark when clicked by the user. A dark status indicates that the person selected at that location is not participating in karaoke. Clicking the seat again will change the status control's color from dark to light, indicating that the person at that location is participating in karaoke. Alternatively, the user can click the clear button to cancel the current setting. This means the user can... Figure 6A Select at least one location from the indicated locations as needed. After completing the settings, the user clicks the save button to store the current settings as secondary settings information.

[0257] For example, suppose the user clicks Figure 6A In the interface shown, the "Driver's Seat" position, in response to this operation, changes its status control from light to dark. The user clicks the save button to generate second setting information, which instructs the person in the driver's seat not to participate in karaoke. It can be understood that the second vocal range determined by the control device based on the second setting information can be found in [reference needed]. Figure 6B As shown, Figure 6B These are schematic diagrams of some second registers provided in embodiments of this application. Figure 6B The dark areas in the text all belong to the second register. Figure 6B In the middle, position 1 corresponds to Figure 6AThe "driver's seat" position, position 2 corresponds to Figure 6A The "passenger" position, position 3 corresponds to Figure 6A The "back left" position, position 4 corresponds to Figure 6A The image shows the "back row right" position. It can be seen that position 1 is where the user selected the person who did not participate in the karaoke session.

[0258] for Figure 6B (1) The corresponding application scenario can be: each seat area in the vehicle cabin forms an independent sound zone. Since there are four seats in the vehicle cabin, four sound zones are pre-divided in the vehicle cabin. These four sound zones correspond to the seat numbers, namely sound zone 1, sound zone 2, sound zone 3 and sound zone 4. Each sound zone is connected by... Figure 6B The ellipse containing the number in (1) represents the response to the user. Figure 6A The second set of input information determines that the second sound zone includes the seat area where position 2 is located (i.e., sound zone 2), the seat area where position 3 is located (i.e., sound zone 3), and the seat area where position 4 is located (i.e., sound zone 4).

[0259] for Figure 6B (2), that is, responding to the user through Figure 6A The second setting information input allows the control device to determine that the second audio zone includes all space areas within the vehicle cabin except for the seating area of ​​seat number 1. The second audio zone is represented by a dark area. It can be seen that, compared to... Figure 6B (1), Figure 6B The second pitch range in (2) includes not only the seating area where position 2 is located, the seating area where position 3 is located, and the seating area where position 4 is located, but also the space between position 2 and position 4, the space between position 3 and position 4, etc.

[0260] here, Figure 6B These are merely examples of determining the second vocal register based on second setting information input by the user. Figure 6A This is merely an example of a display interface provided for users to input secondary settings information, and is not intended to limit the display interface to only one type of information. Figure 6A As shown in the form, Figure 6A The four seats shown in the cockpit are just an example. In some designs, more may be shown. Figure 6A Show more or fewer locations.

[0261] It's understandable that, in most cases, the number of people participating in karaoke (including both singers and listeners) is greater than the number of people not participating. Therefore, through... Figure 6AThe interface shown allows users to select the location of those who do not participate in the karaoke session, reducing the number of clicks required and enabling quick setup, which is beneficial for quickly determining the second vocal range. Figure 6A This setting method is essentially a reverse setting method, which is to remove the spatial area where people who do not participate in karaoke are located from the full coverage sound zone based on the second setting information, thus obtaining the second sound zone.

[0262] In some solutions, a positive setting approach can also be used, for example... Figure 6C As shown, third setting information input by the user is obtained. The third setting information is used to indicate the location of the person participating in the karaoke, and the second vocal range is determined based on the third setting information. The second vocal range includes the spatial area where the person participating in the karaoke is located.

[0263] Figure 6C This is a schematic diagram of a display interface for setting the location of people participating in karaoke in a vehicle-mounted scenario, provided in an embodiment of this application. Figure 6C The displayed interface is a human-computer interaction interface. Figure 6C The interface shown displays status controls for four locations within the vehicle's cabin: the driver's seat, the front passenger seat, the left rear seat, and the right rear seat.

[0264] against Figure 6C The status control for any position shown in the diagram indicates that the person in that position is participating in karaoke when the status control is dark, and that the person in that position is not participating when the status control is light. To minimize user interaction and improve user-friendliness, the status control for each position can be set to dark by default. For example, if a user wants to prevent the person in the "driver's seat" from participating in karaoke, they can click on the "driver's seat," and the status control for that position will change from dark to light. After this operation, the user can click the "save" button. In response to this operation, third setting information is obtained. This third setting information indicates the location of the person participating in karaoke, including the "passenger's seat," "rear left," and "rear right" positions. The second vocal range determined based on this third setting information is described above. Figure 6B As shown, this will not be repeated here. In some solutions, Figure 6C The status control for each position can also be set to a light color by default. In this case, the setting method can be: the user can click on the corresponding position based on the location of the person actually participating in the karaoke, so that the status control of that position changes from full color to dark color.

[0265] For example, the above Figure 6A or Figure 6C It could pop up automatically when the user launches the karaoke app. In some solutions, Figure 6AThe interface shown can also be displayed in response to the user clicking to set the location of people who do not participate in the karaoke session within the karaoke application. Figure 6C It can also be displayed in response to a user clicking to set the location of people participating in the karaoke session in the karaoke app.

[0266] The above M probabilities can be obtained from any of the implementation methods 3 to 5 in S302. Based on these M probabilities, the singer's position is determined from the M positions, including: taking the position corresponding to the probability greater than the probability threshold among these M probabilities as the singer's position, so that the first vocal range can be determined based on the singer's position.

[0267] As an example, among these M probabilities, the probabilities greater than the probability threshold include the first probability and the second probability. Therefore, the singer's position includes the first position corresponding to the first probability and the second position corresponding to the second probability. Both the first and second positions belong to these M positions. The first vocal range includes vocal range A and vocal range B. Vocal range A is associated with the first position (i.e., vocal range A is determined based on the first position), and vocal range B is associated with the second position (i.e., vocal range B is determined based on the second position). The sound pressure level (SPL) of the original vocal audio in vocal range A is the same as the SPL of the original vocal audio in vocal range B. Here, "same SPL" can be understood as the two SPLs being equal, or as the difference between the two SPLs being within an allowable error range.

[0268] In this case, when the first audio is played in the first audio zone, the playback output of the first audio zone can be expressed as shown in the following formula (1).

[0269]

[0270] in, p indicates the playback output of the first register. i Indicates position x i The probability of the presence of singing, a0 represents the probability threshold, M represents the total number of positions, v(p i ) = 1 indicates position x i As for the singer's position, i.e., position x i In active state; v(p) i ) = 0 indicates position x i Not the singer's location, i.e., location x i Currently inactive. 's' represents the original vocal audio; when position x... i When in an active state, Indicates position x iThe output of the independent vocal range mode when the bright area of ​​the s signal is formed and the other positions are dark areas. As can be seen from formula (1), the more active positions there are among these M positions, the larger the range of the first vocal range will be. The control device uses formula (1) to reproduce the output of the first vocal range, which requires high accuracy in detecting the position of the singer, and the probability threshold may also be set relatively high.

[0271] As another example, among these M probabilities, the probabilities greater than the probability threshold include the first probability and the second probability. Then the singer's position includes the first position corresponding to the first probability and the second position corresponding to the second probability. Both the first position and the second position belong to these M positions. Among them, the first vocal range includes vocal range A and vocal range B. Vocal range A is associated with the first position (i.e., vocal range A is determined based on the first position), and vocal range B is associated with the second position (i.e., vocal range B is determined based on the second position). If the first probability is greater than the second probability, then the sound pressure level of the original singer's voice in vocal range A is higher than the sound pressure level of the original singer's voice in vocal range B.

[0272] In this case, when the first audio is played in the first audio zone, the playback output of the first audio zone can be expressed as shown in the following formula (2).

[0273]

[0274] in, This indicates the playback output of the first register, f(p) i ) is the probability p i A mapping process to the corresponding weighted value, for example, the mapping process could be an identity transformation. p i 、v(p i For the parameters a0 and M, please refer to the description of the corresponding parameters in formula (1) above, and they will not be repeated here. As shown in formula (2) As can be seen from the expression, the playback output of the first register is a weighted sum of the independent registers formed at each position of the active state. Therefore, p i The value will affect the original vocal audio at that position x i The sound pressure level at the location, and satisfying p i The larger the value, the higher the original vocal audio at position x. i The higher the sound pressure level, the better. The control device uses formula (2) to reproduce the first sound zone. It is also applicable to scenarios where the same singer moves during the singing process. For example, when a singer is singing karaoke, he may move his body to the music while singing. The movement of his body to the music may cause the singer's position to gradually move from the right position in the back row to the left position in the back row. During the movement, if the first sound zone is reproduced using the method shown in formula (2), the singer can still hear the first audio during the movement. The first audio includes the original vocal audio.

[0275] In one implementation, when applied to a vehicle cabin, a control device controls the playback of a first audio signal in a first audio zone and a second audio signal in a second audio zone. This includes: the control device controlling the playback of the first audio signal in the first audio zone at least by controlling headrest speakers and / or shoulder pillow speakers in the first audio zone, and controlling a speaker group to play the second audio signal in the second audio zone. The speaker group can be a collection of all in-vehicle speakers or all speakers including those in the second audio zone. Playing the first audio signal through the headrest speakers and / or shoulder pillow speakers provides good isolation for people outside the first audio zone; only the singer in the first audio zone can hear the original vocals, and the original vocals will not interfere with people outside the first audio zone. Furthermore, the placement of the headrest speakers and / or shoulder pillow speakers closer to the ears compared to vehicle body speakers or other speakers provides an immersive sound experience for the singer in the first audio zone.

[0276] For example, since the first audio includes the original vocal audio, the first audio can be the audio obtained after applying sound effects processing to the original vocal audio. Since the second audio includes the singer's vocal audio and the accompaniment audio, the second audio can be the audio obtained after applying sound effects processing and / or mixing processing to the singer's vocal audio and the accompaniment audio.

[0277] As described above, the second vocal range includes the first vocal range, and the second audio played in the second vocal range includes the singer's vocal audio and the accompaniment audio. Therefore, at the singer's location, the original vocal audio, the accompaniment audio, and the singer's own vocal audio can be heard. Thus, the first audio can also be the audio obtained by performing sound effects processing and / or mixing processing on the original vocal audio, the accompaniment audio, and the singer's vocal audio. Alternatively, the first audio can be the audio obtained by performing sound effects processing and / or mixing processing on the target song (which includes the original vocal audio and the accompaniment audio) and the singer's vocal audio.

[0278] For example, when mixing the original vocal audio, the accompaniment audio, and the singer's vocal audio to obtain the first audio, the mixing weight of the original vocal audio can be increased. This will relatively increase the volume, prominence, and influence of the original vocal in the first audio during the mixing process, making it more noticeable to the singer when blended with the accompaniment. This not only allows the singer to hear the lyrics more clearly, but also helps the singer to better grasp the details of the performance, such as timbre, pitch, breath control, and articulation.

[0279] Ideally, the singer's location should allow them to hear the original vocals, the accompaniment, and their own vocals, while the listener's location should only hear the vocals and the accompaniment. In some solutions, the sound pressure level (SPL) of the original vocals at the singer's location is higher than that at the listener's location. Conversely, the SPL at the listener's location should be relatively low, perhaps approaching zero, or even zero. This minimizes or eliminates the listener's perception of the original vocals.

[0280] In another implementation, when applied to karaoke rooms or practice rooms, the control device controls the playback of the first audio in the first audio register, including: the control device controls the playback of the first audio in the first audio register using directional sound generation technology. Similarly, the playback of the second audio can also employ directional generation technology.

[0281] Here, directional sound technology is a technology that can focus sound onto a specific area or direction for propagation. Examples of directional sound technology include ultrasonic directional sound technology, loudspeaker array directional sound technology, and reflector directional sound technology.

[0282] Taking ultrasonic directional sound generation technology as an example, the process of controlling the playback of a first audio frequency in a first sound zone using ultrasonic directional sound generation technology can be as follows: Prepare ultrasonic directional sound generation equipment, which typically includes an audio signal generator, an ultrasonic transducer array, etc. First, input the first audio frequency into the audio signal generator, modulate it, and load it onto an ultrasonic carrier wave. Then, transmit the modulated signal through the ultrasonic transducer array. By controlling the parameters and transmission angle of the ultrasonic transducer array, the ultrasonic beam is accurately directed to the first sound zone, and the first audio frequency is reproduced within the first zone.

[0283] In some possible embodiments, during karaoke, a pitch similarity analysis can be performed between the singer's audio track and the original vocal audio. When the pitch similarity between the singer's audio track and the original vocal audio reaches a similarity threshold, the mixing weight of the original vocal audio in the first audio file can be reduced. The singer's audio track carries the singer's vocal audio. This allows the singer to sing more immersively, hear their own voice more clearly, and thus better appreciate their singing ability. In some solutions, when the pitch similarity between the singer's audio track and the original vocal audio reaches a similarity threshold, the playback of the original vocal audio can be stopped, allowing the singer to enjoy a pure musical experience during karaoke and focus on the fusion of their own voice and the accompaniment.

[0284] Implementation Figure 3In this embodiment, a first audio range is determined based on the singer's location, and a second audio range is determined based on both the singer's and the listener's locations. The system controls the playback of a first audio track containing the original singer's vocals within the first audio range, and a second audio track containing both the singer's vocals and the accompaniment within the second audio range. This ensures that the singer can hear the original vocals in a karaoke setting, while the listener is not disturbed by the original vocals, thus improving both the singer's and listener's experience. Furthermore, various methods for detecting the singer's location are provided, such as user-defined location settings, location based on the singer's handheld microphone, and detection via vehicle sensors (e.g., microphone arrays and / or cameras), catering to diverse user needs and enriching application scenarios.

[0285] In some possible embodiments, for scenarios involving multiple singers performing a song together, the original vocal audio can be divided into N separate original vocal tracks, where N is an integer greater than 1. That is, the original vocal audio includes N separate original vocal tracks, namely, original vocal track 1, original vocal track 2, ..., original vocal track N, where each track is a portion of the original vocal audio. In this case, when executing... Figure 3 Before embodiment S304, it is also necessary to determine the original vocal track audio that matches the location of each singer. For example, the division of the original vocal audio can be based on the number of original vocals involved in the original vocal audio, or it can be based on the segmentation of words in a multi-person chorus, etc.

[0286] As an example, when implementing method 3 in S303 above (i.e., detecting the singer's location via a microphone array), if the original vocal audio includes N original vocal tracks, the similarity between the voice track at position i and each of the N original vocal tracks can be calculated. The original vocal track with the highest similarity to the voice track at position i among the N original vocal tracks is then used as the original vocal track matched for position i. In this case, determining the probability of singing at position i based on the voice track and the original vocal audio means: determining the probability of singing at position i based on the voice track at position i and the original vocal track matched for position i.

[0287] It's understandable that in a multi-person chorus, multiple singers may take turns singing, or multiple singers may sing one or more lines of lyrics simultaneously. Therefore, the original vocal tracks matched at different positions may be the same or different.

[0288] See Figure 7A , Figure 7A This is a schematic diagram illustrating a method for original vocal track matching based on a microphone array, as provided in the application embodiment. Figure 7AThe control unit is equipped with an in-vehicle voice processing module, a voice activity and voice energy detection module, and a voice-to-original-singer similarity calculation module. Microphone signals collected by the cockpit microphone are input to the in-vehicle voice processing module for analysis, resulting in voice tracks at M locations. The in-vehicle voice processing module then sequentially inputs each voice track to the voice activity and voice energy detection module and the voice-to-original-singer similarity calculation module. Taking the voice track at the i-th location as an example, the voice activity and voice energy detection module outputs the probability of a voice signal existing at the i-th location and the voice energy of the voice track at the i-th location, thus obtaining the probability of singing at the i-th location based on the probability of a voice signal existing at the i-th location and the voice energy. The voice-to-original-singer similarity calculation module calculates the similarity between the voice track at the i-th location and each original-singer track, and outputs the original-singer track with the highest similarity among the N original-singer tracks as the original-singer track matched at the i-th location. As an example, the voice-original vocal similarity calculation module includes a song feature extraction module, a voice feature extraction module, and a similarity calculation module. The song feature extraction module is used to extract the song features of each original vocal track audio. The voice feature extraction module can be used to extract the voice features (such as the acoustic features mentioned above) of the voice track at the i-th position. The similarity calculation module is used to calculate the similarity between the voice track at the i-th position and each original vocal track audio based on the voice features of the voice track at the i-th position and the song features of each original vocal track audio.

[0289] here, Figure 7A As an example only, the similarity calculation method between the voice track at position i and the original vocal track can refer to the relevant description of the similarity calculation between the voice track at position i and the original vocal audio in implementation method 3 above. In some solutions, Figure 7A The probability that there is singing at the i-th position can also be obtained by weighted summation based on the similarity between the audio track at the i-th position and the original vocal track matched at the i-th position, the probability of the existence of a speech signal at the i-th position, and the speech energy.

[0290] As another example, when implementing method 4 in S303 above (i.e., detecting the singer's location via a camera), if the original vocal audio includes N original vocal tracks, the similarity between the lip movement features of the person at the i-th position and the song features of each of these N original vocal tracks can be calculated. The original vocal track with the highest similarity among these N original vocal tracks is then used as the original vocal track matched at the i-th position. In this case, the probability that singing exists at the i-th position can be the similarity between the lip movement features of the person at the i-th position and the song features of the original vocal track matched at the i-th position.

[0291] See Figure 7B , Figure 7B This is a schematic diagram illustrating a method for original vocal track matching based on a camera, as provided in the application's embodiment. Figure 7B In this system, an artificial intelligence model is deployed on the control device. This model includes a lip movement feature extraction module, a song feature extraction module, and a similarity calculation module. The image data of the lip movements of the person at the i-th position, captured by the camera, and the original vocal audio are input into the AI ​​model to obtain the probability that singing exists at the i-th position. This includes: inputting the image data of the lip movements of the person at the i-th position, captured by the camera, into the lip movement feature extraction module; the lip movement feature extraction module outputs the lip movement features of the person at the i-th position to the similarity calculation module; inputting N original vocal track audios into the song feature extraction module; the song feature extraction module outputs the song features of each original vocal track audio to the similarity calculation module; the similarity calculation module calculates the similarity between the lip movement features of the person at the i-th position and the song features of each original vocal track audio; the original vocal track audio with the highest similarity is used as the matched original vocal track audio for the i-th position; and the probability that singing exists at the i-th position is determined based on the highest similarity. Finally, the probability that singing exists at the i-th position and the matched original vocal track audio for the i-th position are output. Here, Figure 7B This is just one example.

[0292] With the above implementation, after determining the original vocal track audio that matches each position, the singer's position can be determined from the M positions based on the obtained probabilities. Knowing the singer's position, the original vocal track audio that matches the singer's position can also be known.

[0293] See Figure 7C , Figure 7C This is a schematic diagram of a car karaoke scenario provided in an embodiment of this application. Figure 7C As shown in the diagram illustrating the vehicle's cabin seating arrangement, occupant 1 is in position 1 (driver's seat), occupant 2 is in position 2 (front passenger seat), occupant 3 is in position 3 (rear left), and occupant 4 is in position 4 (rear right). Figure 7C (1) or Figure 7C In (2), the left attached diagram uses light colors to represent the first audio range, while the right attached diagram uses dark colors to represent the second audio range. It is assumed that the second audio range is a full-coverage range, encompassing the entire space within the cockpit used for audio playback. This can be understood as... Figure 7C The terms "light" and "dark" are relative; the colored areas in the left image are referred to as light, and the colored areas in the right image are referred to as dark.

[0294] In a multi-person duet scenario, when passenger 3 sings, passenger 3 is the singer, and passengers 1, 2, and 4 are all listeners. The singer's position is determined as position 3 using the method described above, and the original vocal track audio matched to position 3 is original vocal track audio 1. Based on position 3, the first vocal register is determined, and its representation is as follows: Figure 7C As shown in the left figure of (1), the control is to play the first audio in the first register, which includes the original vocal track audio 1; and to play the second audio in the second register, as shown in the figure. Figure 7C As shown in the right-hand image of (1), the second audio includes the accompaniment audio and the singer's vocal audio of crew member 3. That is, for Figure 7C (1) When the crew member 3 is located, the original vocal track 1, the accompaniment audio and the singer's voice audio of the crew member 3 can be heard. When the other crew members (i.e., crew member 1, crew member 2 and crew member 4) are located, only the accompaniment audio and the singer's voice audio of the crew member 3 can be heard.

[0295] In a multi-person duet scenario, when passenger 4 sings, passenger 4 acts as the singer, while passengers 1, 2, and 3 act as listeners. The singer's position is determined as position 4 using the method described above, and the original vocal track audio matched to position 4 is original vocal track audio 2. Based on position 4, the first vocal register is determined, and its representation is as follows: Figure 7C As shown in the left attached figure of (2), the control is to play the first audio in the first register, which includes the original vocal track 2; and to play the second audio in the second register, as shown in the attached figure of (2). Figure 7C As shown in the right-hand image of (2) in the figure, the second audio includes the accompaniment audio and the singer's vocal audio of crew member 4. That is to say, for Figure 7C (2) When the crew member 4 is located, the original vocal track 2, the accompaniment audio and the singer's voice audio of the crew member 4 can be heard. When the crew members (i.e., crew members 1, 2 and 3) are located, only the accompaniment audio and the singer's voice audio of the crew member 4 can be heard.

[0296] here, Figure 7C This is merely one example of a multi-person duet. A multi-person duet is an example of a multi-person chorus, where multiple singers take turns singing in a certain order. In some scenarios, a multi-person chorus may also involve at least two singers singing simultaneously, and / or, the number of singers may exceed two.

[0297] For example, if the original vocal audio of the target song involves multiple original singers, taking two original singers as an example—one male and one female—then the original vocal audio includes two separate tracks: one for the male singer and one for the female singer. Assuming the singer's location is determined to include a first position and a second position, where the first position is the location of the first singer (e.g., a female singer) and the second position is the location of the second singer (e.g., a male singer), the above method might determine that the first position matches the female original vocal track and the second position matches the male original vocal track, or vice versa. In some possible scenarios, the original vocal audio might also include separate tracks for the male and female singers, as well as tracks for a male-female duet.

[0298] For example, when the original vocal audio includes multiple original vocal tracks, the above control, when the first audio range includes the first audio track, controls the playback output of the first audio range. It can also be expressed as shown in the following formula (3).

[0299]

[0300] Here, the difference between formula (3) and formula (1) above is that s i Indicates position x i For the matched original vocal track audio, the other parameters in formula (3) can be found in the description of the corresponding parameters in formula (1). For the sake of brevity, they will not be repeated here. It can be understood that formula (3) is an extension of formula (1).

[0301] For example, when the original vocal audio includes multiple original vocal tracks, the above control, when the first audio range includes the first audio track, controls the playback output of the first audio range. It can also be expressed as shown in the following formula (4).

[0302]

[0303] Here, the difference between formula (4) and formula (2) above is that s i Indicates position x i For the matched original vocal track audio, the other parameters in formula (4) can be found in the description of the corresponding parameters in formula (2). For the sake of brevity, they will not be repeated here. It can be understood that formula (4) is an extension of formula (2).

[0304] In some solutions, in multi-person chorus scenarios, if the original vocal audio of the target song only involves one original vocalist, taking the first singer and the second singer singing as an example, when the first singer and the second singer sing alternately, the original vocal track audio 1 is matched at the position of the first singer at time 1, and the original vocal track audio 2 is matched at the position of the second singer at time 2. The original vocal track audio 1 and the original vocal track audio 2 are different, but both the original vocal track audio 1 and the original vocal track audio 2 are part of the original vocal audio.

[0305] Implementing the above embodiments further enables the playback of different original vocal tracks at the locations of different singers. Thus, in a duet between the first and second singers, assuming the first singer performs the male vocal part and the second singer performs the female vocal part, when the first singer sings, only the first singer hears the male vocal track, and other participants are not disturbed by it; similarly, when the second singer sings, only the second singer hears the female vocal track, and other participants are not disturbed by it. This enhances the listening experience for both singers and listeners.

[0306] In some possible embodiments, the above implementation method 1 (i.e., the user sets the singer's position) or implementation method 2 (i.e., the position of the identified handheld microphone is taken as the singer's position) can also be used in combination with any of the above implementation methods 3 to 5.

[0307] As an example, the first audio zone is first determined based on the first setting information input by the user (i.e., implementation method 1 above) or based on the position of the identified handheld microphone (i.e., implementation method 2 above), and the first audio is controlled to be played in the first audio zone; during the karaoke process, the singer's position is detected by the microphone group and / or camera (i.e., any one of implementation methods 3 to 5 above), and the singer's position detection result is obtained, which is used to indicate the singer's current position; the first audio zone is updated according to the singer's position detection result, and the first audio is controlled to be played in the updated first audio zone.

[0308] For example, if the singer's location obtained through implementation method 1 or implementation method 2 is called the initial location detection result, then the singer's location detection result is different from the initial location detection result. Here, "different" means that the number of singer locations changes (e.g., increases or decreases), or the number of singer locations does not change but the singer's location is different (e.g., combined with...). Figure 6B The initial position detection result is Figure 6B Position 3 in the middle, the singer's position detection result is Figure 6B Position 4 in the middle, etc.

[0309] See Figure 8A , Figure 8A This is a schematic diagram of an in-vehicle karaoke scenario provided in an embodiment of this application. Figure 8A The vehicle cabin positions shown include position 1 (driver's seat), position 2 (front passenger seat), position 3 (rear left), and position 4 (rear right). It can be seen that passenger 1 is in position 1, passenger 2 is in position 2, passenger 3 is in position 3, and passenger 4 is in position 4. Assuming that before the karaoke session begins, the user sets the singer's position (i.e., the initial position detection result) to position 3, then passenger 3 in position 3 is the singer. At the start of the karaoke session, the first audio file is played in the first register. For the representation of the first register, please refer to [link to relevant documentation]. Figure 8A (1), that is, the spatial area where passenger 3 is located, the first audio includes the original vocal audio; as the karaoke continues, passenger 4 is led by passenger 3 to sing along with passenger 3. In this case, the singer position detection result obtained by detecting the singer's position through the microphone group and / or camera includes position 3 and position 4. It can be seen that the number of singer positions increases compared with the initial position detection result. Therefore, the first vocal range is updated according to the singer position detection result, and the above first audio is played in the updated first vocal range. The representation of the updated first vocal range is shown in [reference]. Figure 8A (2) includes the space area where occupant 3 is located and the space area where occupant 4 is located.

[0310] See Figure 8B , Figure 8B This is a schematic diagram of a car karaoke scenario provided in an embodiment of this application. Figure 8B The diagram showing the seating arrangement within the vehicle cabin indicates that occupant 1 is at position 1, occupant 2 is at position 2, and occupant 3 is at position 3. Assuming the user sets the singer's position (i.e., the initial position detection result) to position 3 before the karaoke session begins, then occupant 3 at position 3 is the singer. At the start of the karaoke session, the first audio track is played in the first register. For the representation of the first register, please refer to [link to relevant documentation]. Figure 8B (1) The first audio includes the original vocal audio; as the karaoke continues, passenger 3 sings and moves his / her body to the music, and the position of passenger 3 changes from position 3 to position 4. In this case, the singer's position detection result obtained by the microphone group and / or camera is position 4. It can be seen that the number of singers' positions has not changed compared with the initial position detection result, but the singer's position is different. Therefore, the first vocal range is updated according to the singer's position detection result, and the above first audio is played in the updated first vocal range. For the representation of the updated first vocal range, please refer to Figure 8B (2)

[0311] Understandable. Figure 8A and Figure 8BThese are just some examples of how the first vocal register might be updated in a karaoke setting. In some scenarios, the first vocal register might also be updated in situations such as: a change in the solo singer, singers taking turns singing in a group setting, or alternating between solo and group singing.

[0312] By implementing the above embodiments, the singer's location can be determined before karaoke begins by either setting the singer's position or detecting the position of the handheld microphone. Therefore, when karaoke starts, the first audio file can be played in the first vocal range, providing a better karaoke experience for the singer. As the karaoke progresses, the singer's position can be detected in real time to determine if their location has changed. If the singer's location changes, the first vocal range can be updated promptly to further enhance the singer's karaoke experience.

[0313] In some possible embodiments, the following method can be used to ensure that the singer hears the original vocals without disturbing the listener: First, acquire the original vocal audio and accompaniment audio of the target song (the song sung by the singer); and acquire the singer's vocal audio; then, control the playback of the first target audio in the first target sound region and the playback of the second target audio in the second target sound region. The first target sound region is the spatial region where the singer is located, and the first target audio includes the original vocal audio, the singer's vocal audio, and the accompaniment audio; the second target sound region is the spatial region where the listener is located, and the second target audio includes the singer's vocal audio and the accompaniment audio. Here, the first target sound region is the aforementioned first sound region.

[0314] The following is combined Figure 9 To more clearly illustrate the first and second target pitch regions. See also Figure 9 , Figure 9 This is a schematic diagram illustrating a vocal register division provided in an embodiment of this application. Figure 9 middle, Figure 9 The positions in the vehicle cabin shown include position 1 (driver's seat), position 2 (passenger's seat), position 3 (rear left), and position 4 (rear right). It can be seen that passenger 1 is in position 1, passenger 2 is in position 2, passenger 3 is in position 3, and passenger 4 is in position 4. Figure 9 This is just one example. In some designs, the number of seats in the cabin, the number of occupants in the cabin, etc., can be more or less, and the number of current singers can also be more.

[0315] exist Figure 9 In this context, assuming passenger 2 is the singer and passengers 1, 3, and 4 are all listeners, then the first target sound region is the spatial region where passenger 2 is located. The first target sound region is represented as follows: Figure 9The light-colored area in the middle; the second target sound zone includes the space area in the cockpit excluding the space area where occupant 2 is located, and the second target sound zone is represented as Figure 9 The darker areas in the text. In some schemes, the second target sound region can be similarly referenced to the above. Figure 6B The second target audio region is defined as the spatial region where occupant 1 is located (corresponding to position 1), the spatial region where occupant 3 is located (corresponding to position 3), and the spatial region where occupant 4 is located (corresponding to position 4). The first target audio is played within the first target audio region, and the second target audio is played within the second target audio region.

[0316] For example, since the first target audio includes the original vocal audio, the singer's vocal audio, and the accompaniment audio, the first target audio can be the audio obtained by performing sound effects processing and / or mixing processing on the original vocal audio, the accompaniment audio, and the singer's vocal audio; or, the first target audio can be the audio obtained by performing sound effects processing and / or mixing processing on the target song (which includes the original vocal audio and the accompaniment audio) and the singer's vocal audio. Since the second target audio includes the singer's vocal audio and the accompaniment audio, the second target audio can be the audio obtained by performing sound effects processing and / or mixing processing on the singer's vocal audio and the accompaniment audio.

[0317] Here, the method for determining the first target pitch range can be referred to the description of the method for determining the first pitch range above, and will not be repeated here.

[0318] In one implementation, the second target sound region is determined based on the listener's location.

[0319] As an example, the listener's location can be set by the user. Please refer to the previous description of the user setting the singer's location. For the sake of brevity, it will not be repeated here.

[0320] In some solutions, when applied to a vehicle cabin, a camera can be used to detect the singer's location within the cabin. Based on the location of all positions within the cabin and the singer's location, the listener's location is determined; that is, the listener's location includes all positions within the cabin except the singer's. Alternatively, a camera can be used to detect the locations of all occupants and the singer within the cabin; based on these locations, the listener's location is determined; that is, the listener's location includes all positions of the occupants within the cabin except the singer's.

[0321] To more clearly illustrate Figure 9 The difference between the register division shown and the register division in the above embodiment is that... Figure 10 The division of the sound regions in the above embodiments is illustrated by way of example. Figure 10This is a schematic diagram of another type of vocal register division provided in the embodiments of this application. Figure 10 The vehicle cabin layout shown is similar to Figure 9 The seating arrangement within the vehicle cabins shown is identical. Figure 10 In this scenario, assuming passenger 2 is the singer and passengers 1, 3, and 4 are all listeners, then the first register is the spatial region where passenger 2 is located. For a representation of the first register, please refer to [link to relevant documentation]. Figure 10 As shown in (1); the second sound zone is the full-coverage sound zone, that is, it includes the entire space area of ​​the cockpit. For the representation of the second sound zone, please refer to [reference needed]. Figure 10 As shown in (2). The first audio is played in the first audio range and the second audio is played in the second audio range. The first audio includes the original vocal audio and the second audio includes the singer's vocal audio and the accompaniment audio. For the description of the first audio and the second audio, please refer to the description of the corresponding content above, which will not be repeated here.

[0322] here, Figure 10 This is just one example. In some designs, the number of seats in the cabin, the number of occupants in the cabin, etc., can be more or less, and the number of current singers can also be more.

[0323] See Figure 11 , Figure 11 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application. The audio processing device 30 includes an acquisition unit 310 and a processing unit 312. The audio processing device 30 can be implemented by hardware, software, or a combination of hardware and software.

[0324] The acquisition unit 310 is used to acquire the original vocal audio and accompaniment audio of the target song, which is a song sung by a singer; and to acquire the singer's vocal audio. The processing unit 312 is used to control the playback of the first audio in the first vocal range and the playback of the second audio in the second vocal range. The first vocal range is the spatial area where the singer is located, and the first audio includes the original vocal audio; the second vocal range includes the spatial area where the singer and the listener are located, and the second audio includes the singer's vocal audio and the accompaniment audio.

[0325] Optionally, the audio processing device 30 further includes a display unit 314, which is used to display an interface for setting the location of the singer, and an interface for displaying the location of people who are not participating in karaoke, or an interface for displaying the location of people who are participating in karaoke.

[0326] The audio processing device 30 can be used to achieve Figure 3 The method described in the embodiments. Figure 3 In this embodiment, the acquisition unit 310 can be used to execute S301 and S302, and the processing unit 312 can be used to execute S303 and S304. The display unit 314 can be used to present... Figure 4, Figure 6A as well as Figure 6C The interface shown. The audio processing device 30 can also be used to implement Figures 5A-5D , Figure 7A as well as Figure 7B The methods described in any of the embodiments are not repeated here for the sake of brevity.

[0327] It should be understood that the division of the units in the audio processing device 30 described above is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units in the device can be implemented by a processor calling software; for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each unit in the device. The processor can be, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented as hardware circuits. The functionality of some or all units can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the above units is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through a configuration file, thereby achieving the functionality of some or all of the above units. All units of the above device can be implemented entirely through processor-invoked software, entirely through hardware circuits, or partially through processor-invoked software with the remaining parts implemented through hardware circuits.

[0328] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a type of microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships of hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), or a deep learning processing unit (DPU).

[0329] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0330] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together as a system-on-a-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.

[0331] See Figure 12 , Figure 12 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 12As shown, the computing device 40 includes a processor 401, a communication interface 402, a memory 403, and a bus 404. The processor 401, the memory 403, and the communication interface 402 communicate with each other via the bus 404. It should be understood that this application does not limit the number of processors and memories in the computing device 40.

[0332] In one implementation, the computing device 40 may be a device with computing capabilities, such as the aforementioned control device or a device incorporating the aforementioned control device. For example, when applied to a vehicle, the computing device 40 may be a domain controller within the vehicle; the domain controller is described above. Figure 1A The descriptions of the relevant content in the embodiments will not be repeated here.

[0333] Bus 404 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 12 The bus 404 is represented by only one line, but this does not mean that there is only one bus or one type of bus. The bus 404 may include a path for transmitting information between various components of the computing device 40 (e.g., memory 403, processor 401, communication interface 402).

[0334] The processor 401 can be referred to the relevant description of the processor in the above embodiments, and will not be repeated here.

[0335] Memory 403 provides storage space, which can store data such as the operating system and computer programs. Memory 403 can be one or a combination of several of the following: random access memory (RAM), erasable programmable read-only memory (EPROM), read-only memory (ROM), or compact disc read memory (CD-ROM). Memory 403 can exist alone or be integrated into processor 401.

[0336] The communication interface 402 can be used to provide information input or output to the processor 401. Alternatively, the communication interface 402 can be used to receive and / or send data to externally transmitted data, and can be a wired link interface including an Ethernet cable, or a wireless link interface (such as Wi-Fi, Bluetooth, general wireless transmission, etc.). Alternatively, the communication interface 402 may also include a transmitter (such as an RF transmitter, antenna, etc.) or a receiver coupled to the interface.

[0337] In some possible embodiments, the computing device 40 may also include a display 405. The display 405 is connected or coupled to the processor 401 via a bus 404. The display 405 can be used to present an interface to the user for setting the location of the singer, as well as an interface displaying the location of those not participating in the karaoke session or an interface displaying the location of those participating in the karaoke session. The display 405 can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), etc. The display 405 can also be an in-vehicle tablet, a vehicle-mounted display, or a head-up display (HUD) system, etc.

[0338] The processor 401 in the computing device 40 is used to read the computer program stored in the memory 403 to execute the aforementioned method, for example... Figure 3 , Figures 5A-5D , Figure 7A as well as Figure 7B The method described in [the document / document].

[0339] In one possible design, the computing device 40 may be for execution Figure 3 One or more modules in the execution body of the method shown, the processor 401 can be used to read one or more computer programs stored in memory, for performing the following operations:

[0340] By acquiring the original vocal audio and accompaniment audio of the target song in Unit 310, the target song is the song sung by the singer; and by acquiring the singer's vocal audio.

[0341] The system controls the playback of the first audio in the first vocal register and the second audio in the second vocal register. The first vocal register is the spatial region where the singer is located, and the first audio includes the original vocal audio. The second vocal register includes the spatial region where both the singer and the listener are located, and the second audio includes the singer's vocal audio and the accompaniment audio.

[0342] In the embodiments described above, each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant descriptions in other embodiments. Furthermore, in the embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between the embodiments are consistent and can be mutually referenced. Technical features from different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0343] It should be noted that those skilled in the art will recognize that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0344] The technical solution of this application, in essence, or the part that makes the contribution, or all or part of the technical solution, can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes several instructions to cause a device (which may be a personal computer, server, network device, robot, microcontroller, chip, robot, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

Claims

1. An audio processing method, characterized in that, The method includes: Obtain the original vocal audio and accompaniment audio of the target song, wherein the target song is a song sung by the singer. Obtain the singer's vocal audio; Control the playback of the first audio in the first audio register and the playback of the second audio in the second audio register; The first audio includes the original vocal audio, the second audio includes the singer's vocal audio and the accompaniment audio, the first pitch range is the spatial region where the singer is located, the second pitch range includes the spatial region where the singer and the listener are located, and the sound pressure level of the original vocal audio at the singer's location is higher than the sound pressure level of the original vocal audio at the listener's location.

2. The method according to claim 1, characterized in that, The original vocal audio includes a first original vocal track and a second original vocal track. The singers include a first singer and a second singer. When the first singer sings, the singer's vocal audio is the first vocal audio. When the second singer sings, the singer's vocal audio is the second vocal audio. The first vocal audio matches the first original vocal track, and the second vocal audio matches the second original vocal track. The first vocal range includes a third vocal range and a fourth vocal range, wherein the third vocal range is the spatial region where the first singer is located, and the fourth vocal range is the spatial region where the second singer is located. The control to play the first audio in the first audio zone includes: When the first singer sings, the system controls the playback of the first audio in the third vocal register, the first audio including the first original vocal track; and / or... When the second singer sings, the first audio is played in the fourth vocal register, and the first audio includes the second original vocal track audio.

3. The method according to claim 1, characterized in that, During the singer's performance, the method further includes: The singer's location is detected using a microphone array and / or a camera. The first vocal register is determined based on the singer's location.

4. The method according to claim 1, characterized in that, Before the karaoke session begins, the method also includes: Receive first setting information input by the user, the first setting information being used to indicate the location of the singer; The first audio range is determined based on the first setting information.

5. The method according to claim 4, characterized in that, The first setting information input by the user includes: Displays a first status control for multiple locations; Receive the user's first operation on the first status control of the first target location to obtain the first setting information; Wherein, the first target location belongs to the plurality of locations, and the first operation causes a first state control of the first target location to indicate that the first target location is selected as the singer's location.

6. The method according to claim 1, characterized in that, When the singer uses a handheld microphone, the method further includes: The position of the handheld microphone is identified using a camera; The first audio zone is determined based on the position of the handheld microphone.

7. The method according to any one of claims 4-6, characterized in that, The method further includes: During the karaoke session, the singer's position is detected using a microphone array and / or a camera to obtain the singer's position detection results; The first vocal register is updated based on the singer's position detection results; Control the playback of the first audio in the updated first audio region.

8. The method according to claim 1, characterized in that, The method further includes: Receive second setting information input by the user, the second setting information being used to indicate the location of the target person who is not participating in karaoke; The second sound zone is determined based on the second setting information, and the second sound zone does not include the spatial area where the target person is located.

9. The method according to claim 8, characterized in that, The second setting information received from user input includes: A second status control that displays multiple locations; Receive the user's second operation on the second status control of the second target location to obtain the second setting information; Wherein, the second target location belongs to the plurality of locations, and the second operation causes the second status control of the second target location to indicate that the target person at the second target location does not participate in karaoke.

10. The method according to claim 3, characterized in that, The method of detecting the singer's location via microphone arrays and / or cameras includes: Based on the microphone signals collected by the microphone group, obtain the voice tracks at M positions, where M is a positive integer; Based on the audio track at the i-th position, determine the probability that singing exists at the i-th position, where i is a positive integer less than or equal to M. Based on the obtained M probabilities, the singer's location is determined from the M locations.

11. The method according to claim 10, characterized in that, The step of determining the probability that singing exists at the i-th position based on the audio track at the i-th position includes: Based on the audio track at the i-th position and the original vocal audio, determine the probability that singing exists at the i-th position.

12. The method according to claim 10, characterized in that, When the original vocal audio includes multiple original vocal tracks, the method further includes: Determine the similarity between the voice track at the i-th position and each of the multiple original vocal tracks; The original vocal track with the highest similarity to the voice track at the i-th position among the multiple original vocal tracks is taken as the original vocal track matched at the i-th position.

13. The method according to claim 10, characterized in that, Before obtaining the speech tracks at M positions, the method further includes: The camera identifies the M locations where the person is located.

14. The method according to claim 3, characterized in that, The method of detecting the singer's location via microphone arrays and / or cameras includes: The camera identifies the lip movement characteristics of people at M locations, where M is a positive integer; Based on the lip movement characteristics of the person at the i-th position and the song characteristics of the original singer's voice audio, the probability that there is singing at the i-th position is obtained, where i is a positive integer less than or equal to M. Based on the obtained M probabilities, the singer's location is determined from the M locations.

15. The method according to claim 14, characterized in that, When the original vocal audio includes multiple original vocal tracks, obtaining the probability that singing exists at the i-th position based on the lip movement characteristics of the person at the i-th position and the song characteristics of the original vocal audio includes: Determine the similarity between the lip movement features of the person at the i-th position and the song features of each original vocal track audio; The probability that singing exists at the i-th position is determined based on the maximum similarity.

16. The method according to claim 15, characterized in that, The method further includes: The original vocal track with the highest similarity among the multiple original vocal tracks is taken as the original vocal track matched at the i-th position.

17. The method according to any one of claims 10-16, characterized in that, The step of determining the singer's location from the M locations based on the obtained M probabilities includes: The position corresponding to the probability greater than the probability threshold among the M probabilities is taken as the position of the singer.

18. The method according to claim 17, characterized in that, The probability greater than the probability threshold among the M probabilities includes a first probability and a second probability, and the singer's location includes a first position corresponding to the first probability and a second position corresponding to the second probability. The first vocal range includes a fifth vocal range and a sixth vocal range. The fifth vocal range is associated with the first position, and the sixth vocal range is associated with the second position. The sound pressure level of the original vocal audio in the fifth vocal range is the same as the sound pressure level of the original vocal audio in the sixth vocal range.

19. The method according to claim 17, characterized in that, The probability greater than the probability threshold among the M probabilities includes a first probability and a second probability, and the singer's location includes a first position corresponding to the first probability and a second position corresponding to the second probability. The first vocal range includes a fifth vocal range and a sixth vocal range. The fifth vocal range is associated with the first position, and the sixth vocal range is associated with the second position. The first probability is greater than the second probability, and the sound pressure level of the original vocal audio in the fifth vocal range is higher than the sound pressure level of the original vocal audio in the sixth vocal range.

20. An apparatus for audio processing, characterized in that, The apparatus includes a communication unit and a processing unit, and is used to perform the method as described in any one of claims 1-19.

21. An apparatus for karaoke processing, characterized in that, The device includes a memory and a processor, the memory storing computer program instructions, and the processor executing the computer program instructions to cause the device to perform the method as described in any one of claims 1-19.

22. An audio processing system, characterized in that, The audio processing system includes a control device, a microphone group, and a speaker group, wherein the control device is connected to the microphone group and the speaker group respectively, and the control device is used to perform the method as described in any one of claims 1-19.

23. The system according to claim 22, characterized in that, The audio processing system also includes a camera, which is used to transmit captured image data containing people to the control device, and the image data is used to determine the location of the singer.

24. The system according to claim 22 or 23, characterized in that, The microphone array includes vehicle-mounted microphones and / or handheld microphones.

25. A vehicle, characterized in that, The vehicle includes the device as described in claim 20 or 21, or includes the audio processing system as described in any one of claims 22-24.

26. A computer-readable storage medium containing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method as described in any one of claims 1-19.

27. A computer program product containing instructions, characterized in that, When the instructions are executed by the computing device, the computing device performs the method as described in any one of claims 1-19.

Citation Information

Patent Citations

  • Method and device for controlling karaoke sound equipment through audio playing and system

    CN113470602A