Vehicle-mounted karaoke method and system, controller and vehicle
By using a camera inside the vehicle to capture images of occupants and combining this with voice processing technology to filter human voices, the problem of low voice recognition rate caused by noise interference inside the vehicle is solved, resulting in a better karaoke experience.
Patent Information
- Application Number
- CN202310107268.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-30
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2043-01-30
AI Technical Summary
The in-car karaoke function suffers from low voice recognition rates due to noise interference, resulting in a poor passenger experience.
By acquiring occupant images through in-vehicle cameras, combining voice processing technology to filter out human voices, and then mixing and processing them with accompaniment before outputting, the voice recognition rate is improved.
It improves the accuracy of voice recognition and user experience of the in-car karaoke function, reduces noise interference, and enhances the karaoke experience.
Smart Images

Figure CN118397990B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of vehicles, in particular to a vehicle KTV method, system, controller and vehicle. BACKGROUND
[0002] At present, the potential user number of KTV function is large, and it is expected to become a standard configuration of passenger vehicles. In the related art, the implementation of the vehicle KTV function is through a microphone to collect audio signals in the vehicle, and the audio signals and accompaniment are mixed and played through a loudspeaker. Due to the noise in the vehicle environment and the possibility of more than one person singing in the vehicle, the loudspeaker plays a lot of noise, and the passengers have a poor experience of the KTV function. SUMMARY
[0003] The present disclosure aims to at least partially solve one of the technical problems in the related art. To this end, one object of the present disclosure is to provide a vehicle KTV method, system, controller and vehicle to remove noise well and improve voice recognition rate by visual combination with voice processing capability when implementing the KTV function.
[0004] In a first aspect, the present disclosure provides a vehicle KTV method, a vehicle having a loudspeaker, a microphone and a camera, the method comprising: after the vehicle enters a KTV mode, acquiring an audio signal through the microphone and acquiring a vehicle occupant image through the camera; screening a vocal sound from the audio signal according to the vehicle occupant image; acquiring an accompaniment sound and mixing the accompaniment sound and the vocal sound; and outputting the mixed sound signal to the loudspeaker for playing.
[0005] In a second aspect, the present disclosure provides a controller comprising a memory, a processor and a computer program stored in the memory, wherein the computer program is executed by the processor to implement the vehicle KTV method described above.
[0006] In a third aspect, the present disclosure provides a vehicle-mounted KTV system, which comprises a system on chip (SOC) chip, a digital processor, a loudspeaker, a microphone and a camera; wherein the digital processor is connected with the SOC chip, the loudspeaker and the microphone respectively, and is configured to acquire an audio signal through the microphone after the vehicle enters a KTV mode, and send the audio signal to the SOC chip; the SOC chip is connected with the camera, and is configured to acquire an image of an occupant in the vehicle through the camera after the vehicle enters the KTV mode, filter a vocal sound from the audio signal according to the image of the occupant in the vehicle, acquire an accompaniment sound, and send the accompaniment sound and the vocal sound to the digital processor, so that the digital processor performs mixing processing on the accompaniment sound and the vocal sound, and outputs a sound signal after the mixing processing to the loudspeaker for playing.
[0007] In a fourth aspect, the present disclosure provides a vehicle comprising the vehicle-mounted KTV system in the above embodiment.
[0008] The vehicle-mounted KTV method, system and controller and vehicle of the present disclosure acquire an audio signal through a vehicle-mounted microphone, acquire an image of an occupant in the vehicle through a vehicle-mounted camera, filter a vocal sound from the audio signal according to the image of the occupant in the vehicle, perform mixing processing on an accompaniment sound and the vocal sound, and output a sound signal after the mixing processing to a loudspeaker for playing, so as to realize a vehicle-mounted KTV function, and improve the voice recognition rate through the voice processing capability after the combination of vision and voice, thereby improving the KTV experience effect of a user.
[0009] Additional aspects and advantages of the present disclosure will be made apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a flowchart of a vehicle-mounted KTV method according to an embodiment of the present disclosure;
[0011] Figure 2 is an implementation structural arrangement diagram of the vehicle-mounted KTV method according to an embodiment of the present disclosure;
[0012] Figure 3 is an implementation structural block diagram of the vehicle-mounted KTV method according to an embodiment of the present disclosure;
[0013] Figure 4 is an implementation structural block diagram of the vehicle-mounted KTV method according to a specific embodiment of the present disclosure;
[0014] Figure 5 is a flowchart of filtering a vocal sound according to an embodiment of the present disclosure;
[0015] Figure 6is a structural block diagram of a vehicle KTV system according to an embodiment of the present disclosure;
[0016] Figure 7 is a structural block diagram of a vehicle KTV system according to another embodiment of the present disclosure;
[0017] Figure 8 is a structural block diagram of a vehicle KTV system according to yet another embodiment of the present disclosure;
[0018] Figure 9 is a structural block diagram of a vehicle according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0019] Embodiments of the present disclosure are described in detail below with reference to the accompanying drawings, in which like or similar elements or elements having the same or similar functions are denoted by the same or similar reference numerals throughout the drawings. The embodiments described below by reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and cannot be understood as limiting the present disclosure.
[0020] The vehicle KTV method, system, and controller, vehicle of the embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0021] Figure 1 is a flowchart of a vehicle KTV method according to an embodiment of the present disclosure.
[0022] In embodiments of the present disclosure, a speaker, a microphone, and a camera are provided in the vehicle, wherein the number of the speaker, the microphone, and the camera can each be one or more.
[0023] In some embodiments, the microphone is a built-in microphone, which can make the KTV function not need to purchase an additional microphone and not need to connect the microphone additionally, so that the operation is more simple. As shown in Figure 2 , the number of the microphone is 4, which are respectively marked as 41, 42, 43, and 44, and can correspond to four positions of the ceiling in the vehicle for the front row and the back row of seats respectively. Based on this, the vehicle can be divided into four sound zones, and each sound zone corresponds to a microphone. The number of the camera is 2, which are respectively marked as 51 and 52, and as shown in Figure 2 , one (marked as 51) is provided at the main driver A-pillar or the steering wheel, which can reuse the fatigue detection camera in the vehicle, and the other (marked as 52) is provided at the rearview mirror position in the vehicle, so that the face image of the occupant at each seat can be collected. The number of the speaker can be one or multiple (as shown in Figure 2 , which are respectively marked as 31, 32, 33, 34, 35, and 36), and the setting positions are as shown in Figure 2As shown, 31 is provided in the instrument panel in the vehicle, 32, 33, 34, 35 are provided in the door position, 32, 33 provided in the front position of the door can be high-pitched speakers, 34, 35 provided in the rear position of the door are bass speakers, 36 is provided near the rear trunk position, which can be a subwoofer, through the setting of multiple different speakers in different positions, stereo sound can be realized, and the user K song experience can be improved.
[0024] As shown in Figure 1 The vehicle K song method includes:
[0025] S1, after the vehicle enters the K song mode, the audio signal is acquired through the microphone, and the in-vehicle occupant image is acquired through the camera.
[0026] Specifically, the K song function can be implemented through the structure shown in Figure 3 The implementation process is shown in Figure 4 When the whole vehicle starts OK file, the sensing microphone and the camera are powered on and work, and the intelligent interaction application platform of the vehicle terminal starts normally, if the AISOC (Artificial Intelligence System on Chip) chip detects that the user voice has the command word "I want to sing" or the user manually clicks the K song software installed in the vehicle terminal, the AI SOC chip starts the K song software, and then the user can select the corresponding song to start K song. After the user selects the corresponding song to start K song, the vehicle enters the K song mode. Therefore, the K song software is convenient to start, without additional connection line, switch and the like, and the operation is more convenient, and the K song software can be used in the driving and parking states.
[0027] After the vehicle enters the K song mode, the audio signal collected by the microphone can be received by the digital signal processor (DSP), and the audio signal is obtained, and the in-vehicle occupant image collected by the camera can be acquired by the AI SOC chip.
[0028] S2, according to the in-vehicle occupant image, the vocal sound is selected from the audio signal.
[0029] As an implementation manner, face recognition can be performed on the in-vehicle occupant image to determine the occupant who is K singing, and information such as the in-vehicle position of the K singing occupant and the vocalization state (such as the opening degree, shape, etc. of the mouth) of the K singing occupant can be obtained. According to the in-vehicle position, the distance between the K singing occupant and the built-in microphone can be obtained, and then the human voice can be screened from the audio signal according to the distance and the vocalization state. For example, if the K singing occupant stops vocalizing, but sound is still collected, it can be considered that the sound is noise, and the noise can be identified and deleted when the K singing occupant vocalizes. For another example, under the same volume, the closer to the microphone, the clearer the collected sound. Based on this, when multiple sound sources are obtained according to the audio signal, the distance of the sound sources can be further determined, and the sound source closest to the distance is taken as the screened human voice. When there are multiple human voices closest to the distance, the matching condition of the vocalization state of each sound source and the accompaniment sound can be further screened, such as taking the high matching degree as the screened human voice.
[0030] In some embodiments, the number of microphones is multiple, and the multiple microphones correspond to multiple sound zones in the vehicle. Step S2 can include: determining a target sound zone according to the in-vehicle occupant image; and screening human voice of the target sound zone from the audio signal.
[0031] Specifically, referring to Figure 4 , the AI SOC chip obtains the in-vehicle occupant image, and determines a target sound zone according to the in-vehicle occupant image. The DSP can transmit the audio signal to the AI SOC chip through an I2S (Inter-IC Sound) channel, and the AI SOC chip can screen human voice of the target sound zone from the audio signal.
[0032] In this embodiment, the accompaniment sound and the human voice are mixed, including: mixing the accompaniment sound and the human voice of the target sound zone.
[0033] In some embodiments, determining the target sound zone according to the in-vehicle occupant image includes: performing recognition processing on the in-vehicle occupant image to obtain lip movement information of each in-vehicle occupant; and determining the target sound zone according to the lip movement information.
[0034] Specifically, the vehicle can be divided into four sound zones, corresponding Figure 2The four different positions of the microphone are shown. After the images of the passengers in the vehicle are collected, the face full key data (68 or 106 or other number of key points obtained by face key point positioning technology) extracted by detection processing of each passenger image is obtained. Then, the lip part data is screened from the key data to obtain the lip movement information (i.e. lip movement information). If the passenger is not singing, the amount of lip movement is small, and if the passenger is singing, the amount of lip movement is large, and the lip movement information obtained in the two cases is different, so that the passenger with large amount of lip movement (i.e. singing) can be screened according to the lip movement information, and the sound area where the passenger is located can be taken as the target sound area. Among them, the amount of lip movement can be determined according to the positional relationship of the lip part key points. For example, when the position of the passenger's head changes very little, the amount of movement can be directly determined according to the position information of the key points at different times; when the position of the passenger's head changes greatly, the amount of movement can be determined according to the relative position information of the key points at different times.
[0035] In some embodiments, the target sound area is screened from the audio signal. The voice of the target sound area includes: screening the audio signal collected by the microphone arranged in the target sound area from the audio signal to obtain the target audio signal; obtaining the lip movement shape feature data according to the lip movement information, and using the lip movement shape feature data to process the target audio signal to obtain the voice of the target sound area.
[0036] Specifically, referring to Figure 2 , if the target sound area is the sound area where the microphone 41 is located, it means that the passenger in the driver's seat is singing, and at this time, the audio signal collected by the microphone 41 is selected for processing, which can reduce the voice interference of non-target voice; if the target sound area is the sound area where the multiple microphones are located, it means that multiple passengers in the vehicle are singing, and at this time, the audio signals collected by the multiple microphones are selected for processing, which can reduce the voice interference of non-target voice, and can support KTV function to support chorus, solo and duet. At the same time, the corresponding lip movement shape feature data (which can be obtained according to the shape of multiple lip key points at different times) can be obtained according to the lip movement information of the passengers in the target sound area. Because different lip movement shape feature data can correspond to different phonemes, the phonemes can be distinguished by using the lip movement shape feature data, and the phonemes can be compared with the target audio signal to determine the noise in the target audio signal, and the noise can be removed, that is, the noise reduction processing of the target audio signal is realized, and then the clean voice of the target sound area can be obtained.
[0037] The present disclosure adds face full image detection through camera vision, lip movement information feature fusion arbitration analysis on the basis of pure language recognition, and more accurately performs human voice separation and extraction of target sound sources. Compared with the scheme of using a microphone array to realize noise reduction through the inter-channel difference between microphones, the present disclosure can more accurately analyze how many people are in the in-vehicle scene and the identity of the corresponding person through the visual scheme, so that the blind source task becomes non-blind source, and the sound source positioning capability and accuracy are greatly improved. At the same time, since the visual input is not disturbed by noise, the multi-modal recognition and multi-modal noise reduction of lip fusion speech can realize stable output of clean target human voice under high noise interference conditions, and even for the interference of two pieces of speech with the same voiceprint, accurate noise reduction and recognition can also be realized.
[0038] In specific implementation, a multi-modal algorithm fusion model can be trained in advance by using a deep learning algorithm, as shown in Figure 5 When performing multi-modal fusion of in-vehicle microphone and camera collected data, the data is first synchronized, and then different modal data such as audio signals and video signals can be input into the pre-trained multi-modal algorithm fusion model. The model can enable joint learning and joint prediction of different modal data at the feature level, so that the model can learn the complementary information between the data, make an overall decision, and output an accurate and clean human voice.
[0039] S3, obtaining an accompaniment sound, and mixing the accompaniment sound and the human voice.
[0040] The accompaniment sound is the accompaniment sound of the currently played song, and after the mixing of the accompaniment sound and the human voice, karaoke music is formed.
[0041] In some embodiments, as shown in Figure 3 , Figure 4 The accompaniment sound can be obtained through the main SOC chip, the AI SOC chip transmits the screened human voice to the main SOC chip, the main SOC chip transmits the accompaniment sound and the human voice to the DSP, and the digital processor DSP mixes the accompaniment sound and the human voice. The AI SOC chip and the main SOC chip communicate through I2S, and the main SOC chip and the DSP can communicate through TDM (Time-Division Multiplexing). In some examples, as shown in Figure 3 , the main SOC chip also communicates with the MCU through SPI (Serial Peripheral Interface) to realize volume adjustment of the vehicle terminal and vehicle control, etc. For example, the MCU receives a volume adjustment instruction, sends the instruction to the main SOC chip, and adjusts the volume of the vehicle terminal by the main SOC chip.
[0042] Alternatively, the functions of the AI SOC chip and the main SOC chip can also be implemented through a single SOC chip.
[0043] S4 outputs the mixed audio signal to the speaker for playback.
[0044] Specifically, the system acquires audio signals through the vehicle's built-in microphone and images of the occupants through the vehicle's built-in camera. Based on these images, it then filters out human voices from the audio signal, mixes the accompaniment and vocals, and outputs the mixed audio signal to the speakers for playback. Thus, the karaoke function can be activated simply by speaking, making the operation highly intelligent. Furthermore, the combination of visual and voice processing capabilities enhances the voice recognition rate.
[0045] In some embodiments, outputting the mixed audio signal to a speaker for playback includes: performing sound effect processing on the mixed audio signal (such as adding two-dimensional images, beautifying sound effects, MV recording sound effects, DJ sound effects, etc.), and performing power amplification processing on the sound signal after sound effect processing; and outputting the amplified audio signal to a speaker for playback.
[0046] Specifically, an audio power amplifier can be installed inside the vehicle, such as... Figure 2 As shown, the audio power amplifier (marked as 60) can be installed under the driver's seat to amplify the sound signal. This allows for sound effect processing of the mixed sound signal, followed by amplification and output to the speakers for playback, thus improving the karaoke experience.
[0047] In some embodiments, before selecting human voices from an audio signal based on an image of the in-vehicle occupants, the method further includes: acquiring a reference tone signal and using the reference tone signal to perform Acoustic Echo Chancellor (AEC) processing on the audio signal.
[0048] Specifically, such as Figure 4 As shown, the digital signal processor (DSP) can be used to transmit the reference tone signal and the audio signal collected by the microphone through the audio signal channel to the echo cancellation (AEC) module in the AI SOC chip to perform audio signal echo cancellation and suppress reference tone processing, thereby improving the karaoke effect.
[0049] In some embodiments, the in-vehicle karaoke method further includes: using lip movement information to perform acoustic echo cancellation (AEC) processing on the audio signal.
[0050] Specifically, see Figure 4 Lip movement information can also be provided to the echo cancellation AEC module to improve the echo cancellation effect.
[0051] In some embodiments, the vehicle karaoke method further comprises: taking the sound signal after the sound effect processing as a reference sound signal for the next echo cancellation AEC processing.
[0052] Specifically, the DSP can send the accompaniment sound and the human voice after the sound effect processing back to the AI SOC chip as a reference sound signal for echo suppression to improve the echo cancellation effect.
[0053] In some embodiments, the vehicle karaoke method further comprises: in response to a song ordering instruction, controlling the vehicle terminal to order a song; in response to a song switching instruction, controlling the vehicle terminal to switch a song; wherein the song ordering instruction and the song switching instruction are received by at least one microphone, at least one camera, or the vehicle terminal.
[0054] Specifically, referring to Figure 2 , the vehicle terminal can include a central control display screen (labeled 70), which can display a karaoke interface. The karaoke interface can include different interfaces of the karaoke software installed on the vehicle terminal to display karaoke-related information, such as orderable songs, currently singing songs, song ordering controls, song switching controls, etc. The central control display screen can be a touch screen that can respond to a song ordering instruction for the song ordering control and a song switching instruction for the song switching control.
[0055] In some examples, the song ordering instruction and the song switching instruction can also be input by voice, such as the passenger issuing a voice "order a song", the microphone receiving the voice, and the AI SOC chip recognizing the song ordering instruction from the voice and displaying a song ordering list based on which the passenger can further issue a voice such as "select a song by ** singer" or "select ** song" to order a song; such as the passenger issuing a voice "switch a song", the microphone receiving the voice, and the AI SOC chip recognizing the song switching instruction from the voice and switching to the next song.
[0056] In other examples, the song ordering instruction and the song switching instruction can also be input by gestures, such as the passenger making a song ordering gesture (such as a left swipe gesture), the camera capturing the passenger's gesture image, and the AI SOC chip recognizing the song ordering instruction from the image and displaying a song ordering list based on which the user can order a song; such as the passenger making a song switching gesture (such as a right swipe gesture), the camera capturing the passenger's gesture image, and the AI SOC chip recognizing the song switching instruction from the image and switching to the next song.
[0057] The vehicle karaoke method of the embodiments of the present disclosure is described below through two scenarios: Figure 3 , Figure 4
[0058] Scenario one: when the whole vehicle starts OK file, the perception software starts normally, the vehicle intelligent interaction application platform starts normally; if the AI SOC chip detects that the user voice has the command word "I want to sing", start the KTV software, the vehicle terminal displays the KTV interface, and then the user selects the corresponding song to start KTV. After the user selects the song to start KTV, the AI SOC chip determines that the vehicle enters the KTV mode, at this time the AI SOC chip receives the audio signal transmitted by the DSP through the I2S channel between the AI SOC chip and the DSP and processes it, and then transmits the processed voice to the main SOC chip through the I2S channel between the AI SOC chip and the main SOC chip. The main SOC chip obtains the accompaniment sound in the KTV software, and processes the accompaniment sound in the same frequency. The accompaniment sound after the same frequency processing and the voice output by the AI SOC chip are mixed in the DSP driving layer. After the DSP mixes the voice and the accompaniment sound, the output is sent to the audio power amplifier, the audio power amplifier amplifies the sound signal, and outputs the amplified sound signal to the loudspeaker for playing. And in the KTV mode, gesture recognition song selection, song cutting and the like can be realized.
[0059] Scenario two: when the whole vehicle starts OK file, the perception software starts normally, the vehicle intelligent interaction application platform starts normally; if the AI SOC chip detects that the user clicks the KTV software displayed on the vehicle terminal, start the KTV software, the vehicle terminal displays the KTV interface, and then the user selects the corresponding song to start KTV. After the user selects the song to start KTV, the AI SOC chip determines that the vehicle enters the KTV mode, at this time the AI SOC chip receives the audio signal transmitted by the DSP through the I2S channel between the AI SOC chip and the DSP and processes it, and then transmits the processed voice to the main SOC chip through the I2S channel between the AI SOC chip and the main SOC chip. The main SOC chip obtains the accompaniment sound in the KTV software, and processes the accompaniment sound in the same frequency. The accompaniment sound after the same frequency processing and the voice output by the AI SOC chip are mixed in the DSP driving layer. After the DSP mixes the voice and the accompaniment sound, the output is sent to the audio power amplifier, the audio power amplifier amplifies the sound signal, and outputs the amplified sound signal to the loudspeaker for playing. And in the KTV mode, gesture recognition song selection, song cutting and the like can be realized.
[0060] Based on the above vehicle KTV method, the present disclosure proposes a controller.
[0061] In this embodiment, the controller includes a memory, a processor and a computer program stored on the memory, and the computer program is executed by the processor to implement the vehicle KTV method of the above-mentioned embodiment.
[0062] Figure 6 is the structure block diagram of the vehicle KTV system of the embodiment of the present disclosure.
[0063] As Figure 6As shown, the vehicle-mounted KTV system 100 comprises a system on chip (SOC) chip 10, a digital processor 20, a loudspeaker 30, a microphone 40 and a camera 50.
[0064] Referring to Figure 6 , the digital processor 20 is connected with the system on chip (SOC) chip 10, the loudspeaker 30 and the microphone 40 respectively, for acquiring audio signals through the microphone 40 after the vehicle enters the KTV mode, and sending the audio signals to the system on chip (SOC) chip 10. The system on chip (SOC) chip 10 is connected with the camera 50, for acquiring images of passengers in the vehicle through the camera 50 after the vehicle enters the KTV mode, screening out human voices from the audio signals according to the images of passengers in the vehicle, acquiring accompaniment sounds, and sending the accompaniment sounds and the human voices to the digital processor 20, so that the digital processor 20 performs mixing processing on the accompaniment sounds and the human voices, and outputs the sound signals after the mixing processing to the loudspeaker 30 for playing.
[0065] In some embodiments, as shown in Figure 7 , the system 100 further comprises an audio power amplifier 60, which is connected with the digital processor 20 and the loudspeaker 30 respectively. When the digital processor 20 outputs the sound signals after the mixing processing to the loudspeaker 30 for playing, the digital processor 20 is specifically configured to perform sound effect processing on the sound signals after the mixing processing, and output the sound signals after the sound effect processing to the audio power amplifier 60 for power amplification processing, so that the audio power amplifier 60 outputs the sound signals after the amplification processing to the loudspeaker 30 for playing.
[0066] In some embodiments, as shown in Figure 8 , the number of the system on chip (SOC) chips 10 is two, which are respectively a first SOC chip 11 (i.e. the AI SOC chip mentioned above) and a second SOC chip 12 (i.e. the main SOC mentioned above).
[0067] Referring to Figure 8 , the first SOC chip 11 is connected with the digital processor 20, the second SOC chip 12 and the plurality of cameras 50 respectively, for acquiring images of passengers in the vehicle through the cameras 50 after the vehicle enters the KTV mode, screening out human voices from the audio signals according to the images of passengers in the vehicle, and sending the human voices to the second SOC chip 12. The second SOC chip 12 is configured to acquire accompaniment sounds, and send the accompaniment sounds and the human voices to the digital processor 20.
[0068] In some embodiments, referring to Figure 2The number of the microphone 40 is multiple, and the multiple microphones 40 are respectively arranged at multiple positions of the vehicle roof; the number of the camera 50 is multiple, at least one of the multiple cameras 50 is arranged at the main driver A column or the steering wheel, and at least one of the multiple cameras 50 is arranged at the rearview mirror position; the number of the speaker 30 is multiple, and the multiple speakers 30 are respectively arranged at the instrument desk, multiple door positions, and the rear luggage compartment position in the vehicle; and the system on chip SOC chip 10 is arranged directly below the instrument desk.
[0069] It should be noted that other specific implementations of the vehicle KTV system of the embodiments of the present disclosure can refer to the specific implementations of the vehicle KTV method of the above-mentioned embodiments of the present disclosure.
[0070] Figure 9 is a structural block diagram of the vehicle of the embodiments of the present disclosure.
[0071] As shown in Figure 9 , the vehicle 1000 comprises the vehicle KTV system 100 of the above-mentioned embodiments.
[0072] The vehicle KTV method, system, controller, and vehicle of the embodiments of the present disclosure can directly utilize the built-in microphone in the vehicle to realize the microphone sound area positioning, support the KTV function of multiple people singing together, and realize the functions of voice and gesture recognition ordering and song switching. At the same time, by increasing the visual signal on the basis of the voice signal, the visual signal and the voice signal are combined for processing, which can improve the voice recognition rate in complex scenes.
[0073] It should be noted that the logical and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logic functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions, or a combination of the above. For the purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a product of the manufacturing and / or processing, and / or an article of manufacture. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electronic connection having one or more wires (electronic devices), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via an optical scanner, then compiled, interpreted, or otherwise processed, and stored in a computer memory in a suitable format.
[0074] It should be understood that portions of the present disclosure can be implemented in hardware, software, firmware, or combinations thereof. In the above-described embodiments, a number of steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, or a combination thereof, known in the art can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0075] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples.
[0076] In the description of the present disclosure, it needs to be understood that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present disclosure and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present disclosure.
[0077] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise explicitly specified and limited.
[0078] In the present disclosure, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting", "fixing" and the like should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements, unless otherwise explicitly limited. For those skilled in the art, the specific meaning of the above terms in the present disclosure can be understood according to the specific circumstances.
[0079] In the present disclosure, unless otherwise explicitly specified and limited, the first feature "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature "above", "over" and "on" the second feature can be that the first feature is directly above or obliquely above the second feature, or it can only mean that the horizontal height of the first feature is higher than that of the second feature. The first feature "below", "under" and "under" the second feature can be that the first feature is directly below or obliquely below the second feature, or it can only mean that the horizontal height of the first feature is less than that of the second feature.
[0080] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as a limitation on the present disclosure, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A car-mounted karaoke method, characterized in that, The car is provided with a loudspeaker, a microphone and a camera, and the method comprises: After the vehicle enters the K song mode, an audio signal is acquired through the microphone, and an image of a passenger in the vehicle is acquired through the camera; Human voice is screened from the audio signal according to the image of the passenger in the vehicle; Accompanying sound is acquired, and the accompanying sound and the human voice are mixed; The mixed sound signal is output to the loudspeaker for playing; The number of the microphone is multiple, and the multiple microphones correspond to multiple sound areas in the vehicle, and the human voice is screened from the audio signal according to the image of the passenger in the vehicle, which comprises: A target sound area is determined according to the image of the passenger in the vehicle; Human voice in the target sound area is screened from the audio signal; The accompanying sound and the human voice are mixed, which comprises: The accompanying sound and the human voice in the target sound area are mixed; The target sound area is determined according to the image of the passenger in the vehicle, which comprises: Lip movement information of each passenger in the vehicle is obtained by identifying the image of the passenger in the vehicle; The target sound area is determined according to the lip movement information; The human voice in the target sound area is screened from the audio signal, which comprises: Audio signal collected by the microphone arranged in the target sound area is screened from the audio signal to obtain target audio signal; Lip movement shape feature data is obtained according to the lip movement information, and the target audio signal is de-noised by using the lip movement shape feature data to obtain human voice in the target sound area; wherein the target audio signal is de-noised by using the lip movement shape feature data, which comprises: using the lip movement shape feature data to identify phonemes, and comparing the phonemes with the target audio signal to determine noise in the target audio signal.
2. The car-mounted K song method according to claim 1, characterized in that, The mixed sound signal is output to the loudspeaker for playing, which comprises: Sound effect processing is performed on the mixed sound signal, and power amplification processing is performed on the sound effect processed sound signal; The amplified sound signal is output to the loudspeaker for playing.
3. The car-mounted K song method according to claim 2, characterized in that, Before the human voice is screened from the audio signal according to the image of the passenger in the vehicle, the method further comprises: Reference sound signal is acquired, and the audio signal is AEC processed by using the reference sound signal.
4. The car-mounted K song method according to claim 3, characterized in that, The method further comprises: The sound effect processed sound signal is used as the reference sound signal for the next AEC processing.
5. The car-mounted K song method according to claim 1, characterized in that, The method further comprises: The audio signal is AEC processed by using the lip movement information.
6. A controller characterized by comprising: A system comprising a memory, a processor and a computer program stored on the memory, wherein the computer program is executed by the processor to implement the vehicle-mounted K song method according to any one of claims 1-5.
7. A car-mounted karaoke system, characterized in that, A system for implementing the vehicle-mounted K song method according to any one of claims 1-5, the system comprising a SOC chip, a digital processor, a loudspeaker, a microphone and a camera; wherein The digital processor is connected with the system on chip SOC chip, the loudspeaker and the microphone respectively, and is configured to acquire an audio signal through the microphone after the vehicle enters a KTV mode, and send the audio signal to the system on chip SOC chip. The system on chip SOC chip is connected with the camera, and is configured to acquire a vehicle occupant image through the camera after the vehicle enters the KTV mode, filter a vocal sound from the audio signal according to the vehicle occupant image, acquire an accompaniment sound, and send the accompaniment sound and the vocal sound to the digital processor, so that the digital processor performs mixing processing on the accompaniment sound and the vocal sound, and outputs a sound signal after the mixing processing to the loudspeaker for playing.
8. The car-mounted KTV system according to claim 7, characterized in that, The system further includes an audio power amplifier connected with the digital processor and the loudspeaker respectively, and the digital processor is specifically configured to perform the following when outputting the sound signal after the mixing processing to the loudspeaker for playing: performing sound effect processing on the sound signal after the mixing processing, and outputting the sound signal after the sound effect processing to the audio power amplifier for power amplification processing, so that the audio power amplifier outputs the sound signal after the power amplification processing to the loudspeaker for playing.
9. The car-mounted KTV system according to claim 7, characterized in that, The number of the system on chip SOC chips is two, and the two system on chip SOC chips are respectively a first SOC chip and a second SOC chip; wherein The first SOC chip is connected with the digital processor, the second SOC chip and the camera respectively, and is configured to acquire a vehicle occupant image through the camera after the vehicle enters the KTV mode, filter a vocal sound from the audio signal according to the vehicle occupant image, and send the vocal sound to the second SOC chip. The second SOC chip is configured to acquire an accompaniment sound, and send the accompaniment sound and the vocal sound to the digital processor.
10. The car-mounted KTV system according to claim 7, characterized in that, The number of the microphone is multiple, and multiple microphones are arranged at multiple positions of a vehicle roof; the number of the camera is multiple, and at least one of multiple cameras is arranged at a main driver A-pillar or a steering wheel, and at least one of multiple cameras is arranged at a rearview mirror position in the vehicle; the number of the loudspeaker is multiple, and multiple loudspeakers are arranged at a center console, multiple door positions and a position close to a rear trunk in the vehicle respectively.
11. A vehicle characterized by comprising: The vehicle-mounted KTV system includes the vehicle-mounted KTV system according to any one of claims 7-10.
Citation Information
Patent Citations
Audio processing method and device, terminal equipment and computer storage medium
CN111091845A
Vehicle-mounted KTV control method and device, and vehicle-mounted intelligent network connection terminal
CN113270082A
Audio signal processing method and device, storage medium and electronic equipment
CN114708878A
Method and device for adjusting sound effect of vehicle-mounted sound system
CN114734942A
Vehicle-mounted KTV system and vehicle
CN212073936U