Audio output method, KTV system, electronic device and storage medium

By centrally processing microphone and song audio data through a host server, merging and enhancing the data before splitting it into two recording streams, the system solves the problems of limited recording functionality and high operating costs in KTV systems, thereby enriching entertainment features and saving costs.

CN116758884BActive Publication Date: 2026-04-07VIDEOSTRONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-10
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The existing KTV system has a relatively simple recording function, which is difficult to meet people's diverse needs for entertainment functions, and the separate configuration of a host for each room leads to high operating costs.

Method used

The system uses a host server to uniformly process microphone and song audio data. After being merged and enhanced by timestamps, the data is decomposed into two recording data streams, which are used for singing scoring and singing playback, respectively, thus reducing the host configuration in each room.

Benefits of technology

It enriches the entertainment functions of the KTV system, reduces operating costs and maintenance time, and saves labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758884B_ABST
    Figure CN116758884B_ABST
Patent Text Reader

Abstract

The application discloses a recording output method, a KTV system, an electronic device and a storage medium. The method is applied to the KTV system, and the method comprises the following steps: receiving microphone audio data and song audio data from the same KTV room unit, and generating time stamps for the two data based on the same clock signal; according to the time stamps, the microphone audio data and the song audio data are combined into mixed data, and the mixed data is subjected to enhancement processing to obtain enhanced mixed data; the enhanced mixed data is decomposed into microphone recording data and song recording data; the microphone recording data and the song recording data are transmitted to a singing scoring unit and a singing playback unit respectively, so that the singing scoring unit and the singing playback unit process the microphone recording data and the song recording data respectively, and then return processing information to an audio-video interaction unit of the corresponding KTV room unit. The application can output two-way recording data for singing scoring playback, and enriches the entertainment function of the KTV system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio data processing, and in particular to a recording output method, a KTV system, an electronic device, and a storage medium. Background Technology

[0002] KTV is a venue or space that provides audio-visual equipment for singing. Going to KTV to sing is a common consumer entertainment activity. Among related technologies, the recording function of KTV systems mainly records the audio data (singing) collected by the microphone and outputs it to the singing scoring software, which then scores the singer's singing.

[0003] However, the current entertainment functions of KTV systems based on recording are relatively simple and cannot meet people's diverse needs for KTV entertainment functions. Summary of the Invention

[0004] To address the aforementioned technical problems and deficiencies, the present invention aims to provide a recording output method, a KTV system, an electronic device, and a storage medium that can output two channels of recording data for performance scoring and playback, thereby enriching the entertainment functions of the KTV system.

[0005] To achieve the above objectives, in a first aspect, the present invention provides a recording output method applied to a host server of a KTV system. The KTV system further includes a singing scoring unit, a singing playback unit, and multiple KTV room units. Each KTV room unit includes an audio-visual interaction unit, a first audio processing unit, and a second audio processing unit. The method includes:

[0006] The system receives microphone audio data from a first audio processing unit and song audio data from a second audio processing unit in the same KTV room unit. The microphone audio data and song audio data generate timestamps based on the same clock signal. The microphone audio data is transmitted from the microphone to the first audio processing unit, and the song audio data is transmitted from the audio-visual interaction unit to the second audio processing unit.

[0007] Based on the timestamp, the microphone audio data and the song audio data are merged into mixed data, and the mixed data is enhanced to obtain enhanced mixed data;

[0008] The enhanced mixed data is broken down into microphone recording data and song recording data;

[0009] The microphone recording data and song recording data are transmitted to the singing scoring unit and the singing playback unit, respectively, so that the singing scoring unit and the singing playback unit can process the microphone recording data and the song recording data respectively, and then return the processing information to the audio-visual interaction unit of the corresponding KTV room unit.

[0010] In the above embodiment, the host server can output two channels of recording data: one channel of microphone recording data for the singing scoring function and the other channel of song recording data for the singing playback function, enriching the entertainment functions of the KTV system. Simultaneously, in this embodiment of the KTV system, each KTV room no longer needs a separate host; all KTV rooms share a single host server, saving on KTV operating costs and simplifying operation and maintenance. Only periodic maintenance of the host server is required, saving time and manpower costs compared to related technologies where each KTV room has a separate host.

[0011] Optionally, the step of merging microphone audio data and song audio data into mixed data based on timestamps includes: splicing microphone audio data and song audio data with the same timestamp to obtain mixed data.

[0012] In the above embodiments, enhancing the mixed data can involve amplification and filtering to remove noise and improve the audio quality. Since the mixed data is a single stream, host server 1 can improve processing efficiency by performing filtering, amplification, and other enhancement processes on that single stream.

[0013] Optionally, the enhanced mixed data includes first left channel data, first right channel data, second left channel data, and second right channel data arranged in sequence. The step of decomposing the enhanced mixed data into microphone recording data and song recording data includes:

[0014] The connection point between the first right channel data and the second left channel data is determined as the cutting point;

[0015] The data is cut based on the cut point, and the first left channel data and the first right channel data are used as microphone recording data, while the second left channel data and the second right channel data are used as song recording data.

[0016] In the above embodiment, the first left channel data and the first right channel data correspond to song recording data, and the second left channel data and the second right channel data correspond to microphone recording data. This allows the enhanced mixed data to be precisely divided into two data streams.

[0017] Optionally, before the step of transmitting the microphone recording data and song recording data to the performance scoring unit and performance playback unit respectively, the method further includes:

[0018] Convert the microphone recording data and song recording data into a data format suitable for the performance scoring unit and performance playback unit.

[0019] In the above embodiments, the singing scoring unit can directly process the microphone recording data, and the singing playback unit can directly process the song recording data, thereby improving data processing efficiency.

[0020] Optionally, the KTV system further includes a first relay unit, the input of which is connected to a first audio processing unit and a second audio processing unit respectively, and the output of which is connected to a host server; the step of receiving microphone audio data from the first audio processing unit and song audio data from the second audio processing unit in the same KTV room unit includes:

[0021] It receives microphone audio data and song audio data that have been enhanced and processed by the first relay unit, respectively.

[0022] In the above embodiments, the first relay unit can amplify and filter microphone audio data and song audio data to improve data signal quality. Especially when the distance between the KTV room unit and the host server is far, data is prone to distortion or loss during transmission. Through the enhancement processing of the first relay unit, the microphone audio data and song audio data received by the host server can be complete and accurate.

[0023] Optionally, the KTV system also includes a second relay unit. The input of the second relay unit is connected to the singing playback unit and the singing scoring unit, and the output is connected to the audio-visual interactive device. After transmitting the microphone recording data and song recording data to the singing scoring unit and the singing playback unit respectively, the system further includes:

[0024] The system receives delivery messages from the performance playback unit and the performance scoring unit. These delivery messages are generated by the performance playback unit and the performance scoring unit after transmitting the microphone recording data and song recording data to the second relay unit for enhancement processing.

[0025] In the above embodiments, the enhanced processing of the second relay unit can ensure the data transmission quality between the KTV room unit and the singing scoring unit and singing playback unit.

[0026] Optionally, after the step of decomposing the enhanced mixed data into microphone recording data and song recording data, the method further includes:

[0027] Obtain first volume information and second volume information from the song recording data; the first volume information is used to characterize the volume of the singing voice in the song recording data, and the second information is used to characterize the volume of the accompaniment music in the song recording data.

[0028] Adjust the first volume information and / or the second volume information to match the volume of the vocals in the song recording data with the volume of the accompaniment music.

[0029] This allows the volume of the vocals and accompaniment in the song recording to reach a relatively balanced state, improving the playback effect of the song.

[0030] Secondly, the present invention provides a KTV system, including a host server, a singing scoring unit, a singing playback unit, and multiple KTV room units. Each KTV room unit includes an audio-visual interaction unit, a first audio processing unit, and a second audio processing unit. The host server includes:

[0031] The receiving module is used to receive microphone audio data from the first audio processing unit and song audio data from the second audio processing unit in the same KTV room unit. The microphone audio data and song audio data generate timestamps based on the same clock signal. The microphone audio data is transmitted from the microphone to the first audio processing unit, and the song audio data is transmitted from the audio-visual interaction unit to the second audio processing unit.

[0032] The merging module is used to merge microphone audio data and song audio data into mixed data based on timestamps, and to enhance the mixed data to obtain enhanced mixed data.

[0033] The decomposition module is used to decompose the enhanced mixed data into microphone recording data and song recording data;

[0034] The transmission module is used to transmit microphone recording data and song recording data to the singing scoring unit and singing playback unit respectively, so that the singing scoring unit and singing playback unit can process the microphone recording data and song recording data respectively and return the processing information to the audio-visual interaction unit of the corresponding KTV room unit.

[0035] Thirdly, the present invention provides an electronic device, including a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned recording output method is implemented.

[0036] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described recording output method.

[0037] One or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:

[0038] 1. The host server can output two recording data channels: one microphone recording data channel for the singing scoring function and the other song recording data channel for the singing playback function, which enriches the entertainment functions of the KTV system.

[0039] 2. In the KTV system of this embodiment, each KTV room no longer needs to be configured with a separate host. All KTV rooms share a single host server, which saves the operating costs of the KTV and facilitates the operation and maintenance of the KTV. Only regular maintenance of the host server is required. Compared with the related technologies where each KTV room is configured with a separate host, it saves the time and manpower costs of maintenance.

[0040] 3. First, the microphone audio data and song recording data are merged into one data stream for enhancement processing, and then decomposed into microphone recording data and song recording data. These two data streams are then converted into data formats suitable for the singing scoring unit and singing playback unit, respectively, which improves the efficiency of data processing and transmission. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0042] Figure 1 This is a schematic diagram of the architecture of a KTV system in related technologies;

[0043] Figure 2 This is a schematic diagram of the architecture of a KTV system in an embodiment of the present invention;

[0044] Figure 3 This is the flow chart of the recording output method in this embodiment of the invention. Figure 1 ;

[0045] Figure 4 This is a schematic diagram of the structure for enhancing hybrid data in an embodiment of the present invention. Figure 1 ;

[0046] Figure 5 This is a schematic diagram of the structure for enhancing hybrid data in an embodiment of the present invention. Figure 2 ;

[0047] Figure 6 This is the flow chart of the recording output method in this embodiment of the invention. Figure 2 ;

[0048] Figure 7 This is a schematic diagram of the architecture of a host server in an embodiment of the present invention;

[0049] Figure 8 This is a schematic diagram of the architecture of an electronic device in an embodiment of the present invention. Detailed Implementation

[0050] The terminology used in the following embodiments of the present invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the specification and appended claims of the present invention, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in the present invention refers to and includes any or all possible combinations of one or more of the listed items.

[0051] Hereinafter, the terms "first" and "second" are used for descriptive purposes only to distinguish technical features and should not be construed as implying relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, unless otherwise stated, "multiple" means two or more.

[0052] In related technologies, such as Figure 1 As shown, a traditional KTV system typically includes a service counter and multiple KTV rooms. Each KTV room includes an audio-visual interaction unit, an audio processing unit, a host computer, and a singing scoring unit.

[0053] The audio processing unit transmits the human voice signal captured by the microphone to the audio processing unit. After processing the microphone's human voice signal, the audio processing unit generates microphone audio data and sends it to the audio-visual interaction unit, which plays it in conjunction with the accompaniment music. The host receives the microphone audio data from the audio processing unit, amplifies and filters it, generates microphone recording data, and sends it to the singing scoring unit. The singing scoring unit evaluates and scores the microphone recording data and sends the scoring result to the audio-visual interaction unit.

[0054] Each KTV room's main unit is connected to the service center, which controls and monitors the status of each KTV room.

[0055] It is evident that, among the relevant technologies, the entertainment functions of KTV systems based on microphone recording data are relatively simple, and the high operating costs of KTVs are due to the need for a separate host for each room.

[0056] Therefore, embodiments of the present invention provide a KTV system, such as Figure 2 As shown, it includes a host server 1, a singing scoring unit 2, a singing playback unit 3, and multiple KTV room units 4. Each KTV room unit 4 includes an audio-visual interaction unit 41, a first audio processing unit 42, and a second audio processing unit 43.

[0057] The first audio processing unit 42 processes the human voice signal collected by the microphone mic to generate microphone audio data, and sends it to the audio-visual interaction unit 41. The audio-visual interaction unit 41 plays the audio data in combination with the accompaniment music, and at the same time generates song audio data, which is then transmitted to the second audio processing unit 43. The song audio data is generated by the audio-visual interaction unit 41 by combining the microphone audio data with the accompaniment music.

[0058] The first audio processing unit 42 filters and amplifies the microphone audio data before sending it to the host server 1. The second audio processing unit 43 filters and amplifies the song audio data before sending it to the host server 1. The host server 1 receives the two data streams, merges them into one stream for enhancement, and then decomposes them into two streams: one is microphone recording data, and the other is song audio data. These streams are then sent to the singing scoring unit 2 and the singing playback unit 3 for processing, respectively.

[0059] The first processing unit and the second audio processing unit 43 can be DSP (Digital Signal Processing) hardware.

[0060] The singing scoring unit 2 evaluates and scores the vocals based on the microphone recording data and sends the score to the audio-visual interaction unit 41. The singing playback unit 3 can correct the vocal parts of the song audio data, for example, by correcting off-key pitches to the correct pitch, and then send the corrected song audio data to the audio-visual interaction unit 41 for song playback. The singing scoring unit 2 and the singing playback unit 3 can be software programs or separate hardware modules.

[0061] Thus, the KTV system in this embodiment can output two channels of recording data: one channel of microphone recording data for the singing scoring function and the other channel of song recording data for the singing playback function, enriching the entertainment functions of the KTV system. Simultaneously, in this embodiment, each KTV room no longer needs a separate host computer; all KTV rooms share a single host server 1, saving on KTV operating costs and simplifying operation and maintenance. Only periodic maintenance of host server 1 is required, saving time and manpower costs compared to related technologies where each KTV room has a separate host computer.

[0062] This invention provides a recording output method applied to the host server 1 of the KTV system provided in the above embodiment. The method is executed by the host server 1, such as... Figure 3 As shown, the steps include 101, 102, 103, and 104:

[0063] Step 101: Receive microphone audio data from the first audio processing unit 42 and song audio data from the second audio processing unit 43 within the same KTV room unit 4. The microphone audio data and song audio data generate timestamps based on the same clock signal. The microphone audio data is transmitted from the microphone to the first audio processing unit 42, and the song audio data is transmitted from the audio-visual interaction unit 41 to the second audio processing unit 43.

[0064] The microphone audio data and the song audio data share the same clock signal, so their timestamps are also synchronized. For example, assuming the timestamp is 17:26:30 on May 10, 2023, the corresponding song time in both the microphone audio data and the song audio data is 1 minute and 12 seconds. This ensures that the microphone audio data and the song data are synchronized and avoids timing chaos in subsequent data processing.

[0065] In one embodiment, microphone audio data and song audio data can be in I2S (Inter-IC Sound) format. I2S is a digital audio transmission protocol used to transmit audio data between various digital audio devices, such as computers, mobile devices, Bluetooth speakers, digital audio processors, etc. It uses synchronous data transmission and transmits audio data from multiple channels simultaneously. I2S is commonly used in digital audio processing, such as audio recording, playback, audio encoding / decoding, filtering, etc.

[0066] Step 102: Based on the timestamp, merge the microphone audio data and the song audio data into mixed data, and perform enhancement processing on the mixed data to obtain enhanced mixed data.

[0067] In one embodiment, microphone audio data and song audio data with the same timestamp can be concatenated to obtain mixed data. Specifically, a unit time is set based on the timestamp, and data portions within the same unit time period from the microphone audio data and song audio data are extracted and concatenated together. For example, if the unit time is 0.1 seconds, the portion from 17:26:30.1 to 17:26:30.2 on May 10, 2023 can be extracted and concatenated to obtain the mixed data portion from 17:26:30.1 to 17:26:30.2 on May 10, 2023. Then, according to the timestamp order, the mixed data portions within each 0.1-second interval are concatenated to obtain the overall mixed data.

[0068] This enhancement process for mixed data can involve amplification and filtering to remove noise and improve audio quality. Since the mixed data is processed as a single stream, host server 1 can improve processing efficiency by performing filtering, amplification, and other enhancements on that single stream.

[0069] Step 103: Decompose the enhanced mixed data into microphone recording data and song recording data.

[0070] The enhanced mixed data includes, in sequence, first left channel data, first right channel data, second left channel data, and second right channel data. Specifically, as follows... Figure 4 and Figure 5 As shown, the enhanced mixed data is 32-bit data, where bits 1 to 8 are the first left channel data, bits 9 to 16 are the first right channel data, bits 17 to 24 are the second left channel data, and bits 25 to 32 are the second right channel data.

[0071] In some embodiments, when microphone audio data and song audio data are combined into mixed data, the microphone audio data is spliced ​​before the song audio data in each unit of time. In the enhanced mixed data, the first left channel data and the first right channel data correspond to the microphone recording data, and the second left channel data and the second right channel data correspond to the song recording data.

[0072] Specifically, step 103 may include:

[0073] The junction of the first right channel data and the second left channel data is determined as the cutting point.

[0074] The data is cut based on the cut point, and the first left channel data and the first right channel data are used as microphone recording data, while the second left channel data and the second right channel data are used as song recording data.

[0075] For example, such as Figure 4 As shown, the data is divided into two groups: bits 1 through 16 are separated into one group, and bits 17 through 32 are separated into another group. Specifically, bits 1 through 16 correspond to the first left and first right channel data, which are the microphone recording data. Bits 17 through 32 correspond to the second left and second right channel data, which are the song recording data.

[0076] In some embodiments, when microphone audio data and song audio data are combined into mixed data, the song audio data is spliced ​​before the microphone audio data in each unit of time. In the enhanced mixed data, the first left channel data and the first right channel data correspond to the song recording data, and the second left channel data and the second right channel data correspond to the microphone recording data.

[0077] Specifically, step 103 may also include:

[0078] The junction of the first right channel data and the second left channel data is determined as the cutting point.

[0079] The data is cut based on the cut point, and the first left channel data and the first right channel data are used as the song recording data, while the second left channel data and the second right channel data are used as the microphone recording data.

[0080] For example, such as Figure 5 As shown, the data is divided into two groups: bits 1 through 16 are separated into one group, and bits 17 through 32 are separated into another group. Specifically, bits 1 through 16 correspond to the first left and first right channel data, which is the recording data for the song. Bits 17 through 32 correspond to the second left and second right channel data, which is the microphone recording data.

[0081] Step 104: The microphone recording data and the song recording data are transmitted to the singing scoring unit 2 and the singing playback unit 3 respectively, so that the singing scoring unit 2 and the singing playback unit 3 can process the microphone recording data and the song recording data respectively, and then return the processing information to the audio-visual interaction unit 41 of the corresponding KTV room unit 4.

[0082] The singing scoring unit 2 evaluates and scores the performance based on the microphone recording data and sends the results to the audio-visual interaction unit 41. Specifically, it compares the microphone recording data with the original audio data, analyzing their similarity. The higher the similarity, the higher the score; the lower the similarity, the lower the score. The final score is then returned to the audio-visual interaction unit 41 in the corresponding KTV room unit 4. Users can view their singing scores on the audio-visual interaction unit 41.

[0083] The singing playback unit 3 can correct the vocal part in the song audio data, for example, correcting off-key pitches to the correct pitches, and then send the corrected song audio data to the audio-visual interaction unit 41, which will then play back the song.

[0084] Using the recording output method of this embodiment, the host server 1 can output two channels of recording data: one channel of microphone recording data for the singing scoring function and the other channel of song recording data for the singing playback function, enriching the entertainment functions of the KTV system. Simultaneously, in this embodiment's KTV system, each KTV room no longer needs a separate host; all KTV rooms share a single host server 1, saving on KTV operating costs and simplifying operation and maintenance. Only periodic maintenance of the host server 1 is required, saving time and manpower costs compared to related technologies where each KTV room has a separate host.

[0085] In one embodiment, the KTV system further includes a first relay unit 5, the input of which is connected to the audio processing unit and the output of which is connected to the host server 1.

[0086] The first relay unit 5 is used to enhance the microphone audio data output by the first audio processing unit 42 and the song audio data output by the second audio processing unit 43. Specifically, it can amplify and filter the microphone audio data and song audio data to improve the data signal quality. Especially when the distance between the KTV room unit 4 and the host server 1 is far, data is prone to distortion or loss during transmission. Through the enhancement processing of the first relay unit 5, the microphone audio data and song audio data received by the host server 1 can be complete and accurate.

[0087] In one embodiment, the step of receiving microphone audio data from a first audio processing unit 42 and song audio data from a second audio processing unit 43 in the same KTV room unit 4 includes:

[0088] It receives microphone audio data and song audio data that have been enhanced and processed by the first relay unit 5, respectively.

[0089] After the first relay unit 5 enhances the processing, it can be ensured that the microphone audio data and song audio data of the host server 1 can be accurately and completely transmitted to the host server 1.

[0090] In one embodiment, the KTV system further includes a second relay unit 6, the input of which is connected to the singing playback unit 3 and the singing scoring unit 2, and the output of which is connected to the audio-visual interactive device.

[0091] The second relay unit 6 is used to enhance the microphone recording data and song recording data by amplifying and filtering them. Considering that the singing scoring unit 2 and singing playback unit 3 may be relatively far from the KTV room unit 4, the enhancement processing of the second relay unit 6 can ensure the data transmission quality between the KTV room unit 4 and the singing scoring unit 2 and singing playback unit 3.

[0092] After the steps of transmitting the microphone recording data and the song recording data to the singing scoring unit 2 and the singing playback unit 3 respectively, the following steps are also included:

[0093] The system receives delivery messages from the performance playback unit 3 and the performance scoring unit 2. These delivery messages are generated by the performance playback unit 3 and the performance scoring unit 2 after transmitting the microphone recording data and song recording data to the second relay unit 6 for enhancement processing.

[0094] Specifically, host server 1 sends microphone recording data and song recording data to performance playback unit 3 and performance scoring unit 2, respectively. Performance playback unit 3 and performance scoring unit 2 process the microphone recording data and song recording data, then send the processed data to the second relay unit 6. The relay unit enhances the processed data before sending it to the audio-visual interactive device in KTV room unit 4. Afterwards, performance playback unit 3 and performance scoring unit 2 return a delivery message to host server 1, indicating that the microphone recording data and song recording data have been processed.

[0095] In one embodiment, such as Figure 6 As shown, the recording output method specifically includes the following steps:

[0096] Step 201: Host server 1 receives microphone audio data and song audio data from the same KTV room unit 4.

[0097] The microphone audio data and the song audio data are timestamped based on the same clock signal. The microphone audio data and the song audio data are transmitted by the audio-visual interaction unit 41 to the first audio processing unit 42 and the second audio processing unit 43, respectively.

[0098] Step 202: Based on the timestamp, host server 1 merges the microphone audio data and song audio data into mixed data, and performs enhancement processing on the mixed data to obtain enhanced mixed data.

[0099] Specifically, a time unit based on timestamps is set, and data portions within the same time unit from both microphone audio data and song audio data are extracted and concatenated. For example, if the time unit is 0.2 seconds, the portion from 14:59:30.4 to 14:59:30.6 on May 11, 2023 can be extracted and concatenated to obtain the mixed data portion from 14:59:30.4 to 14:59:30.6 on May 11, 2023. Then, according to the timestamp order, the mixed data portions within each 0.2-second interval are concatenated to obtain the overall mixed data.

[0100] Step 203: Host server 1 decomposes the enhanced mixed data into microphone recording data and song recording data.

[0101] The enhanced mixed data includes first left channel data, first right channel data, second left channel data, and second right channel data arranged sequentially. Specifically, the enhanced mixed data is 32-bit data, where bits 1 to 8 are the first left channel data, bits 9 to 16 are the first right channel data, bits 17 to 24 are the second left channel data, and bits 25 to 32 are the second right channel data.

[0102] When merging microphone audio data and song audio data into mixed data, the microphone audio data is spliced ​​before the song audio data in each unit of time. In the enhanced mixed data, the first left channel data and the first right channel data correspond to the microphone recording data, and the second left channel data and the second right channel data correspond to the song recording data.

[0103] When merging microphone audio data and song audio data into mixed data, in each unit of time, the song audio data is spliced ​​before the microphone audio data. In the enhanced mixed data, the first left channel data and the first right channel data correspond to the song recording data, and the second left channel data and the second right channel data correspond to the microphone recording data.

[0104] Step 204: Host server 1 obtains the first volume information and the second volume information of the song recording data.

[0105] The first volume information is used to characterize the volume of the singing voice in the song recording data, and the second information is used to characterize the volume of the accompaniment music in the song recording data.

[0106] Step 205: Adjust the first volume information and / or the second volume information to match the volume of the vocals in the song recording data with the volume of the accompaniment music.

[0107] Specifically, when the first volume information is too loud, the volume of the singing voice is reduced so that the volume of the singing voice is within the standard range and the difference between the volume of the singing voice and the volume of the accompaniment music is within the set range.

[0108] When the second volume information is too loud, reduce the volume of the accompaniment music so that the volume of the accompaniment music is within the standard range and the difference between the volume of the singing voice and the volume of the singing voice is within the set range.

[0109] When the initial volume information is too low, increase the volume of the singing voice to ensure that the volume of the singing voice is within the standard range and that the volume difference between the singing voice and the accompaniment music is within the set range.

[0110] When the second volume information is too low, increase the volume of the accompaniment music to ensure that the volume of the accompaniment music is within the standard range and that the difference between the volume of the accompaniment music and the volume of the singing voice is within the set range.

[0111] When both the first and second volume information are too low, the volume of the singing voice and the accompaniment music are increased simultaneously, so that the volume of the singing voice and the accompaniment music reaches the set volume value, while the difference between the two is within the set range.

[0112] When both the first and second volume information are too high, the volume of the singing voice and the accompaniment music will be reduced simultaneously so that the volume of the singing voice and the accompaniment music reaches the set volume value, while the difference between the two is within the set range.

[0113] This allows the volume of the vocals in the song recording to match the volume of the accompaniment music, achieving a relatively balanced state and improving the playback effect of the song.

[0114] Step 206: Host server 1 converts the microphone recording data and song recording data into a data format suitable for singing scoring unit 2 and singing playback unit 3.

[0115] This allows the singing scoring unit 2 to directly process the microphone recording data, and the singing playback unit 3 to directly process the song recording data, thus improving data processing efficiency.

[0116] Step 207: The host server 1 transmits the microphone recording data and the song recording data to the singing scoring unit 2 and the singing playback unit 3, respectively.

[0117] Specifically, the singing scoring unit 2 evaluates and scores the performance based on the microphone recording data and sends the score to the audio-visual interaction unit 41. Specifically, it compares the microphone recording data with the original audio data, analyzing their similarity. Higher similarity results in a higher score, and lower similarity results in a lower score. The final score information is then returned to the audio-visual interaction unit 41 in the corresponding KTV room unit 4. Users can view their singing score on the audio-visual interaction unit 41.

[0118] The singing playback unit 3 can correct the vocal part in the song audio data, for example, correcting off-key pitches to the correct pitches, and then send the corrected song audio data to the audio-visual interaction unit 41, which will then play back the song.

[0119] The recording output method provided in this embodiment allows the host server 1 to output two channels of recording data: one channel of microphone recording data for the singing scoring function and the other channel of song recording data for the singing playback function, enriching the entertainment functions of the KTV system. Furthermore, in this embodiment's KTV system, each KTV room no longer needs a separate host; all KTV rooms share a single host server 1, saving on KTV operating costs and simplifying operation and maintenance. Only periodic maintenance of the host server 1 is required, saving time and manpower costs compared to related technologies where each KTV room has a separate host.

[0120] This invention provides a KTV system, such as... Figure 1 As shown, it includes a host server 1, a singing scoring unit 2, a singing playback unit 3, and multiple KTV room units 4. Each KTV room unit 4 includes an audio-visual interaction unit 41, a first audio processing unit 42, and a second audio processing unit 43. The host server 1 applies the recording output method provided in the above embodiment, such as... Figure 7 As shown, it specifically includes a receiving module 11, a merging module 12, a decomposition module 13, and a transmission module 14.

[0121] The receiving module 11 is used to receive microphone audio data from the first audio processing unit 42 and song audio data from the second audio processing unit 43 in the same KTV room unit 4. The microphone audio data and song audio data generate timestamps based on the same clock signal. The microphone audio data is transmitted from the microphone to the first audio processing unit 42, and the song audio data is transmitted from the audio-visual interaction unit 41 to the second audio processing unit 43.

[0122] The merging module 12 is used to merge microphone audio data and song audio data into mixed data based on timestamps, and to enhance the mixed data to obtain enhanced mixed data;

[0123] Decomposition module 13 is used to decompose the enhanced mixed data into microphone recording data and song recording data;

[0124] The transmission module 14 is used to transmit the microphone recording data and the song recording data to the singing scoring unit 2 and the singing playback unit 3 respectively, so that the singing scoring unit 2 and the singing playback unit 3 can process the microphone recording data and the song recording data respectively, and then return the processing information to the audio-visual interaction unit 41 of the corresponding KTV room unit 4.

[0125] In this embodiment, the host server 1 of the KTV system adopts the method provided in the above embodiment, and can output two channels of recording data: one channel of microphone recording data for the singing scoring function, and the other channel of song recording data for the singing playback function, thus enriching the entertainment functions of the KTV system. Simultaneously, in this embodiment, each KTV room no longer needs a separate host; all KTV rooms share a single host server 1, saving on KTV operating costs and simplifying operation and maintenance. Only periodic maintenance of the host server 1 is required, saving time and manpower costs compared to related technologies where each KTV room has a separate host.

[0126] Figure 8 A schematic diagram of a computer system suitable for implementing embodiments of the present invention is shown.

[0127] It should be noted that, Figure 8 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.

[0128] like Figure 8 As shown, the computer system includes a Central Processing Unit (CPU) 1801, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1802 or programs loaded from storage portion 1808 into Random Access Memory (RAM) 1803, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 1803. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via bus 1804. An Input / Output (I / O) interface 1805 is also connected to bus 1804.

[0129] The following components are connected to I / O interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to I / O interface 1805 as needed. Removable media 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1810 as needed so that computer programs read from them can be installed into storage section 1808 as needed.

[0130] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by central processing unit (CPU) 1801, it performs various functions defined in the system of the present invention.

[0131] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0133] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0134] Specifically, the electronic device in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the method provided in the above embodiment.

[0135] The electronic device in this embodiment allows the server to output two recording data streams: one microphone recording data stream for the singing scoring function and the other song recording data stream for the singing playback function, thus enriching the entertainment functions of the KTV system.

[0136] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The storage medium carries one or more computer programs that, when executed by a processor of the electronic device, cause the electronic device to implement the methods provided in the above embodiments.

[0137] It should be noted that although several modules or units of the device for performing actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0138] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, host server, touch terminal, or network device, etc.) to execute the method according to the embodiments of the present invention.

[0139] Specifically, the storage medium of this embodiment can implement the method shown in the above embodiments. It allows the server to output two channels of recording data: one microphone recording data for the singing scoring function and the other song recording data for the singing playback function, thus enriching the entertainment functions of the KTV system.

[0140] In this embodiment, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the above description to give a full understanding of the embodiments of the invention. However, those skilled in the art will recognize that the technical solutions of the invention can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the invention.

[0141] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0142] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0143] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein.

[0144] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A recording output method, characterized in that, A host server is used in a KTV system, which further includes a singing scoring unit, a singing playback unit, and multiple KTV room units. Each KTV room unit includes an audio-visual interaction unit, a first audio processing unit, and a second audio processing unit. The method includes: The system receives microphone audio data from the first audio processing unit and song audio data from the second audio processing unit within the same KTV room unit. The microphone audio data and the song audio data are timestamped based on the same clock signal. The microphone audio data is transmitted from the microphone to the first audio processing unit, and the song audio data is transmitted from the audio-visual interaction unit to the second audio processing unit. The song audio data is generated by the audio-visual interaction unit by combining the microphone audio data with the accompaniment music. Based on the timestamp, the microphone audio data and the song audio data are merged into mixed data, and the mixed data is enhanced to obtain enhanced mixed data; The enhanced mixed data is decomposed into microphone recording data and song recording data; The microphone recording data is transmitted to the singing scoring unit and the song recording data is transmitted to the singing playback unit, so that the singing scoring unit processes the microphone recording data and the singing playback unit processes the song recording data and then returns the processing information to the audio-visual interaction unit of the corresponding KTV room unit.

2. The recording output method according to claim 1, characterized in that, The step of merging the microphone audio data and the song audio data into mixed data based on the timestamp includes: The microphone audio data and the song audio data with the same timestamp are concatenated to obtain mixed data.

3. The recording output method according to claim 2, characterized in that, The enhanced mixed data includes first left channel data, first right channel data, second left channel data, and second right channel data arranged in sequence. The step of decomposing the enhanced mixed data into microphone recording data and song recording data includes: The connection point between the first right channel data and the second left channel data is determined as the cutting point; Based on the cutting point, the data is cut, and the first left channel data and the first right channel data are used as microphone recording data, while the second left channel data and the second right channel data are used as song recording data.

4. The recording output method according to any one of claims 1 to 3, characterized in that, Before the step of transmitting the microphone recording data and the song recording data to the singing scoring unit and the singing playback unit respectively, the method further includes: The microphone recording data and the song recording data are converted into a data format suitable for the singing scoring unit and the singing playback unit.

5. The recording output method according to claim 1, characterized in that, The KTV system further includes a first relay unit, the input of which is connected to the first audio processing unit and the second audio processing unit, and the output of which is connected to the host server; the step of receiving microphone audio data from the first audio processing unit and song audio data from the second audio processing unit in the same KTV room unit includes: The microphone audio data and the song audio data, which have been enhanced and processed by the first relay unit, are received respectively.

6. The recording output method according to claim 1, characterized in that, The KTV system further includes a second relay unit, the input of which is connected to the singing playback unit and the singing scoring unit, and the output of which is connected to the audio-visual interaction unit. After the step of transmitting the microphone recording data and the song recording data to the singing scoring unit and the singing playback unit respectively, the system further includes: The system receives delivery messages returned by the performance playback unit and the performance scoring unit. These delivery messages are generated by the performance playback unit and the performance scoring unit after transmitting the microphone recording data and the song recording data to the second relay unit for enhancement processing.

7. The recording output method according to claim 1, characterized in that, Following the step of decomposing the enhanced mixed data into microphone recording data and song recording data, the method further includes: Obtain first volume information and second volume information from the song recording data; the first volume information is used to characterize the volume of the singing voice in the song recording data, and the second volume information is used to characterize the volume of the accompaniment music in the song recording data. Adjust the first volume information and / or the second volume information to match the volume of the singing voice and the accompaniment music in the song recording data.

8. A KTV system, characterized in that, It includes a host server, a singing scoring unit, a singing playback unit, and multiple KTV room units. The KTV room unit includes an audio-visual interaction unit, a first audio processing unit, and a second audio processing unit. The host server includes: The receiving module is used to receive microphone audio data from the first audio processing unit and song audio data from the second audio processing unit in the same KTV room unit. The microphone audio data and the song audio data generate timestamps based on the same clock signal. The microphone audio data is transmitted from the microphone to the first audio processing unit, and the song audio data is transmitted from the audio-visual interaction unit to the second audio processing unit. The song audio data is generated by the audio-visual interaction unit by combining the microphone audio data with the accompaniment music. The merging module is used to merge the microphone audio data and the song audio data into mixed data according to the timestamp, and to enhance the mixed data to obtain enhanced mixed data; A decomposition module is used to decompose the enhanced mixed data into microphone recording data and song recording data; The transmission module is used to transmit the microphone recording data to the singing scoring unit and the song recording data to the singing playback unit, so that the singing scoring unit processes the microphone recording data and the singing playback unit processes the song recording data and then returns the processing information to the audio-visual interaction unit of the corresponding KTV room unit.

9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the recording output method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the recording output method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Karaoke system based on intelligent terminal and wireless speaker and implementation method of Karaoke system

    CN103874005A

  • Multi-channel audio based real-time scoring method, storage device and application thereof

    CN107221340A