Audio and video data processing method, system, equipment and medium

By receiving and merging audio data groups and video resources from multiple senders, the need for multi-user karaoke is addressed, improving user experience and reducing costs.

CN121644882APending Publication Date: 2026-03-10IMUSIC CULTURE & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient to meet the needs of multiple users singing karaoke simultaneously, resulting in a poor user experience. Furthermore, dedicated microphone equipment is expensive and lacks scalability.

Method used

By receiving user audio data groups from several sending ends, performing audio merging processing, and combining them with video resources for audio-video merging processing, the data is sent to the terminal for playback, enabling multiple users to sing karaoke simultaneously.

Benefits of technology

It fulfills the need for multiple users to sing karaoke simultaneously, improves user experience, reduces user costs, and enhances the scalability and flexibility of audio processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644882A_ABST
    Figure CN121644882A_ABST
Patent Text Reader

Abstract

The invention discloses an audio and video data processing method, system and device and a medium, and the method comprises the steps: receiving user audio data sets from a plurality of transmitting ends, and obtaining video resources corresponding to all user audio data sets; performing audio merging processing on all the user audio data groups to obtain an audio merging signal; according to the video resource, performing audio and video combination processing on the audio combination signal to obtain a signal group, the signal group comprising a target audio stream signal and a target video stream signal; the target audio stream signal comprises the original audio of the video resource and the user audio recorded by each sending end; and sending the signal group to the terminal, so that the terminal plays audios to the outside according to the target audio stream signal in the signal group and plays videos to the outside according to the target video stream signal in the signal group. The method can meet the requirement of multiple users for simultaneously singing karaoke, and the user experience is effectively improved. The invention relates to the technical field of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, system, device and medium for processing audio and video data. Background Technology

[0002] In recent years, with the increasing popularity of audio and video-based karaoke applications in home entertainment scenarios, users' demand for convenient and immersive karaoke has significantly increased.

[0003] Currently, the relevant technologies usually rely on local hardware, specifically by requiring users to purchase dedicated wired or wireless microphone equipment to connect to a TV box to enable karaoke applications. This method typically only supports a single microphone input, making it difficult to meet the needs of multiple users singing karaoke simultaneously, resulting in an unsatisfactory user experience.

[0004] Therefore, the problems existing in the current technology still need to be solved and optimized. Summary of the Invention

[0005] To solve at least one of the above-mentioned technical problems, this application provides a method, system, device, and medium for processing audio and video data, wherein the method can meet the needs of multiple users singing karaoke simultaneously and effectively improve the user experience.

[0006] According to a first aspect of this application, a method for processing audio and video data is provided, applied at a receiving end, the method comprising: Receive user audio data groups from several transmitters and obtain video resources corresponding to all of the user audio data groups; All user audio data groups are subjected to audio merging processing to obtain a merged audio signal; Based on the video resource, the audio merged signal is subjected to audio-video merging processing to obtain a signal group, which includes a target audio stream signal and a target video stream signal; the target audio stream signal includes the original audio of the video resource and the audio merged signal. The signal group is sent to the terminal so that the terminal can play audio to the outside world according to the target audio stream signal in the signal group, and play video to the outside world according to the target video stream signal in the signal group.

[0007] In some embodiments, the audio merging process performed on all the user audio data groups to obtain a merged audio signal includes: Rules for obtaining the array parameters of the user audio data group; According to the array parameter rules, each user audio data group is parsed to obtain the user audio for each user audio data group; All the user audio data are merged to obtain the merged audio signal.

[0008] In some embodiments, the step of performing audio-video merging processing on the audio-video merged signal based on the video resource to obtain a signal group includes: The video resources are subjected to audio and video extraction to obtain a first audio stream signal and an original video stream signal; Based on the first audio stream signal, the audio merge signal is merged to obtain the target audio stream signal; The original video stream signal is hardware decoded to obtain the target video stream signal.

[0009] In some embodiments, the step of combining the audio merged signal based on the first audio stream signal to obtain the target audio stream signal includes: The first audio stream signal is decoded to obtain the second audio stream signal; The second audio stream signal is subjected to channel audio extraction to obtain a third audio stream signal, which is used to characterize part or all of the channel audio of the second audio stream signal; Based on the third audio stream signal, the audio merge signal is merged to obtain the target audio stream signal.

[0010] According to a second aspect of this application, a method for processing audio and video data is provided, applied at a transmitting end, wherein the transmitting end is communicatively connected to a receiving end, the method comprising: Collect the user's raw audio; The original audio is processed to obtain a user audio data set; The user audio data groups are sent to the receiving end so that the receiving end can acquire video resources corresponding to all the user audio data groups; audio merging processing is performed on all the user audio data groups to obtain an audio merged signal; according to the video resources, audio-video merging processing is performed on the audio merged signal to obtain a signal group, the signal group including a target audio stream signal and a target video stream signal; the target audio stream signal includes the original audio of the video resources and the audio merged signal; the signal group is sent to the terminal so that the terminal can play audio to the outside world according to the target audio stream signal in the signal group, and play video to the outside world according to the target video stream signal in the signal group.

[0011] In some embodiments, the audio processing of the original audio to obtain a user audio data set includes: Obtain volume gain configuration data and sound effect gain configuration data; Based on the volume gain configuration data, the original audio is subjected to volume gain to obtain intermediate audio; Based on the sound effect gain configuration data, sound effect gain is applied to the intermediate audio to obtain the user audio; The user audio is encoded to obtain the user audio data group.

[0012] In some embodiments, the step of applying sound effect gain to the intermediate audio based on the sound effect gain configuration data to obtain the user audio includes: The intermediate audio is frequency-domain transformed to obtain the audio frequency domain signal; Based on the sound effect gain configuration data, the audio frequency domain signal is subjected to signal gain to obtain an audio gain signal; The user audio is obtained by performing a time-domain transformation on the audio gain signal.

[0013] According to a third aspect of this application, a method for processing audio and video data is provided, applied to a terminal, the method comprising: Receive a signal group from the receiving end; the signal group includes a target audio stream signal and a target video stream signal; the target audio stream signal includes user audio recorded by each sending end and the original audio of the video resource; Audio is played to the outside world based on the target audio stream signal in the signal group, and video is played to the outside world based on the target video stream signal in the signal group; The signal group is obtained through the following steps: The receiving end receives user audio data groups from several of the sending ends and obtains video resources corresponding to all of the user audio data groups; The receiving end performs audio merging processing on all the user audio data groups to obtain a merged audio signal; The receiving end performs audio-video merging processing on the audio-video merged signal based on the video resource to obtain a signal group.

[0014] According to a fourth aspect of this application, an audio and video data processing system is provided, including a receiving end, a terminal, and a plurality of transmitting ends; The sending end is used to send the user audio data group to the receiving end. The receiving end is configured to acquire video resources corresponding to all the user audio data groups; perform audio merging processing on all the user audio data groups to obtain an audio merged signal; and perform audio-video merging processing on the audio merged signal according to the video resources to obtain a signal group, wherein the signal group includes a target audio stream signal and a target video stream signal; the target audio stream signal includes the original audio of the video resources and the audio merged signal. The terminal is configured to receive a group of signals from the receiving end, and to play audio to the outside world according to the target audio stream signal in the group of signals, and to play video to the outside world according to the target video stream signal in the group of signals.

[0015] According to a fifth aspect of this application, an electronic device is provided, comprising: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described above.

[0016] According to a sixth aspect of this application, a computer-readable storage medium is provided, wherein a processor-executable program is stored, which, when executed by the processor, is used to implement the method as described above.

[0017] According to a seventh aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium, wherein a processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the method described above.

[0018] The beneficial effects of the technical solutions provided in this application are: This application provides a method, system, device, and medium for processing audio and video data. The method receives user audio data groups from several transmitting ends and acquires video resources corresponding to all the user audio data groups. It then performs audio merging processing on all the user audio data groups to obtain an audio merged signal. Based on the video resources, it performs audio-video merging processing on the audio merged signal to obtain a signal group, which includes a target audio stream signal and a target video stream signal. The target audio stream signal includes the original audio of the video resource and the audio merged signal. The signal group is then sent to a terminal, enabling the terminal to play audio and video to the outside world based on the target audio stream signal and the target video stream signal in the signal group. This method, by receiving user audio data groups from several transmitting ends and performing audio merging processing, and then merging the obtained audio merged signal with video resources, can merge the audio input of multiple users into the video resources required for karaoke, thereby meeting the needs of multiple users singing karaoke simultaneously and effectively improving the user experience. Attached Figure Description

[0019] Figure 1 A flowchart illustrating the first method for processing audio and video data provided in this application embodiment; Figure 2 A detailed flowchart of step S120 provided for an embodiment of this application; Figure 3 A detailed flowchart of step S130 provided for an embodiment of this application; Figure 4 A detailed flowchart of step S320 provided for an embodiment of this application; Figure 5 A flowchart illustrating the second audio / video data processing method provided in this application embodiment; Figure 6 A detailed flowchart of step S520 provided for an embodiment of this application; Figure 7 A detailed flowchart of step S630 provided for an embodiment of this application; Figure 8 A flowchart illustrating the third audio / video data processing method provided in this application embodiment. Figure 9 A schematic diagram of the framework of an audio and video data processing system provided in this application embodiment; Figure 10 This is a structural block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0020] The present application will be further described below with reference to the accompanying drawings and specific embodiments. The described embodiments should not be considered as limitations on the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.

[0021] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0023] The technical terms used in the embodiments of this application are explained below: Sampling rate: Also known as sampling speed or sampling frequency, it defines the number of samples extracted from a continuous signal and used to assemble a discrete signal per unit time. It is expressed in Hertz (Hz). The reciprocal of the sampling frequency is the sampling period, or sampling time, which is the time interval between samples.

[0024] Bit depth: Bit depth refers to the number of bits occupied by each sample point in digital audio. Bit depth determines the dynamic range of digital audio, that is, the distance between the maximum and minimum amplitude of the sound signal. The higher the bit depth, the greater the dynamic range of the digital audio, and the more accurately it can display the details and differences of the sound signal.

[0025] Sound Channel: A sound channel refers to the independent audio signals that are collected or played back from different spatial locations during recording or playback. The number of sound channels is the number of sound sources during recording or the corresponding number of speakers during playback.

[0026] UDP: UDP is a protocol that operates at the transport layer of the OSI (Open Systems Interconnection) model. It uses IP as the underlying protocol and provides applications with a minimal protocol mechanism to send messages to other programs.

[0027] The communication connection process between the sending end and the receiving end involved in the embodiments of this application is described below: In this embodiment of the application, taking the local area network IP of the receiving end as IP_RECEIVE and the name as NAME_RECEIVE; the local area network IP of the sending end as IP_SEND; and the sending end having an audio sending APP installed as an example, the specific communication connection process is as follows: 1. After the audio sending app starts, the sending end sends a UDP multicast command DISCOVER_STB in the local area network; 2. The receiving end receives the DISCOVER_STB command in the local area network and obtains the local area network IP address of the sending end, IP_SEND, according to the DISCOVER_STB command; 3. The receiving end sends a UDP unicast command to the sending end based on the obtained IP address IP_SEND. The format of the unicast command STB_RESPONSE is: [IP_RECEIVE]; [NAME_RECEIVE]; [sampling rate]; [sampling bit depth]; [number of channels].

[0028] 4. The sending end obtains the sending end's local area network IP address (IP_RECEIVE), name (NAME_RECEIVE), sampling rate, sampling bit depth, and number of channels based on the UDP unicast command from the receiving end; 5. The sending end will display the names of the received receivers in a list on the interface, and respond to the user's trigger operation (such as the user's click command) to connect with the receiver selected by the user. After the connection is completed, the karaoke interface will be displayed on the audio sending end APP of the sending end.

[0029] Currently, related technologies typically rely on local hardware, requiring users to purchase dedicated wired or wireless microphones and connect them to a TV box to enable karaoke applications. This method usually only supports a single microphone input, making it difficult to meet the needs of multiple users singing simultaneously, resulting in a less than satisfactory user experience. Furthermore, dedicated microphones are relatively expensive, and local hardware methods mostly use dedicated chips or circuits to process audio. This approach often only utilizes the capabilities built into the hardware device, leading to high user costs and poor scalability.

[0030] It should be noted that the aforementioned related technologies are only used to assist in understanding the technical solutions of this application and do not mean that they belong to the publicly disclosed prior art.

[0031] In view of this, embodiments of this application provide a method, system, device, and medium for processing audio and video data. The method receives user audio data groups from several transmitting ends and performs audio merging processing. Then, it performs audio merging processing with video resources, which can merge the audio inputs of multiple users into the video resources required for users to sing karaoke, thereby meeting the needs of multiple users to sing karaoke at the same time and effectively improving the user experience.

[0032] Furthermore, after establishing a communication connection between the sending end and the receiving end via a software app, this method processes the user's original audio based on the sending end app. Specifically, it processes the original audio through volume and sound effect gain configuration data. This software-based audio processing can replace hardware devices, which improves the scalability and flexibility of audio processing. Additionally, the communication connection between the sending and receiving ends eliminates the need for users to purchase additional dedicated hardware, effectively reducing user costs.

[0033] This application provides a method for processing audio and video data, which can be specifically described through the following embodiments. First, an audio and video data processing method in this application is described.

[0034] The audio and video data processing method provided in this application can be applied to home entertainment application scenarios. In home entertainment application scenarios (such as user karaoke scenarios), home entertainment service providers can use the method provided in this application to process the audio of several users in combination with video resources to meet the karaoke needs of several users and effectively improve the user experience.

[0035] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0036] Reference Figure 1 , Figure 1 This is a flowchart illustrating an audio / video data processing method provided in an embodiment of this application. The method is applied at a receiving end and includes, but is not limited to, steps S110 to S140: Step S110: Receive user audio data groups from several sending ends, and obtain video resources corresponding to all of the user audio data groups; In this embodiment, the receiving end can be an IPTV box, and the sending end can be a user's mobile device (such as a mobile phone). The receiving end can communicate with one or more sending ends. The video resource can be a video requested by the user for karaoke, which is typically an MP4 video resource.

[0037] Step S120: Perform audio merging processing on all the user audio data groups to obtain a merged audio signal; In this embodiment of the application, audio data groups provided by the transmitter for each user can be merged to obtain a merged audio signal.

[0038] Reference Figure 2 In some embodiments, step S120, performing audio merging processing on all the user audio data groups to obtain a merged audio signal, includes: Step S210: Obtain the array parameter rules of the user audio data group; Step S220: According to the array parameter rules, perform data parsing on each user audio data group to obtain the user audio for each user audio data group; Step S230: Merge all the user audio data to obtain the merged audio signal.

[0039] In this embodiment, the format of the user audio data group can be a byte array, and the array parameter rules can be the format rules of the user audio data group, which can be [sampling rate; sampling bit depth; number of channels; data length; audio data]. Step S220 can be based on these array parameter rules to parse each user audio data group to extract the audio data (i.e., user audio) of each user audio data group, specifically PCM (Pulse Code Modulation) audio data. Data merging can be performed on all PCM audio format user audio to obtain a PCM audio format merged audio signal.

[0040] For example, the combined audio signal can be represented as:

[0041] in, This is an audio merged signal; The total number of users; , , and The user audio data consists of the first user audio data group (specifically PCM audio data), the second user audio data group, the third user audio data group, and the m-th user audio data group, in that order. The total number of audio frames for the user's audio.

[0042] Step S130: Based on the video resource, perform audio-video merging processing on the audio merged signal to obtain a signal group, wherein the signal group includes a target audio stream signal and a target video stream signal; the target audio stream signal includes the original audio of the video resource and the audio merged signal; In the embodiments of this application, the audio merge signal can be merged into the video resource to obtain a signal group including the target audio stream signal and the target video stream signal.

[0043] Reference Figure 3 In some embodiments, step S130, performing audio-video merging processing on the audio-video merged signal based on the video resource to obtain a signal group, includes: Step S310: Extract audio and video from the video resources to obtain a first audio stream signal and an original video stream signal; In this embodiment of the application, audio and video extraction can be achieved by demultiplexing video resources using FFmpeg, thereby extracting the original audio data of the video resources into a first audio stream signal and extracting the original video data of the video resources into an original video stream signal.

[0044] Step S320: Based on the first audio stream signal, perform audio signal merging on the audio merging signal to obtain the target audio stream signal; Reference Figure 4 Further, step S320, which involves merging the audio merged signal based on the first audio stream signal to obtain the target audio stream signal, includes: Step S410: Decode the first audio stream signal to obtain the second audio stream signal; Step S420: Extract the channel audio from the second audio stream signal to obtain a third audio stream signal, wherein the third audio stream signal is used to characterize part or all of the channel audio of the second audio stream signal; Step S430: Based on the third audio stream signal, perform audio signal merging on the audio merging signal to obtain the target audio stream signal.

[0045] In the embodiments of this application, the first audio stream signal can be decoded into an audio stream signal in PCM audio format, which is referred to as the second audio stream signal. There are various specific audio decoding methods, which will not be described in detail here.

[0046] Understandably, in practical applications, video resources are typically MP4 resources with left-side accompaniment and right-side vocals. This means the second audio stream signal is usually stereo, with the left channel containing the instrumental track and the right channel containing both the instrumental track and the original vocals. Therefore, channel audio extraction can be based on audio filtering algorithms (such as the FFmpeg audio filtering algorithm) to extract the target channel audio from the second audio stream signal. The extracted target channel audios are then merged into a new mono PCM audio data, denoted as the third audio stream signal.

[0047] For example, in a first embodiment, the left channel audio of the second audio stream signal can be extracted and the extracted left channel audio can be determined as the third audio stream signal.

[0048] Alternatively, in the second embodiment, if the video resource is a multi-channel audio resource (i.e., the number of audio channels in the video resource is greater than or equal to 2), its second audio stream signal may include left and right channel audio, as well as surround channel audio, etc. In this case, the audio channel extraction can be to extract the left and right channel audio from the second audio stream signal and merge the extracted left and right channel audio. The specific merging content is similar to the content of the aforementioned step S230, thereby obtaining the third audio stream signal.

[0049] Audio signal merging can be based on the AudioFlinger service built into the transmitter (such as an IPTV TV box), which merges the third audio stream signal and the audio merge signal to obtain the target audio stream signal in PCM audio format.

[0050] Step S330: Perform hardware decoding on the original video stream signal to obtain the target video stream signal.

[0051] In this embodiment, the transmitting end can perform hardware decoding on the original video stream signal based on hardware decoding technology to obtain the decoded original video stream signal, which is denoted as the target video stream signal. Then, the target video stream signal and the target audio stream signal are combined into an HDMI signal by the HDMI driver controller in the transmitting end (such as an IPTV TV box), which is denoted as the signal group.

[0052] Step S140: Send the signal group to the terminal so that the terminal can play audio to the outside world according to the target audio stream signal in the signal group, and play video to the outside world according to the target video stream signal in the signal group.

[0053] In this embodiment, the transmitting end can send a signal group to the terminal, which may be a smart TV, a computer, or a projection TV, etc.

[0054] Reference Figure 5 This application also provides another method for processing audio and video data, applied at a sending end, wherein the sending end and the receiving end are communicatively connected, and the method includes: Step S510: Collect the user's original audio; Step S520: Perform audio processing on the original audio to obtain a user audio data group; Step S530: Send the user audio data groups to the receiving end so that the receiving end can acquire video resources corresponding to all the user audio data groups; perform audio merging processing on all the user audio data groups to obtain an audio merged signal; perform audio-video merging processing on the audio merged signal according to the video resources to obtain a signal group, the signal group including a target audio stream signal and a target video stream signal; the target audio stream signal includes the original audio of the video resources and the audio merged signal; send the signal group to the terminal so that the terminal can play audio to the outside world according to the target audio stream signal in the signal group, and play video to the outside world according to the target video stream signal in the signal group.

[0055] In this embodiment of the application, the sending end (such as a mobile phone) can collect the user's voice and convert it into audio stream data in PCM audio format, which is recorded as the original audio. Then, based on the configuration data set by the user at the sending end, the original audio is processed to obtain the user audio data group, and the user audio data group is sent to the receiving end.

[0056] Reference Figure 6 In some embodiments, step S520, which involves processing the original audio to obtain a user audio data set, includes: Step S610: Obtain volume gain configuration data and sound effect gain configuration data; Step S620: According to the volume gain configuration data, increase the volume of the original audio to obtain the intermediate audio; In this embodiment, the configuration data set by the user at the sending end may include volume gain configuration data and sound effect gain configuration data. The volume gain can be adjusted by adjusting the volume of the original audio according to the user-defined volume gain configuration data to obtain intermediate audio. For example, this intermediate audio can be represented as:

[0057] in, The volume of the i-th audio frame in the middle audio; The volume of the i-th audio frame of the original audio; Configure data for volume gain; This represents the maximum threshold for volume. This is the minimum threshold for volume. It is a minimum value function; It is a function for maximizing the value; This is a rounding function.

[0058] Step S630: According to the sound effect gain configuration data, apply sound effect gain to the intermediate audio to obtain the user audio; Reference Figure 7 Furthermore, step S630, which involves applying sound effect gain to the intermediate audio based on the sound effect gain configuration data to obtain the user audio, includes: Step S710: Perform frequency domain transformation on the intermediate audio to obtain the audio frequency domain signal; Step S720: According to the sound effect gain configuration data, perform signal gain on the audio frequency domain signal to obtain an audio gain signal; Step S730: Perform time-domain transformation on the audio gain signal to obtain the user audio.

[0059] In this embodiment, frequency domain transformation converts intermediate audio in the time domain into a frequency domain signal via Fourier transform (FFT), denoted as the audio frequency domain signal. Then, based on sound effect gain configuration data, the sound effect gain of the audio frequency domain signal is adjusted to obtain an audio gain signal. Exemplarily, this audio gain signal can be represented as:

[0060] in, This is the audio gain signal; Frequency index; The audio frequency domain signal is in complex form and includes the amplitude and phase of the k-th frequency. Configure data for sound effect gain.

[0061] It is understandable that time-domain transformation can convert the audio gain signal in the frequency domain into a time-domain signal to obtain the user audio. Specifically, the audio gain signal in the frequency domain can be converted into the user audio through inverse Fourier transform (IFFT).

[0062] Step S640: Encode the user audio to obtain the user audio data group.

[0063] In this application, the audio encoding can be based on the aforementioned array parameter rules, and the user audio data can be assembled into a byte array through encoding to obtain the user audio data group.

[0064] Reference Figure 8 This application also provides another method for processing audio and video data, applied to a terminal, the method comprising: Step S810: Receive a signal group from the receiving end; the signal group includes a target audio stream signal and a target video stream signal; the target audio stream signal includes user audio recorded by each sending end and the original audio of the video resource; Step S820: Play audio to the outside world according to the target audio stream signal in the signal group, and play video to the outside world according to the target video stream signal in the signal group; The signal group is obtained through the following steps: The receiving end receives user audio data groups from several of the sending ends and obtains video resources corresponding to all of the user audio data groups; The receiving end performs audio merging processing on all the user audio data groups to obtain a merged audio signal; The receiving end performs audio-video merging processing on the audio-video merged signal based on the video resource to obtain a signal group.

[0065] In this embodiment of the application, the terminal can play the screen of a karaoke video to the outside world based on the target audio stream signal in the signal group, and play audio to the outside world based on the target audio stream signal in the signal group. The audio may include the accompaniment of the karaoke video and the voices of several users singing.

[0066] Figure 9 A system block diagram of an audio and video data processing system provided in this application embodiment includes a receiving end, a terminal, and several transmitting ends; The transmitting end 901 is used to send the user audio data group to the receiving end. The receiving end 902 is configured to acquire video resources corresponding to all the user audio data groups; perform audio merging processing on all the user audio data groups to obtain an audio merged signal; and perform audio-video merging processing on the audio merged signal according to the video resources to obtain a signal group, wherein the signal group includes a target audio stream signal and a target video stream signal; the target audio stream signal includes the original audio of the video resources and the audio merged signal. The terminal 903 is used to receive a signal group from the receiving end, and to play audio to the outside world according to the target audio stream signal in the signal group, and to play video to the outside world according to the target video stream signal in the signal group.

[0067] It is worth mentioning that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0068] Figure 10 A schematic diagram of the structure of a computer device provided in this application embodiment includes: At least one processor 980; At least one memory 920 is used to store at least one program; When the at least one program is executed by the at least one processor 980, the at least one processor 980 performs the method as described in the foregoing embodiments.

[0069] This application also provides a computer-readable storage medium storing a processor-executable program, which, when executed by the processor 980, is used to implement the methods described in the foregoing embodiments.

[0070] Specifically, computer equipment can be either a user terminal or a server.

[0071] This application uses a computer device as a user terminal as an example, as detailed below: like Figure 10 As shown, the computer device 900 may include an RF (Radio Frequency) circuit 910, a memory 920 including one or more computer-readable storage media, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a WiFi module 970, a processor 980 including one or more processing cores, and a power supply 990, among other components. Those skilled in the art will understand that... Figure 10The device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. The RF circuit 910 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 980 for processing; additionally, it transmits uplink data to the base station. Typically, the RF circuit 910 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier), a duplexer, etc. Furthermore, the RF circuit 910 can also communicate wirelessly with networks and other devices. Wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System for Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc. The memory 920 can be used to store software programs and modules. The processor 980 executes various functional applications and data processing by running the software programs and modules stored in the memory 920. The memory 920 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device 900 (such as audio data, telephone directory, etc.). In addition, the memory 920 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 920 may also include a memory controller to provide access to the memory 920 by the processor 980 and the input unit 930. Although Figure 10 The RF circuit 910 is shown, but it is understood that it is not a necessary component of the computer device 900 and can be omitted as needed without changing the nature of the invention.

[0072] The input unit 930 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 930 may include a touch-sensitive surface 932 and other input devices 931. The touch-sensitive surface 932, also known as a touch display screen or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 932), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 932 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 980, and can receive and execute commands from the processor 980. In addition, the touch-sensitive surface 932 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 932, the input unit 930 may also include other input devices 931. Specifically, other input devices 931 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc. Display unit 940 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of computer device 900. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 940 may include display panel 941, optionally configured as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc. Further, touch-sensitive surface 932 may cover display panel 941. When touch-sensitive surface 932 detects a touch operation on or near it, it transmits the information to processor 980 to determine the type of touch event. Subsequently, processor 980 provides corresponding visual output on display panel 941 according to the type of touch event. Although in Figure 10 In this embodiment, the touch-sensitive surface 932 and the display panel 941 are implemented as two separate components to realize input and output functions. However, in some embodiments, the touch-sensitive surface 932 and the display panel 941 can be integrated to realize input and output functions.

[0073] The computer device 900 may also include at least one sensor 950, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 941 according to the ambient light level, and the proximity sensor can turn off the display panel 941 and / or backlight when the computer device 900 is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometers, taps), etc. Other sensors that the computer device 900 may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0074] Audio circuitry 960, speaker 961, and microphone 962 provide an audio interface between the user and computer device 900. Audio circuitry 960 converts received audio data into electrical signals, which are then transmitted to speaker 961, where they are converted into sound signals for output. Conversely, microphone 962 converts collected sound signals into electrical signals, which are received by audio circuitry 960, converted back into audio data, and then processed by processor 980 before being transmitted via RF circuitry 910 to another control device, or output to memory 920 for further processing. Audio circuitry 960 may also include an earphone jack to facilitate communication between peripheral headphones and computer device 900.

[0075] Computer device 900 can transmit information with the wireless transmission module set up on the battle equipment via WiFi module 970.

[0076] The processor 980 is the control center of the computer device 900. It connects various parts of the control device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 920, and by calling data stored in the memory 920, it performs various functions of the computer device 900 and processes data, thereby providing overall monitoring of the control device. Optionally, the processor 980 may include one or more processing cores; optionally, the processor 980 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 980.

[0077] The computer device 900 also includes a power supply 990 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 980 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 990 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. Although not shown, the computer device 900 may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0078] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the methods described in the foregoing embodiments.

[0079] This application also discloses a computer program product or computer program, which includes computer instructions stored in the aforementioned computer-readable storage medium; the processor of the aforementioned electronic device can read the computer instructions from the aforementioned computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to perform the aforementioned method embodiment.

[0080] It is understood that the content of the above method embodiments is applicable to this computer program product or computer program embodiment. The specific functions implemented by this computer program product or computer program embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0081] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0082] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0083] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0084] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0085] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0086] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0087] The step numbers in the above method embodiments are set only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0088] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for processing audiovisual data, characterized in that, Applied to a receiving end, the method comprises: Receiving user audio data groups from a plurality of sending ends, and obtaining video resources corresponding to all the user audio data groups; Performing audio merging processing on all the user audio data groups to obtain an audio merging signal; Performing audio-video merging processing on the audio merging signal according to the video resources to obtain a signal group, the signal group comprising a target audio stream signal and a target video stream signal; the target audio stream signal comprising original audio of the video resources and the audio merging signal; Sending the signal group to a terminal, so that the terminal plays audio to the outside world according to the target audio stream signal in the signal group, and plays video to the outside world according to the target video stream signal in the signal group.

2. The method of claim 1, wherein, The audio merging processing on all the user audio data groups to obtain an audio merging signal comprises: Obtaining array parameter rules of the user audio data groups; Performing data analysis on each of the user audio data groups according to the array parameter rules to obtain user audio of each of the user audio data groups; Performing data merging on all the user audio to obtain the audio merging signal.

3. The method of claim 1, wherein, The audio-video merging processing on the audio merging signal according to the video resources to obtain a signal group comprises: Performing audio-video extraction on the video resources to obtain a first audio stream signal and an original video stream signal; Performing audio signal merging on the audio merging signal according to the first audio stream signal to obtain the target audio stream signal; Performing video hard decoding on the original video stream signal to obtain the target video stream signal.

4. The method of claim 3, wherein, The audio signal merging on the audio merging signal according to the first audio stream signal to obtain the target audio stream signal comprises: Performing audio decoding on the first audio stream signal to obtain a second audio stream signal; Performing channel audio extraction on the second audio stream signal to obtain a third audio stream signal, the third audio stream signal being used to represent part or all of channel audio of the second audio stream signal; Performing audio signal merging on the audio merging signal according to the third audio stream signal to obtain the target audio stream signal.

5. A method for processing audiovisual data, characterized in that, Applied to a sending end, the sending end being in communication connection with a receiving end, the method comprises: Collecting original audio of a user; Performing audio processing on the original audio to obtain a user audio data group; Sending the user audio data group to the receiving end, so that the receiving end obtains video resources corresponding to all the user audio data groups; performs audio merging processing on all the user audio data groups to obtain an audio merging signal; performs audio-video merging processing on the audio merging signal according to the video resources to obtain a signal group, the signal group comprising a target audio stream signal and a target video stream signal; the target audio stream signal comprising original audio of the video resources and the audio merging signal; and sends the signal group to a terminal, so that the terminal plays audio to the outside world according to the target audio stream signal in the signal group, and plays video to the outside world according to the target video stream signal in the signal group.

6. The method of claim 5, wherein, The audio processing on the original audio obtains a user audio data set, and the audio processing on the original audio comprises: obtaining volume gain configuration data and sound effect gain configuration data; applying volume gain to the original audio according to the volume gain configuration data to obtain intermediate audio; applying sound effect gain to the intermediate audio according to the sound effect gain configuration data to obtain the user audio; applying audio encoding to the user audio to obtain the user audio data set.

7. The method of claim 6, wherein, The audio processing on the original audio obtains a user audio data set, and the audio processing on the original audio comprises: applying frequency domain transformation to the intermediate audio to obtain an audio frequency domain signal; applying signal gain to the audio frequency domain signal according to the sound effect gain configuration data to obtain an audio gain signal; applying time domain transformation to the audio gain signal to obtain the user audio.

8. A method for processing audiovisual data, characterized in that, The method applied to a terminal comprises: receiving a signal set from a receiving end; the signal set comprises target audio stream signals and target video stream signals; the target audio stream signals comprise user audio recorded by each sending end and original audio of the video resource; playing audio to the outside world according to the target audio stream signals in the signal set and playing video to the outside world according to the target video stream signals in the signal set; wherein the signal set is obtained through the following steps: the receiving end receives user audio data sets from a plurality of sending ends and obtains video resources corresponding to all the user audio data sets; the receiving end performs audio merging processing on all the user audio data sets to obtain an audio merging signal; the receiving end performs audio-video merging processing on the audio merging signal according to the video resources to obtain a signal set.

9. A system for processing audiovisual data, characterized in that comprising a receiving end, a terminal and a plurality of sending ends; the sending end is configured to send the user audio data set to the receiving end the receiving end is configured to obtain video resources corresponding to all the user audio data sets; performing audio merging processing on all the user audio data sets to obtain an audio merging signal; performing audio-video merging processing on the audio merging signal according to the video resources to obtain a signal set, the signal set comprising target audio stream signals and target video stream signals; the target audio stream signals comprise original audio of the video resource and the audio merging signal; the terminal is configured to receive the signal set from the receiving end, play audio to the outside world according to the target audio stream signals in the signal set and play video to the outside world according to the target video stream signals in the signal set.

10. An electronic device, comprising: comprise: at least one processor; at least one memory configured to store at least one program; when the at least one program is executed by the at least one processor, the at least one processor implements the method of any one of claims 1-4, the method of any one of claims 5-7 or the method of claim 8.