Audio processing method and apparatus, device, and storage medium
Patent Information
- Application Number
- PCT/CN2026/078932
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-19
- Filing Date
- 2026-02-12
- Publication Date
- 2026-08-27
Smart Images

Figure CN2026078932_27082026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices and storage media for audio processing
[0001] This application claims priority to Chinese Patent Application No. 202510188455.9, filed on February 19, 2025, entitled "Method, Apparatus, Device and Storage Medium for Audio Processing", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices and computer-readable storage media for audio processing. Background Technology
[0003] With the development of computer technology, more and more applications are able to provide live streaming functionality. During a live stream, participants can engage in the event and output audio content; for example, participants can sing together during the live stream. How to process the audio content to improve the user's interactive experience is a key concern. Summary of the Invention
[0004] In a first aspect of this disclosure, an audio processing method is provided. The method includes: acquiring audio content associated with a live event; processing the audio content using a first model to determine a first song recognition result; processing text content using a second model to determine a second song recognition result, the text content being determined by recognizing the audio content; and triggering the presentation of the song's lyrics in a live streaming interface associated with the live event in response to the first song recognition result and the second song recognition result indicating the same song.
[0005] In a second aspect of this disclosure, an apparatus for audio processing is provided. The apparatus includes: a first acquisition module configured to acquire audio content associated with a live event; a first processing module configured to process the audio content using a first model to determine a first song recognition result; a second processing module configured to process text content using a second model to determine a second song recognition result, the text content being determined by recognizing the audio content; and a first triggering module configured to trigger the presentation of the song's lyrics in a live streaming interface associated with the live event in response to the first song recognition result and the second song recognition result indicating the same song.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] Figure 2 shows a flowchart of an audio processing procedure according to some embodiments of the present disclosure;
[0013] Figure 3 shows a schematic diagram of a live streaming interface according to some embodiments of the present disclosure;
[0014] Figure 4 shows an example flowchart of audio processing according to some embodiments of the present disclosure;
[0015] Figure 5 shows a schematic structural block diagram of an apparatus for audio processing according to certain embodiments of the present disclosure;
[0016] Figure 6 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0022] Traditionally, when participants in a live streaming event sing during the live stream, they can use a third-party player and sound card to ensure the singing effect. At this time, the users watching the live stream can only hear the sound of the song being sung and cannot see the lyrics of the song being sung from the live stream interface.
[0023] Alternatively, participants can use a personal computer as the main device and use pre-defined software to project the lyrics from the third player onto the live stream to achieve the effect of displaying the lyrics. However, this method cannot guarantee the clarity of the lyrics displayed on the live stream interface, and the process of displaying the lyrics is relatively cumbersome and requires a high level of skill.
[0024] Embodiments of this disclosure propose an audio processing scheme. The scheme includes: acquiring audio content associated with a live event; processing the audio content using a first model to determine a first song recognition result; processing text content using a second model to determine a second song recognition result, wherein the text content is determined by recognizing the audio content; and triggering the presentation of the song's lyrics in a live streaming interface associated with the live event in response to the first song recognition result and the second song recognition result indicating the same song.
[0025] Based on this approach, the embodiments of this disclosure can perform song recognition using both audio and text modalities, reducing the misjudgment problems caused by song recognition based on a single modal and improving the reliability of the song recognition results. Furthermore, when the song recognition results based on the two modalities are consistent, the corresponding lyrics can be displayed on the live streaming interface. This provides users with real-time lyrics while avoiding the problems of low lyrics clarity and high operational requirements associated with displaying lyrics via screen projection.
[0026] Example Environment
[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include a terminal device 110.
[0028] In this example environment 100, terminal device 110 may run an application 120 for providing and watching live streams. Application 120 may be any suitable type of application for providing and watching live streams, examples of which may include, but are not limited to, online video applications and live streaming applications. User 140 may interact with application 120 via terminal device 110 and / or its attached devices. User 140 may be any suitable user, such as a participant in the live stream event, or a user who is not participating in the live stream event but only watching the live stream (e.g., a viewer, listener, or spectator).
[0029] In environment 100 of Figure 1, if application 120 is active, terminal device 110 can present live interface 150 through application 120.
[0030] Taking terminal device 110 as an example, the server 130 can obtain the audio content sent by terminal device 110. The audio content can be audio captured by an audio acquisition device deployed on terminal device 110, such as audio captured by a microphone. Furthermore, the server 130 can determine the song being sung by the target participant in the live stream based on the audio content. Further, the server 130 can display the lyrics of the song sung by the target participant on terminal device 110 and other electronic devices (not shown in the figure) corresponding to other participants or viewers.
[0031] In some embodiments, both the terminal device 110 and the electronic device can communicate with the server 130 to provide services to the application 120.
[0032] Terminal device 110 or electronic device can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of user-facing interface (such as "wearable" circuitry).
[0033] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 supporting virtual scenarios in terminal device 110.
[0034] A communication connection can be established between server 130 and terminal device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and terminal device 110 can achieve signaling interaction through the communication connection between them.
[0035] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0036] The various example implementations of this disclosure will be described in detail below.
[0037] Example process
[0038] Figure 2 shows a flowchart of an audio processing procedure 200 according to some embodiments of the present disclosure. Procedure 200 may be implemented at server 130. Procedure 200 is described below with reference to Figure 1.
[0039] In box 210, server 130 retrieves audio content associated with the live event.
[0040] In some embodiments, a live streaming event may include interactive events initiated by at least one participant joining a predetermined live streaming room during the live stream. These events could include a live stream by the host, a multi-person chat event within the live streaming room, or a competition event within the live streaming room. Competition events can be of any suitable type, such as a singing competition. Such at least one participant may include one host user, or multiple host users. Alternatively, multiple participants may include a single host user and multiple guest users, etc.
[0041] In some embodiments, the audio content may be audio data associated with the at least one participant. In some embodiments, the audio content may be any suitable type of content, such as singing, chatting, noise from the participant's surroundings, etc. It should be noted that the audio content may also be referred to as the first audio content in subsequent embodiments.
[0042] As an example, the audio content could be the audio of user A singing during a live stream while participating in a live chat event.
[0043] In box 220, server 130 uses the first model to process the audio content in order to determine the recognition result of the first song.
[0044] In some embodiments, the first model can be any suitable machine learning model, such as an audio recognition model. The first model can be deployed on server 130, or on other devices besides server 130, which will not be elaborated here.
[0045] In some embodiments, server 130 may provide audio content to a first model, enabling the first model to process the audio content and output a first song recognition result. In some embodiments, the first song recognition result may indicate whether the audio content is associated with a certain song. In response to the audio content being associated with a certain song, the first song recognition result may also indicate a first identifier of the song associated with the audio content, wherein the first identifier may be a name or any other suitable information that can characterize the song.
[0046] Since audio content can only be associated with a song if it includes singing content, and will not correspond to a song if it does not (e.g., chat content or ambient noise), server 130 can first determine whether the audio content includes singing content to determine if further song identification information is needed. Specifically, server 130 can use an audio energy detection model to process the audio content to determine whether it includes singing content. The audio energy detection model can be any suitable machine learning model, which mainly judges whether the audio content contains singing content based on the audio energy corresponding to the audio content. Further, in response to determining that the audio content includes singing content, server 130 can trigger the first model to process the audio content and trigger the second model to process the text content. The process of the second model processing the text content is described in detail in box 230 and will not be repeated here. Further, in response to determining that the audio content does not include singing content, server 130 can not trigger the first model to process the audio content and trigger the second model to process the text content, i.e., not perform subsequent audio processing operations to reduce additional audio processing steps.
[0047] To ensure the accuracy of song recognition results, server 130 can perform the operation of determining the first song recognition result based on the audio content only after collecting audio content of a certain duration. Specifically, server 130 can provide audio content to the first model in response to the audio content reaching a preset duration to determine the first song recognition result. The preset duration can be any appropriate length, such as 10 seconds, etc.
[0048] In some embodiments, the first model may further be configured to output at least one of the following: lyrics of a candidate song corresponding to the first song identification result; and progress information of the audio content in the candidate song. The candidate song is a song predicted by the first model and associated with the audio content. The progress information may indicate the start and end times of the lyrics segment corresponding to the audio content, etc. For each song, each lyrics segment or each word included in each lyrics segment corresponds to a predetermined time in the song. The start time of a lyrics segment corresponds to the predetermined performance time of the first word in the lyrics segment in the song, and the end time of a lyrics segment corresponds to the predetermined performance time of the last word in the lyrics segment in the song.
[0049] As an example, the first model can output audio content corresponding to lyric fragment A, with a singing time of 10 to 15 seconds. Specifically, the first word in lyric fragment A is scheduled to be sung starting at the 10-second mark of the song, and the last word in lyric fragment A is scheduled to be sung starting at the 15-second mark of the song.
[0050] To improve the efficiency of determining the first song recognition result, the progress information of the lyrics and / or audio content of the candidate songs corresponding to the first song recognition result can be output from the first model along with the first song recognition result. Specifically, the server 130 can use the first model to process the audio content to obtain the target recognition result output by the first model. In addition to indicating the first song recognition result, the target recognition result can also indicate the progress information of the lyrics and / or audio content of the candidate songs corresponding to the first song recognition result.
[0051] In box 230, server 130 uses a second model to process text content to determine the second song recognition result, the text content being determined by recognizing audio content.
[0052] To improve the accuracy of song recognition results, this disclosure can determine the second song recognition result based on text content, in addition to determining the first song recognition result based on audio content, so as to further ensure that the song recognition result can be double-verified based on multimodal content.
[0053] In some embodiments, server 130 may utilize a speech recognition model to process audio content to determine the speech recognition result. The speech recognition model can be any suitable machine learning model that can convert speech into text. As an example, server 130 may utilize a language recognition model based on Automatic Speech Recognition (ASR) technology to convert speech into text.
[0054] In some embodiments, the speech recognition result may include text converted from the singing content, as well as text converted from dialogue content, text converted from environmental noise, etc. To improve the accuracy of the song recognition result, in some embodiments, the server 130 can determine the text content corresponding to a preset grammatical structure from the speech recognition result to remove interference from other forms of content besides the singing content on the second song recognition result. The preset grammatical structure indicates any appropriate text format or pattern, which conforms to the format corresponding to the lyrics.
[0055] Furthermore, server 130 can provide the text content to the second model to determine the second song recognition result. In some embodiments, the second song recognition result can indicate whether the text content is associated with a certain song. In response to the text content being associated with a certain song, the second song recognition result can also indicate a second identifier of the song associated with the text content, where the second identifier can be a name or other information that can characterize the song. The second model can be any suitable machine learning model, such as a language model. The second model can be deployed on server 130, or it can be deployed on other devices besides server 130, which will not be elaborated here.
[0056] It should be noted that the execution order of boxes 230 and 220 can be set according to requirements, and this disclosure does not limit this. To improve the efficiency of audio processing, as an example, server 130 can execute the operations corresponding to boxes 230 and 220 in parallel.
[0057] In box 240, server 130 responds to the first song identification result and the second song identification result indicating the same song, triggering the presentation of the song's lyrics in the live broadcast interface associated with the live broadcast event.
[0058] In some embodiments, server 130 may determine whether the first song recognition result and the second song recognition result indicate the same song based on whether the first identifier of the song indicated by the first lyrics recognition result and the second identifier of the song indicated by the second lyrics recognition result are consistent. As an example, server 130 may determine that the first song recognition result and the second song recognition result indicate the same song in response to the first identifier and the second identifier being consistent. As another example, server 130 may determine that the first song recognition result and the second song recognition result indicate different songs in response to the first identifier and the second identifier being inconsistent.
[0059] In some embodiments, server 130 may trigger the presentation of the lyrics of a song in a live streaming interface associated with a live streaming event in response to the first song identification result and the second song identification result indicating the same song. The live streaming interface can be an interface presented to the terminals of all users who can join the live streaming room. These users may include multiple participants in the live streaming event, as well as viewers who are not participating in the live streaming event. The lyrics can be presented at any location on the live streaming interface, such as in the target area corresponding to the participant associated with the audio content.
[0060] Referring to Figure 3 as an example, terminal device 110 can present interface 300. Terminal device 110 can present multiple content areas in interface 300, these multiple content areas corresponding to multiple participants in the target activity of the live broadcast event, with one participant corresponding to one content area. Referring to Figure 3 as an example, the live broadcast interface 300 may include content area 301 corresponding to user A, content area 302 corresponding to user B, content area 303 corresponding to user C, and content area 304 corresponding to user D, etc. In some embodiments, terminal device 110 can present the lyrics 310 corresponding to song 1 in content area 301 corresponding to user A, where user A can be a participant singing in the live broadcast room.
[0061] In some embodiments, the terminal device 110 may display interface content associated with the lyrics on the live streaming interface to indicate the singing progress corresponding to the song. The singing progress does not represent the actual progress of the participant singing the song, but rather the expected progress of the participant singing the song. For ease of description, in subsequent embodiments, "sung to completion" indicates that it is expected to be sung to completion, and "not sung to completion" indicates that it is not expected to be sung to completion. The interface content may be a progress bar corresponding to the song, the currently sung lyric fragment, etc., to indicate which line of lyrics the song has currently been sung. As an example, the terminal device 110 may display the third line of lyrics corresponding to the song on the live streaming interface to indicate that the song has currently reached the third line of lyrics.
[0062] Furthermore, the terminal device 110 can also use predetermined dynamic effects to indicate which parts of the lyrics in the currently sung segment have been sung and which parts have not yet been sung, in order to further indicate the singing progress corresponding to the lyrics. The predetermined dynamic effects can be any appropriate effects that can represent the difference between sung and unsung words. For example, the sung words in the currently sung lyrics can be presented in a first style, and the unsung words in the currently sung lyrics can be presented in a second style.
[0063] Using Figure 3 as an example, server 130 can present the first four words of lyrics 310 in a first form in content area 301 to indicate that the four words have been sung. Server 130 can also present the last seven words of lyrics 310 in a second form in content area 301 to indicate that the seven words have not been sung.
[0064] In some embodiments, to ensure that viewers can follow the song's progress in real time, the live streaming interface is also configured to periodically update its content to indicate the song's progress. As an example, the terminal device 110 can display each line of lyrics in the live streaming interface at a predetermined initial and end time, thus indicating the song's progress.
[0065] Furthermore, the terminal device 110 can set the presentation format of each word in the lyrics based on the predetermined initial time and end time of each word in the lyrics, so as to indicate which words in the lyrics have been sung and which words have not yet been sung.
[0066] The process of determining the concert schedule is explained below.
[0067] In some embodiments, server 130 can obtain progress information generated by the first model based on audio content, the progress information indicating the start and end times of a song segment corresponding to the audio content. Further, server 130 can determine the singing progress of the song based on the progress information. For example, if the start time of a song segment is 10 seconds and the end time is 15 seconds, then the singing progress of the song can indicate that the current song has reached the lyrics segment scheduled to be sung between 10 and 15 seconds, and the lyrics corresponding to 15 seconds have already been sung.
[0068] Since obtaining the song recognition result takes a certain amount of time, to avoid the problem of lyrics or progress bars being out of sync with the audio, in some embodiments, the server 130 can determine a first duration for the first model to generate the first song recognition result and a second duration for the second model to generate the second song recognition result. It should be noted that, to improve the efficiency of determining the song recognition result, in some embodiments, the process of the first model generating the first song recognition result and the process of the second model generating the second song recognition result can be performed in parallel.
[0069] Furthermore, server 130 can determine the processing time based on the first duration and the second duration. As an example, server 130 can determine the longer of the first duration and the second duration as the processing time.
[0070] Furthermore, server 130 can determine the singing progress of the song based on the progress information and processing time. As an example, server 130 can determine the sum of the end time indicated in the progress information and the processing time. Further, server 130 can determine the singing progress of the song based on this sum. For example, if the start time of a song segment is 10 seconds, the end time of the song segment is 15 seconds, and the processing time is 3 seconds, then the singing progress of the song can indicate that the current song has reached the lyrics segment scheduled to be sung between 10 seconds and 18 seconds, and the lyrics corresponding to 18 seconds have already been sung.
[0071] In some embodiments, after triggering the presentation of song lyrics in a live streaming interface associated with a live event, server 130 may acquire second audio content associated with the live streaming event. The second audio content may be audio data acquired from the live stream and associated with at least one of the plurality of participants. In some embodiments, the second audio content may be any suitable type of content, such as singing, chat, or ambient noise from the participant's surroundings.
[0072] As an example, the second audio content could be the audio of a conversation between user A and user B during the live stream, where user A is participating in the live stream event.
[0073] Furthermore, server 130 can use an audio energy detection model to detect whether the second audio content includes singing content. The audio energy detection model can be any suitable machine learning model, which is mainly used to analyze the energy of the audio signal and determine whether the audio content contains singing content.
[0074] Furthermore, server 130 can trigger the live streaming interface to stop displaying lyrics if no singing content is detected from the second audio content within a preset time period. This indicates that the participant has stopped singing, so terminal device 110 can not display lyrics on the live streaming interface. The preset time period can be set according to needs, such as 6 seconds, 4 seconds, etc.
[0075] Figure 4 shows an example flowchart of audio processing according to some embodiments of the present disclosure, and will now be described with reference to Figure 4.
[0076] In box 401, server 130 retrieves the audio content from the live stream.
[0077] As an example, audio content can be added to the audio output of at least one of the multiple participants in the live stream. As an example, the audio content can be singing, chat, ambient noise from the participant's surroundings, etc. As an example, the audio content can be collected by the voice capture device of the terminal device used by the participant outputting audio in the live stream and sent to server 130.
[0078] In box 402, server 130 performs audio event detection.
[0079] In some embodiments, audio event detection is used to detect whether a singing event has occurred in the live stream. As an example, server 130 can use an audio energy detection model to process audio content to determine whether the audio content includes singing content (i.e., to determine whether a singing event has occurred in the live stream). As an example, the audio energy detection model can determine whether the audio content contains singing content based on the audio energy corresponding to the audio content.
[0080] In box 403, server 130 periodically determines whether the audio content includes singing content based on the event detection results.
[0081] As an example, server 130 can periodically check whether the audio content includes singing content, with a cycle length of two seconds.
[0082] Furthermore, server 130 may perform operations in boxes 404 and 407 in response to determining that the audio content includes vocal content. Server 130 may also perform operation in box 414 in response to determining that the audio content does not include vocal content.
[0083] In box 404, server 130 cached 10s of audio.
[0084] As an example, server 130 may wait for a predetermined time to cache audio content with a duration of not less than 10 seconds in response to the audio content obtained based on box 401 having a duration of less than 10 seconds.
[0085] In box 405, server 130 provides audio content to the first model for song recognition.
[0086] As an example, the first model can be any suitable machine learning model, such as an audio recognition model. The first model can be deployed on server 130, or on other devices besides server 130, which will not be elaborated here.
[0087] In box 406, server 130 obtains the first song information output by the first model, which includes the song title, lyrics and progress information.
[0088] In some embodiments, the first song information may include the song title associated with the audio content, the lyrics corresponding to the song region, and the progress information corresponding to the lyric segment of the audio content. The progress information may indicate the start and end times of the lyric segment of the audio content, etc.
[0089] As an example, the first model can output audio content corresponding to lyric fragment A, with the corresponding singing time of 10 to 15 seconds. Specifically, the first word in lyric fragment A is scheduled to be sung at the 10th second of the song, and the last word in lyric fragment A is scheduled to be sung at the 15th second of the song.
[0090] In box 407, server 130 identifies the text content corresponding to the audio content based on automatic speech recognition technology.
[0091] As an example, the server can use automatic speech recognition technology to identify the speech recognition results corresponding to audio content. These results may include not only the text converted from the singing content, but also the text converted from dialogue, environmental noise, and so on. As another example, server 130 can determine the text content corresponding to the singing content from the speech recognition results to remove interference from other forms of content on the accuracy of song recognition.
[0092] In box 408, server 130 provides the text content to the second model for song recognition.
[0093] As an example, the second model can be any suitable machine learning model, such as a language model. The second model can be deployed on server 130, or on other devices besides server 130, which will not be elaborated upon here.
[0094] In box 409, server 130 obtains the second song information output by the second model, where the second song information includes the song title.
[0095] As an example, the second song information could include the name of the song associated with the text content. Of course, besides the name, it could also be other forms of information that can identify the song.
[0096] In box 410, server 130 determines whether the song titles match.
[0097] As an example, server 130 can determine whether the name of the song included in the first song information is the same as the name of the song included in the second song information.
[0098] Furthermore, server 130 may execute the operation in box 411 in response to determining that the song titles match. Server 130 may execute the operation in box 402 in response to determining that the song titles do not match.
[0099] In box 411, server 130 determines the actual progress information of the song by detecting the process time.
[0100] As an example, server 130 can determine a first duration for the first model to generate the first song recognition result and a second duration for the second model to generate the second song recognition result. Further, server 130 can determine the longer of the first and second durations as the processing time. Further, server 130 can determine the actual progress information corresponding to the song based on progress information and processing time. For example, server 130 can determine the sum of the end time indicated in the progress information and the processing time. Further, server 130 can determine the words in the song that are scheduled to be sung at the time corresponding to the sum as the words currently sung by the participant. For example, if the sum is 18 seconds, server 130 can determine the words scheduled to be sung at 18 seconds as the words currently sung by the participant in the live stream.
[0101] In box 412, terminal device 110 can start a lyrics rendering timer to render the lyrics.
[0102] As an example, server 130 can trigger terminal device 110 to display the lyrics of the song on the live streaming interface. As an example, terminal device 110 can display lyrics fragments of the song that have not yet been sung on the live streaming interface, wherein each line of lyrics in the lyrics fragment can be rendered separately based on the timer's corresponding time interval, and after rendering each line of lyrics, the operation of box 413 is executed.
[0103] In frame 413, terminal device 110 displays the lyrics on the live streaming interface.
[0104] As an example, terminal device 110 can present each line of lyrics in the unperformed lyrics segment according to the corresponding presentation time of the lyrics. For each line of lyrics presented in the live broadcast interface, the display format of each word in the lyrics can be determined based on the predetermined presentation time of that word, wherein the display style of words that are predetermined to be presented is different from the display style of words that are predetermined to not be presented.
[0105] In box 414, server 130 determines whether lyrics are displayed in the current live stream interface.
[0106] As an example, server 130 may execute the operation of box 402 in response to determining that lyrics are not displayed in the current live stream interface. Server 130 may execute the operation of box 415 in response to determining that lyrics are displayed in the current live stream interface.
[0107] In box 415, server 130 triggers terminal device 110 to hide lyrics in the live broadcast interface.
[0108] As an example, server 130 can send a lyrics hiding instruction to terminal device 110, so that terminal device 110 stops displaying lyrics on the live broadcast interface after receiving the lyrics hiding instruction.
[0109] Based on this approach, the embodiments of this disclosure can perform song recognition using both audio and text modalities, reducing the misjudgment problems caused by song recognition based on a single modal and improving the reliability of the song recognition results. Furthermore, when the song recognition results based on the two modalities are consistent, the corresponding lyrics can be displayed on the live streaming interface. This provides users with real-time lyrics while avoiding the problems of low lyrics clarity and high operational requirements associated with displaying lyrics via screen projection.
[0110] Example devices and equipment
[0111] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows a schematic structural block diagram of an apparatus 500 for audio processing according to certain embodiments of this disclosure. The apparatus 500 may be implemented as or included in the server 130 discussed above. The various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0112] As shown in Figure 5, the device 500 includes a first acquisition module 510 configured to acquire audio content associated with a live event; a first processing module 520 configured to process the audio content using a first model to determine a first song recognition result; a second processing module 530 configured to process text content using a second model to determine a second song recognition result, wherein the text content is determined by recognizing the audio content; and a first trigger module 540 configured to trigger the presentation of the song lyrics in the live interface associated with the live event in response to the first song recognition result and the second song recognition result indicating the same song.
[0113] In some embodiments, the first processing module 520 is further configured to provide audio content to the first model in response to the audio content reaching a preset duration, so as to determine the first song recognition result.
[0114] In some embodiments, the first model is further configured to output at least one of the following: lyrics of a candidate song corresponding to the first song identification result; and progress information of the audio content in the candidate song.
[0115] In some embodiments, the device 500 further includes a third processing module configured to: process audio content using a speech recognition model to determine a speech recognition result; and a first determining module configured to: determine text content corresponding to a preset grammatical structure from the speech recognition result.
[0116] In some embodiments, the device 500 further includes a fourth processing module configured to: process audio content using an audio energy detection model to determine whether the audio content includes singing content; and a second triggering module configured to: trigger a first model to process the audio content and trigger a second model to process the text content in response to determining that the audio content includes singing content.
[0117] In some embodiments, the audio content is first audio content, and the device 500 further includes: a second acquisition module configured to acquire second audio content associated with a live event; a detection module configured to: use an audio energy detection model to detect whether the second audio content includes singing content; and a third trigger module configured to: trigger the live interface to stop displaying lyrics content in response to the fact that no singing content is detected from the second audio content within a preset time period.
[0118] In some embodiments, the live streaming interface is also triggered to display interface content associated with the lyrics to indicate the singing progress corresponding to the song.
[0119] In some embodiments, the live streaming interface is also configured to periodically update the interface content to indicate the update progress of the corresponding song.
[0120] In some embodiments, the device 500 further includes a third acquisition module configured to: acquire progress information generated by the first model based on audio content, the progress information indicating the start and end times of a song segment corresponding to the audio content; and a second determination module configured to: determine the singing progress corresponding to the song based on the progress information.
[0121] In some embodiments, the second determining module is further configured to: determine a first duration for the first model to generate the first song recognition result, and determine a second duration for the second model to generate the second song recognition result; determine the processing time based on the first duration and the second duration; and determine the singing progress corresponding to the song based on the progress information and the processing time.
[0122] The units included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 500 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.
[0123] Figure 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 600 shown in Figure 6 can be used to implement the server 130 shown in Figure 1.
[0124] As shown in Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage devices 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0125] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 600.
[0126] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0127] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0128] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0129] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0130] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0131] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0132] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0134] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An audio processing method, comprising: Retrieve audio content associated with the live event; The audio content is processed using a first model to determine the recognition result of the first song; The second model is used to process the text content in order to determine the second song recognition result, wherein the text content is determined by recognizing the audio content; as well as In response to the first song identification result and the second song identification result indicating the same song, the lyrics of the song are triggered to be displayed in the live broadcast interface associated with the live broadcast event.
2. The method according to claim 1, wherein processing the audio content using a first model to determine the first song recognition result includes: In response to the audio content reaching a preset duration, the audio content is provided to the first model to determine the recognition result of the first song.
3. The method according to any one of claims 1 to 2, wherein the first model is further configured to output at least one of the following: The lyrics of the candidate songs corresponding to the first song identification result; The audio content contains progress information within the candidate songs.
4. The method according to any one of claims 1 to 3, wherein before processing the text content using the second model to determine the second song recognition result, the method further comprises: The audio content is processed using a speech recognition model to determine the speech recognition result; as well as The text content corresponding to the preset grammatical structure is determined from the speech recognition results.
5. The method according to any one of claims 1 to 4, further comprising: The audio content is processed using an audio energy detection model to determine whether the audio content includes singing content; as well as In response to determining that the audio content includes singing content, the first model is triggered to process the audio content and the second model is triggered to process the text content.
6. The method of claim 5, wherein the audio content is first audio content, and after triggering the presentation of the song's lyrics in a live streaming interface associated with the live streaming event, the method further comprises: Obtain the second audio content associated with the live event; The audio energy detection model is used to detect whether the second audio content includes singing content; as well as In response to the failure to detect singing content in the second audio content within a preset time period, the live streaming interface is triggered to stop displaying the lyrics content.
7. The method according to any one of claims 1 to 6, wherein the live streaming interface is further triggered as follows: The interface content associated with the lyrics is displayed to indicate the singing progress corresponding to the song.
8. The method according to claim 7, wherein the live streaming interface is further configured to periodically update the interface content to indicate the update performance progress of the song.
9. The method according to claim 7, further comprising: Obtain progress information generated by the first model based on the audio content, wherein the progress information indicates the start and end times of the song segment corresponding to the audio content; as well as Based on the progress information, the singing progress corresponding to the song is determined.
10. The method according to claim 9, wherein determining the singing progress corresponding to the song based on the progress information includes: Determine a first duration for the first model to generate the first song recognition result, and determine a second duration for the second model to generate the second song recognition result; Based on the first duration and the second duration, the processing time is determined; as well as Based on the progress information and the processing time, the singing progress corresponding to the song is determined.
11. An apparatus for audio processing, comprising: The first acquisition module is configured to acquire audio content associated with the live event; The first processing module is configured to process the audio content using a first model to determine the first song recognition result; The second processing module is configured to process the text content using a second model to determine the second song recognition result, wherein the text content is determined by recognizing the audio content; as well as The first triggering module is configured to trigger the presentation of the lyrics of the song in the live broadcast interface associated with the live broadcast event in response to the first song identification result and the second song identification result indicating the same song.
12. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processor.
13. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, implement the method according to any one of claims 1 to 10.
14. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.