Multimedia recording method, device, equipment and readable storage medium

By displaying text prompts and receiving selections during multimedia content playback, custom recording of multimedia segments can be achieved, solving the problems of low flexibility in song recording and low efficiency in human-computer interaction in existing technologies, and improving the freedom of multimedia recording and the efficiency of chorus performances.

CN115701713BActive Publication Date: 2025-12-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110879698.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-02
Publication Date
2025-12-12
Estimated Expiration
2041-08-02

AI Technical Summary

Technical Problem

Existing song recording applications lack flexibility and have low human-computer interaction efficiency, requiring users to use pre-provided unvoiced accompaniment in specific applications to record songs.

Method used

During multimedia content playback, a text prompt screen is displayed, selection operations are received, and a multimedia recording screen is displayed in response to the selection. Text segments are recorded and integrated with the multimedia content to achieve customized recording of multimedia segments.

Benefits of technology

It increases the freedom and diversity of multimedia recording, enhances the efficiency of users singing along with the original singers, and improves human-computer interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115701713B_ABST
    Figure CN115701713B_ABST
Patent Text Reader

Abstract

The application discloses a multimedia recording method and device, equipment and readable storage medium, and relates to the field of interface interaction. The method comprises the following steps: displaying a text prompt picture, the text prompt content is synchronized with multimedia content; receiving a selection operation on the text prompt content; displaying a multimedia recording picture; in response to the end of recording of a text segment, displaying a recording result, the recording result is a result obtained by fusing the multimedia content and a multimedia segment. In the process of multimedia playing, the selection of the recording segment is performed through the text prompt content, the recording of the multimedia segment is performed in the multimedia recording picture, the segment obtained by recording is fused with the multimedia content, and thus a customized multimedia result is obtained, the freedom and diversity of the multimedia recording are improved, the efficiency of the user in singing along with the original singer of a song is improved, and the human-computer interaction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of interface interaction, and in particular, to a multimedia recording method and device, equipment and readable storage medium. BACKGROUND

[0002] A singing recording application is an application that provides a song accompaniment to a user for singing recording, wherein the accompaniment of the song is provided in the application and does not contain the singing voice of the original singer or contains part of the singing voice of the original singer.

[0003] In related technologies, in a song recording application, a user first selects a song accompaniment to be sung, and then the selected song accompaniment is played, and the user sings the song when the accompaniment is played to the singing segment, and the recording is completed after the accompaniment is played.

[0004] However, when recording a song in the above manner, it needs to be implemented in a specific singing recording application, and the song accompaniment is a same pre-provided accompaniment without a singing voice, and the flexibility in the song recording process is low, and the human-computer interaction efficiency is low. SUMMARY

[0005] Embodiments of the present application provide a multimedia recording method, device, equipment and readable storage medium, which can improve the flexibility in the song recording process. The technical solution is as follows:

[0006] On the one hand, a multimedia recording method is provided, and the method comprises:

[0007] In the process of playing multimedia content, a text prompt picture is displayed, the text prompt picture comprises text prompt content arranged and displayed in sequence, and the text prompt content is synchronized with the multimedia content;

[0008] A selection operation on the text prompt content is received, and the selection operation is used to select a text segment in the text prompt content;

[0009] A multimedia recording picture is displayed in response to the selection operation, and the multimedia recording picture is used to indicate recording of a multimedia segment for the text segment;

[0010] A recording result is displayed in response to the end of recording for the text segment, and the recording result is a result obtained by fusing the multimedia content and the multimedia segment.

[0011] On the other hand, a multimedia recording device is provided, and the device comprises:

[0012] The display module is configured to display a text prompt picture during playing of the multimedia content, the text prompt picture comprising text prompt content arranged in sequence, the text prompt content being synchronized with the multimedia content.

[0013] The receiving module is configured to receive a selection operation on the text prompt content, the selection operation being configured to select a text segment in the text prompt content.

[0014] The display module is further configured to display a multimedia recording picture in response to the selection operation, the multimedia recording picture being configured to indicate recording of a multimedia segment for the text segment.

[0015] The display module is further configured to display a recording result in response to the recording of the text segment being completed, the recording result being a result of fusing the multimedia content and the multimedia segment.

[0016] In another aspect, a computer device is provided, the computer device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the multimedia recording method according to any one of the above embodiments of the present application.

[0017] In another aspect, a computer readable storage medium is provided, the storage medium storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by a processor to implement the multimedia recording method according to any one of the above embodiments of the present application.

[0018] In another aspect, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the multimedia recording method according to any one of the above embodiments.

[0019] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:

[0020] During playing of the multimedia, a recording segment is selected through the text prompt content, a multimedia segment is recorded in the multimedia recording picture, the recorded segment is fused with the multimedia content, and thus a customized multimedia result is obtained, the freedom and diversity of multimedia recording are improved, the efficiency of a user singing along with an original singer is improved, and the human-computer interaction efficiency is improved. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the interface interaction provided in an exemplary embodiment of this application;

[0023] Figure 2 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application;

[0024] Figure 3 This is a flowchart of a multimedia recording method provided in an exemplary embodiment of this application;

[0025] Figure 4 Based on Figure 3 The illustrated embodiment provides a schematic diagram of the lyrics prompt screen display process;

[0026] Figure 5 Based on Figure 3 The illustrated embodiment provides a schematic diagram of the display process of the dialogue prompt screen;

[0027] Figure 6 Based on Figure 3 The illustrated embodiment provides a schematic diagram of the lyrics selection process interface;

[0028] Figure 7 This is a flowchart of a multimedia recording method provided in another exemplary embodiment of this application;

[0029] Figure 8 Based on Figure 7 A schematic diagram of the interface for multimedia recording provided in the illustrated embodiment;

[0030] Figure 9 This is a flowchart of a multimedia recording method provided in another exemplary embodiment of this application;

[0031] Figure 10 Based on Figure 9 A schematic diagram of the processing candidates provided in the illustrated embodiment;

[0032] Figure 11 This is a schematic diagram of the overall flow of a multimedia recording method provided in an exemplary embodiment of this application;

[0033] Figure 12is a structural block diagram of a multimedia recording device provided by an example embodiment of the present application.

[0034] Figure 13 is a structural block diagram of a multimedia recording device provided by another example embodiment of the present application.

[0035] Figure 14 is a structural block diagram of a terminal provided by an example embodiment of the present application. DETAILED DESCRIPTION

[0036] For the purpose, technical solutions and advantages of the present application to be clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0037] Singing application is a kind of application that is currently favored by users. Singing application provides online solo singing and online chorus and other functions.

[0038] When a user wants to perform online chorus, the song chorus method in the related art includes: a first user sings a first part in a target song, and a first client records to obtain a first audio file; the first client uploads the first audio file to a server; a second client downloads the first audio file, and in the process of playing the first audio file, a second user sings a second part in the target song, and the second client records to obtain a second audio file; the first audio file and the second audio file are merged into a chorus audio file. The second client uploads the chorus audio file to the server.

[0039] However, when the song chorus is performed in the above manner, the flexibility is poor, and the human-computer interaction efficiency in the chorus process is low.

[0040] First, the application scenario of the multimedia recording method provided by the embodiments of the present application is introduced, including at least one of the following application scenarios:

[0041] First, the multimedia recording method is implemented as an audio recording method and applied to a song chorus scenario.

[0042] The user selects the music to be played in the music playing application program, selects the part to be sung in the lyrics interface of the music, enters the recording interface, performs voice muting on the selected part, and performs audio recording, so that the background combines the recorded audio and the music after voice muting to obtain chorus audio. The effect of removing the voice of the selected part in the music and embedding the user recorded audio is realized, so that the effect of the original singing of the music and the chorus of the user is realized.

[0043] Second, the multimedia recording method is implemented as an audio recording method and applied to a dialogue dubbing scenario.

[0044] The user selects a video to be played in a video playing application, selects a script part to be dubbed in a script interface of the video, enters a recording interface, performs voice de-noising on the selected part, and records audio, so that the background combines the recorded audio with the video audio track after voice de-noising to obtain a dubbed video.

[0045] Thirdly, the multimedia recording method is implemented as a video method and applied to a music video (MV).

[0046] The user selects a music video to be played in a music playing application, selects a script part to be recorded in a lyrics interface corresponding to the music video, enters a recording interface, starts a camera, and records video content, so that the background combines the recorded video content with the music video to obtain a combined video.

[0047] It is worth noting that the above application scenarios are only illustrative examples, and the multimedia recording method is not limited to the specific application scenarios.

[0048] The above is described by taking the song duet scenario as an example, Figure 1 is an interface interaction diagram provided by an example embodiment of the present application, as Figure 1 shown, a music playing interface 100 is first displayed, which is a lyrics display interface of the music and includes lyrics content of the music. After receiving a selection operation on a lyrics segment 110 in the lyrics content, a singing control 130 is clicked to trigger the display of a recording interface 140. In the recording interface 140, audio recording corresponding to the lyrics segment 110 is performed according to the song playing progress, wherein the original singing corresponding to the lyrics segment 110 is de-noised in real time. After the audio recording is completed, a recording result 150 is displayed, which is a result obtained by combining the recording audio of the user with the original singing of the song.

[0049] Figure 2 is an implementation environment diagram provided by an example embodiment of the present application, as Figure 2 shown, the implementation environment includes a terminal 210 and a server 220, and the terminal 210 and the server 220 are connected through a communication network 230.

[0050] The terminal 210 is installed with an application program for multimedia content playing, such as a music playing application program, a video playing application program, and the like. During playing of the multimedia content, text prompt content, such as lyrics content, dialogue content, and the like, is displayed. After the user selects a text segment on the terminal 210 that needs to be recorded, the user triggers multimedia content recording for the text segment, and a multimedia recording screen is displayed. When the terminal 210 receives the trigger operation of the multimedia content recording, the terminal 210 sends a recording signal to the server 220.

[0051] After receiving the recording signal, the server 220 sends the multimedia content that needs to be played to the terminal 210 in real time, wherein the selected text segment is sent to the terminal 210 after being subjected to voice de-noising by the server 220. It should be noted that in the above embodiment, the server 220 is taken as an example to complete voice de-noising, and in some embodiments, the voice de-noising process can also be completed by the terminal 210, or by the server 220 and the terminal 210 together, such as after the server 220 performs voice de-noising, the terminal 210 performs fine de-noising processing, and the embodiments of the present application are not limited in this regard.

[0052] Since it is necessary to maintain synchronization between the terminal 210 and the server 220, the terminal 210 and the server 220 realize network relay acceleration through a software defined real-time network (SD-RTN). The SD-RTN provides a real-time data transmission cloud service based on a user datagram protocol (UDP) and with end-to-end network delay of milliseconds. The SD-RTN is a service architecture that can carry any point-to-point real-time data transmission requirement: as long as an open application programming interface (API) is called.

[0053] The terminal described above can be a mobile phone, a tablet computer, a desktop computer, a portable notebook computer, or the like, and the embodiments of the present application are not limited in this regard.

[0054] It should be noted that the server described above can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.

[0055] The cloud technology refers to a hosting technology of unifying a series of resources such as hardware, software, network, etc. in a wide area network or a local area network to realize data calculation, storage, processing and sharing. The cloud technology is a general term of network technology, information technology, integration technology, management platform technology, application technology and the like applied based on a cloud computing business model, can form a resource pool, and is used on demand, flexibly and conveniently. The cloud computing technology will become an important support. The background service of a technical network system needs a large amount of calculation and storage resources, such as a video website, a picture website and more portal websites. With the high development and application of the Internet industry, in the future, each item is likely to have its own identification mark and needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and various industry data need strong system support, which can only be realized through cloud computing.

[0056] In some embodiments, the server described above can also be implemented as a node in a blockchain system. The blockchain is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. The blockchain is essentially a decentralized database, which is a series of data blocks associated using cryptography. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-fake) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer and an application service layer.

[0057] In combination with the above application scenarios and implementation environments, the multimedia recording method provided by the present application is described, Figure 3 is a flowchart of the multimedia recording method provided by an exemplary embodiment of the present application. Taking the case that the method is applied in a terminal, as shown in Figure 3 , the method comprises:

[0058] Step 301, in the process of playing multimedia content, a text prompt picture is displayed.

[0059] The text prompt picture includes text prompt content displayed in sequence, wherein the text prompt content is synchronized with the multimedia content. Illustratively, the text prompt content includes at least one of lyric content, dialogue content and narration content.

[0060] Illustratively, in the process of playing music, lyric content is displayed, wherein the lyric content is displayed in a lyric prompt picture, and the lyric prompt picture is displayed according to a trigger operation.

[0061] In some embodiments, a multimedia playing picture is first displayed in a playing process of the multimedia content, the multimedia information of the multimedia content is included in the multimedia playing picture, and the text prompt picture is displayed in response to receiving the text prompt operation. Illustratively, when the user selects to play a song A, a song playing picture is first displayed, the song playing picture includes a cover image, a name, a singer of the playing song, and a control playing control, such as pause, next, previous, and the like. When the lyrics prompt operation on the song playing picture is received, the lyrics prompt picture is displayed, which includes the lyrics content of the playing song, and the display of the lyrics content is synchronized with the progress of the song playing.

[0062] Illustratively, Figure 4 is a display process diagram of the lyrics prompt picture provided by an exemplary embodiment of the present application, as shown in Figure 4 A song playing picture 400 is first displayed, the song playing picture 400 includes a song cover 401, a song name 402, a song playing control 403, and the like, when the left swipe operation on the song playing picture 400 is received, a lyrics prompt picture 410 is displayed, the lyrics prompt picture 410 includes the lyrics content arranged and displayed in sequence.

[0063] The above Figure 4 example is directed to the display of the lyrics content in the song playing process, in some embodiments, the multimedia content can also be implemented as video content, Figure 5 is a display process diagram of the dialogue prompt picture provided by an exemplary embodiment of the present application, as shown in Figure 5 A video playing picture 500 is first displayed, the video playing picture 500 includes a video picture 501, a dialogue display control 502, when the selection operation on the dialogue display control 502 is received, a dialogue display picture 510 is displayed, the dialogue display picture 510 includes the dialogue content arranged and displayed in pairs in sequence.

[0064] The process of playing the multimedia content refers to the state that the video picture or the audio of the multimedia content is in the playing state, or the state that the video picture or the audio of the multimedia content is in the pause state after the multimedia content is selected.

[0065] In step 302, the selection operation on the text prompt content is received, the selection operation is used to select the text segment in the text prompt content.

[0066] The selection operation is used to select one or more text sentences from the text prompt content as the selected text segment. Illustratively, the selection operation is used to select a song lyric from the song lyric prompt content as the text segment; or the selection operation is used to select multiple consecutive or non-consecutive song lyrics from the song lyric prompt content as the text segment.

[0067] The selection manner of the text segment includes at least one of the following manners:

[0068] Firstly, since the text prompt content is displayed in sequence, the long press operation on a single text content is used as the selection operation of the text segment. When the text segment includes multiple text contents, the first selected text content needs to be selected by the long press operation, and the subsequent text content can be directly selected by the click operation.

[0069] That is, the long press operation on the text segment in the text prompt content is received, and the long press operation is used as the selection operation. For the first selected text content, the long press operation is received as the selection operation of the first selected text content, and for the nth selected text content, the click operation is received as the selection operation of the nth selected text content, n≥2.

[0070] Illustratively, please refer to Figure 6 which shows a song lyric segment selection process interface diagram provided by an example embodiment of the present application, as shown in Figure 6 In the song lyric prompt interface 600, the song lyric content is displayed in sequence, and no control is displayed on the right side of the current song lyric content. When the long press operation on the song lyric content 610 is received, the selection operation of the song lyric content 610 is triggered, and the highlighted selection control is displayed on the right side of the song lyric content 610. The selection controls on the right side of other song lyrics are not highlighted, indicating that the current song lyric content 610 is selected, and other song lyric contents are not selected. The user can select other song lyric contents by clicking the selection control, or quickly select other song lyric contents in batch by sliding operation. In some embodiments, all text prompt contents can also be selected by the select-all control.

[0071] Secondly, the recording trigger operation is first triggered, and the selection control is displayed on the corresponding position of the text prompt content based on the recording trigger operation, and the trigger operation on the selection control is used as the selection operation.

[0072] Illustratively, the text prompt screen also includes a recording trigger control. When the trigger operation on the recording trigger control is received, the selection control to be selected is displayed on the right side of the text prompt content. When the selection operation on the selection control is received, it means that the corresponding text content is selected.

[0073] It is worth noting that the selection of the text segment is only an illustrative example, and the determination of the text segment is not limited in the embodiments of the present application.

[0074] In step 303, a multimedia recording interface is displayed in response to the selection operation.

[0075] The multimedia recording interface is used to indicate the recording of the multimedia segment for the text segment.

[0076] In some embodiments, the multimedia recording interface is played from the initial position of the multimedia content, and when the multimedia content corresponding to the selected text segment is played, it is prompted that the multimedia recording is being performed.

[0077] In other embodiments, the multimedia recording interface is played from the specified position of the multimedia content, and when the multimedia content corresponding to the selected text segment is played, it is prompted that the multimedia recording is being performed. The specified position is determined according to the first selected text content, such as a specified number of text contents before the first selected text content, or a last segment node position before the first selected text content, or a specified time length before the first selected text content. The implementation of the specified position is not limited in the embodiments of the present application.

[0078] In some embodiments, the multimedia recording component is started in response to the selection operation, and the multimedia recording interface is displayed, such as starting the microphone or starting the camera in response to the selection operation, and displaying the multimedia recording interface.

[0079] In step 304, the recording result is displayed in response to the end of the recording of the text segment.

[0080] The recording result is the result of fusing the multimedia content and the multimedia segment. The fusion manner of the multimedia content and the multimedia segment includes at least one of the superimposition manner, the overlay manner, and the partial overlay manner.

[0081] The superimposition manner means that the multimedia segment is superimposed at the position corresponding to the selected text segment based on the original multimedia content. Illustratively, for audio content, the audio segment is superimposed and fused with the originally played audio content to obtain an audio with both the original singing and the user's recording sound at the position corresponding to the text segment. For video content, the video segment is combined with the originally played video content, that is, the video segment is superimposed and displayed at the specified position of the video screen corresponding to the text segment in the originally played video.

[0082] The covering manner refers to deleting the segment content corresponding to the text segment in the original multimedia content, and covering with the recorded multimedia segment. Illustratively, for audio content, the selected text segment in the audio content corresponds to 2:00:15 to 2:01:01, the audio content in this time period is deleted, and the recorded audio segment is inserted into this position.

[0083] The partial covering refers to partially deleting the segment content corresponding to the text segment in the original multimedia content, and superimposing the recorded multimedia segment with the non-deleted part. Illustratively, for audio content, the audio segment corresponding to the selected text segment in the audio content is voice-removed, and the recorded audio segment is superimposed with the non-removed accompaniment sound.

[0084] The fusion manner of the multimedia content and the multimedia segment can also include other fusion manners, which are not limited by the embodiments of the present application.

[0085] In summary, the method provided by the embodiments of the present application can select the recording segment through the text prompt content in the process of multimedia playing, record the multimedia segment in the multimedia recording picture, fuse the recorded segment with the multimedia content, and thus obtain a customized multimedia result, thereby improving the freedom and diversity of multimedia recording, and improving the efficiency of user singing with the original singer and the efficiency of human-computer interaction.

[0086] In some embodiments, the multimedia segment includes an audio segment, and the multimedia content includes audio content, Figure 7 is a flowchart of a multimedia recording method provided by another exemplary embodiment of the present application. The method is taken as an example applied to a terminal, as shown in Figure 7 The method comprises the following steps:

[0087] Step 701, displaying a text prompt picture in the process of playing multimedia content.

[0088] The text prompt picture includes text prompt content displayed in sequence, wherein the text prompt content is synchronized with the multimedia content. Illustratively, the text prompt content includes at least one of lyric content, dialogue content, and narration content.

[0089] In some embodiments, the multimedia playing picture is first displayed in the process of playing the multimedia content, the multimedia playing picture includes multimedia information of the multimedia content, and the text prompt picture is displayed in response to receiving a text prompt operation.

[0090] Step 702, receiving a selection operation on the text prompt content, the selection operation being used for selecting a text segment in the text prompt content.

[0091] The selecting operation is used to select one or more text sentences from the text prompt content as the selected text segment. Illustratively, the selecting operation is used to select a song lyric from the song lyric prompt content as the text segment; or the selecting operation is used to select multiple song lyrics, which are continuous or discontinuous, from the song lyric prompt content as the text segment.

[0092] In step 703, an audio collection component is started in response to the selecting operation.

[0093] In response to the selecting operation, a microphone is started, and audio is collected through the microphone.

[0094] In some embodiments, when the multimedia segment includes a video segment, a video stream collection component, such as a camera, is started in response to the selecting operation, and the video stream collection component is used to collect a video picture.

[0095] In step 704, a multimedia recording interface is displayed.

[0096] For the audio segment, the multimedia recording interface is used to indicate the recording progress of the audio segment, that is, the multimedia recording interface includes the playing progress of the current multimedia content, and when the multimedia content plays to the part corresponding to the text segment, the user needs to perform audio recording. The selected text segment is marked and displayed in the multimedia recording interface.

[0097] Optionally, the multimedia recording interface further displays first prompt information. For example, since the audio segment recording needs a relatively quiet environment, the first prompt information “singing, a quiet environment helps audio collection” is displayed.

[0098] In some embodiments, in the multimedia recording interface, if the recording for the text segment has not been completed and an exit recording operation is received, second prompt information is displayed, which is used to prompt that the recording is in process. For example, the second prompt information “still in the singing process, do you want to leave?” is displayed.

[0099] Illustratively, please refer to Figure 8 which shows the interface of the multimedia recording interface provided by an example embodiment of the present application, as shown in Figure 8 In the multimedia recording interface 800, first prompt information 810 is displayed, which is used to prompt that the user is in the audio recording process. The multimedia recording interface 800 further includes song lyric prompt content. According to the audio playing, the song lyrics are synchronously displayed by scrolling. The selected song lyric 820 is marked and highlighted, which is used to prompt the user to record the singing of the song lyric 820. When the recording is not completed and an exit recording operation is received, second prompt information 830 is displayed, which is used to indicate that the user needs to confirm again whether to exit the recording.

[0100] In the recording process of audio, the server real-time processes the part corresponding to the text segment by human voice noise reduction, wherein the human voice noise reduction is achieved by audio human voice elimination technology. The sound wave form of human voice is the same or similar in two channels of a song, and the method of subtracting two channels can be adopted to eliminate human voice in a stereo song. The human voice elimination process includes the following steps: step one, high-pass filter; step two, channel mixer; and step three, low-pass filter. The channel mixer is generally divided into four parameters, which are respectively: the percentage a1 of the original left channel Left in the new left channel newLeft; the percentage a2 of the original right channel Right in the new left channel newLeft; the percentage b1 of the original left channel Left in the new right channel newRight; and the percentage b2 of the original right channel Right in the new right channel newRight. The values of the four numbers a1, a2, a3 and a4 are between -100 and 100. Then the new left channel sample value newLeft = a1×Left / 100+a2×Right / 100, and the new right channel sample value newRight = b1×Left / 100+b2×Right / 100.

[0101] In order to realize the subtraction of left and right channels, the four values of the channel mixer are respectively 100, -100, -100 and 100, to generate a stereo waveform with opposite left and right channel waveforms; however, in this case, no sound will appear if the sound is played on a single-channel speaker. In order to prevent this error from occurring, the right channel should be flipped, that is, the four values of the channel mixer are 100, -100, 100 and -100.

[0102] The cutoff frequency of the low-pass filter and the pass frequency of the high-pass filter should be the same (below 400 Hz). For a small part of songs, the human voice elimination will also eliminate the high frequency part of the music. In this case, the process of human voice elimination should be modified as follows:

[0103] The audio track 1 of the original song is input into the band-stop filter for channel mixing. The audio track 2 is input into the band-pass filter, wherein the settings of the channel mixing are still the same as above, and the two frequencies of the band-pass filter and the band-stop filter are still the same, that is, the cutoff frequency (frequency 1) of the band-pass filter and the pass frequency (frequency 1) of the band-stop filter are set below 400 Hz, and the pass frequency (frequency 2) of the band-pass filter and the cutoff frequency (frequency 2) of the band-stop filter are set between 2000 Hz and 8000 Hz.

[0104] When the system eliminates the human voice, the user can directly sing the pressed lyrics when pressing the lyrics for a long time, triggering the recording function to start. The part of the human voice audio will be automatically recorded and uploaded to the cloud according to the SD-RTN technology. The SD-RTN can be automatically converted into a low-latency Internet Transmission Layer Protocol (Quick UDP Internet Connection, QUIC) to reduce the generation of delay accumulation, thereby realizing the real-time chorus mode. The audio information is stored in the cloud at the same time. During the process, the API can be called to notify the access point of the SD-RTN to send data to the specified IP and port. At the same time, the chorus background music data will also be uploaded to the cloud in real time. After the whole song is played or the user automatically intervenes to stop playing, the final audio synthesized in the cloud will prompt the user to download / save to the local.

[0105] In some embodiments, when the multimedia segment includes a video segment, the multimedia recording picture is used to indicate the recording progress of the video segment, such as indicating the playing progress of the current video through the scrolling display of the text prompt content, thereby indicating the recording progress.

[0106] Step 705, in response to the end of recording of the text segment, embedding the audio segment into the part of the audio content corresponding to the text segment.

[0107] In some embodiments, the way of embedding the audio segment into the audio content includes at least one of the following ways:

[0108] 1. performing human voice sound elimination on the part of the audio content corresponding to the text segment to obtain accompaniment audio, and adding the audio segment to the part of the accompaniment audio corresponding to the text segment;

[0109] Wherein, the human voice sound elimination is realized by the above-mentioned audio human voice elimination technology.

[0110] 2. covering the audio segment on the part of the audio content corresponding to the text segment;

[0111] 3. superimposing the audio segment on the part of the audio content corresponding to the text segment.

[0112] It is worth noting that the above-mentioned embedding way of the audio segment is only an illustrative example, and the embedding of the audio segment by the embodiments of the present application is not limited.

[0113] In some embodiments, the fusion manner of the audio segment and the audio content is selected by the user; or the fusion manner of the audio segment and the audio content is default; or the fusion manner of the audio segment and the audio content is determined according to the user-selected chorus form, for example, when the user selects to sing together with the original singer, the superposition manner is adopted, i.e., the above-mentioned manner 3; when the user selects to sing with the original singer in sections, the voice sound elimination and covering manner is adopted, i.e., the above-mentioned manner 1.

[0114] In step 706, the recording result is displayed.

[0115] The recording result is the result obtained by fusing the multimedia content and the multimedia segment.

[0116] In summary, the method provided by the embodiments of the present application can select the recording segment through the text prompt content in the process of multimedia playing, record the multimedia segment in the multimedia recording picture, fuse the recorded segment with the multimedia content, and thus obtain a self-defined multimedia result, thereby improving the freedom and diversity of multimedia recording, and the user does not need to download a specific chorus accompaniment to sing a song, which improves the efficiency of the user singing with the original singer and improves the human-computer interaction efficiency.

[0117] The method provided by the embodiments can provide various singing manners of the user and the original singer through different embedding manners, improve the diversity of multimedia processing, and improve the human-computer interaction efficiency.

[0118] In some embodiments, after the recording of the text segment corresponding part is completed, the multimedia content can be continuously played according to the selection, or the recording result can be selected to be viewed. Figure 9 FIG. 7 is a flowchart of a multimedia recording method provided by another exemplary embodiment of the present application, which is applied to a terminal for example, as shown in FIG. 7, the method comprises the following steps. Figure 9

[0119] In step 901, a text prompt picture is displayed in the process of playing the multimedia content.

[0120] The text prompt picture comprises text prompt content arranged and displayed in sequence, wherein the text prompt content is synchronized with the multimedia content. Illustratively, the text prompt content comprises at least one of lyric content, dialogue content, and narration content.

[0121] In some embodiments, a multimedia playing picture is first displayed in the process of playing the multimedia content, the multimedia playing picture comprises multimedia information of the multimedia content, and the text prompt picture is displayed in response to receiving a text prompt operation.

[0122] ​Step 902, receiving a selection operation on the text prompt content, the selection operation being used to select a text segment in the text prompt content.

[0123] The selection operation is used to select one or more text sentences from the text prompt content as the selected text segment. Illustratively, the selection operation is used to select a song lyric from the song lyric prompt content as the text segment; or the selection operation is used to select multiple continuous or discontinuous song lyrics from the song lyric prompt content as the text segment.

[0124] Step 903, displaying a multimedia recording screen in response to the selection operation.

[0125] The multimedia recording screen is used to indicate the recording of a multimedia segment for the text segment.

[0126] In some embodiments, the multimedia recording screen is used to play the multimedia content from the initial position of the multimedia content, and to prompt that the multimedia recording is being performed when the multimedia content corresponding to the selected text segment is played.

[0127] In some embodiments, the multimedia recording component is started in response to the selection operation, and the multimedia recording screen is displayed, such as starting the microphone or starting the camera in response to the selection operation, and displaying the multimedia recording screen.

[0128] Step 904, displaying a processing candidate in response to the end of the recording for the text segment.

[0129] The processing candidate includes a result candidate and a play candidate. The result candidate is used to indicate the display of the multimedia recording result, and the play candidate is used to indicate the continuous playing of the current multimedia content.

[0130] Illustratively, please refer to Figure 10 which shows a schematic diagram of the processing candidate provided by one exemplary embodiment of the present application, as shown in Figure 10 When the recording for the text segment ends, the processing candidate 1010 is displayed on the multimedia recording screen 1000, which includes the result candidate 1011 and the play candidate 1012. A prompt message 1013 “recording success, the complete chorus file can be viewed after the song ends” is also displayed.

[0131] Step 905, receiving a first selection operation on the result display candidate.

[0132] Optionally, the first selection operation includes a click operation on the display candidate.

[0133] Step 906, displaying the recording result based on the first selection operation.

[0134] The recording result is a result of fusing the multimedia content and the multimedia segment.

[0135] Optionally, a recording result picture is displayed, and when the multimedia segment includes an audio segment, audio obtained by playing the audio content combined with the audio segment based on the recording result picture is played.

[0136] Step 907: receiving a second selection operation on the playing candidate.

[0137] Optionally, the first selection operation includes a click operation on the playing candidate.

[0138] Step 908: continuing to play the multimedia content based on the second selection operation.

[0139] In summary, the method provided by the embodiments of the present application enables the selection of a recording segment through a text prompt during multimedia playing, thereby recording a multimedia segment in a multimedia recording picture, fusing the recorded segment with the multimedia content, and obtaining a customized multimedia result, which improves the freedom and diversity of multimedia recording, and enables a user to sing a song without downloading a specific accompaniment, thereby improving the efficiency of singing a song with an original singer and the efficiency of human-computer interaction.

[0140] Figure 11 is a schematic diagram of the overall flow of a multimedia recording method provided by an exemplary embodiment of the present application, as shown in the figure, the method includes an original audio end 1110, i.e., an original singer end, and a new audio end 1120, i.e., a singing end. The original audio end 1110 can be a terminal or a server end. Figure 11

[0141] First, the original audio end 1110 generates an online audio file, so that the new audio end 1120 obtains the online audio file from the original audio end 1110 for playing.

[0142] When the new audio end 1120 receives a long press operation of the lyrics to enter a singing mode, the original audio end 1110 automatically eliminates the voice of the long-pressed selected lyrics part and retains the background accompaniment, and sends the audio file with the eliminated voice to the new audio end 1120 for playing. The new audio end 1120 sings the selected lyrics part, thereby generating a recording audio, and finally synthesizes the accompaniment with the eliminated voice and the recording audio to obtain a singing audio file 1130.

[0143] Figure 12 is a structural block diagram of a multimedia recording device provided by an exemplary embodiment of the present application, as shown in the figure, the device includes: Figure 12

[0144] ​​The display module 1210 is configured to display a text prompt picture during playing of the multimedia content, the text prompt picture comprising text prompt content arranged in sequence, the text prompt content being synchronized with the multimedia content.

[0145] The receiving module 1220 is configured to receive a selection operation on the text prompt content, the selection operation being used for selecting a text segment in the text prompt content.

[0146] The display module 1210 is further configured to display a multimedia recording picture in response to the selection operation, the multimedia recording picture being used for indicating recording of a multimedia segment for the text segment.

[0147] The display module 1210 is further configured to display a recording result in response to the recording of the text segment being ended, the recording result being a result of fusing the multimedia content and the multimedia segment.

[0148] In an optional embodiment, the multimedia segment comprises an audio segment.

[0149] As shown in Figure 13 The apparatus further comprises:

[0150] The starting module 1230 is configured to start an audio acquisition component in response to the selection operation, the audio acquisition component being used for acquiring an audio signal.

[0151] The display module 1210 is further configured to display the multimedia recording picture, the multimedia recording picture being used for indicating a recording progress of the audio segment.

[0152] In an optional embodiment, the multimedia content comprises audio content.

[0153] The apparatus further comprises:

[0154] The editing module 1240 is configured to embed the audio segment into a part of the audio content corresponding to the text segment in response to the recording of the text segment being ended.

[0155] The display module 1210 is further configured to display the recording result.

[0156] In an optional embodiment, the editing module 1240 is further configured to perform voice de-noising processing on the part of the audio content corresponding to the text segment to obtain accompaniment audio, and add the audio segment to a part of the accompaniment audio corresponding to the text segment.

[0157] Or,

[0158] The editing module 1240 is further configured to overlay the audio segment in the audio content corresponding to the text segment.

[0159] Or,

[0160] The editing module 1240 is further configured to overlay the audio segment in the audio content corresponding to the text segment.

[0161] In an optional embodiment, the multimedia segment comprises a video segment.

[0162] The apparatus further comprises:

[0163] The starting module 1230 is configured to start a video stream acquisition component in response to the selection operation, the video stream acquisition component being configured to acquire a video picture.

[0164] The display module 1210 is further configured to display the multimedia recording picture, the multimedia recording picture being configured to indicate a recording progress of the video segment.

[0165] In an optional embodiment, the display module 1210 is further configured to display a recording control in response to the selection operation, the recording control being configured to indicate multimedia recording based on the text segment.

[0166] The display module 1210 is further configured to display the multimedia recording picture in response to receiving a triggering operation on the recording control.

[0167] In an optional embodiment, the display module 1210 is further configured to display processing candidates in response to recording of the text segment being completed, the processing candidates comprising a result display candidate and a play candidate.

[0168] The receiving module 1220 is further configured to receive a first selection operation on the result display candidate.

[0169] The display module 1210 is further configured to display the recording result based on the first selection operation.

[0170] In an optional embodiment, the receiving module 1220 is further configured to receive a second selection operation on the play candidate.

[0171] The apparatus further comprises:

[0172] The play module 1250 is configured to continue playing the multimedia content based on the second selection operation.

[0173] In an optional embodiment, the display module 1210 is further configured to display a multimedia playing picture during playing of the multimedia content, and the multimedia playing picture comprises multimedia information of the multimedia content.

[0174] The display module 1210 is further configured to display the text prompt picture in response to receiving a text prompt operation.

[0175] In an optional embodiment, the receiving module 1220 is further configured to receive a long-press operation on the text segment in the text prompt content, and take the long-press operation as the selection operation.

[0176] Alternatively,

[0177] The receiving module 1220 is further configured to receive a recording trigger operation, display a selection control at a position corresponding to the text prompt content based on the recording trigger operation, and take a trigger operation on the selection control as the selection operation.

[0178] To sum up, the device provided by the embodiments of the present application can select a segment for recording through text prompt content during multimedia playing, thereby recording a multimedia segment in a multimedia recording picture, fusing the recorded segment with the multimedia content, and obtaining a customized multimedia result, which improves the freedom and diversity of multimedia recording, and improves the efficiency of song chorus by users without downloading a specific chorus accompaniment and improves the efficiency of human-computer interaction.

[0179] It should be noted that: the multimedia recording device provided by the above embodiments is only exemplified by the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the multimedia recording device and the multimedia recording method provided by the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.

[0180] Figure 14A structure block diagram of a terminal 1400 is shown, which is provided by an example embodiment of the present application. The terminal 1400 can be a smart phone, a tablet computer, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 1400 can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other names.

[0181] Generally, the terminal 1400 includes a processor 1401 and a memory 1402.

[0182] The processor 1401 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1401 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 1401 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1401 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by a display screen. In some embodiments, the processor 1401 can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.

[0183] The memory 1402 can include one or more computer-readable storage media, which can be non-transitory. The memory 1402 can also include a high-speed random access memory, and a non-volatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1402 is used to store at least one instruction for being executed by the processor 1401 to implement a multimedia recording method provided by a method embodiment of the present application.

[0184] In some embodiments, the terminal 1400 can further optionally include a peripheral device interface 1403 and at least one peripheral device. The processor 1401, the memory 1402 and the peripheral device interface 1403 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 1403 through a bus, a signal line or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1404, a display screen 1405, a camera 1406, an audio circuit 1407 and a power supply 1409.

[0185] The peripheral device interface 1403 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402 and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402 and the peripheral device interface 1403 can be implemented on a separate chip or circuit board, and the present embodiments are not limited in this regard.

[0186] The radio frequency circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1404 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 1404 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 1404 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1404 can also include NFC (Near Field Communication) related circuitry, and the present application is not limited in this regard.

[0187] The display screen 1405 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 1405 is a touch display screen, the display screen 1405 is further configured to capture touch signals on or above the surface of the display screen 1405. The touch signals can be input to the processor 1401 as control signals for processing. In this case, the display screen 1405 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 1405 can be one, arranged on the front panel of the terminal 1400; in other embodiments, the display screen 1405 can be at least two, arranged on different surfaces of the terminal 1400 or in a folding design; in still other embodiments, the display screen 1405 can be a flexible display screen, arranged on a curved surface or a folding surface of the terminal 1400. Even, the display screen 1405 can also be arranged in an irregular shape other than a rectangle, i.e., a special-shaped screen. The display screen 1405 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0188] The camera assembly 1406 is configured to capture images or videos. Optionally, the camera assembly 1406 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is arranged on the front panel of the terminal, and the rear-facing camera is arranged on the back of the terminal. In some embodiments, the rear-facing camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 1406 can further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0189] The audio circuit 1407 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into an electrical signal input to the processor 1401 for processing, or input to the radio frequency circuit 1404 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, respectively arranged at different parts of the terminal 1400. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signal from the processor 1401 or the radio frequency circuit 1404 into sound waves. The speaker can be a traditional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, not only can it convert electrical signals into sound waves that humans can hear, but it can also convert electrical signals into sound waves that humans cannot hear for ranging purposes. In some embodiments, the audio circuit 1407 can also include a headphone jack.

[0190] The power supply 1409 is used to supply power to each component in the terminal 1400. The power supply 1409 can be alternating current, direct current, disposable battery or rechargeable battery. When the power supply 1409 includes a rechargeable battery, the rechargeable battery can be a wired charging battery or a wireless charging battery. The wired charging battery is a battery charged through a wired line, and the wireless charging battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0191] In some embodiments, the terminal 1400 further includes one or more sensors 1410. The one or more sensors 1410 include, but are not limited to, an acceleration sensor 1411, a gyroscope sensor 1412, a pressure sensor 1413, an optical sensor 1415, and a proximity sensor 1416.

[0192] The acceleration sensor 1411 can detect the acceleration magnitude in three coordinate axes of the coordinate system established by the terminal 1400. For example, the acceleration sensor 1411 can be used to detect the components of the gravitational acceleration in three coordinate axes. The processor 1401 can control the touch display screen 1405 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1411. The acceleration sensor 1411 can also be used for game or user motion data collection.

[0193] The gyroscope sensor 1412 can detect the body direction and rotation angle of the terminal 1400, and the gyroscope sensor 1412 can collect 3D actions of the user on the terminal 1400 in cooperation with the acceleration sensor 1411. The processor 1401 can realize the following functions according to the data collected by the gyroscope sensor 1412: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization when shooting, game control, and inertial navigation.

[0194] The pressure sensor 1413 can be disposed on the side bezel of the terminal 1400 and / or on the lower layer of the touch display screen 1405. When the pressure sensor 1413 is disposed on the side bezel of the terminal 1400, it can detect the user's grip signal on the terminal 1400, and the processor 1401 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1413. When the pressure sensor 1413 is disposed on the lower layer of the touch display screen 1405, the processor 1401 can control the operable controls on the UI interface based on the user's pressure operation on the touch display screen 1405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0195] An optical sensor 1415 is used to collect ambient light intensity. In one embodiment, the processor 1401 can control the display brightness of the touch screen 1405 based on the ambient light intensity collected by the optical sensor 1415. Specifically, when the ambient light intensity is high, the display brightness of the touch screen 1405 is increased; when the ambient light intensity is low, the display brightness of the touch screen 1405 is decreased. In another embodiment, the processor 1401 can also dynamically adjust the shooting parameters of the camera assembly 1406 based on the ambient light intensity collected by the optical sensor 1415.

[0196] The proximity sensor 1416, also known as a distance sensor, is typically located on the front panel of the terminal 1400. The proximity sensor 1416 is used to detect the distance between the user and the front of the terminal 1400. In one embodiment, when the proximity sensor 1416 detects that the distance between the user and the front of the terminal 1400 is gradually decreasing, the processor 1401 controls the touchscreen display 1405 to switch from a screen-on state to a screen-off state; when the proximity sensor 1416 detects that the distance between the user and the front of the terminal 1400 is gradually increasing, the processor 1401 controls the touchscreen display 1405 to switch from a screen-off state to a screen-on state.

[0197] Those skilled in the art will understand that Figure 14 The structure shown does not constitute a limitation on terminal 1400 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0198] Optionally, the computer readable storage medium can include a read only memory (ROM), a random access memory (RAM), a solid state disk (SSD), an optical disk, etc. Among them, the random access memory can include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM). The above-mentioned application embodiment serial number is only for description, not representing the pros and cons of the embodiment.

[0199] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read only memory, a magnetic disk or an optical disk, etc.

[0200] The above is only an optional embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multimedia recording method, characterized by, The method comprises: In the process of playing multimedia content, synchronously display text prompt content corresponding to the multimedia content and sequentially rolling in time order; During the playing of the multimedia content, receive a selection operation on any one or more text segments in the rolling text prompt content, the text segments comprising text prompt content discontinuous in time; In response to the selection operation, display a multimedia recording picture, the multimedia recording picture being used to indicate recording of a multimedia segment for each selected text segment; In response to the end of recording for the text segment, overlay or partially overlay or superimpose each multimedia segment recorded to a part of the multimedia content corresponding to the corresponding text segment in time to generate a fusion recording result; and display the recording result.

2. The method of claim 1, wherein, The multimedia segment comprises an audio segment; The response to the selection operation to display a multimedia recording picture comprises: In response to the selection operation, start an audio acquisition component, the audio acquisition component being used to acquire an audio signal; Display the multimedia recording picture, the multimedia recording picture being used to indicate the recording progress of the audio segment.

3. The method of claim 2, wherein, The multimedia content comprises audio content; The method further comprises: In response to the end of recording for the text segment, embed the audio segment into a part of the audio content corresponding to the text segment; Display the recording result.

4. The method of claim 3, wherein, The embedding of the audio segment into a part of the audio content corresponding to the text segment comprises: Perform voice de-noising processing on a part of the audio content corresponding to the text segment to obtain a sound track audio; and add the audio segment to a part of the sound track audio corresponding to the text segment; Or, Overlay the audio segment on a part of the audio content corresponding to the text segment; Or, Superimpose the audio segment on a part of the audio content corresponding to the text segment.

5. The method of claim 1, wherein, The multimedia segment comprises a video segment; The response to the selection operation to display a multimedia recording picture comprises: In response to the selection operation, start a video stream acquisition component, the video stream acquisition component being used to acquire a video picture; Display the multimedia recording picture, the multimedia recording picture being used to indicate the recording progress of the video segment.

6. The method according to any one of claims 1 to 5, characterized in that, The response to the selection operation to display a multimedia recording picture comprises: In response to the selection operation, display a recording control, the recording control being used to indicate multimedia recording based on the text segment; In response to receiving a trigger operation on the recording control, display the multimedia recording picture.

7. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: In response to the end of recording for the text segment, display processing candidates, the processing candidates comprising a result display candidate and a play candidate; Receive a first selection operation on the result display candidate; Based on the first selection operation, display the recording result.

8. The method of claim 7, wherein, The method further comprises: Receive a second selection operation on the play candidate; Based on the second selection operation, continue to play the multimedia content.

9. The method according to any one of claims 1 to 5, characterized in that, The method comprises the following steps: In the process of playing the multimedia content, the text prompt content corresponding to the multimedia content is displayed synchronously and scrolls in time sequence. During the playing of the multimedia content, a multimedia playing picture is displayed, and the multimedia playing picture comprises multimedia information of the multimedia content.

10. The method according to any one of claims 1 to 5, characterized in that, In response to receiving a text prompt operation, a text prompt picture is displayed, and the text prompt picture comprises the text prompt content. The receiving of the selection operation on one or more text segments in the scrolling text prompt content comprises: Receiving a long press operation on the text segment in the text prompt content, and taking the long press operation as the selection operation. Or, 11. A multimedia recording device, characterized by Receiving a recording trigger operation, displaying a selection control at a position corresponding to the text prompt content based on the recording trigger operation, and taking a trigger operation on the selection control as the selection operation. The device comprises: A display module configured to display, synchronously with playing of multimedia content, text prompt content corresponding to the multimedia content and scrolling in time sequence. A receiving module configured to receive, during playing of the multimedia content, a selection operation on one or more text segments in the scrolling text prompt content, the text segments comprising text prompt content that is not continuous in time. The display module is further configured to display, in response to the selection operation, a multimedia recording picture, the multimedia recording picture being configured to indicate recording of a multimedia segment for each selected text segment.

12. The apparatus of claim 11, wherein, The display module is further configured to, in response to the recording of the text segment ending, overlay or partially overlay or superimpose each multimedia segment obtained by recording to a portion of the multimedia content corresponding in time to the corresponding text segment, to generate a fused recording result, and display the recording result. The multimedia segment comprises an audio segment. The device further comprises: An opening module configured to open, in response to the selection operation, an audio acquisition component configured to acquire an audio signal.

13. The apparatus of claim 12, wherein, The display module is further configured to display the multimedia recording picture, the multimedia recording picture being configured to indicate a recording progress of the audio segment. The multimedia content comprises audio content. The device further comprises: An editing module configured to, in response to the recording of the text segment ending, embed the audio segment in a portion of the audio content corresponding to the text segment.

14. A computer device, comprising: The display module is further configured to display the recording result. The computer device comprises a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the multimedia recording method according to any one of claims 1 to 10.

15. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set or an instruction set, which are loaded and executed by the processor to implement the multimedia recording method according to any one of claims 1 to 10.

16. A computer program product, characterised in that, The computer program product comprises computer instructions, which are executed by the processor to implement the multimedia recording method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Video generating method and video generating device

    CN104967900A

  • Karaoke processing method and apparatus

    CN105006234A

  • Voice synthesis method and system

    CN108269560A