Audio and video processing method and device, electronic equipment and storage medium
By sending frame data to the server for storage during audio and video acquisition, and displaying local frames on the client and replacing server frames when publishing conditions are detected, the problem of storage and transmission lag during video publishing is solved, and efficient audio and video publishing is achieved.
Patent Information
- Application Number
- CN202311257480.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-26
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-09-26
AI Technical Summary
In existing technologies, audio and video data need to be stored both locally and on the server during the video publishing process, which leads to increased encoding time, increased storage space consumption, and insufficient data transmission lag and real-time performance.
During the audio and video acquisition process, frame data is sent to the server for storage. When the publishing conditions are detected, the locally stored audio and video frames are displayed on the client. At the same time, the client receives audio and video frames sent by the server to replace the locally displayed frames.
It enables rapid processing and efficient publishing of audio and video data, reducing user waiting time and improving user experience.
Smart Images

Figure CN119729064B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of audio and video processing, and particularly relate to an audio and video processing method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the development of computer technology, more and more application software emerges as the times require, or corresponding application functions are added in the application software. Optionally, the application software or the added application function can be video shooting, video publishing and the like.
[0003] At present, the main process of video publishing is to trigger the camera function of the application software to shoot corresponding audio and video data, and store the audio and video data locally. Based on the triggering operation of the user on the publishing control, the audio and video data stored locally is uploaded to the server, so that the server distributes the audio and video data; or the user can edit the audio and video data, and send the edited audio and video data to the server, so that the server publishes the audio and video data.
[0004] However, the above-mentioned method has the following problems: the audio and video data needs to be stored locally and on the server, which increases the encoding time and storage space occupation, and further, the uploading of the audio and video data is limited by the network and the device, resulting in a certain lag in data transmission, and accordingly, the real-time performance of video publishing is also a problem. SUMMARY
[0005] The present disclosure provides an audio and video processing method and device, electronic equipment and storage medium, to realize fast and effective processing of audio and video data, so as to achieve the technical effect of real-time video publishing.
[0006] In a first aspect, the embodiments of the present disclosure provide an audio and video processing method applied in a client, comprising:
[0007] In the process of collecting audio and video, the audio and video frame is sent to the server, so that the server stores the audio and video frame;
[0008] When it is detected that the audio and video publishing condition is met, the locally stored audio and video frame is displayed on the audio and video collection client;
[0009] When the audio and video frame issued by the server is received, the audio and video frame displayed on the audio and video collection client is replaced based on the audio and video frame, and the audio and video frame is published.
[0010] In a second aspect, the embodiments of the present disclosure provide an audio and video processing method applied in a server, comprising:
[0011] receive and store the audio and video frames sent by the audio and video collection client; wherein the audio and video frames are sent in the process of collecting audio and video;
[0012] when receiving the audio and video publishing instruction, issue the stored audio and video frames to the audio and video collection client, to replace the locally stored audio and video frames displayed by the audio and video collection client based on the audio and video frames;
[0013] wherein the audio and video publishing instruction is determined based on the audio and video publishing condition.
[0014] In a third aspect, the embodiments of the present disclosure further provide an audio and video processing device, which is configured in a client, and the device comprises:
[0015] an audio and video collection module, configured to send audio and video frames to a server in the process of collecting audio and video, so that the server stores the audio and video frames;
[0016] an audio and video first playing module, configured to display locally stored audio and video frames in the audio and video collection client when detecting that the audio and video publishing condition is met;
[0017] an audio and video second playing module, configured to replace the audio and video frames displayed by the audio and video collection client based on the audio and video frames and publish the audio and video frames when receiving the audio and video frames issued by the server.
[0018] In a fourth aspect, the embodiments of the present disclosure further provide an audio and video processing device, which is configured in a server, and the device comprises:
[0019] an audio and video receiving module, configured to receive and store the audio and video frames sent by the audio and video collection client; wherein the audio and video frames are sent in the process of collecting audio and video;
[0020] an audio and video issuing module, configured to issue the stored audio and video frames to the audio and video collection client when receiving the audio and video publishing instruction, to replace the locally stored audio and video frames displayed by the audio and video collection client based on the audio and video frames;
[0021] wherein the audio and video publishing instruction is determined based on the audio and video publishing condition.
[0022] In a fifth aspect, the embodiments of the present disclosure further provide an electronic device, which comprises:
[0023] one or more processors;
[0024] a storage device, configured to store one or more programs,
[0025] The one or more programs, when executed by the one or more processors, cause the one or more processors to implement the audio and video processing method according to any of the embodiments of the present disclosure.
[0026] In a sixth aspect, the embodiments of the present disclosure further provide a storage medium containing computer executable instructions for executing the audio and video processing method according to any of the embodiments of the present disclosure when executed by a computer processor.
[0027] The technical solution provided by the embodiments of the present disclosure can send the audio and video frames to the server in the process of collecting the audio and video, so that the server stores the audio and video frames. Further, in order to avoid the problem that the user needs to wait for a certain period of time due to the delay of the server in issuing the audio and video frames, and thus causing poor user experience, the locally stored audio and video frames can be displayed in the audio and video collection client, and the audio and video frames displayed by the audio and video collection client can be replaced by the audio and video frames issued by the server, and the audio and video frames issued by the server can be published, thereby solving the problem of certain delay and hysteresis when the audio and video frames are uploaded at the time of publishing and then pushed to other clients by the server, and achieving the technical effect of efficient publishing of audio and video. BRIEF DESCRIPTION OF DRAWINGS
[0028] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0029] Figure 1 is a flowchart of an audio and video processing method provided by an embodiment of the present disclosure;
[0030] Figure 2 is a flowchart of an audio and video processing method provided by an embodiment of the present disclosure;
[0031] Figure 3 is a flowchart of an interaction between a server and a client provided by an embodiment of the present disclosure;
[0032] Figure 4 is a structural diagram of an audio and video processing device provided by an embodiment of the present disclosure;
[0033] Figure 5 is a structural diagram of an audio and video processing device provided by an embodiment of the present disclosure;
[0034] Figure 6 is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood.
[0036] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0037] The term “comprising” and variations thereof as used herein are open-ended, and mean “including but not limited to”. The term “based on” means “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related terms are defined in the description that follows.
[0038] It should be noted that the terms “first”, “second”, and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0039] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that “one or more” should be understood unless otherwise explicitly stated in the context.
[0040] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely used for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0041] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.
[0042] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware, such as an electronic device, an application program, a server or a storage medium, performing the operation of the technical solution of the present disclosure according to the prompt information.
[0043] As an optional but non-limiting implementation, in response to receiving an active request of a user, the prompt information can be sent to the user in the form of a pop-up window, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0044] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0045] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0046] Before introducing the technical solution, an exemplary application scenario can be described. The technical solution provided by the embodiments of the present disclosure can be applied to a scenario in which any client collects audio and video data and publishes the audio and video data so that other users can see the audio and video data, that is, as long as the interaction between the client and the server is involved in the audio and video processing process, the solution provided by the embodiments of the present disclosure can be used to implement it, for example, it can be applied to a scenario in which the client collects audio and video and the server distributes the audio and video data. Optionally, the application scenario can be a short video creation and publishing scenario, a special effect video shooting and publishing scenario, etc.
[0047] It should be noted that the above described application scenario is only a simple explanation of the application scenario of the solution, and is not a limitation of the application scenario, and can also be applied to other application scenarios.
[0048] It should also be noted that the audio and video processing apparatus provided by the embodiments of the present disclosure can be integrated in an application software supporting video and audio processing function, and the software can be installed in an electronic device. Optionally, the electronic device can be a mobile terminal or a PC terminal, etc. The application software can be a software for audio and video processing, and the specific application software is not described here, as long as it can realize audio and video processing. Optionally, the audio and video processing can be audio and video collection or adding corresponding content to the audio and video, which can be text, stickers, music, filters, etc.
[0049] Figure 1 is a flowchart of an audio and video processing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to a case where any client collects audio and video and a server distributes the audio and video. The method can be executed by an audio and video processing device. The device can be implemented in the form of software and / or hardware. Optionally, the device is implemented by an electronic device, which can be a mobile terminal, a PC terminal, or a server.
[0050] As shown in Figure 1 , the method comprises the following steps.
[0051] In S110, in the process of collecting audio and video, an audio and video frame is sent to a server, so that the server stores the audio and video frame.
[0052] The collection of the audio and video can be implemented by a client, which can be an installed application software. The application software can call a camera and a microphone array of a terminal device to collect corresponding audio and video. The server corresponding to the client can be understood as a server for processing the audio and video.
[0053] Generally, the processing and synthesis of the audio and video are performed by the server. Therefore, when the client completes the collection of the video, the audio and video needs to be sent to the server so that the server processes the audio and video. However, when the audio and video is collected locally and then sent to the server, it is affected by the network, the device to which the client belongs, and the size of the video. Therefore, the collected audio and video frame data can be encoded and sent to the server in the process of collecting the audio and video, that is, the collected audio and video frame data is sent to the server in the process of collecting the audio and video, so as to avoid the problem that when the audio and video is published, it needs to be uploaded to the server, and when the data amount is large, the problem of large lag.
[0054] Specifically, the corresponding camera and / or microphone array can be called based on a target function in the client to collect corresponding audio and video frames, and the collected audio and video frames are sent to the server so that the server stores the audio and video frames, thereby quickly distributing the audio and video frames when the audio and video frames are published.
[0055] For example, the embodiment can be applied in a short video shooting or special effect video shooting scene. A user can open an application software A and trigger a video shooting control in the application software A to start a camera and a microphone array. The audio and video frames including a target object are collected based on the camera and the microphone array, and the audio and video frames are sent to the server in the process of shooting.
[0056] The audio and video frames can be directly sent to the server, or the audio and video frames can be uploaded to the server after being stream encapsulated.
[0057] In this embodiment, the audio and video frames can be processed in a streaming encapsulation manner. This has the advantage that the video can be uploaded while being captured. Furthermore, when the video is published or played, it can be played while being downloaded, so that the user can continue to download the subsequent video frames while watching the previous video frames, thereby avoiding the problem that the video cannot be played until all the video frames are downloaded.
[0058] S120, when it is detected that the audio and video publishing condition is met, the locally stored audio and video frames are displayed on the audio and video collection client.
[0059] The audio and video publishing condition includes triggering the audio and video publishing control, or the shooting time of the audio and video reaches a preset time threshold, or a gesture for triggering the audio and video publishing is detected, etc. The specific publishing condition is not limited in this embodiment, as long as it can inform the server that the audio and video needs to be published. The audio and video collection client is a client for collecting audio and video frames.
[0060] It should be noted that during the audio and video collection process, it is usually necessary to store it locally first in order to edit the audio and video frames, or to preview the audio and video based on the locally stored audio and video frames. That is, when collecting audio and video frames, it is necessary to store them locally first.
[0061] Specifically, when it is detected that the audio and video publishing condition is met, the locally stored audio and video frames can be played on the audio and video collection client. This has the advantage that the server needs a certain waiting time to issue the audio and video frames. During the waiting process, the locally stored audio and video can be displayed, thereby achieving the effect of no waiting for the user.
[0062] In this embodiment, if the user does not add a corresponding special effect element to the collected audio and video frames, the steps of S110 to S120 can be directly executed to realize the collection and publishing of the audio and video frames. If a corresponding special effect element needs to be added to the audio and video, the following steps can be performed before detecting that the audio and video publishing condition is met: adding at least one special effect element to the locally stored audio and video frames to update the locally stored audio and video frames; recording element association information of the at least one special effect element and adding attribute information in the audio and video frames, so as to fuse the at least one special effect element into the audio and video frames stored in the server based on the element association information and the adding attribute information; wherein the element association information includes a resource package identifier of the special effect element, and the adding attribute information includes at least one of position information and timestamp information.
[0063] The number of the at least one special effect element can be one or more, and the number of the specific special effect element can be matched with the actual demand of the user. The at least one special effect element can be a map, text, music, filter, or beautifying, etc. That is, the special effect element can be understood as any element that is different from the original audio / video frame. In the process of adding the special effect element, the element association information corresponding to each special effect element can be recorded. The element association information can include a resource package identifier, that is, an identifier for calling the special effect element. It can be understood that all special effect elements can be stored in a database, and the resource package corresponding to each special effect element exists in a unique identifier corresponding thereto. The unique identifier is the resource package identifier. At the same time, the adding attribute of each special effect element in the audio / video frame can also be recorded. The adding attribute information can include position information and / or timestamp information. The position information can be understood as the specific adding position of the special effect element in the audio / video frame, and the specific adding position can be represented by a pixel coordinate. The adding timestamp information can be the playback timestamp of the special effect element in the audio / video frame, that is, the time information of the application of the special effect element in the audio / video frame. The time information is a relative time, that is, the application timestamp of the special effect element is counted from the start frame. That is, the adding attribute information is the application information of the special effect element in the audio / video frame.
[0064] It also needs to be noted that the adding attribute information can be one or both of the timestamp information and the position information, because the element attributes corresponding to different special effect elements are different, and therefore the above adding attribute also has certain differences. For example, if the music is added, there can be no position information.
[0065] Specifically, if it is necessary to edit the operation of the audio / video frame, and the editing operation is to add the special effect element, the locally stored audio / video frame can be called, and the special effect element can be added in the audio / video frame. In the process of adding the special effect element, the resource package identifier corresponding to each special effect element can be recorded, that is, the resource for calling the special effect element is recorded. At the same time, the use position and / or use timestamp of the special effect element in the audio / video frame can also be recorded.
[0066] For example, if the special effect element added at the playback timestamp of the audio and video frame is 00:05-00:10 is background music A, then the element association information corresponding to background music A can be recorded as resource package identifier ID1, and the added attribute information includes timestamp information, and the timestamp information is 00:05-00:10; or, for example, if the special effect element added at the playback timestamp of the audio and video frame is 00:15 and the pixel coordinates are (500,500) (the pixel coordinates at this time are the center point of the text box) is text box B, and text content XXX is entered, then the resource package identifier ID5 of text box B can be recorded, and the added attribute information includes timestamp information 00:15 and position information (500,500). It should also be noted that multiple special effect elements can be added to an audio and video frame, and they only need to be recorded separately when recording.
[0067] Typically, after adding special effect elements, the video can be published. If the video publishing conditions are met, the server storage identifier of the audio and video frame, as well as the element association information and corresponding added attribute information of the at least one special effect element are sent to the server, so that the server retrieves the corresponding audio and video frame based on the server storage identifier, and adds special effect elements to the audio and video frame based on the element association information and the added attributes.
[0068] During the audio and video capture process, the captured audio and video can be uploaded to the server while being captured, so that the captured audio and video can be stored on the server. In order to facilitate the synchronization of special effect elements to the server when the user edits the audio and video, a storage identifier, i.e., a server storage identifier, can be fed back to the client when the server first receives and stores the audio and video frame. It can be understood that the server storage identifier is the storage address identifier of the stored audio and video frame, that is, the address of the stored audio and video frame. Based on this storage address identifier, the stored audio and video frame can be retrieved from the corresponding location.
[0069] In this embodiment, the element association information and corresponding added attribute information for all special effect elements added to the audio and video frame can be packaged into a single data packet and sent to the server. That is, only the incremental information that differs from the captured audio and video frame needs to be packaged and sent to the client, avoiding the current practice of sending all audio and video frames with added special effect elements to the server, which results in large data volumes, low transmission efficiency, and long latency. The server can then fuse the special effect elements into the audio and video frames based on the received data packets to obtain the audio and video frames with added special effect elements.
[0070] S130: When receiving the audio and video frames sent by the server, replace the audio and video frames displayed by the audio and video acquisition client based on the audio and video frames and publish the audio and video frames.
[0071] It can be understood that in the case of playing the locally stored audio and video frames, if the server-side issued audio and video frames are received, the server-side issued audio and video frames can replace the audio and video frames currently played by the audio and video frame collection client. At the same time, the server-side issued audio and video frames are published, that is, the server-side issued audio and video frames are visible to other clients.
[0072] That is, there is a certain time length in the process of the server issuing the audio and video frames. In order to improve the user experience, that is, the user does not need to wait for a certain time length, the locally stored audio and video frames can be preferentially displayed in the audio and video collection client. When the server issues the audio and video frames, they can be replaced to achieve a user-perception-free effect, thereby improving the user experience. At this time, the audio and video has been published for the user corresponding to the audio and video collection client, but the published audio and video frames have not been seen by the users of other clients.
[0073] It should be noted that if the server receives the data packet, the server needs to obtain the corresponding special effect resource package according to the content in the data packet and fuse the special effect resource package with the stored audio and video frames, and then issue them, which requires a certain waiting time length, resulting in poor user experience. The advantage of displaying the locally stored audio and video frames in the audio and video collection client is that the audio and video frames can be quickly previewed and played. For the users of the audio and video collection client, the audio and video frames have been successfully published and are in a browsable state, which weakens the waiting time length of the user, thereby improving the user experience.
[0074] In this embodiment, the way of replacing the audio and video frames currently displayed by the audio and video collection client based on the server-side issued audio and video frames can be: based on the display timestamp of the audio and video collection client, the server-side issued audio and video frames are displayed.
[0075] That is, in order to improve the smoothness of video playback and the user's perception-free, after the server issues the audio and video, the display timestamp corresponding to the currently played audio and video frame by the audio and video collection client can be obtained, that is, the audio and video frame corresponding to which time point is currently displayed. The issued audio and video frames can also be played at the corresponding display timestamp to achieve the effect of smooth playback.
[0076] The technical scheme provided by the embodiments of the present disclosure can send the audio and video frames to the server in the process of collecting the audio and video, so that the server stores the audio and video, and further, when the audio and video publishing condition is detected, in order to avoid the delay of the server in issuing the audio and video frames, causing the user to wait for a certain period of time, and thus causing the user to have a poor user experience, the audio and video frames stored locally can be displayed in the audio and video collection client, and the audio and video frames displayed by the audio and video collection client can be replaced in the audio and video frames issued by the server, and the audio and video frames issued by the server can be published, solving the problem of a certain delay and hysteresis when the audio and video frames are uploaded at the time of publishing and then pushed to other clients by the server, and achieving the technical effect of efficient publishing of audio and video.
[0077] Figure 2 is a flow diagram of an audio and video processing method provided by the embodiments of the present disclosure. The embodiments of the present disclosure are applicable to any situation where a client collects audio and video and a server distributes audio and video. The implementation of the technical scheme of the present embodiment can be performed by the server. The method can be performed by an audio and video processing device. The device can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, a PC terminal, or a server, etc. The same or corresponding technical terms as described above will not be described here.
[0078] As shown in Figure 2 , the method comprises:
[0079] S210, receiving and storing the audio and video frames sent by the audio and video collection client; wherein the audio and video frames are sent in the process of collecting the audio and video.
[0080] It can be understood that the collected audio and video frames can be sent to the server in the process of collecting the audio and video. The server can receive the audio and video frames sent by the audio and video collection client.
[0081] When the server receives the audio and video frames, it can store the audio and video frames. In order to facilitate subsequent tracking of the audio and video frames, the server storage identifier can be determined based on the storage address of the audio and video frames, and the server storage identifier can be fed back to the audio and video collection client, so that when the audio and video collection client sends the audio and video frames, the server storage identifier can be carried, so that the server stores the continuously collected audio and video frames in the corresponding storage address based on the server storage identifier. At the same time, if there is a corresponding editing operation in the client, in order to achieve the effect of synchronization, the corresponding audio and video frames can be retrieved based on the server storage identifier, and the same processing as the editing operation of the client can be performed. Optionally, the server storage identifier for storing the audio and video frames is fed back to the audio and video collection client, so that the corresponding audio and video frames can be retrieved from the server based on the server storage identifier.
[0082] S220, in response to receiving the audio / video publishing instruction, the stored audio / video frame is sent to the audio / video collection client to replace the locally stored audio / video frame displayed by the audio / video collection client based on the audio / video frame.
[0083] The audio / video publishing instruction is generated when the audio / video publishing condition is met. That is, when the audio / video publishing condition is met, the audio / video publishing instruction can be generated and sent to the server, and the server can receive the audio / video publishing instruction.
[0084] Specifically, when the server receives the audio / video publishing instruction, the server can send the stored audio / video frame to the audio / video collection client and push the audio / video frame to other clients. The specific pushing method is not limited in this embodiment. When the audio / video collection client receives the audio / video frame sent by the server, the currently displayed audio / video frame of the audio / video collection client can be replaced by the audio / video frame sent by the server. At the same time, in order to ensure the continuity of the audio / video frame playing, the display timestamp of the played audio / video frame of the audio / video collection client can be determined, and the audio / video frame sent by the server can be played based on the display timestamp.
[0085] It should be noted that before the server sends the audio / video frame, the audio / video frame displayed by the audio / video collection client is the audio / video frame stored locally by the audio / video collection client.
[0086] Optionally, the server storage identifier, the element association information of the at least one special effect element and the corresponding addition attribute information sent by the audio / video collection client are received; based on the server storage identifier, the element association information and the corresponding addition attribute information, at least one special effect element is added to the audio / video frame; wherein the at least one special effect element is generated by the editing operation of the audio / video frame on the audio / video collection client, the element association information includes the resource package identifier of the special effect element, and the addition attribute information is the position information and / or timestamp information of the special effect element in the corresponding audio / video frame.
[0087] The editing operation on the audio / video frame can be an operation of adding a special effect element to the audio / video frame, an operation of deleting one or more frames, an operation of editing the audio / video frame, or an operation of removing a certain object P in the audio / video frame.
[0088] Specifically, the server storage identifier, the element association information of the at least one special effect element and the corresponding addition attribute information can be sent to the server. After the server receives such information, the corresponding audio / video frame can be called and the corresponding special effect element can be added.
[0089] In the embodiment, the corresponding special effect element is added to the audio and video frame in the server, which can be based on the server storage identifier, element association information and corresponding addition attribute information, at least one special effect element is added to the audio and video frame, including: calling the audio and video frame based on the server storage identifier; calling the corresponding special effect resource package based on the resource package identifier in the element association information; determining the display time and / or display position of the corresponding special effect element in the audio and video frame based on the timestamp information and / or position information of the at least one special effect element; based on the display time and / or display position, adding the special effect element in the special effect resource package to the audio and video frame to obtain the audio and video frame added with at least one special effect element.
[0090] The technical scheme provided by the embodiments of the present disclosure can send the audio and video frame to the server in the process of collecting the audio and video, so that the server stores the audio and video, and further, when the audio and video publishing condition is detected, in order to avoid the delay of the server issuing the audio and video frame, causing the user to wait for a certain period of time, and thus causing the user to have a poor user experience, the locally stored audio and video frame can be displayed in the audio and video collection client, and the audio and video frame displayed by the audio and video collection client can be replaced in the received audio and video frame issued by the server, and the audio and video frame issued by the server can be published, solving the problem of certain delay and hysteresis when the audio and video frame is uploaded at the time of publishing and the server pushes it to other clients, and achieving the technical effect of efficient publishing of audio and video.
[0091] Figure 3 The flowchart of the interaction between the server and the client provided by the embodiments of the present disclosure can be explained based on the order processing between the server and the client, and the specific implementation can be referred to the detailed description of the embodiments, wherein the same or corresponding technical terms as the above embodiments are not repeated here.
[0092] As Figure 3 shown, the method comprises:
[0093] S310, in the process of collecting the audio and video frame in the audio and video collection client, storing the audio and video frame locally and sending the audio and video frame to the server.
[0094] It can be understood that in the process of shooting the audio and video frame, a real-time encoding output of one audio and video stream push (TS stream) to the server can be understood as a stream encapsulation of the audio and video, which supports playing only by downloading a part. At the same time, the shot audio and video frame is stored locally. This way is different from the current way of only storing the collected audio and video frame locally and uploading it to the server when publishing.
[0095] S320: The server stores the received audio and video frames and feeds back a server storage identifier to the audio and video acquisition client.
[0096] It can be understood that after the server receives the audio and video frame, it can store it on the server and generate a server storage identifier corresponding to the audio and video frame. The server storage identifier can be the storage address of the audio and video frame or the storage connection corresponding to the audio and video frame. Furthermore, after generating the server storage identifier, the server storage identifier can be sent to the audio and video capture client.
[0097] S330: Add at least one special effect element to the locally stored audio and video frames, and determine element association information and added attribute information corresponding to the at least one special effect element.
[0098] It can be understood that on the editing page, at least one special effect element can be added to at least one locally stored audio or video frame. At the same time, the element association information and added attribute information corresponding to the special effect element can be recorded. Based on the element association information and added attribute information of each special effect element, a corresponding data packet is generated.
[0099] S340: When it is detected that the audio and video publishing conditions are met, the locally stored audio and video frames are played on the audio and video acquisition client, and at the same time, the server storage identifier, element association information of at least one special effect element and corresponding added attribute information are sent to the server.
[0100] It can be understood that when it is detected that the audio and video publishing conditions are met, the locally stored audio and video frames can be played on the audio and video acquisition client. If a special effect element is added, the audio and video frame played may be an audio and video frame with the special effect element added. If no special effect element is added, the audio and video frame played may be an audio and video frame without the special effect element added. Furthermore, the server storage identifier and the data packet generated based on the element association information and corresponding added attribute information of at least one special effect element can be sent to the server. This method is different from the prior art, in which the captured audio and video frames are stored locally, and the editing process uses this audio and video frame for editing. After the editing is completed, a new video is synthesized locally. When the newly synthesized video is uploaded to the server during publishing, there is a problem of a large amount of uploaded data, which in turn causes a publishing delay. This method achieves uploading only incremental data (element association information and added attribute information of special effect elements), greatly reducing the amount of data uploaded and thus improving data processing efficiency.
[0101] S350: The server retrieves the stored audio and video frames based on the server storage identifier, and adds corresponding special effect elements to the audio and video based on the element association information and addition attributes of at least one special effect element.
[0102] It can be understood that the server can call the stored audio and video frames based on the received server storage identifier, call the corresponding special effect resource package based on the resource package identifier in the element association information of the special effect element, and add the special effect element corresponding to the special effect resource package to the audio and video frames based on the position information and / or timestamp information in the addition attribute information. Since the performance of the server is good and the processing efficiency is high, the special effect element can be quickly fused into the audio and video frames to obtain the published audio and video, greatly shortening the publishing time and saving time.
[0103] S360, based on the server, the audio and video frames with added special effect elements are sent to the audio and video collection client to replace the audio and video frames played by the audio and video collection client.
[0104] It can be understood that after the server is rendered, the audio and video frames with added special effect elements can be sent to the audio and video collection client. After the client receives the audio and video frames sent by the server, the audio and video frames played by the audio and video collection client can be replaced. When replacing, the audio and video frames can be played according to the playback timestamp of the audio and video collection client to jump to the playback position corresponding to the audio and video frames sent to continue playing the audio and video frames.
[0105] S370, publish the audio and video.
[0106] It can be understood that when the server sends the audio and video frames to the audio and video collection client, it can be understood that the audio and video frames are published, at which time other clients can browse the audio and video frames.
[0107] The technical scheme provided by the embodiments of the present disclosure can send the audio and video frames to the server during the collection of the audio and video, so that the server stores the audio and video. Further, when the audio and video publishing condition is detected, in order to avoid the delay of the server sending the audio and video frames, causing the user to wait for a certain period of time, and thus causing the user to have a poor user experience, the local stored audio and video frames can be displayed in the audio and video collection client, and the audio and video frames displayed by the audio and video collection client can be replaced in the received audio and video frames sent by the server, and the audio and video frames sent by the server can be published, solving the problem of certain delay and hysteresis when the audio and video frames are uploaded during publishing and the server pushes the audio and video frames to other clients, and achieving the technical effect of efficient publishing of audio and video.
[0108] Figure 4 is a schematic diagram of an audio and video processing device provided by an embodiment of the present disclosure. The device is configured in a client, as shown in Figure 4 The device includes an audio and video collection module 410, an audio and video first playback module 420, and an audio and video second playback module 430.
[0109] The audio and video collection module 410 is configured to send the audio and video frames to the server during the process of collecting the audio and video, so that the server stores the audio and video frames; the audio and video first playing module 420 is configured to display the locally stored audio and video frames on the audio and video collection client when it is detected that the audio and video publishing condition is met; and the audio and video second playing module 430 is configured to replace the audio and video frames displayed on the audio and video collection client based on the audio and video frames sent by the server and publish the audio and video frames when the audio and video frames sent by the server are received.
[0110] On the basis of the above technical solutions, the device comprises a special effect element editing module, which comprises:
[0111] The special effect adding unit is configured to add at least one special effect element to the locally stored audio and video frames to update the locally stored audio and video frames; the information recording unit is configured to record element association information of the at least one special effect element and adding attribute information in the audio and video frames, so as to fuse the at least one special effect element into the audio and video frames stored by the server based on the element association information and the adding attribute information; wherein the element association information comprises a resource package identifier of the special effect element, and the adding attribute information comprises at least one of position information and a timestamp.
[0112] On the basis of the above technical solutions, the device further comprises:
[0113] The data sending module is configured to send the server storage identifier of the audio and video frames, and the element association information and the corresponding adding attribute of the at least one special effect element to the server, so that the server calls the corresponding audio and video frames based on the server storage identifier, and adds special effect elements to the audio and video frames based on the element association information and the adding attribute.
[0114] On the basis of the above technical solutions, the audio and video frames sent by the server are audio and video frames to which at least one special effect element is added, and the audio and video second playing module is further configured to display the audio and video frames sent by the server based on a display timestamp of the audio and video collection client.
[0115] The technical solution provided by the embodiments of the present disclosure can send the audio and video frames to the server in the process of collecting the audio and video, so that the server stores the audio and video frames. Further, in order to avoid the problem that the user needs to wait for a certain period of time due to the delay of the server in issuing the audio and video frames, thereby causing poor user experience, the audio and video frames stored locally can be displayed in the audio and video collection client, and the audio and video frames displayed by the audio and video collection client can be replaced by the audio and video frames issued by the server, and the audio and video frames issued by the server can be published, thereby solving the problem of certain delay and hysteresis when the audio and video frames are uploaded at the time of publishing and then pushed to other clients by the server, and achieving the technical effect of efficient publishing of audio and video.
[0116] The audio and video processing device provided by the embodiments of the present disclosure can execute the audio and video processing method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.
[0117] It should be noted that each unit and module included in the above device is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for convenient mutual distinction, and does not limit the protection scope of the embodiments of the present disclosure.
[0118] Figure 5 is a structural schematic diagram of an audio and video processing device provided by the embodiments of the present disclosure, which is configured in a server, as shown in Figure 5 The device includes an audio and video receiving module 510 and an audio and video issuing module 520.
[0119] The audio and video receiving module 510 is configured to receive and store the audio and video frames sent by the audio and video collection client; wherein the audio and video frames are sent in the process of collecting the audio and video; the audio and video issuing module 520 is configured to issue the stored audio and video frames to the audio and video collection client when receiving the audio and video publishing instruction, so as to replace the locally stored audio and video frames displayed by the audio and video collection client based on the audio and video frames; wherein the audio and video publishing instruction is determined based on the audio and video publishing condition.
[0120] On the basis of the above technical solutions, the device further includes a data feedback module configured to feed back the server storage identifier for storing the audio and video to the audio and video collection client, so as to call the corresponding audio and video frames from the server based on the server storage identifier.
[0121] On the basis of the above technical solutions, the device further includes:
[0122] The attribute information receiving module is configured to receive the server storage identifier sent by the audio / video collection client, element association information of at least one special effect element, and corresponding adding attribute information. The special effect synthesizing module is configured to add at least one special effect element to the audio / video frame based on the server storage identifier, the element association information, and the corresponding adding attribute information. The at least one special effect element is an element added to the audio / video frame on the audio / video collection client. The element association information includes a resource package identifier of the special effect element. The adding attribute information is position information and / or timestamp information of the special effect element in the corresponding audio / video frame.
[0123] On the basis of the above technical solutions, the special effect synthesizing module comprises:
[0124] The audio / video calling unit is configured to call the audio / video frame based on the server storage identifier. The resource package calling unit is configured to call a corresponding special effect resource package based on the resource package identifier in the element association information. The display information determining unit is configured to determine display time and / or display position of a corresponding special effect element in the audio / video frame based on timestamp information and / or position information of at least one special effect element. The special effect adding unit is configured to add a special effect element in the special effect resource package to the audio / video frame based on the display time and / or display position, so as to obtain an audio / video frame to which at least one special effect element is added.
[0125] The technical solution provided by the embodiments of the present disclosure can send an audio / video frame to a server in the process of collecting an audio / video, so that the server stores the audio / video. Further, in order to avoid a certain delay in the server issuing an audio / video frame, which causes a user to wait for a certain length of time and thus causes a poor user experience, the embodiments of the present disclosure can display a locally stored audio / video frame in an audio / video collection client, replace the audio / video frame displayed by the audio / video collection client in the audio / video frame issued by the server, and publish the audio / video frame issued by the server, thereby solving the problem of a certain delay and hysteresis when an audio / video frame is uploaded at the time of publishing and the server pushes the audio / video frame to other clients, and achieving the technical effect of efficiently publishing an audio / video.
[0126] The audio / video processing apparatus provided by the embodiments of the present disclosure can execute the audio / video processing method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.
[0127] It should be noted that each unit and module included in the above apparatus is only divided according to a function logic, but is not limited to the above division, as long as the corresponding function can be implemented; in addition, the specific names of each functional unit are only for convenient mutual distinction, and do not serve to limit the protection scope of the embodiments of the present disclosure.
[0128] Figure 6 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 6 , which shows an electronic device (eg Figure 6 The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0129] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An edit / output (I / O) interface 605 is also connected to the bus 604.
[0130] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0131] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication apparatus 609, or installed from the storage apparatus 608, or installed from the ROM 602. When the computer program is executed by the processing apparatus 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
[0132] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0133] The electronic device provided by the embodiments of the present disclosure and the audio and video processing method provided by the above-mentioned embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above-mentioned embodiments, and the present embodiment has the same beneficial effects as the above-mentioned embodiments.
[0134] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the audio and video processing method provided by the above-mentioned embodiments.
[0135] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, in which the computer-readable program code is contained. Such a propagated data signal can take any of a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the foregoing.
[0136] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0137] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and can be accessed via the electronic device.
[0138] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to:
[0139] In the process of collecting audio and video, the audio and video frames are sent to the server, so that the server stores the audio and video frames.
[0140] When it is detected that the audio and video publishing condition is met, the locally stored audio and video frames are displayed on the audio and video collection client.
[0141] When the audio and video frames issued by the server are received, the audio and video frames displayed on the audio and video collection client are replaced based on the audio and video frames and the audio and video frames are published.
[0142] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device:
[0143] Receive the audio and video frames sent by the audio and video collection client and store them; wherein the audio and video frames are sent in the process of collecting audio and video;
[0144] When the audio and video publishing instruction is received, the stored audio and video frames are issued to the audio and video collection client, so as to replace the locally stored audio and video frames displayed on the audio and video collection client based on the audio and video frames.
[0145] The audio and video publishing instruction is determined based on the audio and video publishing condition.
[0146] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as "C" or similar programming languages. Program code can execute entirely on a user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0147] The computer program product of the present disclosure can be a computer program embodied on a non-transitory computer readable medium. The body of computer program instructions can be read and executed by a processor to perform methods in accordance with the present disclosure.
[0148] The above description is only preferred embodiments of the present disclosure and the technical principles of the application. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0149] In addition, although each operation is described in a specific order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination.
[0150] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An audio-video processing method, characterized in that, Applied in the client, the method comprises: In the process of collecting audio and video, the audio and video frames are sent to the server to make the server store the audio and video frames; When detecting that the audio and video publishing condition is met, the locally stored audio and video frames are displayed on the audio and video collection client; When receiving the audio and video frames issued by the server, the audio and video frames displayed on the audio and video collection client are replaced based on the audio and video frames and the audio and video frames are published.
2. The method of claim 1, wherein, The audio and video frames are sent to the server, comprising: After the audio and video frames are stream encapsulated, they are uploaded to the server.
3. The method of claim 1, wherein, Before the detection that the audio and video publishing condition is met, it further comprises: At least one special effect element is added to the locally stored audio and video frames to update the locally stored audio and video frames; The element association information of the at least one special effect element and the adding attribute information in the audio and video frames are recorded to fuse the at least one special effect element into the audio and video frames stored in the server based on the element association information and the adding attribute information; The element association information comprises the resource package identifier of the special effect element, and the adding attribute information comprises at least one of the position information and the timestamp information.
4. The method of claim 3, wherein, After the detection that the audio and video publishing condition is met, it further comprises: The server storage identifier of the audio and video frames, and the element association information and the corresponding adding attribute information of the at least one special effect element are sent to the server to make the server call the corresponding audio and video frames based on the server storage identifier, and add special effect elements to the audio and video frames based on the element association information and the adding attribute information.
5. The method of claim 4, wherein, The audio and video frames issued by the server are the audio and video frames added with at least one special effect element, and when the audio and video frames displayed on the audio and video collection client are replaced based on the audio and video frames, it further comprises: Based on the display timestamp of the audio and video collection client, the audio and video frames issued by the server are displayed.
6. An audio-video processing method, characterized in that, Applied in the server, the method comprises: The audio and video frames sent by the audio and video collection client are received and stored; wherein the audio and video frames are sent in the process of collecting audio and video; When receiving the audio and video publishing instruction, the stored audio and video frames are issued to the audio and video collection client to replace the locally stored audio and video frames displayed on the audio and video collection client based on the audio and video frames; The audio and video publishing instruction is generated when detecting that the audio and video publishing condition is met.
7. The method of claim 6, wherein, After storing the audio and video frames, it further comprises: The server storage identifier of storing the audio and video frames is fed back to the audio and video collection client to call the corresponding audio and video frames from the server based on the server storage identifier.
8. The method of claim 6, wherein, It further comprises: The server storage identifier sent by the audio and video collection client and the element association information and the corresponding adding attribute information of at least one special effect element are received; At least one special effect element is added to the audio and video frames based on the server storage identifier, the element association information and the corresponding adding attribute information; The at least one special effect element is generated by an editing operation on an audio / video frame on the audio / video collection client, and the element association information includes a resource package identifier of the special effect element. The addition attribute information is position information and / or timestamp information of the special effect element in the corresponding audio / video frame.
9. The method of claim 8, wherein, The adding of the at least one special effect element to the audio / video frame based on the server storage identifier, the element association information and the corresponding addition attribute information includes: The audio / video frame is called based on the server storage identifier. The corresponding special effect resource package is called based on the resource package identifier in the element association information. The display time and / or display position of the corresponding special effect element in the audio / video frame are determined based on the timestamp information and / or position information of the at least one special effect element. The special effect element in the special effect resource package is added to the audio / video frame based on the display time and / or display position, so as to obtain the audio / video frame to which the at least one special effect element is added.
10. An audio-video processing apparatus, characterized by comprising: The device is configured in a client and includes: An audio / video collection module configured to send an audio / video frame to a server in the process of collecting an audio / video, so that the server stores the audio / video frame. An audio / video first playing module configured to display a locally stored audio / video frame on an audio / video collection client when it is detected that an audio / video publishing condition is met. An audio / video second playing module configured to replace the audio / video frame displayed on the audio / video collection client with the audio / video frame based on the audio / video frame received from the server and publish the audio / video frame.
11. An audio-video processing apparatus, characterized by comprising: The device is configured in a server and includes: An audio / video receiving module configured to receive and store an audio / video frame sent by an audio / video collection client, wherein the audio / video frame is sent in the process of collecting an audio / video. An audio / video issuing module configured to issue the stored audio / video frame to the audio / video collection client based on an audio / video publishing instruction, so as to replace the locally stored audio / video frame displayed on the audio / video collection client with the audio / video frame. The audio / video publishing instruction is determined based on an audio / video publishing condition.
12. An electronic device, comprising: The electronic device includes: One or more processors; A storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the audio / video processing method of any one of claims 1-5 or 6-9.
13. A storage medium containing computer executable instructions for performing the audio / video processing method of any one of claims 1-5 or 6-9 when executed by a computer processor.
Citation Information
Patent Citations
Issuing method, system and client of video microblog
CN102938775A
Video playing method and device, electronic equipment and storage medium
CN110446083A