Video processing method, device, electronic device and storage medium

By extracting highlight time materials from historical video clips in related live broadcast rooms, and using the start and end time to generate and push video clips, the problems of poor time-consuming and high cost of generating highlight time-sensitive video clips in live broadcast rooms are solved, real-time display and cost reduction are achieved.

CN117156224BActive Publication Date: 2025-07-08BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311155211.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2025-07-08
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

In the prior art, the generation of video clips at high-light moments in the live broadcast room is poor in time, resulting in poor short video push effect and high storage and transmission costs.

Method used

By extracting the highlight time video clip material from the historical video clip of the second live broadcast room where the conditions associated with the first live broadcast room are present, only the start and end times are stored and transmitted, and the highlight time video clip material is generated using the start and end times and pushed.

Benefits of technology

It realizes the display of instant highlight moment video clips, reduces storage and transmission costs, increases the number of visits in the live broadcast room, and improves the timeliness and coverage of video clip generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117156224B_ABST
    Figure CN117156224B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video processing method, apparatus, electronic device, and storage medium. The method includes: determining a first video segment corresponding to a first live broadcast room, determining start and end times of a second video segment corresponding to the first video segment, where the second video segment is highlight moment video segment material of the first video segment, and the second video segment is selected from a third video segment corresponding to a second live broadcast room as highlight moment video segment material that can improve the access volume of the second live broadcast room, and the live broadcast contents of the first live broadcast room and the second live broadcast room meet a preset association condition; and pushing the first video segment, the start and end times of the second video segment, and the third video segment to the first live broadcast room. This solution can realize the instant application of highlight moment video segment material in the live broadcast room for highlight display, and ensure the timeliness of the highlight moment video segment material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to data processing technologies, and in particular, to a video processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the continuous development of live broadcast technologies, many business parties choose short videos to push live broadcast rooms. To provide a better video push effect, some exciting moments or highlight video clips are extracted from the video. The highlight video clips can be relatively exciting or attractive clips in the video. However, usually, editors watch the video to manually edit the highlight video clips in the video, resulting in poor timeliness of the highlight video clips and inability to be used immediately, which leads to a poor improvement effect on the access volume of the live broadcast room. Summary of the Invention

[0003] The present disclosure provides a video processing method, apparatus, electronic device, and storage medium to quickly edit and generate appropriate highlight video clip materials to form short videos and improve the access volume of the live broadcast room.

[0004] In a first aspect, embodiments of the present disclosure provide a video processing method, the method including:

[0005] Determine a first video clip corresponding to a first live broadcast room, where the first video clip is real-time live video data of the first live broadcast room;

[0006] Determine start and end times of a second video clip corresponding to the first video clip, where the second video clip is highlight video clip material of the first video clip, and the second video clip is selected from a third video clip corresponding to a second live broadcast room and can improve the access volume of the second live broadcast room. The third video clip is historical live video data of the second live broadcast room, and the live broadcast contents of the first live broadcast room and the second live broadcast room meet a preset association condition;

[0007] Push the first live broadcast room based on the first video clip, the start and end times of the second video clip, and the third video clip.

[0008] In a second aspect, embodiments of the present disclosure further provide a video processing apparatus, the apparatus including:

[0009] A first determination module, configured to determine a first video clip corresponding to a first live broadcast room, where the first video clip is real-time live video data of the first live broadcast room;

[0010] A second determination module, configured to determine the start and end times of a second video segment corresponding to the first video segment, where the second video segment is a highlight video segment material of the first video segment, and the second video segment is a highlight video segment material selected from a third video segment corresponding to a second live broadcast room and capable of increasing the access volume of the second live broadcast room. The third video segment is historical live video data of the second live broadcast room, and the live broadcast contents of the first live broadcast room and the second live broadcast room satisfy a preset association condition;

[0011] A video processing module, configured to push to the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment.

[0012] In a third aspect, an electronic device is further provided in the embodiments of the present disclosure. The electronic device includes:

[0013] At least one processor; and

[0014] A memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores a computer program executable by the at least one processor. When the computer program is executed by the at least one processor, the at least one processor is enabled to execute the video processing method according to any one of the foregoing embodiments.

[0016] In a fourth aspect, a computer-readable medium is further provided in the embodiments of the present disclosure. The computer-readable medium stores computer instructions for causing a processor to execute the video processing method according to any one of the foregoing embodiments when executed.

[0017] In the embodiments of the present disclosure, when using short videos to increase the viewership of a live broadcast room, the first video segment corresponding to the first live broadcast room for which the viewership needs to be increased and the start and end times of the second video segment corresponding to the first video segment are determined. The second video segment is the highlight moment video segment material for increasing the viewership of the second live broadcast room in the third video segment corresponding to the second live broadcast room. The third video segment is the historical live video data of the second live broadcast room. The live broadcast content of the first live broadcast room and the second live broadcast room satisfies a preset association condition. Then, the first video segment, the start and end times of the second video segment, and the third video segment are used to push to the first live broadcast room. Since the second video segment used to increase the viewership of the first live broadcast room is derived from the third video segment of the second live broadcast room that has a preset association condition with the first live broadcast room, it can be ensured that the highlight moment video segment material indicated by the second video segment can fully reflect the core of the live broadcast content of the first live broadcast room, realizing the ability to immediately apply the second video segment for high-light display in the first live broadcast room, ensuring the timeliness of the highlight moment video segment material, and generating appropriate highlight moment video segment material without having to wait until the live broadcast of the first live broadcast room ends and review the live broadcast content. At the same time, when determining the second video segment, only the start and end times of the second video segment are generated and stored, and the second video segment is not directly generated, which can reduce the storage cost of the highlight moment video segment material and the transmission traffic cost of the highlight moment video segment material.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0020] Figure 1 is a schematic flowchart of a video processing method provided by an embodiment of the present disclosure;

[0021] Figure 2 is an interactive schematic diagram of a video processing process applicable to an embodiment of the present disclosure;

[0022] Figure 3 is a schematic flowchart of a video splicing process applicable to an embodiment of the present disclosure;

[0023] Figure 4 is another schematic flowchart of a video processing method provided by an embodiment of the present disclosure;

[0024] Figure 5 It is a schematic diagram for extracting the highlight video clip material in a video processing process applicable to the embodiments of the present disclosure;

[0025] Figure 6 It is a schematic structural diagram of a video processing device provided by the embodiments of the present disclosure;

[0026] Figure 7 It is a schematic structural diagram of an electronic device for implementing a video processing method provided by the embodiments of the present disclosure. Specific Embodiments

[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not used to limit the protection scope of the present disclosure.

[0028] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit steps shown. The scope of the present disclosure is not limited in this regard.

[0029] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0030] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0031] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0033] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.

[0034] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0035] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0036] It is understandable that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0037] It is understandable that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0038] Figure 1 The figure is a schematic flowchart of a video processing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of using short videos to increase the access volume of a live broadcast room. The method can be executed by a video processing device, and the video processing device can be implemented in the form of software and / or hardware and is generally integrated in any electronic device with network communication functions. The electronic device can be a mobile terminal, a PC terminal or a server, etc.

[0039] As shown in Figure 1 the figure, the video processing method of the embodiment of the present disclosure may include the following processes:

[0040] S110. Determine a first video segment corresponding to a first live broadcast room, where the first video segment is the real-time live video data of the first live broadcast room.

[0041] The business forms of live streaming include short videos for increasing the visit volume of the live streaming room and directly promoting the live streaming room. Among them, the short video enters the live streaming room by using live streaming advertisements in the recommended stream in the form of short videos to increase the visit volume of the live streaming room, with a superimposed conversion component and a breathing light. The main entrance to the live streaming room is the conversion component + clicking on the live streaming room icon; directly promoting the live streaming room means promoting the visit volume of the live streaming room in the form of directly promoting the live streaming room in the recommended stream. The whole screen is clickable, and the information stream directly pulls the live painting of the live streaming room in real time for placement, with a superimposed conversion card style to increase the click interest in the live streaming room and guide users to enter the live streaming room.

[0042] The first live streaming room can be a live streaming room that needs to increase the visit volume of the live streaming room in the form of short videos. The first video segment can be the video stream data generated by the first live streaming room during real-time live streaming.

[0043] S120. Determine the start and end time of the second video segment corresponding to the first video segment. The second video segment is the highlight moment video segment material of the first video segment. The second video segment is selected from the third video segment corresponding to the second live streaming room as the highlight moment video segment material that can increase the visit volume of the second live streaming room. The third video segment is the historical live streaming video data of the second live streaming room. The live streaming content of the first live streaming room and the second live streaming room meets the preset association conditions.

[0044] Generally, while using the method of directly promoting the live streaming room for placement, short videos are also made to guide the increase of the visit volume of the live streaming room. Many of these short videos are made by reprocessing the highlight moment video segment materials of the live streaming content, so as to better increase the visit volume of the live streaming room. However, when using short videos to increase the visit volume of the live streaming room, the following problems are found: the timeliness of short videos is relatively low, the generation of short videos requires a large time cost and it is difficult to be immediately available for use, and the storage cost after the generation of short videos is relatively large.

[0045] See Figure 2 , when generating the highlight moment video segment material for the first live streaming room, a second live streaming room that meets the preset association conditions with the first live streaming room in terms of live streaming content will be determined from multiple live streaming rooms that have conducted live streaming to form historical live streaming video data. Then, the highlight moment video segment material that can increase the visit volume of the second live streaming room can be extracted from the third video segment corresponding to the second live streaming room, and this video segment material is determined as the highlight moment video segment material corresponding to the first video segment. In this way, the second video segment directly extracted from the historical live streaming video data corresponding to the second live streaming room can be used as the highlight moment video segment material for the first live streaming room to be pushed for the first live streaming room. In this way, when the first live streaming room is in live streaming, the highlight moment video segment can be directly displayed on the spot, without waiting to watch the whole live stream after the live stream ends and then selecting frame by frame to generate the highlight moment video segment material for display, ensuring the timeliness of the highlight moment video segment material.

[0046] See Figure 2 When generating the second video segment containing the highlight video segment material for the first live broadcast room, it is necessary to store the second video segment corresponding to the first video segment so that the highlight video segment material can be sent to the client when requested by the client corresponding to the first live broadcast room. However, storing the second video segment will occupy a certain amount of storage resources. As the demand for highlight video segment material increases, the storage resources required will also increase. Moreover, as time goes by, the highlight video segment material required by the first live broadcast room will also change dynamically, further leading to a large amount of storage resources being occupied. Therefore, when determining the second video segment corresponding to the first video segment, only the start and end times of the second video segment in the third video segment need to be determined. In this way, there will be no additional large amount of storage resources occupied for the highlight video segment material, which can reduce the storage cost, effectively solve the storage resource pressure brought by the server generating the highlight video segment material, and since the start and end times of the second video segment are used, the transmission traffic cost for transmitting this part of the video segment material can be saved, and when the storage resource cost is reduced, the coverage rate of the highlight video segment material can be expanded.

[0047] S130. Push to the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment.

[0048] Since the second video segment is the highlight video segment material that can improve the access volume of the second live broadcast room selected from the third video segment corresponding to the second live broadcast room, and the first live broadcast room and the second live broadcast room meet the preset association conditions, the second video segment can be used as the highlight video segment material of the first video segment in the first live broadcast room. Therefore, the second video segment can be extracted from the third video segment through the start and end times of the second video segment, and then the extracted second video segment and the first video segment can be used to push to the first live broadcast room.

[0049] As an optional but non-limiting implementation manner, pushing to the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment includes the following steps A1 - A2:

[0050] Step A1. Send at least one start and end time of the second video segment to the client corresponding to the first live broadcast room, so that the client corresponding to the first live broadcast room extracts the second video segment from the third video segment based on the start and end time of the second video segment.

[0051] Step A2: Send the first video clip to the client corresponding to the first live stream, so that the client corresponding to the first live stream preprocesses the second video clip and the first video clip and displays them on the client corresponding to the first live stream.

[0052] See Figure 2 , when the client corresponding to the first live stream sends a request for the highlight moment video clip material of the first live stream to the server, the server can send the start and end times of the second video clip extracted from the third video clip corresponding to the second live stream to the client corresponding to the first live stream. The start and end times of the second video clip are used to describe the start time and end time of the highlight moment video clip material in the third video clip that can increase the access volume of the second live stream in the third video clip. For the server, only storing the start and end times of the highlight moment video clip material can represent the extracted video clip, without the need to store the video clip itself, reducing the occupancy of storage resources.

[0053] See Figure 2 and Figure 3 , sending at least one start and end time of the second video clip to the client. The client corresponding to the first live stream can extract and download the second video clip from the third video clip provided by the server according to the received start and end times of the second video clip, which can facilitate the client to extract the second video clip that meets its own needs from the third video clip, making the generation of the highlight moment video clip material more flexible and dynamic. The highlight moment video clip material can be continuously adjusted by modifying the start and end times of the second video clip or updating to obtain new start and end times of the second video clip. At the same time, the transmission traffic cost of pulling the highlight moment video clip material on the client can be reduced.

[0054] See Figure 3 , the second video clip is the highlight moment video clip material that is extracted from the third video clip corresponding to the second live stream, increases the access volume of the second live stream, and belongs to the first video clip. When the client corresponding to the first live stream receives the first video clip sent by the server and at least one start and end time of the second video clip corresponding to the first video clip, it extracts and downloads the video clip according to the start and end times in the third video clip indicated by at least one start and end time of the second video clip. Among them, at least one start and end time of the second video clip can be recorded as follows: Highlight 1 (start time 1, end time 1), Highlight 2 (start time 2, end time 2), Highlight 3 (start time 3, end time 3),..., Highlight N (start time N, end time N).

[0055] See Figure 2 and Figure 3, the client corresponding to the first live broadcast room can receive the first video segment sent by the server, and then preprocess the second video segment and the first video segment before rendering and displaying the second video segment and the first video segment. Among them, the video segment preprocessing may include at least one of the following: video segment splicing, adding special effects to the video segment, adding stickers to the video segment, adding transition animations to the video segment, and adding watermarks to the video segment.

[0056] See Figure 3 , optionally, after the splicing of the first video segment and the second video segment corresponding to the first video segment is completed, the spliced video segment is configured to preload the first video in the video list before the spliced video segment is triggered to be displayed, and preload the remaining videos in the video list during the process of triggering the display of the first video in the spliced video segment. Among them, the splicing of the first video segment and the second video segment corresponding to the first video segment includes video display styles (including vertical video styles and horizontal video styles), highlight moment judgment, highlight moment duration, and splicing methods (such as setting the playback progress, and setting the highlight moment video segment material to play from the Nth second (highlight moment) to the end, and playing from the beginning the next time).

[0057] The technical solution of the embodiment of the present disclosure can use the first video segment, the start and end times of the second video segment, and the third video segment to increase the access volume of the first live broadcast room when using short videos to increase the access volume of the live broadcast room. Since the second video segment used to increase the access volume of the first live broadcast room is from the third video segment of the second live broadcast room that has a preset association condition with the first live broadcast room, it can ensure that the highlight moment video segment material indicated by the second video segment can fully reflect the core of the live broadcast content of the first live broadcast room, realize the instant application of the second video segment for highlight display in the first live broadcast room, ensure the timeliness of the highlight moment video segment material, and do not need to wait until the live broadcast of the first live broadcast room ends to review the live broadcast content to generate appropriate highlight moment video segment material; at the same time, when determining the second video segment, only the start and end times of the second video segment are generated and stored without directly generating the second video segment, which can reduce the storage cost of the highlight moment video segment material and the transmission traffic cost of the highlight moment video segment material, and thus facilitate the expansion of the coverage rate of the highlight moment video segment material when the storage resource cost is reduced.

[0058] Figure 4 It is a schematic flowchart of another video processing method provided by the embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of determining the start and end times of the second video segment corresponding to the first video segment in the foregoing embodiment on the basis of the foregoing embodiment. This embodiment can be combined with each optional solution in the above one or more embodiments.

[0059] As Figure 4 shown, the video processing method of the embodiments of the present disclosure may include the following processes:

[0060] S410. Determine a first video segment corresponding to a first live broadcast room, where the first video segment is real-time live video data of the first live broadcast room.

[0061] S420. Determine at least two third video segments in a second live broadcast room and corresponding interaction parameter values, where the interaction parameter is a detection index for detecting whether the video segment contains video segment material of a highlight moment that can guide an increase in the access volume of the second live broadcast room.

[0062] The third video segment is historical live video data of the second live broadcast room. The live broadcast content of the first live broadcast room and the second live broadcast room satisfies a preset association condition. Therefore, the core of the live broadcast content of the second live broadcast room can, to a certain extent, represent the core of the live broadcast content of the first live broadcast room. Optionally, the live broadcast content of the first live broadcast room and the second live broadcast room satisfying the preset association condition includes that the product identification information indicated by the live broadcast content of the first live broadcast room and the product identification information indicated by the live broadcast content of the second live broadcast room satisfy a preset similarity. Among them, the product identification information includes product type, product name, etc.

[0063] Optionally, the interaction parameter may be the video duration of the video segment, the click-through rate of the video segment, the conversion rate of the video segment, the product of the click-through rate and the conversion rate of the video segment, the total playback volume of the video segment, the instantaneous playback rate of the video segment (such as the 3s playback rate), the completion rate of the video segment, the average playback duration of the video segment, the like rate of the video segment, the comment rate of the video segment, the activation rate of the video segment, the feedback rate of the video segment, the number of paid users of the video segment, the payment rate of the video segment, etc. The interaction parameter may also be the number of plays of the video segment, the attention of the video segment, the penetration rate of the video segment, etc.

[0064] As an optional but non-limiting implementation manner, determining at least two third video segments in the second live broadcast room and corresponding interaction parameter values includes the following steps B1 - B2:

[0065] Step B1. Determine a second live broadcast room associated with the first live broadcast room, where the product identification information pushed in the live broadcast of the first live broadcast room and the product identification information pushed in the live broadcast of the second live broadcast room satisfy a preset similarity condition.

[0066] Step B2. Obtain at least two third video segments in the second live broadcast room and the corresponding interaction parameter values of the third video segments from the historical live video data corresponding to the second live broadcast room.

[0067] See Figure 2 And Figure 5, when the client corresponding to the first live broadcast room sends a request for the highlight video clip material of the first live broadcast room to the server, the second live broadcast room that meets the preset association condition with the first live broadcast room in terms of live content will be determined from multiple live broadcast rooms that have conducted live broadcasts to form historical live video data, and it is ensured that the product identification information pushed in the first live broadcast room and the product identification information pushed in the second live broadcast room meet the preset similarity condition. Furthermore, at least two third video clips in the second live broadcast room can be obtained from the historical live video data corresponding to the second live broadcast room, and at the same time, the value of the interaction parameter corresponding to the third video clip needs to be marked. For example, see Figure 5 , for 6 1-minute third video clips, the value of the interaction parameter corresponding to each third video clip is marked.

[0068] S430. According to the values of the interaction parameters corresponding to at least two third video clips, at least one fourth video clip is screened out from the at least two third video clips.

[0069] See Figure 5 , according to the values of the interaction parameters corresponding to each third video clip, it can be judged whether each third video clip contains highlight video clip material that can guide the increase of the access volume of the second live broadcast room, realizing a rough screening of each third video clip, and the remaining third video clips are determined as the fourth video clips. The fourth video clips are likely to contain highlight video clip material that can guide the increase of the access volume of the second live broadcast room. Optionally, if it is detected that the value of the interaction parameter corresponding to the third video clip is greater than or equal to the preset interaction parameter threshold, it is determined that the third video clip contains highlight video clip material that can guide the increase of the access volume of the second live broadcast room; if it is detected that the value of the interaction parameter corresponding to the third video clip is less than the preset interaction parameter threshold, it is determined that the third video clip does not contain highlight video clip material that can guide the increase of the access volume of the second live broadcast room.

[0070] Exemplarily, see Figure 5 , based on a series of strategies, according to the values of the interaction parameters corresponding to each third video clip and the corresponding interaction parameter thresholds, it is judged whether the third video clip is a video clip containing highlight video clip material. For example, the number of plays of the video clip, the attention of the video clip, and the penetration rate of the video clip can be used as the interaction parameter indicators for defining the highlight video clip material, and relatively rich live broadcast forms (such as flash sales, lotteries, reviews of popular products, live broadcast previews, etc.) can be located.

[0071] S440. Identify the start and end times of the second video clip corresponding to the first video clip from at least one fourth video clip.

[0072] See Figure 5, for six third video segments each lasting 1 minute, the interaction parameter values of each third video segment are marked. In this way, through speech recognition text technology, the words, punctuation characters, and their respective start and end times in the fourth video segment can be recognized, so that the start and end times of at least one second video segment included in the fourth video segment can be recognized.

[0073] As an optional but non-limiting implementation, identifying the start and end times of the second video segment corresponding to the first video segment from at least one fourth video segment includes the following steps C1 - C3:

[0074] Step C1: Identify at least three clause texts from the fourth video segment and predict the keywords corresponding to the clause texts.

[0075] Step C2: Generate at least two candidate complex sentence texts based on the keywords corresponding to at least three clause texts. Each candidate complex sentence text is associated with at least one start and end time of a fifth video segment, and the fifth video segment is a local video segment in the fourth video segment that is associated with the keyword corresponding to the clause text.

[0076] Step C3: Determine the target complex sentence text from at least two candidate complex sentence texts, and determine the start and end times of the second video segment corresponding to the first video segment based on the at least one start and end time of the fifth video segment associated with the target complex sentence text.

[0077] See Figure 5 , for the fourth video segment, at least three clause texts can be recognized from the fourth video segment through speech recognition text technology, and the interface service can be called to perform word segmentation label prediction on each clause text to obtain the keyword corresponding to each clause text. The keywords of the clause text can be product names, product functions, product promotion information, live broadcast speed, and the applicable objects of the product, etc.

[0078] See Figure 5 , in combination with the keywords corresponding to the clause texts, while ensuring the integrity of the sentence, at least two candidate complex sentence texts can be produced (for example, by combining the clause texts with product or selling point information labels to form candidate complex sentence texts. Each candidate complex sentence text is associated with at least one start and end time of a fifth video segment, and the fifth video segment is a local video segment in the fourth video segment that is associated with the keyword corresponding to the clause text. For a general 2 - minute video, 2 - 5 complex sentences can be produced.

[0079] See Figure 5, optionally, determining a target complex sentence text from at least two candidate complex sentence texts, including: calculating weights for the keywords included in the candidate complex sentence texts, and determining the target complex sentence text from the at least two candidate complex sentence texts based on the weighted calculation results corresponding to each candidate complex sentence text. For example, calculating a weighted score based on the keywords included in the candidate complex sentence text, if the score is greater than a threshold, determining the candidate complex sentence text as the target complex sentence text, otherwise discarding it. Further, at least one start and end time of the fifth video segment associated with the target complex sentence text can be determined as the start and end time of the second video segment corresponding to the first video segment, realizing the output of the highlight video segment material of the first live broadcast room from the second live broadcast room.

[0080] S450. Push to the first live broadcast room based on the start and end time of the first video segment, the second video segment, and the third video segment.

[0081] The technical solution of the embodiments of the present disclosure can use the start and end time of the first video segment, the second video segment, and the third video segment to increase the access volume of the first live broadcast room when using short videos to increase the access volume of the live broadcast room. Since the second video segment used to increase the access volume of the first live broadcast room is from the third video segment of the second live broadcast room that has a preset association condition with the first live broadcast room, it can ensure that the highlight video segment material indicated by the second video segment can fully reflect the core of the live broadcast content of the first live broadcast room, realizing the instant application of the second video segment for highlight display in the first live broadcast room, ensuring the timeliness of the highlight video segment material, and generating appropriate highlight video segment material without waiting to review the live broadcast content after the live broadcast of the first live broadcast room ends, which can reduce the skill requirements for editors, reduce the time-consuming of the editing process, improve the editing efficiency and quality, and enhance the effect of increasing the access volume of the live broadcast room; at the same time, when determining the second video segment, only the start and end time of the second video segment is generated and stored without directly generating the second video segment, which can reduce the storage cost of the highlight video segment material and the transmission traffic cost of the highlight video segment material. Furthermore, when the storage resource cost is reduced, it is conducive to expanding the coverage rate of the highlight video segment material.

[0082] Figure 6 FIG. is a schematic structural diagram of a video processing device provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of using short videos to increase the access volume of a live broadcast room. The video processing device can be implemented in the form of software and / or hardware, and is generally integrated on any electronic device with network communication functions. The electronic device can be a mobile terminal, a PC, or a server, etc.

[0083] As Figure 6As shown in the figure, the video processing device according to the embodiments of the present disclosure may include the following: a first determination module 610, a second determination module 620, and a video processing module 630. Among them:

[0084] The first determination module 610 is configured to determine a first video segment corresponding to a first live broadcast room, where the first video segment is real-time live video data of the first live broadcast room;

[0085] The second determination module 620 is configured to determine the start and end times of a second video segment corresponding to the first video segment. The second video segment is highlight moment video segment material of the first video segment. The second video segment is selected from a third video segment corresponding to a second live broadcast room as highlight moment video segment material that can increase the access volume of the second live broadcast room. The third video segment is historical live video data of the second live broadcast room, and the live content of the first live broadcast room and the second live broadcast room satisfies a preset association condition;

[0086] The video processing module 630 is configured to push the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment.

[0087] Based on the above embodiments, optionally, determining the start and end times of the second video segment corresponding to the first video segment includes:

[0088] Determine at least two third video segments in the second live broadcast room and corresponding interaction parameter values. The interaction parameter is a detection index for detecting whether the video segment contains highlight moment video segment material that can guide an increase in the access volume of the second live broadcast room;

[0089] According to the interaction parameter values corresponding to the at least two third video segments, screen out at least one fourth video segment from the at least two third video segments;

[0090] Identify the start and end times of the second video segment corresponding to the first video segment from the at least one fourth video segment.

[0091] Based on the above embodiments, optionally, determining at least two third video segments in the second live broadcast room and corresponding interaction parameter values includes:

[0092] Determine a second live broadcast room associated with the first live broadcast room, where the product identification information pushed in the live broadcast of the first live broadcast room and the product identification information pushed in the live broadcast of the second live broadcast room satisfy a preset similarity condition;

[0093] Obtain at least two third video segments in the second live broadcast room and the corresponding interaction parameter values of the third video segments from the historical live video data corresponding to the second live broadcast room.

[0094] Based on the above embodiments, optionally, identifying the start and end times of the second video segment corresponding to the first video segment from the at least one fourth video segment includes:

[0095] Identifying at least three clause texts from the fourth video segment and predicting the keywords corresponding to the clause texts;

[0096] Generating at least two candidate complex sentence texts based on the keywords corresponding to the at least three clause texts, each candidate complex sentence text associated with at least one start and end time of a fifth video segment, where the fifth video segment is a local video segment in the fourth video segment associated with the keywords corresponding to the clause texts;

[0097] Determining a target complex sentence text from the at least two candidate complex sentence texts and determining the start and end times of the second video segment corresponding to the first video segment based on the at least one start and end time of the fifth video segment associated with the target complex sentence text.

[0098] Based on the above embodiments, optionally, determining a target complex sentence text from the at least two candidate complex sentence texts includes:

[0099] Performing weighted calculation on the keywords included in the candidate complex sentence texts and determining the target complex sentence text from the at least two candidate complex sentence texts based on the weighted calculation results corresponding to each candidate complex sentence text.

[0100] Based on the above embodiments, optionally, pushing to the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment includes:

[0101] Sending at least one start and end time of the second video segment to the client corresponding to the first live broadcast room, so that the client corresponding to the first live broadcast room extracts the second video segment from the third video segment based on the start and end time of the second video segment;

[0102] Sending the first video segment to the client corresponding to the first live broadcast room, so that the client corresponding to the first live broadcast room performs video segment preprocessing on the second video segment and the first video segment and displays them on the client corresponding to the first live broadcast room.

[0103] Based on the above embodiments, optionally, the video segment preprocessing includes at least one of the following: video segment splicing, adding special effects to the video segment, adding stickers to the video segment, adding transition animations to the video segment, and adding watermarks to the video segment.

[0104] When using short videos to increase the viewership of a live streaming room, the technical solution provided by the embodiments of the present disclosure can use the start and end times of the first video segment and the second video segment and the third video segment to increase the viewership of the first live streaming room. Since the second video segment used to increase the viewership of the first live streaming room is derived from the third video segment of the second live streaming room that has a preset association condition with the first live streaming room, it can ensure that the highlight moment video segment material indicated by the second video segment can fully reflect the core of the live streaming content of the first live streaming room, realizing the ability to immediately apply the second video segment for highlight display in the first live streaming room, ensuring the timeliness of the highlight moment video segment material, and generating appropriate highlight moment video segment material without having to review the live streaming content after the live streaming of the first live streaming room ends; at the same time, when determining the second video segment, only the start and end times of the second video segment are generated and stored without directly generating the second video segment, which can reduce the storage cost of the highlight moment video segment material and the transmission traffic cost of the highlight moment video segment material. Furthermore, when the storage resource cost is reduced, it is conducive to expanding the coverage rate of the highlight moment video segment material.

[0105] The video processing device provided by the embodiments of the present disclosure can execute the video processing method provided by any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution of the method.

[0106] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0107] Figure 7 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Referring below to Figure 7 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure (such as Figure 7 the terminal device or server in). The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0108] As Figure 7As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An editing / output (I / O) interface 505 is also connected to the bus 504.

[0109] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 an electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0110] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0111] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0112] The electronic device provided by the embodiment of the present disclosure and the video processing method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0113] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the video processing method provided by the above embodiment is implemented.

[0114] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0115] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (for example, the Internet), and end-to-end networks (for example, ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0116] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.

[0117] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: determine a first video segment corresponding to a first live broadcast room, the first video segment being real-time live video data of the first live broadcast room; determine start and end times of a second video segment corresponding to the first video segment, the second video segment being highlight moment video segment material of the first video segment, the second video segment being selected from a third video segment corresponding to a second live broadcast room as highlight moment video segment material that can increase the number of visits to the second live broadcast room, the third video segment being historical live video data of the second live broadcast room, and the live broadcast content of the first live broadcast room and the second live broadcast room satisfying a preset association condition; and perform a push on the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment.

[0118] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0120] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of a unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".

[0121] The functions described above in this article can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0122] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] The above description is only for the preferred embodiments of the present disclosure and the illustration of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0124] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0125] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video processing method, characterized in that, The method includes: Determine a first video segment corresponding to a first live streaming room, where the first video segment is real-time live streaming video data of the first live streaming room; Determine the start and end time of a second video segment corresponding to the first video segment, where the second video segment is high-light moment video segment material corresponding to the first video segment. The second video segment is high-light moment video segment material selected from a third video segment corresponding to a second live streaming room and capable of increasing the access volume of the second live streaming room. The third video segment is historical live streaming video data of the second live streaming room, and the live streaming content of the first live streaming room and the second live streaming room meets a preset association condition; Push to the first live streaming room based on the first video segment, the start and end time of the second video segment, and the third video segment; Pushing to the first live streaming room based on the first video segment, the start and end time of the second video segment, and the third video segment includes: determining the second video segment according to the third video segment and the start and end time of the second video segment; performing video segment preprocessing on the second video segment and the first video segment, and pushing to the first live streaming room to achieve instant application of the second video segment for high-light display in the first live streaming room.

2. The method according to claim 1, wherein Determining the start and end time of the second video segment corresponding to the first video segment includes: Determine at least two third video segments in the second live streaming room and corresponding interaction parameter values, where the interaction parameter is a detection index for detecting whether the video segment contains high-light moment video segment material capable of increasing the access volume of the second live streaming room; Screen at least one fourth video segment from the at least two third video segments according to the interaction parameter values corresponding to the at least two third video segments; Identify the start and end time of the second video segment corresponding to the first video segment from the at least one fourth video segment.

3. The method according to claim 2, wherein Determining at least two third video segments in the second live streaming room and corresponding interaction parameter values includes: Determine a second live streaming room associated with the first live streaming room, where the product identification information live-streamed and pushed in the first live streaming room and the product identification information live-streamed and pushed in the second live streaming room meet a preset similarity condition; Obtain at least two third video segments in the second live streaming room and the corresponding interaction parameter values of the third video segments from the historical live streaming video data corresponding to the second live streaming room.

4. The method according to claim 2, wherein Identifying the start and end time of the second video segment corresponding to the first video segment from the at least one fourth video segment includes: Identify at least three clause texts from the fourth video segment and predict the keywords corresponding to the clause texts; Generate at least two candidate complex sentence texts based on the keywords corresponding to the at least three clause texts, and each candidate complex sentence text is associated with at least one start and end time of a fifth video segment, where the fifth video segment is a local video segment in the fourth video segment associated with the keywords corresponding to the clause texts; Determine a target complex sentence text from the at least two candidate complex sentence texts, and determine the start and end times of the second video segment corresponding to the first video segment based on the start and end times of at least one fifth video segment associated with the target complex sentence text.

5. The method according to claim 4, wherein Determining a target complex sentence text from the at least two candidate complex sentence texts includes: Performing weighted calculation on the keywords included in the candidate complex sentence texts, and determining the target complex sentence text from the at least two candidate complex sentence texts based on the weighted calculation results corresponding to each candidate complex sentence text.

6. The method according to any one of claims 1 to 5, characterized in that Pushing to the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment includes: Sending at least one start and end time of the second video segment to the client corresponding to the first live broadcast room, so that the client corresponding to the first live broadcast room extracts the second video segment from the third video segment based on the start and end time of the second video segment; Sending the first video segment to the client corresponding to the first live broadcast room, so that the client corresponding to the first live broadcast room performs video segment preprocessing on the second video segment and the first video segment and displays them on the client corresponding to the first live broadcast room.

7. The method according to claim 6, wherein The video segment preprocessing includes at least one of the following: video segment splicing, adding special effects to the video segment, adding stickers to the video segment, adding transition animations to the video segment, and adding watermarks to the video segment.

8. A video processing device, characterized in that, The device includes: A first determination module, configured to determine a first video segment corresponding to the first live broadcast room, where the first video segment is real-time live video data of the first live broadcast room; A second determination module, configured to determine the start and end times of the second video segment corresponding to the first video segment, where the second video segment is a highlight moment video segment material corresponding to the first video segment, and the second video segment is a highlight moment video segment material selected from the third video segment corresponding to the second live broadcast room that can increase the access volume of the second live broadcast room, and the third video segment is historical live video data of the second live broadcast room, and the live broadcast contents of the first live broadcast room and the second live broadcast room meet a preset association condition; A video processing module, configured to push to the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment; pushing to the first live broadcast room based on the first video segment, the start and end times of the second video segment, and the third video segment includes: determining the second video segment according to the third video segment and the start and end times of the second video segment; performing video segment preprocessing on the second video segment and the first video segment, and pushing to the first live broadcast room to achieve instant application of the second video segment for highlight display in the first live broadcast room.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the video processing method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method, device and system for video live broadcast

    CN107172443A

  • Live streaming processing method and device, electronic equipment and computer readable storage medium

    CN111918085A