Live broadcast fragment generation method and device, medium, electronic equipment and product

CN121509783APending Publication Date: 2026-02-10BYTEDANCE TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511785661.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-10

Smart Images

  • Figure CN121509783A_ABST
    Figure CN121509783A_ABST
Patent Text Reader

Abstract

The invention discloses a live broadcast fragment generation method and device, a medium, electronic equipment and a product. The method comprises the following steps: receiving a live video stream of a target live room in real time; extracting a video clip to be processed from the live video stream; performing content understanding on the video clip, and determining at least one object appearing in the video clip and a video sub-clip corresponding to each object; extracting an initial slice of the target object from an aggregation result by aggregating a plurality of video sub-fragments corresponding to the same target object; highlight fragment identification is carried out on the initial slice of the target object to extract a highlight fragment of the target object; and during the live broadcast period of the target live broadcast room, outputting the highlight segment of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a method, apparatus, medium, electronic device, and product for generating live broadcast segments. Background Technology

[0002] Current live streaming content processing solutions require extracting highlights from the complete live stream content. Typically, this involves manually watching and editing the entire stream, or extracting highlight segments from replays of completed streams using recognition methods (e.g., identifying the most frequently interacted parts by viewers). However, manual editing is extremely inefficient and cannot meet the needs of a massive number of live streams. Extracting highlight segments using recognition methods suffers from significant time lag, and the extracted segments may be incomplete and of poor quality. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, this disclosure provides a method for generating live stream segments, the method comprising: Receive live video streams from the target live streaming room in real time; Extract video segments to be processed from the live video stream; Perform content understanding on the video segment to identify at least one object appearing in the video segment and the video sub-segment corresponding to each object; By aggregating multiple video segments corresponding to the same target object, the initial slice of the target object is extracted from the aggregation result; Spectral fragment identification is performed on the initial slice of the target object to extract the spectral fragments of the target object; During the live broadcast in the target live room, the highlight segment of the target object is output.

[0005] Secondly, this disclosure provides a live broadcast segment generation apparatus, the apparatus comprising: The receiving module is used to receive the live video stream from the target live streaming room in real time. The first extraction module is used to extract video segments to be processed from the live video stream; The determination module is used to perform content understanding on the video segment and determine at least one object appearing in the video segment and the video sub-segment corresponding to each object. The second extraction module is used to extract the initial slice of the target object from the aggregation result by aggregating multiple video sub-segments corresponding to the same target object; The third extraction module is used to identify highlight segments in the initial slice of the target object in order to extract the highlight segments of the target object; An output module is used to output the highlight fragment of the target object during the live broadcast in the target live room.

[0006] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of this disclosure.

[0007] Fourthly, this disclosure provides an electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect of this disclosure.

[0008] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect of this disclosure.

[0009] The above technical solution receives real-time live video streams and continuously extracts video segments for processing. For each video segment, content understanding is used to analyze the objects within it in real time, and segments are aggregated according to the target object to extract initial slices. This allows for the accurate extraction of high-quality slices related to the target object from lengthy and complex live streams. Highlight segment recognition is then performed on the initial slices to extract more essential highlight segments, ensuring high quality. Therefore, by constructing an automated processing flow for real-time live content recognition, high-quality slice extraction, and highlight segment extraction, high-quality highlight segments for specific objects in the current live stream can be automatically generated with extremely high processing efficiency. Furthermore, during the live stream of the target live stream, highlight segments of the target object can be automatically output to attract potential users to the target live stream, thus improving the viewing experience.

[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart of a live stream segment generation method according to one embodiment of the present disclosure; Figure 2 This is a flowchart and architecture diagram of the live stream segment generation method provided in this public document; Figure 3 This is a block diagram of a live broadcast segment generation apparatus according to one embodiment of the present disclosure; Figure 4 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0012] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0013] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0022] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0023] Figure 1 This is a flowchart of a live stream segment generation method according to one embodiment of this disclosure. Figure 1 As shown, the method provided in this disclosure may include steps 11 to 16.

[0024] In step 11, the live video stream of the target live room is received in real time.

[0025] In step 12, the video segments to be processed are extracted from the live video stream.

[0026] For a target live stream that is currently broadcasting, you can obtain the live video stream of that target live stream in real time.

[0027] For the acquired live video stream, video segments to be processed can be continuously extracted. Optionally, video segments of a preset duration can be continuously extracted from the live video stream as video segments to be processed. For example, if the preset duration is set to 60 seconds, for the received live video stream, a 60-second segment can be extracted every 60 seconds as a video segment to be processed. For example, for a 2-minute live video stream, the first 60 seconds can be extracted as one video segment, and the next 60 seconds can be extracted as the next video segment, so that each video segment can be processed separately.

[0028] In step 13, content understanding is performed on the video segment to identify at least one object appearing in the video segment and the corresponding video sub-segment for each object.

[0029] It should be noted that in the actual processing, multiple video segments to be processed can be extracted from the live video stream. Each video segment needs to be processed. This disclosure will provide a detailed description of the processing method for a video segment, and the processing method for each video segment can refer to this process.

[0030] In this disclosure, the aforementioned objects represent the entity objects that need to be focused on, and can be set according to actual needs. For example, if the theme of the live stream is item sharing, then the object is the item shared in the live stream.

[0031] In one possible implementation, step 13 may include the following steps: The video segment is identified by a preset object detection model, which identifies at least one object in the video segment and the time period of each object, so as to determine the video sub-segment corresponding to each object. And / or, The speech corresponding to the video segment is converted into text by a preset speech recognition model, and at least one object being explained and the time period of each object are identified based on the text, so as to determine the video sub-segment corresponding to each object.

[0032] Optionally, an object list can be pre-set, which may include image information or image features of different objects. The object detection model can identify whether an object from the object list exists in a video clip, and mark the time period in which the object appears if an object is detected. Different objects can correspond to different identifiers. By associating video frames with object identifiers, it can be shown that the video frame contains the object indicated by the identifier. Thus, based on the object identifier associated with each video frame, the time period in which the object appears can be located, and the corresponding time period can be extracted as the video sub-segment corresponding to that object.

[0033] Optionally, a pre-defined object list can be set, which may include text information (e.g., keywords, synonyms, etc.) or text features of different objects. After converting the audio in the video clip into text using a pre-defined speech recognition model, a natural language understanding model can be used to recognize the text to determine whether an object from the object list exists, and if an object is recognized, the time period in which that object appears can be marked. Different objects can correspond to different identifiers. By associating audio frames with the representations of objects, it can be ensured that the text of the audio frame contains the object indicated by the identifier. Thus, based on the object identifier associated with each audio frame, the time period in which the object appears can be located, and the corresponding time period in the video clip can be extracted as the video sub-segment corresponding to that object.

[0034] For example, suppose in a 60-second video clip, the first 20 seconds show object 1, the next 20 seconds show object 2, the following 10 seconds are spent explaining object 1, and the last 10 seconds show neither object nor are explained. After understanding the content in step 13, the first 20 seconds can be extracted as video segment 1, the next 20 seconds as video segment 2, and the last 10 seconds as video segment 3. Video segment 1 and video segment 3 correspond to object 1 (i.e., are associated with it), and video segment 2 corresponds to object 2. The last 10 seconds, since they are not associated with any object, will not be extracted as a video segment.

[0035] This method allows for the continuous acquisition of video clips from a constant stream of live video, and the extraction of objects and their corresponding video segments from the live stream, extracting as many relevant and valid segments as possible. Furthermore, by combining image recognition and speech recognition during the segment extraction process, it ensures that no potentially valid segments are missed.

[0036] In step 14, the initial slice of the target object is extracted from the aggregation result by aggregating multiple video sub-segments corresponding to the same target object.

[0037] After step 13, extract all video segments of the object of interest up to the present. Based on these video segments, aggregate all video segments of the same object to obtain all the content related to the object in this live broadcast up to the present. Based on this, select the more core and high-quality initial slices.

[0038] In one possible implementation, step 14 may include the following steps: For a target object, the corresponding video segments are aggregated to obtain the aggregation result; According to the preset screening criteria, the initial slices that meet the screening criteria are selected from the aggregation results.

[0039] For at least one object identified in step 13, each object can be treated as a target object for subsequent processing. For example, in the previous example, if the relevant content of object 1 and object 2 is extracted from the 60s clip, then object 1 and object 2 can be treated as target objects for subsequent processing steps, ultimately generating highlight clips of object 1 and object 2.

[0040] For a target object, you can first extract all video segments of that target object from the live video stream up to the present (or a certain period of time before the current moment, which can be set as needed), and then splice and aggregate them in chronological order to obtain the aggregated result, which is the aggregated video.

[0041] For the aggregated results, segments that meet the preset filtering criteria can be selected as initial slices. The filtering criteria can be flexibly set according to actual needs. For example, the filtering criteria can be set to the continuous appearance of the target object. Another example is that the filtering criteria can be set to the target object being located in a specific area of ​​the image and not stationary.

[0042] In one possible implementation, initial slices that meet the selection criteria can be selected from the aggregation results in the following way: In the aggregation results, the longest consecutive segment of the target object in time is selected as the initial slice.

[0043] Optionally, for the aggregation result, it is possible to identify whether a target object exists in each video frame of the aggregation result, and mark the video frames in which the target object exists. This allows for the determination of the set of video frames in which the target object appears consecutively, identifying the set of video frames with the most consecutive frames, and then extracting the portion corresponding to this set of video frames from the aggregation result to obtain the initial slice. For example, if the aggregation result has 120 frames, and the consecutively appearing video frames include frames 3-33 and frames 50-120, then the video segment corresponding to frames 50-120 in the aggregation result can be extracted as the initial slice.

[0044] In another possible implementation, initial slices that meet the selection criteria can be selected from the aggregation results in the following way: Based on the first preset evaluation dimension, determine the score of each video frame in the aggregation result in the first preset evaluation dimension; In the aggregation results, segments with scores continuously greater than the first score threshold are selected as the initial slices.

[0045] Optionally, the first preset evaluation dimension may include at least one of the image quality dimension and the interaction dimension.

[0046] The image quality dimension can be related to at least one of the following: image sharpness, whether the anchor is on screen, and whether the image is still. For example, a higher image sharpness can be rated higher than a lower image sharpness. Another example is that an image with an anchor on screen can be rated higher than an image without an anchor on screen. Yet another example is that a still image can be rated lower than a changing image.

[0047] Interaction dimensions can be related to the interaction data corresponding to the live stream segment featuring the target object. Interaction data generally refers to data generated by users within the live stream that can quantify user attention or interest, such as traffic data, interaction data, and conversion data. Traffic data may include, but is not limited to, the speed at which users enter the live stream, the number of users entering the live stream at the same time (or time period), etc. Interaction data may include, but is not limited to, the number of likes, comments, and shares (or, frequency) of users towards the target object. Conversion data can include data on users' pre-selected behaviors. If the target object is a product, conversion data can be determined through user behaviors such as adding the target object to favorites, purchasing, or adding it to the shopping cart.

[0048] For each of the first preset evaluation dimensions, the method for determining the score can be implemented by those skilled in the art by selecting a mature algorithm model or custom rules, and can be flexibly set according to actual needs. For example, the score for the picture quality dimension can be determined based on visual features extracted from video frames. These visual features may include, but are not limited to, picture sharpness (e.g., determined by sharpness algorithms such as gradient functions), the presence of the anchor's face on camera (e.g., determined by the confidence level of face recognition), and non-still frames (e.g., determined by the differences between consecutive frames). By assigning weights to these visual features, the score corresponding to the picture quality dimension can be calculated comprehensively. As another example, for the interaction dimension, various types of interaction data can be mapped to weighted scores according to preset rules, and then weighted summed to calculate the interaction dimension score; alternatively, the interaction data can be input into a pre-trained machine learning model. This machine learning model has the basic ability to determine scores based on interaction data and can be trained by collecting interaction data and the corresponding scores from historical live broadcast segments.

[0049] In step 15, highlight fragment identification is performed on the initial slice of the target object to extract the highlight fragments of the target object.

[0050] After obtaining the initial slices of the target object, to improve their quality, highlight segment recognition can be performed on the initial slices to obtain the optimal highlight segments. In this disclosure, highlight segment recognition is used to identify and locate the relatively optimal segment that best meets the expected target from a relatively long initial slice. It is a quantitative analysis and screening process. The output result obtained through highlight segment recognition is the highlight segment, which represents a series of consecutive video frames that, after quantitative evaluation, extract the best performance in terms of comprehensive evaluation dimensions from the initial slices. The determination of the highlight segment depends on its evaluation criteria, which can be flexibly set according to actual needs (such as subsequent second and third preset evaluation dimensions). For example, a highlight segment can be the segment with the highest information density in the initial slice (e.g., the explanation content describes the characteristics of the target object most completely and is attractive), the segment with the best visual display effect (e.g., the target object is displayed most clearly and from the most comprehensive angle), or the segment most relevant to positive user feedback (e.g., the segment with the most users entering the live broadcast room and the most users liking).

[0051] In one possible implementation, step 15 may include the following steps: Based on the second preset evaluation dimension, a comprehensive score is determined for each processing unit in the initial slice according to the second preset evaluation dimension. Each processing unit includes at least one video frame. Based on the comprehensive score of each processing unit, one or more consecutive processing units are extracted as the highlight segments of the target object.

[0052] The initial slice can be divided into finer-grained processing units, and each unit can be comprehensively scored. Based on the comprehensive score, the highest quality highlight segment is obtained. Each processing unit can include at least one video frame. For example, each video frame can be divided into one processing unit. Alternatively, every 10 video frames can be divided into one processing unit.

[0053] Optionally, determining the comprehensive score of each processing unit in the initial slice based on the second preset evaluation dimension may include the following steps: Identify the first type of segments in the initial slice that contain explanatory content and the second type of segments that do not contain explanatory content; For the first type of segment, the segment is divided into multiple semantic units, and each semantic unit is scored based on the second preset scoring dimension; For the second type of segment, the segment is divided into multiple time units, and each time unit is scored based on a third preset scoring dimension; Based on the scores of semantic units and temporal units, a comprehensive score is determined for each processing unit in the initial slice.

[0054] As mentioned earlier, during content understanding in step 13, the audio of the video clip can be converted into text using a speech recognition model, and the object being explained can be determined using a natural language understanding model. Therefore, each object being explained will have corresponding explanation content, and its timing can also be determined. Based on this, the initial slice typically contains the explanation content corresponding to the target object.

[0055] Since the initial slices may contain both parts that include explanatory content and parts that do not, and the presence or absence of explanatory content affects the semantic richness of the text, they can be categorized and scored using different criteria. The initial slices can be classified into two categories based on whether they contain explanatory content: Category 1 slices containing explanatory content and Category 2 slices that do not (e.g., slices that only show the target object without explanation).

[0056] For the first type of fragment, it can be divided into multiple semantic units.

[0057] Optionally, the first type of segment can be divided into multiple semantic units based on semantic boundaries, such as specific punctuation marks (e.g., period, question mark, exclamation mark, etc.). For example, sentences can be used as the unit of division, with each sentence being divided into a semantic unit.

[0058] For the second type of segment, it can be divided into multiple time units.

[0059] Optionally, the second type of segment can be divided into multiple time units at certain time intervals. For example, the second type of segment can be divided into multiple 10-second time units, with each unit being 10 seconds.

[0060] By using the above methods, different types of live streaming content can be divided at the most reasonable level, which lays a solid foundation for accurate scoring in the future.

[0061] After dividing the data into multiple processing units, that is, after dividing it into the aforementioned semantic units or time units, scoring can be performed separately. For semantic units, scoring can be based on a second preset evaluation dimension, and for time units, scoring can be based on a third preset evaluation dimension. The second preset evaluation dimension may include at least one of the following: visual dimension, text dimension, and interaction dimension; the third preset evaluation dimension may include at least one of the following: visual dimension and interaction dimension.

[0062] The visual dimension can be related to the appearance of the target object and / or the anchor. The appearance of the target object can include, but is not limited to, the size, position, completeness, and clarity of the target object in the frame. The appearance of the anchor can include, but is not limited to, whether the anchor's face is in the frame, the proportion of the anchor's face in the frame, and whether the anchor's face is in the same frame as the target object. In addition, the visual dimension can also include the visual quality dimension mentioned above.

[0063] The interaction dimension is related to the interaction data corresponding to the live stream segment where the target object is located. Explanations regarding the interaction dimension have already been provided in the previous summary and will not be repeated here.

[0064] Textual dimensions can be related to the semantic quality of the content. Textual dimensions are specifically designed for semantic units to assess the information content and appeal of the content. Optionally, factors that can be considered for textual dimensions include, but are not limited to, the following: keyword density, the completeness of the description of the target object's characteristics, semantic coherence, and whether it encourages interaction. Target object characteristics are the features that best represent the target object. For example, if the target object is a product, its characteristics may include its selling points; similarly, if the target object is something to share, its characteristics may include the reasons for sharing it.

[0065] For each of the second or third preset evaluation dimensions, the method for determining the score can be implemented by a person skilled in the art by selecting a mature algorithm model or custom rules, and can be flexibly set according to actual needs. The scoring of the visual dimension and the scoring of the interaction dimension are similar to those of the visual quality dimension and the interaction dimension mentioned above, and will not be repeated here.

[0066] The following is a brief example of scoring for the text dimension. For the text dimension, for example, a keyword list can be pre-set, and scoring can be achieved based on the frequency of keywords appearing in semantic units and the number of keywords in semantic units that cover the keyword list. For instance, the more keywords in a semantic unit that cover the keyword list, the higher the score. Another example is checking the grammatical structure to determine the coherence of the meaning, thus achieving a score. Yet another example is identifying whether the content contains words or phrases that encourage user interaction; more interactive expressions will result in a higher score.

[0067] After obtaining the scores for each semantic unit and temporal unit, each frame within a semantic unit shares the score for that semantic unit, and each frame within a temporal unit shares the score for that temporal unit. Thus, a comprehensive score can be obtained for each processing unit. For example, the comprehensive score for each processing unit can be obtained by weighted averaging the scores of each frame within that processing unit. Alternatively, a machine learning model for score fusion can be pre-trained, and the frame scores of each processing unit can be input into the machine learning model to obtain the final comprehensive score.

[0068] After obtaining the comprehensive score of each processing unit, the processing units can be arranged in chronological order to form a sequence of comprehensive scores. By analyzing this sequence, one or more consecutive processing units with the highest comprehensive score can be identified, or one or more consecutive processing units with a comprehensive score higher than the second score threshold can be identified. One or more such consecutive processing units can be extracted to form the highlight fragments of the target object.

[0069] For example, the process and architecture for generating highlight fragments of a target object in this disclosure can be as follows: Figure 2 As shown, in Figure 2 middle: As the live video stream plays, video segments to be processed can be extracted sequentially. For each video segment, the objects appearing in the video segment are identified through content understanding (such as object subject detection model and speech recognition model), and the video sub-segment corresponding to each object is stored in the database shown in the figure. In addition, during the content understanding stage, other tests disclosed herein can also be performed simultaneously (such as whether the host appears on camera, the content of the explanation and the corresponding time, whether the host and the product are in the same frame, etc.), and the relevant data is also directly stored in the database shown in the figure so that it can be called up when data is needed later; At the same time, based on the data stored in the database, initial slices can be generated periodically, that is, multiple video sub-segments of the same target object are aggregated, and then each video frame in the aggregated result is scored from the dimensions of picture quality and interaction to select the initial slices. For the extracted initial slices, highlight fragment recognition is performed. Semantic units and temporal units are divided according to whether there is narration content. Semantic units are scored from the dimensions of image, text and interaction. Temporal units are scored from the dimensions of image and interaction. The comprehensive score of each processing unit in the initial slice is then obtained. The final highlight fragment is extracted based on the comprehensive score.

[0070] Back Figure 1 In step 16, during the live broadcast in the target live room, the highlight clips of the target object are output.

[0071] Once you have obtained the highlight clips of the target object, you can output those clips during the target live stream to attract more potential viewers to the target live stream.

[0072] Optionally, step 16 may include the following steps: Generate an access link for accessing the target live stream; Publish highlight clips and access links of the target object to the video platform's content recommendation stream.

[0073] Based on this, the access link to the target live stream can be pushed together with the highlight clips of the target, making it easier for interested users to directly enter the target live stream and improving the viewing speed.

[0074] In addition to outputting highlight clips of the target object, it can also generate titles or highlight descriptions based on the content of the highlight clips and the explanatory content, so as to more intuitively reflect the core content of the highlight clips and make it easier for users to understand quickly.

[0075] The above technical solution receives real-time live video streams and continuously extracts video segments for processing. For each video segment, content understanding is used to analyze the objects within it in real time, and segments are aggregated according to the target object to extract initial slices. This allows for the accurate extraction of high-quality slices related to the target object from lengthy and complex live streams. Highlight segment recognition is then performed on the initial slices to extract more essential highlight segments, ensuring high quality. Therefore, by constructing an automated processing flow for real-time live content recognition, high-quality slice extraction, and highlight segment extraction, high-quality highlight segments for specific objects in the current live stream can be automatically generated with extremely high processing efficiency. Furthermore, during the live stream of the target live stream, highlight segments of the target object can be automatically output to attract potential users to the target live stream, thus improving the viewing experience.

[0076] Figure 3 This is a block diagram of a live stream segment generation apparatus according to one embodiment of the present disclosure. Figure 3 As shown, the device 30 may include: Receiver module 31 is used to receive the live video stream of the target live room in real time; The first extraction module 32 is used to extract video segments to be processed from the live video stream; The determination module 33 is used to perform content understanding on the video segment and determine at least one object appearing in the video segment and the video sub-segment corresponding to each object. The second extraction module 34 is used to extract the initial slice of the target object from the aggregation result by aggregating multiple video sub-segments corresponding to the same target object; The third extraction module 35 is used to identify highlight segments in the initial slice of the target object in order to extract the highlight segments of the target object; Output module 36 is used to output the highlight segment of the target object during the live broadcast in the target live room.

[0077] Optionally, the output module 36 includes: A generation submodule is used to generate access links for accessing the target live streaming room; The sending submodule is used to publish the highlight clip of the target object and the access link to the content recommendation stream of the video platform.

[0078] Optionally, the determining module 33 includes: The first determining submodule is used to identify at least one object appearing in the video segment and the time period of each object by using a preset object subject detection model, so as to determine the video sub-segment corresponding to each object. And / or, The second determining submodule is used to convert the speech corresponding to the video segment into text using a preset speech recognition model, and to identify at least one object being explained and the time period of each object based on the text, so as to determine the video sub-segment corresponding to each object.

[0079] Optionally, the second extraction module 34 includes: The aggregation submodule is used to aggregate the video sub-segments corresponding to the target object to obtain the aggregation result. The first filtering submodule is used to filter out initial slices that meet the preset filtering criteria from the aggregation results.

[0080] Optionally, the first filtering submodule includes: The second filtering submodule is used to filter the longest segment in time that the target object appears consecutively in the aggregation result as the initial slice.

[0081] Optionally, the first filtering submodule includes: The third determining submodule is used to determine the score of each video frame of the aggregation result in the first preset evaluation dimension based on the first preset evaluation dimension. The third filtering submodule is used to filter the segments whose scores are continuously greater than the first score threshold from the aggregation results as the initial slices.

[0082] Optionally, the first preset evaluation dimension includes at least one of the image quality dimension and the interaction dimension; The image quality dimension is related to at least one of the following: image clarity, on-screen appearance of the anchor, and stillness of the image. The interaction dimension is related to the interaction data corresponding to the live stream segment where the target object is located.

[0083] Optionally, the third extraction module 35 includes: The fourth determining submodule is used to determine the comprehensive score of each processing unit in the initial slice based on the second preset evaluation dimension, wherein the processing unit includes at least one video frame: An extraction submodule is used to extract one or more consecutive processing units as highlight segments of the target object based on the comprehensive score of each processing unit.

[0084] Optionally, the fourth determining submodule includes: The fifth determining submodule is used to determine the first type of segments that contain explanatory content and the second type of segments that do not contain explanatory content in the initial slice; The first scoring submodule is used to divide the first type of fragments into multiple semantic units and score each semantic unit based on a second preset scoring dimension. The second scoring submodule is used to divide the second type of segment into multiple time units and score each time unit based on a third preset scoring dimension. The third scoring submodule is used to determine the comprehensive score of each processing unit in the initial slice based on the scores of the semantic units and the scores of the time units.

[0085] Optionally, the second preset evaluation dimension includes at least one of the visual dimension, text dimension, and interaction dimension; the third preset evaluation dimension includes at least one of the visual dimension and interaction dimension. The image dimensions are related to the on-screen presence of the target object and / or the anchor; The text dimension is related to the semantic quality of the content being explained; The interaction dimension is related to the interaction data corresponding to the live stream segment where the target object is located.

[0086] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0087] The following is for reference. Figure 4 The diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0088] like Figure 4 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0089] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0090] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0091] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0092] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0093] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0094] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Receive live video streams from the target live streaming room in real time; Extract video segments to be processed from the live video stream; Perform content understanding on the video segment to identify at least one object appearing in the video segment and the video sub-segment corresponding to each object; By aggregating multiple video segments corresponding to the same target object, the initial slice of the target object is extracted from the aggregation result; Spectral fragment identification is performed on the initial slice of the target object to extract the spectral fragments of the target object; During the live broadcast in the target live room, the highlight segment of the target object is output.

[0095] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0097] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, the first receiving module can also be described as "a module for receiving live video streams from a target live streaming room in real time".

[0098] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0099] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0100] According to one or more embodiments of this disclosure, a method for generating live stream segments is provided, the method comprising: Receive live video streams from the target live streaming room in real time; Extract video segments to be processed from the live video stream; Perform content understanding on the video segment to identify at least one object appearing in the video segment and the video sub-segment corresponding to each object; By aggregating multiple video segments corresponding to the same target object, the initial slice of the target object is extracted from the aggregation result; Spectral fragment identification is performed on the initial slice of the target object to extract the spectral fragments of the target object; During the live broadcast in the target live room, the highlight segment of the target object is output.

[0101] According to one or more embodiments of this disclosure, a method for generating live stream clips is provided, wherein outputting the highlight clip of the target object includes: Generate an access link for accessing the target live stream room; The highlight clip of the target object and the access link are published to the content recommendation stream of the video platform.

[0102] According to one or more embodiments of this disclosure, a method for generating live video clips is provided, wherein performing content understanding on the video clips to determine at least one object appearing in the video clips and a video sub-segment corresponding to each object includes: The video segment is identified by a preset object detection model, which identifies at least one object in the video segment and the time period of each object, so as to determine the video sub-segment corresponding to each object. And / or, The speech corresponding to the video segment is converted into text using a preset speech recognition model, and at least one object being explained and the time period of each object are identified based on the text, so as to determine the video sub-segment corresponding to each object.

[0103] According to one or more embodiments of this disclosure, a method for generating live stream segments is provided, wherein the method extracts an initial slice of the target object from the aggregation result by aggregating multiple video sub-segments corresponding to the same target object, including: For the target object, the corresponding video segments are aggregated to obtain the aggregation result; According to the preset screening criteria, the initial slices that meet the screening criteria are selected from the aggregation results.

[0104] According to one or more embodiments of this disclosure, a method for generating live stream segments is provided, wherein selecting initial segments that meet the preset filtering criteria from the aggregated results includes: In the aggregation result, the longest segment in which the target object appears consecutively in time is selected as the initial slice.

[0105] According to one or more embodiments of this disclosure, a method for generating live stream segments is provided, wherein selecting initial segments that meet the preset filtering criteria from the aggregated results includes: Based on the first preset evaluation dimension, determine the score of each video frame of the aggregation result in the first preset evaluation dimension; In the aggregation results, segments whose scores are continuously greater than a first score threshold are selected as the initial slices.

[0106] According to one or more embodiments of this disclosure, a method for generating live broadcast segments is provided, wherein the first preset evaluation dimension includes at least one of a picture quality dimension and an interaction dimension; The image quality dimension is related to at least one of the following: image clarity, on-screen appearance of the anchor, and stillness of the image. The interaction dimension is related to the interaction data corresponding to the live stream segment where the target object is located.

[0107] According to one or more embodiments of this disclosure, a method for generating live stream segments is provided, wherein the highlight segment identification is performed on an initial slice of the target object to extract the highlight segment of the target object, comprising: Based on a second preset evaluation dimension, a comprehensive score is determined for each processing unit in the initial slice according to the second preset evaluation dimension, wherein the processing unit includes at least one video frame: Based on the comprehensive score of each processing unit, one or more consecutive processing units are extracted as the highlight segments of the target object.

[0108] According to one or more embodiments of this disclosure, a method for generating live broadcast segments is provided, wherein determining the comprehensive score of each processing unit in the initial slice based on the second preset evaluation dimension includes: The initial slice is divided into a first category of segments that contain explanatory content and a second category of segments that do not contain explanatory content. For the first type of fragment, the first type of fragment is divided into multiple semantic units, and each semantic unit is scored based on a second preset scoring dimension; For the second type of segment, the second type of segment is divided into multiple time units, and each time unit is scored based on a third preset scoring dimension; Based on the scores of the semantic units and the scores of the time units, a comprehensive score is determined for each of the processing units in the initial slice.

[0109] According to one or more embodiments of this disclosure, a method for generating live broadcast segments is provided, wherein the second preset evaluation dimension includes at least one of a visual dimension, a text dimension, and an interaction dimension; and the third preset evaluation dimension includes at least one of a visual dimension and an interaction dimension. The image dimensions are related to the on-screen presence of the target object and / or the anchor; The text dimension is related to the semantic quality of the content being explained; The interaction dimension is related to the interaction data corresponding to the live stream segment where the target object is located.

[0110] According to one or more embodiments of this disclosure, a live stream segment generation apparatus is provided, the apparatus comprising: The receiving module is used to receive the live video stream from the target live streaming room in real time. The first extraction module is used to extract video segments to be processed from the live video stream; The determination module is used to perform content understanding on the video segment and determine at least one object appearing in the video segment and the video sub-segment corresponding to each object. The second extraction module is used to extract the initial slice of the target object from the aggregation result by aggregating multiple video sub-segments corresponding to the same target object; The third extraction module is used to identify highlight segments in the initial slice of the target object in order to extract the highlight segments of the target object; An output module is used to output the highlight fragment of the target object during the live broadcast in the target live room.

[0111] According to one or more embodiments of the present disclosure, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the live broadcast segment generation method provided in any embodiment of the present disclosure.

[0112] According to one or more embodiments of this disclosure, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device is configured to execute the computer program in the storage device to implement the steps of the live segment generation method provided in any of the disclosed embodiments.

[0113] According to one or more embodiments of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the live segment generation method provided in any of the disclosed embodiments.

[0114] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0115] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0116] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A method for generating live stream segments, characterized in that, The method includes: Receive live video streams from the target live streaming room in real time; Extract video segments to be processed from the live video stream; Perform content understanding on the video segment to identify at least one object appearing in the video segment and the video sub-segment corresponding to each object; By aggregating multiple video segments corresponding to the same target object, the initial slice of the target object is extracted from the aggregation result; Spectral fragment identification is performed on the initial slice of the target object to extract the spectral fragments of the target object; During the live broadcast in the target live room, the highlight segment of the target object is output.

2. The method according to claim 1, characterized in that, The output of the highlight fragment of the target object includes: Generate an access link for accessing the target live stream room; The highlight clip of the target object and the access link are published to the content recommendation stream of the video platform.

3. The method according to claim 1, characterized in that, The step of performing content understanding on the video segment to determine at least one object appearing in the video segment and the corresponding video sub-segment for each object includes: The video segment is identified by a preset object detection model, which identifies at least one object in the video segment and the time period of each object, so as to determine the video sub-segment corresponding to each object. And / or, The speech corresponding to the video segment is converted into text using a preset speech recognition model, and at least one object being explained and the time period of each object are identified based on the text, so as to determine the video sub-segment corresponding to each object.

4. The method according to claim 1, characterized in that, The step of extracting the initial slice of the target object from the aggregation result by aggregating multiple video sub-segments corresponding to the same target object includes: For the target object, the corresponding video segments are aggregated to obtain the aggregation result; According to the preset screening criteria, the initial slices that meet the screening criteria are selected from the aggregation results.

5. The method according to claim 4, characterized in that, The step of selecting initial slices that meet the preset screening criteria from the aggregation results includes: In the aggregation result, the longest segment in which the target object appears consecutively in time is selected as the initial slice.

6. The method according to claim 4, characterized in that, The step of selecting initial slices that meet the preset screening criteria from the aggregation results includes: Based on the first preset evaluation dimension, determine the score of each video frame of the aggregation result in the first preset evaluation dimension; In the aggregation results, segments whose scores are continuously greater than a first score threshold are selected as the initial slices.

7. The method according to claim 6, characterized in that, The first preset evaluation dimension includes at least one of the image quality dimension and the interaction dimension; The image quality dimension is related to at least one of the following: image clarity, on-screen appearance of the anchor, and stillness of the image. The interaction dimension is related to the interaction data corresponding to the live stream segment where the target object is located.

8. The method according to claim 1, characterized in that, The step of identifying highlight fragments in the initial slice of the target object to extract highlight fragments of the target object includes: Based on a second preset evaluation dimension, a comprehensive score is determined for each processing unit in the initial slice according to the second preset evaluation dimension, wherein the processing unit includes at least one video frame: Based on the comprehensive score of each processing unit, one or more consecutive processing units are extracted as the highlight segments of the target object.

9. The method according to claim 8, characterized in that, The determination of the comprehensive score of each processing unit in the initial slice based on the second preset evaluation dimension includes: The initial slice is divided into a first category of segments that contain explanatory content and a second category of segments that do not contain explanatory content. For the first type of fragment, the first type of fragment is divided into multiple semantic units, and each semantic unit is scored based on a second preset scoring dimension; For the second type of segment, the second type of segment is divided into multiple time units, and each time unit is scored based on a third preset scoring dimension; Based on the scores of the semantic units and the scores of the time units, a comprehensive score is determined for each of the processing units in the initial slice.

10. The method according to claim 8, characterized in that, The second preset evaluation dimension includes at least one of the visual dimension, text dimension, and interaction dimension; the third preset evaluation dimension includes at least one of the visual dimension and interaction dimension. The image dimensions are related to the on-screen presence of the target object and / or the anchor; The text dimension is related to the semantic quality of the content being explained; The interaction dimension is related to the interaction data corresponding to the live stream segment where the target object is located.

11. A live broadcast segment generation device, characterized in that, The device includes: The receiving module is used to receive the live video stream from the target live streaming room in real time. The first extraction module is used to extract video segments to be processed from the live video stream; The determination module is used to perform content understanding on the video segment and determine at least one object appearing in the video segment and the video sub-segment corresponding to each object. The second extraction module is used to extract the initial slice of the target object from the aggregation result by aggregating multiple video sub-segments corresponding to the same target object; The third extraction module is used to identify highlight segments in the initial slice of the target object in order to extract the highlight segments of the target object; An output module is used to output the highlight fragment of the target object during the live broadcast in the target live room.

12. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program performs the steps of the method described in any one of claims 1-10.

13. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-10.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-10.

Citation Information

Cited By

  • A commodity correlation analysis method and system for live monitoring

    CN122617497A