A method and system for intelligent video analysis and synchronous streaming playback

CN117459754BActive Publication Date: 2026-09-01HUIZHIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311262281.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-09-01
Estimated Expiration
2043-09-27

AI Technical Summary

Benefits of technology

[0033]本发明提供的技术方案,通过识别视频流的语义,可以将视频流进行聚合。同一个集合中的视频流,可以具备相同的语义,后续,可以将具备相同语义的视频流的视频画面进行拼接,从而得到能够表征视频语义的合成视频画面。后续只需要向用户推送该合成视频画面对应的视频流即可。这样,既能够节省数据传输的流量,同时也不会遗漏视频流集合中的主要内容。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117459754B_ABST
    Figure CN117459754B_ABST
Patent Text Reader

Abstract

This invention provides an intelligent video analysis and synchronous streaming playback method and system. The method includes: acquiring multiple raw video streams and parsing the image information of the raw video streams to determine the semantics represented by the raw video streams; aggregating the multiple video streams according to the semantics, and for a set of video streams with the same semantics, acquiring the video frames of each video stream in the video stream set after alignment with the time axis; cropping and splicing the acquired video frames according to the semantics corresponding to the video stream set to integrate the acquired multiple video frames into a single frame of composite video; and using the video stream composed of the composite video frames as the integrated video stream corresponding to the video stream set, and pushing the integrated video stream to the client. The technical solution provided by this invention can meet user needs while saving data traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to an intelligent video analysis and synchronous streaming playback method and system. Background Technology

[0002] Currently, when video servers push streams to clients, they may push multiple video streams of the same type to the client. However, in reality, depending on the client's needs, usually only one type of video stream needs to be pushed. But if only one type of video stream is pushed, information from other video streams of the same type will be lost.

[0003] Therefore, there is a need for a video streaming method that can save bandwidth while meeting user needs. Summary of the Invention

[0004] This invention provides an intelligent video analysis and synchronous streaming playback method and system that can meet user needs while saving data traffic.

[0005] In view of this, the present invention provides an intelligent video analysis and synchronous streaming playback method, the method comprising:

[0006] Acquire multiple raw video streams and parse the image information of the raw video streams to determine the semantics represented by the raw video streams;

[0007] The multiple video streams are aggregated according to semantics, and for a set of video streams with the same semantics, the video frames of each video stream in the set are obtained after aligning the time axis.

[0008] Based on the semantics corresponding to the video stream set, the acquired video frames are cropped and spliced ​​to integrate multiple acquired video frames into a single frame for composite video display.

[0009] The video stream composed of the composite video frames is used as the integrated video stream corresponding to the video stream set, and the integrated video stream is pushed to the client.

[0010] In one embodiment, cropping and splicing the acquired video footage includes:

[0011] The main subjects in each video frame are analyzed, and the analyzed main subjects are classified according to the main subjects required to form a composite video frame.

[0012] Select a representative subject from each category of subject matter, and then crop the representative subject out of the original video footage;

[0013] The various representative images obtained from the cropping are then spliced ​​together according to their original positions in the video frame.

[0014] In one implementation, selecting a representative subject from each category of subject areas includes:

[0015] Generate a sharpness score for each subject in each category of images, and select the subject with the highest sharpness score as the representative subject in each category.

[0016] In one embodiment, if integrating multiple video frames fails to yield a complete composite video frame, the method further includes:

[0017] A masked video frame is obtained after integrating multiple video frames, and the masked video frame is input into a content prediction model so as to output the inference image of the masked video frame through the content prediction model.

[0018] The inference image output by the content prediction model is filled into the mask portion of the masked video frame to form a complete synthetic video image.

[0019] In one implementation, outputting the masked inference frame from the masked video frame through the content prediction model includes:

[0020] The content prediction model divides the known screen content in the masked video frame into multiple video blocks, and queries the content index of each video block according to a preset content index directory.

[0021] The retrieved content index is filled into the corresponding position of the video block to obtain the index sequence of the masked video frames;

[0022] Predict the content index at the missing position in the index sequence, and restore the predicted content index to a video block according to the content index directory. The restored video block is used as the inference screen that is masked.

[0023] In another aspect, the present invention provides an intelligent video analysis and synchronous streaming playback system, the system comprising:

[0024] A semantic parsing unit is used to acquire multiple raw video streams and parse the image information of the raw video streams to determine the semantics represented by the raw video streams;

[0025] The image acquisition unit is used to aggregate the multiple video streams according to semantics, and for a set of video streams with the same semantics, acquire the video images of each video stream in the set after aligning the time axis.

[0026] The integration unit is used to crop and splice the acquired video frames according to the semantics corresponding to the video stream set, so as to integrate the acquired multiple video frames into a single frame of composite video frame;

[0027] The push unit is used to take the video stream composed of the composite video frames of each frame as the integrated video stream corresponding to the video stream set, and push the integrated video stream to the client.

[0028] In one embodiment, the integration unit is specifically used to: analyze the main subjects in each of the video frames, and classify the analyzed main subjects according to the main subjects required to form a composite video frame; select a representative main subject from each category of main subjects, and crop the representative main subject from the original video frame; and splice the cropped representative main subjects according to their original positions in the video frame.

[0029] In one embodiment, the integration unit is further configured to generate a sharpness score for each subject in each category of subjects, and to use the subject with the highest sharpness score as the representative subject in each category of subjects.

[0030] In one embodiment, the system further includes:

[0031] The inference unit is configured to, if integrating multiple video frames fails to produce a complete composite video frame, acquire a masked video frame after integrating the multiple video frames, input the masked video frame into a content prediction model, and output the inference frame that has been masked through the content prediction model; and fill the masked portion of the masked video frame with the inference frame output by the content prediction model to form a complete composite video frame.

[0032] In one embodiment, the inference unit is specifically used for: the content prediction model dividing the known screen content in the masked video frame into multiple video blocks, and querying the content index of each video block according to a preset content index directory; filling the queried content index into the corresponding position of the video block to obtain the index sequence of the masked video frame; predicting the content index at the missing position in the index sequence, and restoring the predicted content index into a video block according to the content index directory, wherein the restored video block is used as the inference screen to be masked.

[0033] The technical solution provided by this invention aggregates video streams by recognizing their semantics. Video streams within the same set can possess the same semantics. Subsequently, video frames from these semantically identical streams can be stitched together to obtain a composite video frame that represents the video's semantics. Then, only the video stream corresponding to this composite video frame needs to be pushed to the user. This saves data transmission bandwidth while ensuring that the main content of the video stream set is not missed.

[0034] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0035] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0036] Figure 1 This is a schematic diagram illustrating the steps of an intelligent video analysis and synchronous streaming playback method in an embodiment of the present invention;

[0037] Figure 2 This is a schematic diagram of the functional modules of an intelligent video analysis and synchronous streaming playback system according to an embodiment of the present invention. Detailed Implementation

[0038] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0039] Please see Figure 1 One embodiment of this application provides an intelligent video analysis and synchronous streaming playback method, the method comprising:

[0040] S1: Acquire multiple raw video streams and parse the image information of the raw video streams to determine the semantics represented by the raw video streams;

[0041] S2: Aggregate the multiple video streams according to semantics, and for a set of video streams with the same semantics, obtain the video frames of each video stream in the set after aligning the time axis;

[0042] S3: Based on the semantics corresponding to the video stream set, the acquired video frames are cropped and spliced ​​to integrate the acquired multiple video frames into a single frame to synthesize a video frame;

[0043] S4: Take the video stream composed of the synthesized video frames of each frame as the integrated video stream corresponding to the video stream set, and push the integrated video stream to the client.

[0044] In some application scenarios, servers receive numerous video streams forwarded by users, which often exhibit high repetition. Directly pushing these streams to other users would be a waste of bandwidth and would also result in users receiving highly repetitive video streams. Therefore, this application employs semantic recognition of the video streams to obtain their semantic representation.

[0045] Specifically, an attention-based LSTM model can be used to analyze video frames in a video stream to obtain semantic information. The combination of these semantic information frames constitutes the semantic information of the entire video stream. This semantic information can be quantized into vector parameters. By calculating the distance between these vector parameters, it can be determined whether two video streams are semantically similar enough. Subsequently, video streams with sufficiently similar semantics (e.g., over 80%) can be aggregated into the same video stream set.

[0046] In one embodiment, cropping and splicing the acquired video footage includes:

[0047] The main subjects in each video frame are analyzed, and the analyzed main subjects are classified according to the main subjects required to form a composite video frame.

[0048] Select a representative subject from each category of subject matter, and then crop the representative subject out of the original video footage;

[0049] The various representative images obtained from the cropping are then spliced ​​together according to their original positions in the video frame.

[0050] In practical applications, a composite video frame can contain multiple desired subjects. For example, a composite video frame can include a 16x16 area, each representing a different subject. Some areas might represent close-up environmental scenes, others distant environmental scenes, some close-up scenes of people, and others distant scenes of people. By parsing the subjects in the video stream and classifying them according to the subjects needed for the composite video frame, different sets of subjects can be obtained. Only one representative subject is needed from each set. Subsequently, these representative subjects can be stitched together to obtain the composite video frame.

[0051] Since video streams in the same set have sufficiently similar semantics, the content represented by the video frames is also basically similar. Therefore, after cropping and splicing, the synthesized video frames will not look out of place and may also contain some detailed information.

[0052] In one implementation, selecting a representative subject from each category of subject areas includes:

[0053] Generate a sharpness score for each subject in each category of images, and select the subject with the highest sharpness score as the representative subject in each category.

[0054] In one embodiment, if integrating multiple video frames fails to yield a complete composite video frame, the method further includes:

[0055] A masked video frame is obtained after integrating multiple video frames, and the masked video frame is input into a content prediction model so as to output the inference image of the masked video frame through the content prediction model.

[0056] The inference image output by the content prediction model is filled into the mask portion of the masked video frame to form a complete synthetic video image.

[0057] In one implementation, outputting the masked inference frame from the masked video frame through the content prediction model includes:

[0058] The content prediction model divides the known screen content in the masked video frame into multiple video blocks, and queries the content index of each video block according to a preset content index directory.

[0059] The retrieved content index is filled into the corresponding position of the video block to obtain the index sequence of the masked video frames;

[0060] Predict the content index at the missing position in the index sequence, and restore the predicted content index to a video block according to the content index directory. The restored video block is used as the inference screen that is masked.

[0061] In this embodiment, a content prediction model can be constructed using the transformer architecture. This content prediction model can encode video blocks into content indices, and then infer unknown (masked) content indices based on the existing content indices. The inferred content indices can then be restored to video blocks, thereby filling in the unknown parts in the masked video frames and forming a complete synthetic video image.

[0062] Please see Figure 2 In another aspect, the present invention provides an intelligent video analysis and synchronous streaming playback system, the system comprising:

[0063] A semantic parsing unit is used to acquire multiple raw video streams and parse the image information of the raw video streams to determine the semantics represented by the raw video streams;

[0064] The image acquisition unit is used to aggregate the multiple video streams according to semantics, and for a set of video streams with the same semantics, acquire the video images of each video stream in the set after aligning the time axis.

[0065] The integration unit is used to crop and splice the acquired video frames according to the semantics corresponding to the video stream set, so as to integrate the acquired multiple video frames into a single frame of composite video frame;

[0066] The push unit is used to take the video stream composed of the composite video frames of each frame as the integrated video stream corresponding to the video stream set, and push the integrated video stream to the client.

[0067] In one embodiment, the integration unit is specifically used to: analyze the main subjects in each of the video frames, and classify the analyzed main subjects according to the main subjects required to form a composite video frame; select a representative main subject from each category of main subjects, and crop the representative main subject from the original video frame; and splice the cropped representative main subjects according to their original positions in the video frame.

[0068] In one embodiment, the integration unit is further configured to generate a sharpness score for each subject in each category of subjects, and to use the subject with the highest sharpness score as the representative subject in each category of subjects.

[0069] In one embodiment, the system further includes:

[0070] The inference unit is configured to, if integrating multiple video frames fails to produce a complete composite video frame, acquire a masked video frame after integrating the multiple video frames, input the masked video frame into a content prediction model, and output the inference frame that has been masked through the content prediction model; and fill the masked portion of the masked video frame with the inference frame output by the content prediction model to form a complete composite video frame.

[0071] In one embodiment, the inference unit is specifically used for: the content prediction model dividing the known screen content in the masked video frame into multiple video blocks, and querying the content index of each video block according to a preset content index directory; filling the queried content index into the corresponding position of the video block to obtain the index sequence of the masked video frame; predicting the content index at the missing position in the index sequence, and restoring the predicted content index into a video block according to the content index directory, wherein the restored video block is used as the inference screen to be masked.

[0072] The technical solution provided by this invention aggregates video streams by recognizing their semantics. Video streams within the same set can possess the same semantics. Subsequently, video frames from these semantically identical streams can be stitched together to obtain a composite video frame that represents the video's semantics. Then, only the video stream corresponding to this composite video frame needs to be pushed to the user. This saves data transmission bandwidth while ensuring that the main content of the video stream set is not missed.

[0073] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for intelligent video analysis and synchronous streaming playback, characterized in that, The method includes: Acquire multiple raw video streams and parse the image information of the raw video streams to determine the semantics represented by the raw video streams; The multiple original video streams are aggregated according to semantics, and for a set of video streams with the same semantics, the video frames of each video stream in the set are obtained after aligning the time axis. Based on the semantics corresponding to the video stream set, the acquired video frames are cropped and spliced ​​to integrate multiple acquired video frames into a single frame for composite video display. The video stream composed of the composite video frames is used as the integrated video stream corresponding to the video stream set, and the integrated video stream is pushed to the client.

2. The method according to claim 1, characterized in that, Cropping and splicing the acquired video footage includes: The main subjects in each video frame are analyzed, and the analyzed main subjects are classified according to the main subjects required to form a composite video frame. Select a representative subject from each category of subject matter, and then crop the representative subject out of the original video footage; The various representative images obtained from the cropping are then spliced ​​together according to their original positions in the video frame.

3. The method according to claim 2, characterized in that, Select one representative subject from each category of subject matter, including: Generate a sharpness score for each subject in each category of images, and select the subject with the highest sharpness score as the representative subject in each category.

4. The method according to claim 1, characterized in that, If integrating multiple video frames fails to yield a complete composite video frame, the method further includes: A masked video frame is obtained after integrating multiple video frames, and the masked video frame is input into a content prediction model so as to output the inference image of the masked video frame through the content prediction model. The inference image output by the content prediction model is filled into the mask portion of the masked video frame to form a complete synthetic video image.

5. The method according to claim 4, characterized in that, The content prediction model outputs the inference image of the masked video frame that has been masked, including: The content prediction model divides the known screen content in the masked video frame into multiple video blocks, and queries the content index of each video block according to a preset content index directory. The retrieved content index is filled into the corresponding position of the video block to obtain the index sequence of the masked video frames; Predict the content index at the missing position in the index sequence, and restore the predicted content index to a video block according to the content index directory. The restored video block is used as the inference screen that is masked.

6. An intelligent video analysis and synchronous streaming playback system, characterized in that, The system includes: A semantic parsing unit is used to acquire multiple raw video streams and parse the image information of the raw video streams to determine the semantics represented by the raw video streams; The image acquisition unit is used to aggregate the multiple original video streams according to semantics, and for a set of video streams with the same semantics, acquire the video images of each video stream in the set after aligning the time axis. The integration unit is used to crop and splice the acquired video frames according to the semantics corresponding to the video stream set, so as to integrate the acquired multiple video frames into a single frame of composite video frame; The push unit is used to take the video stream composed of the composite video frames of each frame as the integrated video stream corresponding to the video stream set, and push the integrated video stream to the client.

7. The system according to claim 6, characterized in that, The integration unit is specifically used to: analyze the main subjects in each video frame and classify the analyzed main subjects according to the main subjects required to form a composite video frame; select a representative main subject from each category and cut out the representative main subject from the original video frame; and splice the cut-out representative main subjects according to their original positions in the video frame.

8. The system according to claim 7, characterized in that, The integration unit is further configured to generate a sharpness score for each subject in each category of subjects, and to use the subject with the highest sharpness score as the representative subject in each category of subjects.

9. The system according to claim 6, characterized in that, The system also includes: The inference unit is configured to, if integrating multiple video frames fails to produce a complete composite video frame, acquire a masked video frame after integrating the multiple video frames, input the masked video frame into a content prediction model, and output the inference frame that has been masked through the content prediction model; and fill the masked portion of the masked video frame with the inference frame output by the content prediction model to form a complete composite video frame.

10. The system according to claim 9, characterized in that, The inference unit is specifically used for the following: the content prediction model divides the known screen content in the masked video frame into multiple video blocks, and queries the content index of each video block according to a preset content index directory; fills the queried content index into the corresponding position of the video block to obtain the index sequence of the masked video frame; predicts the content index at the missing position in the index sequence, and restores the predicted content index into a video block according to the content index directory, and the restored video block serves as the inference screen to be masked.

Citation Information

Patent Citations

  • Multi-video-stream playing method and device thereof

    CN104822070A

  • Multipath multi-image video splicing device

    CN105578129A