Video polling playing method and device and storage medium

By analyzing the priority of target videos and dynamically adjusting the playback order, the problems of low-value information and delayed observation caused by fixed rotation rules are solved, realizing the flexibility and timely response of video rotation playback.

CN121887940APending Publication Date: 2026-04-17HANGZHOU HUACHENG SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU HUACHENG SOFTWARE TECH CO LTD
Filing Date
2025-12-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing video loop systems rely on fixed loop rules, resulting in a delay in the provision of a large amount of low-value information and the observation of target events, which affects the video loop playback effect.

Method used

By analyzing the target videos to be played and the target events, the target priority is determined, and compared with the video information of the currently playing videos in the polling playlist. The playback order is dynamically adjusted, and low-priority videos are replaced to play high-priority videos.

Benefits of technology

It improves the flexibility and effectiveness of video loop playback, transforming it from passive loop playback to proactive understanding and response, ensuring timely observation of important events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887940A_ABST
    Figure CN121887940A_ABST
Patent Text Reader

Abstract

The invention discloses a video polling playing method and device and a storage medium, the method is applied to a video polling system, and the method comprises the steps: responding to a detected target event in a target to-be-played video, analyzing the target to-be-played video and the target event, and obtaining a target priority of the target event; analyzing the current video information of each currently played video in the polling playlist, and determining a to-be-replaced video in each currently played video; comparing the current priority corresponding to the video to be replaced with the target priority to obtain a priority comparison result; and if the priority comparison result represents that the target priority is greater than the current priority, performing replacement playing processing on the to-be-replaced video according to the target to-be-played video. According to the scheme, the video polling playing effect can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of target detection technology, and in particular to a video loop playback method, device, and storage medium. Background Technology

[0002] In the field of target detection technology, multiple image acquisition devices are typically deployed in the scene to be detected to record video, and then the target event is analyzed based on the multi-channel video recordings.

[0003] Currently, in order to facilitate users to watch multiple video recordings, a video looping system is usually provided, which can automatically cycle through and play the video footage captured by multiple image acquisition devices.

[0004] However, the current video looping process mainly relies on fixed looping rules. Fixed looping rules may not only provide a lot of low-value information, but may also cause users to observe the target event in a delayed manner, affecting the effect of video looping playback. Summary of the Invention

[0005] This application provides at least one video loop playback method, apparatus, device, and computer-readable storage medium.

[0006] This application provides a video loop playback method, applied to a video loop system, comprising: in response to detecting a target event in a target video to be played, analyzing the target video to be played and the target event to obtain a target priority of the target video to be played; analyzing the current video information of each currently playing video in the loop playback list to determine a video to be replaced in each currently playing video; comparing the current priority corresponding to the video to be replaced with the target priority to obtain a priority comparison result; if the priority comparison result indicates that the target priority is greater than the current priority, performing replacement playback processing on the video to be replaced according to the target video to be played.

[0007] In one embodiment, the current video information includes the current priority. Analyzing the current video information of each currently playing video in the rotating playlist to determine the video to be replaced in each currently playing video includes: in response to the existence of multiple currently playing videos in the rotating playlist, obtaining the current priority of the multiple currently playing videos; and determining the currently playing video with the lowest current priority among the multiple currently playing videos as the video to be replaced based on the current priority.

[0008] In one embodiment, analyzing the current video information of each currently playing video in the rotating playlist to determine the video to be replaced in each currently playing video includes: determining the video overlap of each currently playing video based on the current video information of each currently playing video; and determining the video to be replaced from the currently playing videos based on the video overlap.

[0009] In one embodiment, the current video information includes at least one of current event semantics, current visual features, and current spatiotemporal features. Analyzing the current video information of each currently playing video in the rotating playlist to determine the video to be replaced in each currently playing video includes: determining the semantic similarity between each currently playing video based on the current event semantics; determining the visual similarity between each currently playing video based on the current visual features; determining the spatiotemporal similarity between each currently playing video based on the current spatiotemporal features; and determining the video overlap based on at least one of the semantic similarity, the visual similarity, and the spatiotemporal similarity.

[0010] In one embodiment, the step of replacing the video to be replaced based on the target video to be played if the priority comparison result indicates that the target priority is greater than the current priority includes: obtaining the event confidence of the target event; if the event confidence is greater than the confidence threshold and the priority comparison result indicates that the target priority is greater than the current priority, replacing the video to be replaced based on the target video to be played.

[0011] In one embodiment, after performing the replacement playback process on the video to be replaced based on the target video to be played, the method further includes: obtaining the waiting playback time of the video to be replaced; determining the waiting priority of the video to be replaced based on the waiting playback time; determining the comprehensive priority of the video to be replaced based on the waiting priority and the current priority of the video to be replaced; and determining the playback order of the video to be replaced in the round-robin playlist based on the comprehensive priority.

[0012] In one embodiment, the step of replacing the video to be replaced with the target video to be played includes: performing semantic analysis on the target video to be played to obtain a semantic summary of the target video to be played; and performing replacement playback on the video to be replaced with the target video to be played based on the semantic summary of the target video to be played.

[0013] In one embodiment, the step of performing replacement playback processing on the video to be replaced based on the semantic summary of the target video to be played includes: playing the semantic summary in a sub-region of the playback area where the video to be replaced is located; after performing replacement playback processing on the video to be replaced based on the semantic summary of the target video to be played, the method further includes: in response to the summary playback time of the semantic summary being greater than the summary playback threshold, or receiving a user's confirmation playback instruction for the target video to be played, playing the target video to be played in the playback area.

[0014] A second aspect of this application provides a video carousel playback device, comprising: an acquisition module, configured to, in response to detecting a target event in a target video to be played, analyze the target video to be played and the target event to obtain a target priority of the target video to be played; a video analysis module, configured to analyze the current video information of each currently playing video in the carousel playback list to determine a video to be replaced in each currently playing video; a priority comparison module, configured to compare the current priority corresponding to the video to be replaced with the target priority to obtain a priority comparison result; and a replacement playback module, configured to, if the priority comparison result indicates that the target priority is greater than the current priority, perform replacement playback processing on the video to be replaced according to the target video to be played.

[0015] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-described video rotation playback method.

[0016] The fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the above-described video rotation playback method.

[0017] The above scheme, by acquiring the target detection results of the video to be played, and when the target detection results indicate the presence of a target event in the video to be played, can determine the target priority of the video to be played by performing in-depth understanding and analysis of the target video and the target event. Then, it acquires the current playback list and analyzes the current video information of each currently playing video in the playlist to determine the video to be replaced. Next, it compares the current priority of the video to be replaced with the target priority of the video to be played to obtain a priority comparison result. If the priority comparison result indicates that the target priority is greater than the current priority, it means that the video to be played is more worthy of attention (play) than the video to be replaced, and therefore, the video to be replaced can be replaced based on the target video. This improves the flexibility of video playback, optimizes the effect of video playback, and realizes the transformation of the video playback system from the traditional "passive playback" to the "active understanding and response" rule-based approach of this application.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0020] Figure 1 This is a flowchart illustrating an exemplary embodiment of the video rotation playback method of this application; Figure 2 This is a schematic diagram of an exemplary edge-cloud collaborative architecture in the video polling playback method of this application; Figure 3 This is a block diagram illustrating a video loop playback device according to an exemplary embodiment of this application; Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device of this application; Figure 5 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0021] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0022] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0023] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0024] To facilitate understanding, one of the applicable scenarios of this application will be illustrated by example.

[0025] In the field of target detection technology, multiple image acquisition devices are typically deployed in the scene to be detected to record video separately, and then the target event is analyzed based on the acquired multi-channel video recordings.

[0026] Currently, in order to facilitate users to watch multiple video recordings, a video looping system is usually provided, which can automatically cycle through and play the video footage captured by multiple image acquisition devices.

[0027] However, the current video looping process mainly relies on fixed looping rules. Fixed looping rules may not only provide a lot of low-value information, but may also cause users to observe the target event in a delayed manner, affecting the effect of video looping playback.

[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating an exemplary embodiment of the video loop playback method of this application. The video loop playback method of this application can be applied to a video loop system, which may include a video playback management device that can communicate with multiple image acquisition devices (such as IP cameras, IPCs).

[0029] The video playback management device can determine the number of video playback splits N based on the currently available display resources of the video loop system. The number of video playback splits refers to the number of times each video can be played in a split-screen format; N is usually an integer greater than or equal to 1. Splitting the video can be done within a single display screen or within a display area (including multiple display screens); this is not limited here. The video content played in different splits can be the same or different (usually different); this is not limited here either.

[0030] Specifically, the video rotation playback method of this application may include at least the following steps: Step S110: In response to detecting a target event in the target video to be played, analyze the target video to be played and the target event to obtain the target priority of the target video to be played.

[0031] It should be noted that since video recording and patrol systems typically connect to multiple image acquisition devices (e.g., the number of image acquisition devices is greater than the number of screens), the video played on the display screen each time may only be a portion of all the captured videos.

[0032] Therefore, among the multiple video recordings captured, some videos may belong to the currently playing video (the video currently playing on the display screen), while other videos may belong to the currently waiting video (the video currently waiting to be played on the display screen).

[0033] For example, target detection processing can be performed on all or part of the acquired video recordings; this is not limited here. If a target event is detected in a video to be played, that is, the video to be played is a target video, the priority information (target priority) of the target video can be determined based on the target event.

[0034] The target event can refer to the detection of certain behaviors or events in the video recording, or the detection of the appearance, existence, or disappearance of a target object in the video recording, etc., without any limitation here.

[0035] Furthermore, the target event can refer to all events detectable in the video recording, or events further determined after filtering all detected events in the video recording according to certain rules; there is no limitation here. For example, filtering can be performed by event priority and event priority threshold, and / or by event confidence and event confidence threshold, and / or by the quality priority and quality priority threshold of the video corresponding to the event, etc., without limitation here.

[0036] Methods for determining the target priority of a target video to be played based on a target event may include, but are not limited to: determining the corresponding event priority based on the target event and setting the event priority as the target priority.

[0037] Alternatively, the quality priority of the target video to be played can be determined based on its video quality parameters, and the target priority can be determined based on event priority and / or quality priority. The video quality parameters may include, but are not limited to, one or more of the following: video brightness, video color, video signal-to-noise ratio, video dynamic range, and video bitrate.

[0038] There are various methods for determining quality priorities, and no specific method is specified here. For example, the quantized video quality parameters can be weighted and summed (or processed using other pre-defined mathematical functions), or the parameters can be input into a pre-trained video quality assessment neural network model to obtain the corresponding output quality priorities.

[0039] In addition, there are several methods to determine the target priority of a video to be played based on event priority and / or quality priority. For example, the quantized event priority and quality priority of the same target video can be weighted and summed (or processed by other preset mathematical functions), or the data can be input into a pre-trained priority evaluation neural network model to obtain the corresponding output target priority (which can also be referred to as video priority).

[0040] It should also be noted that, in addition to the target video to be played, this application may contain other types of videos, each with corresponding priority information. The method for determining the priority of other videos can be similarly described in the foregoing embodiments, and will not be repeated hereafter.

[0041] Optionally, the method for performing target detection processing on video recordings can be executed by the video recording loop system or by each image acquisition device separately; this is not limited here. Similarly, the method for analyzing the target priority of target events can be executed by the video recording loop system or by each image acquisition device separately; this is not limited here.

[0042] This application primarily uses the example of each image acquisition device performing target detection processing on its own acquired video recordings, and the video recording rotation system performing priority analysis processing on target events, to illustrate the application.

[0043] In this regard, a cloud-edge collaborative architecture can be constructed, as shown in the following example. Figure 2 As shown, Figure 2 This is an exemplary edge-cloud collaborative architecture diagram of the video rotation playback method in this application. Each image acquisition device is treated as an edge (edge ​​node), or the edge server to which each image acquisition device is communicatively connected is treated as an edge (the edge server can also have a direct or indirect communication connection with the video rotation system), and the video rotation system is treated as the cloud (central cloud node). This allows for asynchronous collaborative implementation of the video rotation playback method.

[0044] For example, configure the communication protocol between the central cloud and edge nodes. Deploy lightweight object detection models (such as the YOLO series) on the edge node side for object detection processing, so that object detection processing can be performed on the video collected by each edge node separately.

[0045] Firstly, if an edge node detects a target event in a video to be played, it can upload the video to the central cloud, where the central cloud will analyze the target priority of the video. Alternatively, secondly, if an edge node detects a target event in a video to be played, it can analyze the target priority of the video and then send the video and its priority to the central cloud.

[0046] Specifically, let's take the example scenario from the first aspect as an example for illustration: 1. Lightweight preprocessing at edge nodes: Deploy a lightweight object detection model (such as YOLO) on each IPC camera or the edge server to which the IPC camera communicates. The main task of this object detection model is to perform high-speed filtering on the acquired video recordings.

[0047] This means that the target video corresponding to the target event will only be uploaded to the central cloud when the target detection model detects a preset target event (such as the appearance of a person, a vehicle, or abnormal behavior).

[0048] Since video recordings are typically recorded continuously over a period of time, a complete video may contain segments where the target event does not occur, while other segments do. Therefore, when edge nodes upload the target video corresponding to the target event to the central cloud, they can either upload the complete video of the target event or a portion of the complete video (or extract keyframe data corresponding to the target event). In other words, the target video can refer to a complete video recording, a portion of a video segment, or a portion of video keyframes; there is no specific limitation here.

[0049] The preprocessing method described above directly reduces the upload of a large amount of meaningless data, greatly alleviating the network bandwidth and computing pressure on the central cloud.

[0050] 2. Central Cloud Deep Semantic Analysis: The central cloud is primarily used to receive and process high-value video clips (target videos to be played) uploaded after filtering by edge nodes. Then, pre-trained neural network models (such as multimodal large models) can be used for deep semantic understanding and semantic summarization generation. Multimodal large models can fuse and analyze image data, audio data, and other sensor data from the video to achieve comprehensive semantic understanding.

[0051] Due to the effectiveness of preprocessing, the amount of data in the semantic analysis process is significantly reduced, allowing central cloud resources to more centrally and quickly complete the analysis of high-value content.

[0052] 3. Asynchronous Edge-Cloud Operation: The processing procedures handled by each edge node and the central cloud can be parallel and asynchronous. Each edge node can continuously collect and filter video recordings and upload them immediately. The central cloud's large-scale model stores the received video recordings in a queue for analysis.

[0053] In other words, once the analysis results of a high-priority target event are completed, the transient response of the central cloud can be triggered immediately without waiting for all image acquisition devices in the current acquisition cycle to finish their acquisition processes.

[0054] In summary, the edge-cloud collaborative architecture provided in this application can achieve a "near real-time response" effect during video rotation playback. This overcomes the contradiction between the computationally intensive and time-consuming analysis of neural network models (especially large multimodal models) and the real-time requirements of video rotation playback systems.

[0055] Step S120: Analyze the current video information of each currently playing video in the rotating playlist to determine the video to be replaced in each currently playing video.

[0056] Based on the steps described above, the central cloud can store each video recording obtained from each edge node into a rotating playback list. This rotating playback list can include the currently playing video and / or videos awaiting playback.

[0057] Since a target video with a target event has been detected, and the priority of the target video may be higher than the priority of one or more currently playing videos, this application needs to analyze whether to replace a video in the currently playing video with the target video.

[0058] If the captured video does not contain the target event, a rotation playlist can be generated periodically according to a preset rotation playback strategy, and each video in the rotation playlist can be played in rotation. The rotation playback strategy can be set as needed, referring to relevant rotation strategies in this technical field (e.g., rotation by priority order, preset order rotation, random order rotation, rotation by capture time, rotation by capture location, etc.), and is not limited here.

[0059] The video to be replaced refers to the video currently playing that may be replaced by the target video to be played.

[0060] It should be noted that video information can include one or more elements (such as priority, event type, video frame, video quality, video duration, video position, etc.), and there are no restrictions here. There can also be one or more methods for analyzing the current video information of each currently playing video, and there are no restrictions here.

[0061] For example, similarly, the priority comparison method can be used to determine the video to be replaced among the currently playing videos based on their priority (e.g., determining the currently playing video with the lowest priority as the video to be replaced). It is also possible to analyze the overlap between the currently playing videos and determine the currently playing video with the highest overlap as the video to be replaced, etc., which will not be elaborated on here.

[0062] It should also be noted that there may be one or more target videos to be played, so there may also be one or more currently playing videos to be replaced (videos to be replaced), which is not limited here.

[0063] Step S130: Compare the current priority of the video to be replaced with the target priority to obtain the priority comparison result.

[0064] To illustrate the steps outlined above, after identifying the video to be replaced in the currently playing video, it is necessary to compare the current priority of the video to be replaced with the target priority of the target video to be played, and obtain the priority comparison result.

[0065] Based on the priority comparison results, it is determined which video to be replaced has a higher priority (or higher value) than the target video to be played, and then it is determined whether to replace the video to be replaced with the target video to be played.

[0066] Step S140: If the priority comparison result indicates that the target priority is greater than the current priority, the replacement video is played according to the target video to be played.

[0067] Based on the steps described above, if the priority comparison result indicates that the target priority is greater than the current priority, it means that the target video to be played has a higher priority (higher value) than the video to be replaced. Therefore, the video to be replaced can be replaced based on the target video to be played.

[0068] Conversely, if the priority comparison result indicates that the target priority is less than or equal to the current priority, it means that the target video to be played has a lower priority (lower value) than the video to be replaced. Therefore, it is not necessary to replace the video to be replaced with the target video to be played, and the video to be replaced can still be played.

[0069] As can be seen, this application obtains the target detection results of the video to be played. When the target detection results indicate that a target event exists in the target video to be played, the target priority of the target video to be played can be determined through in-depth understanding and analysis of the target video to be played and the target event. Then, the current looping playback list is obtained, and the current video information of each currently playing video in the looping playback list is analyzed to determine the video to be replaced in each currently playing video. Then, by comparing the current priority of the video to be replaced with the target priority of the target video to be played, a priority comparison result is obtained. If the priority comparison result indicates that the target priority is greater than the current priority, it means that the target video to be played is more worthy of attention (play) than the video to be replaced. Therefore, the video to be replaced can be replaced based on the target video to be played. This can improve the flexibility of video looping playback, optimize the effect of video looping playback, and realize the transformation of the video looping playback system from the traditional "passive looping" to the "active understanding and response" rule of this application.

[0070] Based on the above embodiments, this embodiment describes one of the feasible methods for analyzing the current video information of each currently playing video in the round-robin playlist in step S120 to determine the video to be replaced in each currently playing video.

[0071] The current video information includes its current priority. Essentially, each currently playing video has a corresponding current priority. Additionally, each video awaiting playback also has a corresponding priority, which will not be elaborated upon here.

[0072] If it is necessary to determine whether to replace the currently playing video with the target video to be played, you can first determine the video that is relatively worth replacing (i.e., obtain the video to be replaced) from the currently playing videos based on their current priority.

[0073] Specifically, in this embodiment, the method for analyzing the current video information of each currently playing video in the rotating playlist and determining the video to be replaced in each currently playing video may include at least the following steps: In response to the presence of multiple currently playing videos in the rotating playlist, the current priority of each video is obtained. Based on the current priority, the video with the lowest current priority among the multiple currently playing videos is identified as the video to be replaced.

[0074] As described in conjunction with the foregoing embodiments, the rotating playlist may contain one or more currently playing videos.

[0075] If there are multiple currently playing videos in the rotating playlist, the method for selecting the video to be replaced from each currently playing video can include at least: obtaining the current priority of the multiple currently playing videos; comparing the current priorities of each currently playing video by priority comparison, and selecting the currently playing video with the lowest current priority as the video to be replaced.

[0076] If there is a currently playing video in the playlist, then the currently playing video can be directly selected as the video to be replaced.

[0077] Based on the above embodiments, this embodiment should also explain that: On the one hand, when the video priority of each video is determined based on the event priority, if there are multiple currently playing videos with the lowest current priority, the quality priority of these currently playing videos with the lowest current priority can be obtained. Similarly, the currently playing video with the lowest quality priority is selected from these currently playing videos with the lowest current priority as the video to be replaced (that is, the currently playing video with the lowest event priority and the lowest quality priority is selected as the video to be replaced).

[0078] On the other hand, when the video priority of each video is determined based on event priority and quality priority, if there are multiple currently playing videos with the lowest current priority, their event priorities can be compared. Similarly, the currently playing video with the lowest current priority and the lowest event priority is selected from these videos as the video to be replaced (i.e., the currently playing video with the lowest current priority and the lowest event priority is selected as the video to be replaced).

[0079] Furthermore, regardless of how the video priority is determined using the methods described above, the event priority of each currently playing video can be compared with a preset event priority threshold during the process of determining which currently playing video needs to be replaced.

[0080] The event priority threshold is primarily used to measure the importance of events occurring in a video. If the event priority of all currently playing videos is greater than the event priority threshold, the video to be replaced can be determined similarly using the example method described above. If there are multiple currently playing videos with event priorities less than or equal to the event priority threshold, the currently playing video with the lowest quality priority among those with event priorities less than or equal to the threshold can be selected as the video to be replaced. Therefore, when the event is not important, videos with lower quality can be prioritized for replacement.

[0081] Based on the above embodiments, this embodiment describes another feasible method for analyzing the current video information of each currently playing video in the round-robin playlist in step S120 to determine the video to be replaced in each currently playing video.

[0082] The foregoing embodiments illustrated how to determine the video to be replaced from each currently playing video based on priority comparison. This embodiment mainly focuses on determining the video to be replaced from each currently playing video based on video overlap.

[0083] The method for determining video overlap can be adaptively and selectively executed. For example, if there are multiple currently playing videos, the overlap between each currently playing video can be determined according to the example method in this application.

[0084] Alternatively, if multiple currently playing videos exist, first determine whether they include the same or highly similar events (this can be determined by one or more of the following: the same (or similar) event type, the same (or similar) target object, the same (or similar) video time, or the same (or similar) video location; this is not limited here). If at least two currently playing videos share the same or highly similar events, then calculate their video overlap for subsequent processing. Thus, the above method can be used to filter out videos of the same event from different perspectives for video overlap analysis and redundancy removal, reducing the computational load and improving processing efficiency.

[0085] Specifically, the method of this embodiment may include at least the following steps S121 to S122: Step S121: Determine the video overlap of each currently playing video based on the current video information of each currently playing video.

[0086] The current video information can include one or more of the following (e.g., priority, event type, video frame, video quality, video duration, video location, etc.), and there are no restrictions here.

[0087] Video overlap refers to the degree to which different videos have the same, similar, or nearly identical parts. To some extent, video overlap can also be understood as video similarity. For example, it can be used to compare whether any two currently playing videos have the same or similar event types, whether any two currently playing videos have the same all or part of the same video footage, or whether the capture locations of two currently playing videos are the same or close, etc. There are no restrictions here.

[0088] Methods for determining video overlap based on current video information may include, but are not limited to: determining the degree of overlap (video overlap) between two videos based on one or more types of video information as described in the above examples.

[0089] For example, to facilitate understanding, a simpler approach is used in the embodiments of this application. For instance, the video overlap of the two currently playing videos is determined based on the proportion of the same frame portion to the overall frame portion in the video frames of the two currently playing videos (specifically, the calculation principle of intersection-union ratio can be referred to, where the overall frame (overall field of view) formed by the two currently playing videos is determined as the union, and the same frame (overlapping field of view) formed by the two currently playing videos is determined as the intersection).

[0090] In addition, if we consider the above-mentioned multiple video information to calculate the video overlap, we can obtain it through techniques such as weighted summation or inputting a pre-trained video overlap evaluation neural network model, which will not be elaborated here.

[0091] Step S122: Determine the video to be replaced from the currently playing video based on the video overlap.

[0092] Based on the steps described above, after obtaining the video overlap between each currently playing video, one of the currently playing videos with the highest video overlap (video similarity) can be selected as the video to be replaced.

[0093] For example, a preset video overlap threshold can be used for filtering. If there are at least two currently playing videos with an overlap greater than the threshold, it indicates that these at least two currently playing videos are highly similar. Therefore, at least one currently playing video can be selected as the video to be replaced.

[0094] This allows for the removal of redundancy in the currently playing video on the display screen, resulting in richer information presented by the video.

[0095] It should be noted that if multiple currently playing videos have an event priority greater than the event priority threshold, even if these currently playing videos are highly similar (video overlap greater than the video overlap threshold), these currently playing videos can continue to play (without replacing them with the target video to be played), so as not to ignore important events and ensure that multiple perspectives are provided when important events occur.

[0096] It should also be noted that one or more event priority thresholds can be set. Setting multiple different event priority thresholds can better distinguish the importance of each event.

[0097] For example, when an event priority threshold is set, if the event priority of the currently playing video is greater than the event priority threshold, the video overlap threshold can be increased to ensure that when observing videos of important events, multiple perspectives of video playback are retained as much as possible, unless multiple videos are indeed highly similar (in which case one of them will be replaced).

[0098] The methods for increasing the video overlap threshold may include, but are not limited to: increasing the threshold based on a preset increase coefficient (such as multiplication); or increasing the threshold based on a preset increase value (such as addition).

[0099] Alternatively, it can be adjusted using a pre-defined adjustment function. It's understood that there are multiple ways to set this adjustment function as needed, and we won't limit it here. The following examples illustrate the main implementation principles of the adjustment function: The main characteristic of this adjustment function is that the video overlap threshold can be positively correlated with the event priority (or, under the condition that the event priority is greater than the event priority threshold, the video overlap threshold is positively correlated with the event priority). Therefore, when the event priority is higher (the corresponding video is more worthy of attention), the video overlap threshold can be increased to preserve as many viewing angles as possible.

[0100] Similarly, the video overlap threshold can be reduced by referring to the aforementioned example method, which will not be elaborated here.

[0101] For example, when multiple event priority thresholds are set, the event priority thresholds can include at least a first event priority threshold and a second event priority threshold, where the first event priority threshold is less than the second event priority threshold.

[0102] The first event priority threshold is mainly used to determine the video to be replaced from the currently playing videos, while the second event priority threshold is mainly used to determine whether to perform redundancy processing. When there is a high degree of similarity among multiple currently playing videos, if the event priorities of these currently playing videos are all greater than the second event priority threshold, then these currently playing videos will continue to play. Conversely, the method described in the previous embodiment can be used to determine the video to be replaced based on event priority (or video priority, quality priority), which will not be elaborated here.

[0103] Furthermore, the method for determining the video to be replaced from at least two currently playing videos with a video overlap greater than the video overlap threshold may include, but is not limited to: random selection; or selecting the currently playing video with the lowest quality priority as the video to be replaced.

[0104] Furthermore, the method for determining the video to be replaced based on video overlap in this embodiment and the method for determining the video to be replaced based on video priority in the aforementioned embodiments can be implemented in any one of them or in combination.

[0105] For example, in a combined implementation scenario, if there is only one video with the lowest video priority among all currently playing videos, it can be directly used as the video to be replaced. If there are multiple videos with the lowest video priority among all currently playing videos, the video to be replaced can be determined based on the video overlap among the multiple videos with the lowest video priority.

[0106] Based on the above embodiments, this embodiment describes the method for determining the video overlap of each currently playing video in step S121 according to the current video information of each currently playing video.

[0107] The current video information may include at least one of the following: current event semantics, current visual features, and current spatiotemporal features.

[0108] Current event semantics is obtained by extracting features from the semantic information of events that occur in the video (such as text features related to the event, such as event type and event content).

[0109] The current visual features are obtained by extracting features from the video images (such as image features related to the video, such as image texture and image color).

[0110] The current spatiotemporal characteristics are the relevant features of the acquisition time and acquisition space (or acquisition location) when the video is acquired.

[0111] Specifically, the method for determining the video overlap of each currently playing video based on the current video information of each currently playing video in this embodiment may include at least the following steps: Based on the current event semantics of each currently playing video, determine the semantic similarity between each currently playing video; based on the current visual features of each currently playing video, determine the visual similarity between each currently playing video; based on the current spatiotemporal features of each currently playing video, determine the spatiotemporal similarity between each currently playing video; determine the video overlap based on at least one of semantic similarity, visual similarity, and spatiotemporal similarity.

[0112] In conjunction with the foregoing embodiments, when calculating the video overlap between any two currently playing videos, at least one of the following data can be obtained: current event semantics, current visual features, and current spatiotemporal features of the two currently playing videos.

[0113] Then, feature similarity between the two currently playing videos can be calculated based on the selected data. For example, semantic similarity can be calculated based on the current event semantics of the two currently playing videos; visual similarity can be determined based on the current visual features of the two currently playing videos; and spatiotemporal similarity can be determined based on the current spatiotemporal features of the two currently playing videos.

[0114] The method for calculating similarity can refer to relevant technologies in this technical field, including but not limited to: cosine similarity, cosine distance, Euclidean distance, Manhattan distance, etc. The specific method can be set as needed and is not limited here.

[0115] Furthermore, video overlap can be determined based on at least one of semantic similarity, visual similarity, and spatiotemporal similarity. If only one type of similarity is chosen to represent video overlap, it can be directly assigned a value. If multiple similarity methods are chosen, the video overlap can be determined by weighted summation.

[0116] Based on the above embodiments, this embodiment describes the method of replacing the video to be replaced with the target video in step S140 if the priority comparison result indicates that the target priority is greater than the current priority.

[0117] It should be noted that when the target detection model detects a target event, it can also simultaneously output the detection confidence level of the target event (also known as the event confidence level). For specific methods, please refer to the target detection technology in this technical field, which will not be elaborated here.

[0118] Therefore, this embodiment provides a method to determine whether to perform replacement playback processing based on event confidence.

[0119] Specifically, the method of this embodiment may include at least the following steps: Obtain the event confidence score of the target event; if the event confidence score is greater than the confidence score threshold and the priority comparison result indicates that the target priority is greater than the current priority, replace the video to be replaced with the target video to be played.

[0120] Among them, event confidence reflects the confidence level of the image acquisition device when detecting a target event, and also characterizes whether the target event (or the detection result of the target video to be played) is credible. Therefore, the event confidence can be compared with a preset confidence threshold, and then it can be determined whether to perform replacement playback processing based on the comparison.

[0121] For example, the event confidence level may be compared with a preset confidence threshold in step S140, or the event confidence level may be compared with a preset confidence threshold before step S110; there is no limitation here.

[0122] For example, in step S140, the event confidence level is compared with a preset confidence threshold. If the event confidence level is greater than the confidence threshold and the priority comparison result indicates that the target priority is greater than the current priority, the video to be replaced is played according to the target video to be played. Otherwise, no replacement playback is performed.

[0123] For example, before step S110, the event confidence score is compared with a preset confidence score threshold. If the event confidence score is greater than the confidence score threshold, subsequent steps (such as step S110) can be executed. Otherwise, the edge node corresponding to the target event can be requested to re-process the target detection in the video to be played.

[0124] Alternatively, before step S110, the event confidence level can be compared with a preset confidence threshold. If the event confidence level is greater than the confidence threshold, subsequent steps can be executed. However, when executing step S110, it is necessary to first determine the initial priority of the target video to be played based on the target event (but the initial priority at this time may be unreliable due to the low event confidence level). Therefore, the initial priority can be downgraded to obtain the target priority.

[0125] There are various methods for downgrading, which can be configured as needed and will not be elaborated here. For example, in applications where priorities are divided by stages, the initial priority can be downgraded in stages. In applications where priorities are divided by numerical values, the quantified value of the initial priority can be downgraded.

[0126] Based on the above embodiments, this embodiment describes the method after step S140.

[0127] In the video rotation playback system, when the video is played periodically on the display screen according to a preset rotation strategy (which can be called steady-state mode), if a target video to be played is suddenly detected to have a target event and replaces the currently playing video (which can be called transient mode), the replaced currently playing video (the video to be replaced) can be temporarily stored in the rotation playback list to wait for playback. Generally, it will wait until the target video to be played finishes playing, or until the video to be replaced is rotated to again, and then continue playing the video to be replaced.

[0128] In this embodiment, after the video to be replaced is temporarily stored in the rotating playback list and waits, its priority can be flexibly adjusted according to its waiting time.

[0129] Specifically, the method in this embodiment after replacing the target video with the replacement video according to the target video to be played may include at least the following steps: Obtain the waiting playback time of the video to be replaced; determine the waiting priority of the video to be replaced based on the waiting playback time; determine the overall priority of the video to be replaced based on the waiting priority and the current priority of the video to be replaced; determine the playback order of the video to be replaced in the rotating playlist based on the overall priority.

[0130] The waiting time refers to the time it takes for a video to be replaced to enter the rotation playlist and wait to be played again.

[0131] It should be noted that when generating the playlist based on the videos captured by each image acquisition device during each polling cycle, the videos can be sorted from highest to lowest priority. This allows for the determination that higher-priority video recordings will be played first.

[0132] After the currently playing video is replaced, the scheduler of the video rotation playback system can assign a waiting priority to the replaced (interrupted) video item that increases (positively correlated) over time (waiting time for playback).

[0133] When generating the video carousel playlist for the next cycle, the current priority (which could be event priority and / or quality priority, etc.) and waiting priority of the videos to be played can be comprehensively calculated to form a "comprehensive priority" as the video priority, which is then used for sorting. This determines the playback order of the video to be replaced in the video carousel playlist.

[0134] This mechanism ensures that the video content that is replaced (interrupted) is not left unattended indefinitely, guaranteeing the system's "global coverage without omissions".

[0135] Based on the above embodiments, this embodiment describes the method of replacing the video to be replaced in step S140 according to the target video to be played.

[0136] Specifically, the method for replacing the target video with the replacement video in this embodiment may include at least the following steps: Semantic analysis is performed on the target video to be played to obtain a semantic summary of the target video to be played; the video to be replaced is then played based on the semantic summary of the target video to be played.

[0137] Referring to the foregoing embodiments, during the replacement playback process, the playback of the pane containing the video to be replaced can be interrupted, and the video stream of the target video to be played can be connected for playback.

[0138] Alternatively, you can choose to obtain a semantic summary of the target video to be played, for example, by generating a semantic summary based on the semantic features of the target event in the target video. Then, use the semantic summary of the target video to replace the original video for playback.

[0139] This allows for a clearer, more concise, and faster acquisition of relevant events occurring in the target video to be played.

[0140] Based on the above embodiments, this embodiment describes a method for replacing and playing a video based on the semantic summary of the target video to be played.

[0141] Specifically, the method of this embodiment may include at least the following steps: The method further includes playing a semantic summary in a sub-region of the playback area where the video to be replaced is located; after replacing the video to be replaced based on the semantic summary of the target video to be played, the method also includes playing the target video to be played in the playback area in response to the semantic summary playback time being greater than the summary playback threshold, or receiving a user's confirmation playback instruction for the target video to be played.

[0142] In conjunction with the foregoing embodiments, during the process of replacing the video to be played based on the semantic summary of the target video to be played, it is also possible to first display the semantic summary of the target video to be played (target event) in the pane of the split screen where the video to be replaced is located (the playback area where the video to be replaced is located), or in a sub-area of ​​the playback area where the video to be replaced is located, in a picture-in-picture format, or in a floating notification bar on the side of the screen (and it is also possible to play its keyframes).

[0143] Furthermore, the playback time of the semantic summary can be recorded and compared with a playback threshold (e.g., 3 seconds). Interactive buttons such as "Full-screen Playback" and "Cancel Playback" can be displayed to the user as needed. If the playback time exceeds the playback threshold, or if the user interacts with the interactive buttons (corresponding to a confirmation or cancellation command), the user can choose to play the target video stream (split-screen or full-screen playback), or cancel playback and switch to another video recording in the rotating playlist for replacement.

[0144] Based on the above embodiments, in order to facilitate a deeper understanding of the application scenarios of this application, this embodiment describes some exemplary settings of this application.

[0145] System Initialization: During initialization, device grouping can be performed: multiple IPC devices can be added to the management device and the number of screens N can be determined; Parameter Settings: a fixed acquisition period T (e.g., 10 seconds) can be set for video acquisition, and a correlation model between the number of screens, the number of devices, and the playback duration can be established; Multimodal Large Model Loading: pre-trained multimodal large models (e.g., GPT-4V, Gemini, etc.) can be loaded for video analysis; Module Deployment: lightweight object detection models (e.g., YOLO series) can be deployed on the edge side, and the communication protocol between the central cloud and the edge nodes can be configured.

[0146] Steady-state mode: As the basic mode for the system's polling playback of video recordings, it can acquire video recordings from all image acquisition devices synchronously or asynchronously within a fixed preset time period. It also utilizes a multimodal large model for event analysis and priority assessment. Furthermore, it allows the option to play semantic summaries of the video recordings from all image acquisition devices when switching to the next polling cycle. Steady-state mode ensures comprehensive and in-depth observation and understanding of the entire video recording system.

[0147] Transient Mode: In normal steady-state operation, once a high-priority event, as determined by the method of the aforementioned embodiments and verified by confidence, is ready, the system can immediately interrupt the current scheduled playback (one or more of the currently playing videos), occupy the resources originally belonging to the currently playing video to prioritize the playback of the target event, and switch to the real-time video stream of the target event for continuous detection. This mode ensures zero-latency response to important events.

[0148] Edge-cloud collaborative architecture: To reduce the computing load and transmission latency of the central cloud, the system can adopt an edge-central cloud collaborative architecture. Lightweight detection models are deployed on cameras or edge servers to filter the collected video data; the central cloud's multimodal large model only performs in-depth analysis on the filtered video segments.

[0149] Confidence assessment and feedback: The large model can output a confidence score for detected events (such as target events). The dynamic scheduler combines this confidence score with other relevant information about the event to make a comprehensive decision, avoiding misscheduling of video playback due to model "illusions". The system can also introduce a human feedback mechanism to continuously optimize the large model and scheduling strategy.

[0150] Progressive replacement playback and state management: To avoid frequent playback changes that could disrupt the user experience, the system supports non-intrusive prompts (such as floating windows, picture-in-picture, etc.) and progressive replacement playback of semantic summaries to the video stream. Simultaneously, a priority guarantee algorithm with time decay ensures that replaced video recordings can still be reviewed promptly, maintaining system task integrity.

[0151] Specifically, the method provided in the above embodiments can be used not only to replace the video in the steady-state mode when triggering transient mode, but also during the playback process in steady-state mode.

[0152] For example: All image acquisition devices synchronously acquire video recordings within the time period [t0, t0+T], perform edge preprocessing, and then send the recordings to the central cloud video playback system.

[0153] The video carousel playback system analyzes the events of each acquired video and sorts them according to a preset carousel playback strategy (e.g., priority sorting) to generate a carousel playlist for the current period. It can also cycle through the videos in the playlist in a steady-state mode.

[0154] To improve information richness and avoid having videos of the same event occupy multiple screens simultaneously, redundancy detection and de-redundancy processing can be performed on the videos displayed on each screen at the same time.

[0155] The methods for redundancy detection and deredundancy processing can be similarly described in the foregoing embodiments. For example, calculating video overlap (e.g., determining whether the event types are the same, geographical locations are adjacent, video times overlap, event semantics are similar, or the shooting perspective is repeated). These will not be elaborated upon here.

[0156] Alternatively, based on redundancy and video priority (event importance), it can be determined whether the same (or similar) videos on the display screen are played from a single perspective, multiple perspectives, or in a main-supplementary linkage.

[0157] Single-viewpoint playback can be achieved by automatically selecting the video stream with the best picture quality and / or the most complete semantic description (most semantic information) from multiple videos if their viewing angles are highly repetitive. It can also be noted that the video is actually "visible from multiple perspectives".

[0158] Multi-view playback: If the videos from different perspectives are complementary, multiple videos can be played simultaneously. Alternatively, if the videos from different perspectives have high redundancy but also high priority, multiple videos can be played simultaneously. Or, the system can generate a fused semantic summary from the semantic summaries of multiple videos. This fused semantic summary can extract several key frames from each perspective's video (obtained through keyframes of event occurrence segments or analysis based on video attention, video heatmaps, etc.) to form a comprehensive "event graph."

[0159] Main and auxiliary video playback: When multiple videos are played simultaneously in split-screen mode, one video can be selected to play in the larger main pane, while other videos play in the smaller auxiliary panes. The display size of the main pane is larger than that of the auxiliary panes on the same screen. Alternatively, semantic summaries of other videos can be played in the auxiliary panes; this is not limited here. The method for selecting the video to play in the main pane may include, but is not limited to, filtering based on quality and / or semantic information as described in the preceding embodiments; these methods will not be elaborated upon here.

[0160] Therefore, the system can generate a rotating playlist for the video recordings captured by the image acquisition device within the time period [t0, t0+T], and can select to rotate and play the video according to the generated rotating playlist within the time period [t0+T, t0+2T].

[0161] If, while the system is playing a looped playlist, the image acquisition device captures a recording (target video to be played) containing a target event within the time period [t0+T, t0+2T], then a transient mode is triggered. The system can then determine, using the methods described in the aforementioned embodiments, whether the target video to be played needs to replace the currently playing video; this will not be elaborated upon here.

[0162] Each image acquisition device may capture multiple recordings containing various events within the time period [t0+T, t0+2T]. All captured events can be used as the target event. Alternatively, to avoid frequent replacement playback of recordings of unimportant events, each event can be filtered and judged, and only after determining that an important target event has occurred can subsequent comparison with the currently playing video be performed.

[0163] Specifically, in traffic detection scenarios, traffic intersections in cities can be detected. The scenario can be a 4-screen split display, managing the access of 16 high-definition IPCs.

[0164] System parameters may include at least: The data collection time is T = 10 seconds.

[0165] Event priority: P1 (traffic accident / severe congestion) > P2 (vehicle violation) > P3 (normal traffic flow) > P4 (static scene).

[0166] Event confidence threshold = 0.85.

[0167] Lightweight edge-side model: YOLOv5s, which can be used to detect people and vehicles.

[0168] Workflow, for example: (1) [t0, t0+10 seconds] – Steady-state mode acquisition and analysis: Synchronous acquisition: 16-channel image acquisition devices simultaneously acquire and record video.

[0169] Edge preprocessing: The edge servers at each intersection can run YOLOv5s. When multiple vehicles stop and gather at t0+2 seconds in the recording of channel B3, they are marked as high-value video clips (target videos of the target event to be played) and immediately uploaded to the central cloud.

[0170] Central cloud in-depth analysis: At t0+4 seconds, the video clip of channel B3 was analyzed by the multimodal large model and the analysis result was generated: "A rear-end collision occurred between two vehicles in the southwest corner, and congestion occurred behind it", priority P1, confidence level 0.93.

[0171] Steady-state playback: If the screen is currently displaying a playlist generated within the [t0-10 seconds, t0] cycle, and is displaying a P2 level event ("Vehicle illegally changing lanes") in channel A1.

[0172] (2) [t0+4 seconds] – Transient mode triggered: Triggering and Decision: The dynamic scheduler receives the P1 level ready signal (confidence 0.93>0.85) from channel B3. Compared with the currently playing P2 level content, P1>P2, and the replacement playback condition is met.

[0173] Progressive prompt (optional): The system pops up a picture-in-picture window in the lower right corner of the screen, displays the keyframes and semantic summary of the rear-end collision in channel B3, and begins a 3-second countdown.

[0174] Replacement Playback: If the user does not cancel the playback of the video on channel B3 within the countdown, the system will automatically interrupt the playback of the current content on channel A1 (P2) and start playing the semantic summary of the P1 event and the video stream on channel B3 (which may include multi-angle keyframes and detailed text descriptions).

[0175] Switching to Real-Time Stream: After the semantic summary of channel B3 finishes playing, you can seamlessly switch to the real-time video stream of channel B3, allowing traffic management personnel to continuously observe the accident handling and traffic control situation.

[0176] (3) [t0+10 seconds, t0+12 seconds] – Steady-state mode recovery and new cycle of round-robin playback scheduling: After the aforementioned cycle of playback concludes, the system analyzes the remaining unplayed video recordings. Furthermore, the scheduler generates a new playlist. This list takes into account: The interrupted A1 channel P2 event has had its "wait priority" increased, so it can be added to the current cycle's playlist with priority.

[0177] By calculating the overall priority, it can be ensured that the video feeds from the remaining image acquisition devices are all arranged reasonably.

[0178] In some special cases, it can be considered that channel B3 has already been played in advance and is being monitored in real time, so its recent video content may not be scheduled again in this cycle.

[0179] (4) [t0+12 seconds, t0+20 seconds] – A new round of steady-state playback: The system plays a new rotating playlist. Meanwhile, the previously used split screen for channel B3 can either stop playing or continue displaying the live video stream from channel B3; this is not limited to either option. Alternatively, after a period of time, the B3 pane can be switched back to the steady-state rotating playlist via user manual operation or by the scheduler according to rules (such as after an incident has been resolved).

[0180] Based on the above embodiments, this embodiment should be noted that various weighted summation calculation methods may be involved in the above examples. The weighted summation calculation methods in different examples may be the same or different, and no limitation is made here. Furthermore, the specific weighted summation method and weight coefficients can be set as needed, and will not be elaborated here.

[0181] It should be further noted that the execution entity of the video rotation playback method can be a video rotation playback device. For example, the video rotation playback method can be executed by a terminal device, a server, or other processing devices. The terminal device can be a user equipment (UE), computer, mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, the video rotation playback method can be implemented by a processor calling computer-readable instructions stored in memory.

[0182] Figure 3 This is a block diagram illustrating a video loop playback device as shown in an exemplary embodiment of this application. Figure 3 As shown, the exemplary video loop playback device 300 includes: an acquisition module 310, a video analysis module 320, a priority comparison module 330, and a replacement playback module 340. Specifically: The acquisition module 310 is used to analyze the target video and the target event in response to the detection of a target event in the target video to be played, and to obtain the target priority of the target video to be played.

[0183] The video analysis module 320 is used to analyze the current video information of each currently playing video in the rotating playlist and determine the video to be replaced in each currently playing video.

[0184] The priority comparison module 330 is used to compare the current priority of the video to be replaced with the target priority to obtain the priority comparison result.

[0185] The replacement playback module 340 is used to replace the video to be replaced if the priority comparison result indicates that the target priority is greater than the current priority.

[0186] In this exemplary video carousel playback device, by acquiring the target detection results of the video to be played, and when the target detection results indicate the presence of a target event in the target video to be played, the target priority of the target video to be played can be determined through in-depth understanding and analysis of the target video to be played and the target event. Then, the current carousel playback list is acquired, and the current video information of each currently playing video in the carousel playback list is analyzed to determine the video to be replaced in each currently playing video. Then, by comparing the current priority corresponding to the video to be replaced with the target priority corresponding to the target video to be played, a priority comparison result is obtained. If the priority comparison result indicates that the target priority is greater than the current priority, it means that the target video to be played is more worthy of attention (play) than the video to be replaced, and therefore the video to be replaced can be replaced based on the target video to be played. This can improve the flexibility of video carousel playback, optimize the effect of video carousel playback, and realize the transformation of the video carousel playback system from the traditional "passive carousel" to the "active understanding and response" rule of this application.

[0187] It should be noted that the apparatus and method provided in the above embodiments belong to the same concept, and the specific ways in which each module and unit performs operations have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the apparatus provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above, and this is not a limitation.

[0188] The functions of each module can be found in the video rotation playback method implementation example, and will not be repeated here.

[0189] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 100 includes a memory 101 and a processor 102. The processor 102 is used to execute program instructions stored in the memory 101 to implement the steps in any of the above-described video loop playback method embodiments. In a specific implementation scenario, the electronic device 100 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 100 may also include mobile devices such as laptops and tablets, which are not limited here.

[0190] Specifically, processor 102 controls itself and memory 101 to implement the steps in any of the above-described video loop playback method embodiments. Processor 102 can also be referred to as a CPU (Central Processing Unit). Processor 102 may be an integrated circuit chip with signal processing capabilities. Processor 102 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 102 can be implemented using integrated circuit chips.

[0191] In this exemplary electronic device, by acquiring the target detection results of the video to be played, and when the target detection results indicate the presence of a target event in the target video to be played, the target priority of the target video to be played can be determined through deep understanding and analysis of the target video to be played and the target event. Then, the current rotating playback list is acquired, and the current video information of each currently playing video in the rotating playback list is analyzed to determine the video to be replaced in each currently playing video. Then, by comparing the current priority corresponding to the video to be replaced with the target priority corresponding to the target video to be played, a priority comparison result is obtained. If the priority comparison result indicates that the target priority is greater than the current priority, it means that the target video to be played is more worthy of attention (play) than the video to be replaced, and therefore the video to be replaced can be replaced based on the target video to be played. This can improve the flexibility of video rotating playback, optimize the effect of video rotating playback, and realize the transformation of the video rotating playback system from the traditional "passive rotation" to the "active understanding and response" rule of this application.

[0192] Please see Figure 5 , Figure 5 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 110 stores program instructions 111 that can be executed by a processor. The program instructions 111 are used to implement the steps in any of the above-described video loop playback method embodiments.

[0193] In this exemplary storage medium, by running program instructions within the storage medium, the target detection results of the video to be played are obtained. When the target detection results indicate the presence of a target event in the target video to be played, the target priority of the target video to be played can be determined through deep understanding and analysis of the target video to be played and the target event. Then, the current rotating playback list is obtained, and the current video information of each currently playing video in the rotating playback list is analyzed to determine the video to be replaced in each currently playing video. The current priority of the video to be replaced is then compared with the target priority of the target video to be played to obtain a priority comparison result. If the priority comparison result indicates that the target priority is greater than the current priority, it means that the target video to be played is more worthy of attention (play) than the video to be replaced, and therefore, the video to be replaced can be replaced based on the target video to be played. This improves the flexibility of video rotating playback, optimizes the effect of video rotating playback, and realizes the transformation of the video rotating playback system from the traditional "passive rotation" to the "active understanding and response" rule of this application.

[0194] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0195] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0196] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0197] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for video loop playback, characterized in that, The method is applied to a video recording patrol system, and the method includes: In response to detecting a target event in a target video to be played, the target video to be played and the target event are analyzed to obtain the target priority of the target video to be played; Analyze the current video information of each currently playing video in the rotating playlist to determine the video to be replaced in each currently playing video; The current priority of the video to be replaced is compared with the target priority to obtain the priority comparison result; If the priority comparison result indicates that the target priority is greater than the current priority, the replacement video is replaced and played according to the target video to be played.

2. The method according to claim 1, characterized in that, The current video information includes the current priority. The analysis of the current video information of each currently playing video in the polling playlist to determine the video to be replaced in each currently playing video includes: In response to the presence of multiple currently playing videos in the polling playlist, the current priority of the multiple currently playing videos is obtained; Based on the current priority, the currently playing video with the lowest current priority among multiple currently playing videos is determined as the video to be replaced.

3. The method according to claim 1, characterized in that, The step of analyzing the current video information of each currently playing video in the rotating playlist to determine the video to be replaced in each currently playing video includes: Based on the current video information of each currently playing video, determine the video overlap of each currently playing video; The video to be replaced is determined from the currently playing video based on the video overlap.

4. The method according to claim 3, characterized in that, The current video information includes at least one of current event semantics, current visual features, and current spatiotemporal features. The step of analyzing the current video information of each currently playing video in the rotating playlist to determine the video to be replaced in each currently playing video includes: Determine the semantic similarity between the currently playing videos based on the current event semantics of each video. Determine the visual similarity between the currently playing videos based on their current visual features. Based on the current spatiotemporal characteristics of each currently playing video, determine the spatiotemporal similarity between each currently playing video; The video overlap is determined based on at least one of the semantic similarity, the visual similarity, and the spatiotemporal similarity.

5. The method according to claim 1, characterized in that, If the priority comparison result indicates that the target priority is greater than the current priority, the replacement video is replaced and played according to the target video to be played, including: Obtain the event confidence level of the target event; If the event confidence level is greater than the confidence threshold, and the priority comparison result indicates that the target priority is greater than the current priority, the replacement video is replaced and played according to the target video to be played.

6. The method according to claim 1, characterized in that, After performing the replacement playback process on the video to be replaced based on the target video to be played, the method further includes: Obtain the waiting playback time of the video to be replaced; The waiting priority of the video to be replaced is determined based on the waiting playback time; The overall priority of the video to be replaced is determined based on the waiting priority and the current priority of the video to be replaced. The playback order of the video to be replaced in the rotating playlist is determined based on the overall priority.

7. The method according to claim 1, characterized in that, The process of replacing the video to be replaced with the target video to be played includes: Semantic analysis is performed on the target video to be played to obtain a semantic summary of the target video to be played. The video to be replaced is replaced and played based on the semantic summary of the target video to be played.

8. The method according to claim 7, characterized in that, The step of replacing the video to be played based on the semantic summary of the target video to be played includes: Play the semantic summary in a sub-region of the playback area where the video to be replaced is located; After performing replacement playback processing on the video to be replaced based on the semantic summary of the target video to be played, the method further includes: In response to the semantic summary playback time being greater than the summary playback threshold, or upon receiving a user's confirmation playback instruction for the target video to be played, the target video to be played is played in the playback area.

9. An electronic device, characterized in that, The method includes a memory and a processor, the processor being configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the method described in any one of claims 1 to 8.