Video tracing display method and device based on visual large model
Through the video tracing and display method based on the visual big model, the problems of low timeliness of event tracing and single information presentation in the existing video surveillance system are solved, the collaborative visualization display of panoramic and close-up videos is realized, and the retrieval efficiency and user experience are improved.
Patent Information
- Application Number
- CN202511152134.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing video surveillance systems have low timeliness in event tracing and single information presentation. They are unable to effectively combine panoramic and close-up videos for spatiotemporal collaborative presentation, resulting in low retrieval efficiency and poor user experience.
A method based on visual big models is used to extract static and dynamic targets in recorded videos, generate target relationship chains, and quickly filter and display panoramic and surveillance recorded videos based on user query information. Collaborative visualization of multi-source videos is performed by enlarging the display area, panoramic display area, and surveillance display area.
It achieves the rapid and accurate extraction and presentation of relevant video recordings, improves the timeliness of event tracing and user experience, and enhances the intelligent association and integrated display capabilities of multi-source videos.
Smart Images

Figure CN120640026A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video surveillance technology, and in particular to a method and device for retrospectively displaying recorded videos based on a visual macro model. Background Art
[0002] Current video surveillance systems generally use a time-based video retrieval mechanism. For example, Chinese patent CN116866534A discloses a processing method and device for digital video surveillance systems. Users must preconfigure a monitoring time period before they can retrieve and display the corresponding surveillance video. This approach presents the following problems: First, retrieval efficiency is low. When users cannot determine the specific time of an event (for example, only behavioral characteristics are known), they must manually segment and screen massive amounts of video footage, which is time-consuming and prone to missing key information, seriously affecting the timeliness of event tracing. Second, the information presented is monotonous. Existing technologies only play extracted video clips linearly in time, lacking intelligent correlation and integrated display of multi-source videos. In particular, it is unable to present the macro scenes of wide-angle panoramic videos in a spatiotemporal coordinated manner with the detailed behaviors of close-up surveillance videos. This forces users to repeatedly switch and compare, making it difficult to quickly construct a complete event chain.
[0003] Therefore, there is an urgent need for a large visual model-based method that can directly parse semantic query requests, automatically locate and extract target video clips, and achieve collaborative visualization of panoramic and close-up videos, fundamentally solving the dual bottlenecks of retrieval efficiency and presentation effect. Summary of the Invention
[0004] Based on this, it is necessary to provide a video tracing display method and device based on a visual large model that can quickly and accurately extract the corresponding video according to the user's question information and clearly present it to the user for playback, in order to address the above technical problems.
[0005] In a first aspect, the present invention provides a method for retrospectively displaying recorded videos based on a visual macro model, the method comprising:
[0006] Obtaining video footage, wherein the video footage includes panoramic video footage and surveillance video footage;
[0007] Extract static targets and dynamic targets in each frame of video, generate corresponding static target relationship chains and dynamic target relationship chains, and store the static target relationship chains and dynamic target relationship chains in text form;
[0008] Obtain the user's question information, filter out the corresponding feature information in the question information, and compare the feature information with the stored static target relationship chain and dynamic target relationship chain to filter out the target panoramic video and target surveillance video in the video;
[0009] The target panoramic video and the target monitoring video are displayed with preset effects on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area and a monitoring display area.
[0010] Optionally, extracting static targets and dynamic targets from each frame of video, generating corresponding static target relationship chains and dynamic target relationship chains, and storing the static target relationship chains and dynamic target relationship chains in text form, including:
[0011] Split the panoramic video and the surveillance video into corresponding panoramic video frames and surveillance video frames according to their frame rates;
[0012] Identify static targets and dynamic targets in each panoramic video frame and each surveillance video frame, and obtain corresponding static target attribute information and dynamic target attribute information;
[0013] Generate corresponding static target video recording sets and dynamic target video recording sets according to static target attribute information or dynamic target attribute information;
[0014] Based on the static target attribute information, the dynamic target attribute information, the static target video set and the dynamic target video set, a corresponding static target relationship chain or a dynamic target relationship chain is generated, and the static target relationship chain and the dynamic target relationship chain are stored in text form.
[0015] Optionally, the static target relationship chain includes: the name of the static target, the time when the static target appears in the recorded video, the position of the static target and the status of the static target; and / or, the dynamic target relationship chain includes: the name of the dynamic target, the time when the dynamic target appears in the recorded video, the motion trajectory of the dynamic target, the motion status of the dynamic target and the ID of the associated surveillance camera.
[0016] Optionally, the method further includes:
[0017] According to the motion trajectory of the dynamic target, determine whether the dynamic target is out of the target panoramic video;
[0018] If the dynamic target leaves the target panoramic video, the associated surveillance camera is determined according to the dynamic target relationship chain, and the target surveillance video after the dynamic target leaves the target panoramic video is found.
[0019] Optionally, the target panoramic video and the target surveillance video are displayed on the display interface of the client with a preset effect, including:
[0020] The enlarged display area is set at the upper part of the display interface, and its width is equal to the width of the display interface, and is used to play the enlarged target panoramic video or target monitoring video;
[0021] The panoramic display area is set in the middle of the display interface, between the magnified display area and the monitoring display area, and is used to play the target panoramic video;
[0022] The monitoring display area is set at the lower part of the display interface, below the panoramic display area, and is used to play and display the target monitoring video.
[0023] Optionally, the panoramic display area further includes a plurality of window playback containers and video axes of the same size;
[0024] The sum of the widths of the multiple window playback containers is equal to the width of the panoramic display area. After the target panoramic video is screened out, the target panoramic video is loaded into the window playback container for video playback;
[0025] The video axis is located below the multiple window playback containers, and the video axis correspondingly displays the playback time information of the target panoramic video in the multiple window playback containers.
[0026] Optionally, when the number of target panoramic video recordings is greater than the number of window playback containers, a temporary video storage container is established, and the target panoramic video recordings exceeding the number of window playback containers are stored in the temporary video storage container in chronological order as candidate target panoramic video recordings.
[0027] Optionally, detecting whether there is a candidate target panoramic video in the temporary video storage container that is played earlier than the target panoramic video played in the first window playback container in the panoramic display area;
[0028] If it exists, a first detection control is set with the left window boundary of the first window playback container as a reference point, wherein the width of the first detection control is less than the width of the first window playback container, and the height of the first detection control is greater than or equal to the height of the first window playback container;
[0029] When an event trigger operation is received on the first detection control, a candidate target panoramic video that is adjacent to the target panoramic video played in the first window playback container and precedes the target panoramic video in the first window playback container is extracted from the temporary video storage container, and loaded into the first window playback container. The target panoramic video in the last window playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is sequentially loaded into the next window playback container for playback;
[0030] and / or detecting whether there exists a candidate target panoramic video in the temporary video storage container whose playing time is later than the target panoramic video played in the last window playback container in the panoramic display area;
[0031] If it exists, a second detection control is set with the right window boundary line of the last window playback container as the reference point, wherein the width of the second detection control is less than the width of the last window playback container, and the height of the second detection control is greater than or equal to the height of the last window playback container;
[0032] When an event trigger operation is received for the second detection control, a candidate target panoramic video that is adjacent to and later than the target panoramic video played in the last window playback container is extracted from the temporary video storage container and loaded into the last window playback container, and the target panoramic video in the first window playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is loaded into the previous window playback container in sequence for playback.
[0033] Optionally, the monitoring display area further includes a monitoring window playback container corresponding to the window playback container; when a single target panoramic video is associated with multiple target monitoring video videos, the first target monitoring video video among the multiple target monitoring video videos is loaded into the monitoring window playback container for playback, and the remaining target monitoring video videos are displayed in the form of a time list below the corresponding monitoring window playback container.
[0034] In a second aspect, the present invention provides a video tracing and display device based on a visual macro model, the device comprising:
[0035] An acquisition module is used to acquire video recordings, wherein the video recordings include panoramic video recordings and surveillance video recordings;
[0036] The visual large model is connected to the acquisition module and is used to extract static and dynamic targets in each frame of video, generate corresponding static target relationship chains and dynamic target relationship chains, and store the static target relationship chains and dynamic target relationship chains in text form; and obtain user question information, filter out corresponding feature information in the question information, and compare the feature information with the stored static target relationship chains and dynamic target relationship chains to filter out target panoramic video and target monitoring video in the video;
[0037] The computing and display module is connected to the visual large model and the client respectively, and is used to display the target panoramic video and the target monitoring video with preset effects on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area and a monitoring display area.
[0038] The present invention provides a method and device for tracing and displaying recorded videos based on a visual macromodel. After acquiring a recorded video, the method and device extract static and dynamic targets from each frame of the recorded video, generate corresponding static and dynamic target relationship chains, and store the static and dynamic target relationship chains in text form. When a user enters a question, the method first filters out the corresponding feature information in the question, then compares the feature information with the stored static and dynamic target relationship chains to filter out the target panoramic and surveillance videos in the recorded video. Finally, the target panoramic and surveillance videos are displayed with a preset effect on the client's display interface. The method and device of the present invention can quickly and accurately extract the corresponding recorded video based on the user's question information and clearly present it to the user for playback. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1a A schematic diagram of a process for tracing and displaying recorded videos based on a visual macro model according to an embodiment of the present invention;
[0040] Figure 1b A schematic diagram showing a preset effect of a display interface of a video tracing display method based on a visual macro model provided by an embodiment of the present invention;
[0041] Figure 1c Another display diagram of a preset effect of the display interface of the video tracing display method based on the visual macro model provided by an embodiment of the present invention;
[0042] Figure 2 A schematic diagram of a circuit module structure of a video tracing and display device based on a visual macro model provided by an embodiment of the present invention;
[0043] Figure 3 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0045] like Figure 1a As shown, the present invention provides a method for retrospective display of recorded videos based on a visual macro model, the method comprising:
[0046] Step S11: Acquire recorded video, wherein the recorded video includes panoramic video and surveillance video;
[0047] Optionally, the panoramic video is obtained by capturing and storing panoramic videos of the corresponding scene with a panoramic camera, and the surveillance video is obtained by capturing and storing surveillance videos of the corresponding scene with surveillance cameras installed at different locations. The panoramic camera may be an LC150 panoramic megapixel camera, an NV180E panoramic megapixel camera, etc., and the surveillance camera may be a DS-2CD5A4XYZ-LS full-color intelligent security camera, a DH-IPC-HFW3449DM-A-LED camera, etc.
[0048] Step S12: extracting static targets and dynamic targets from each frame of video, generating corresponding static target relationship chains and dynamic target relationship chains, and storing the static target relationship chains and dynamic target relationship chains in text form;
[0049] The static object relationship chain includes: the static object's name, the time the static object appears in the recorded video, the static object's location, and the static object's status; and / or the dynamic object relationship chain includes: the dynamic object's name, the time the dynamic object appears in the recorded video, the dynamic object's motion trajectory, the dynamic object's motion status, and the ID of the associated surveillance camera. It should be noted that static objects can include streetlights, buildings, and other objects that are generally immobile; dynamic objects can include pedestrians, pets, bicycles, cars, and other objects.
[0050] The static target relationship chain stored in text form is: [static target name, time when the static target appears in the video, static target location, static target status] = [street lamp, panoramic view - 2012.5.1 12:30:30 - 2013.9.10 13:00:29, left middle of the screen - next to the bench, solar and wind powered street lamp].
[0051] The dynamic target relationship chain stored in text form is: [dynamic target name, time when the dynamic target appears in the video, dynamic target motion trajectory, dynamic target motion state, associated surveillance camera ID] = [pedestrian, panoramic-2012.5.1 12:30:30-2013.9.10 13:00:29, coordinates (54,67)-(54,90)-(54,120), pedestrian in red shirt jogging, 10256], or [dynamic target name, time when the dynamic target appears in the video, dynamic target motion trajectory, dynamic target motion state, associated surveillance camera ID] = [pedestrian, surveillance-2015.5.1 12:30:30-2017.9.10 13:00:29, coordinates (54,67)-(54,90)-(54,120), a pedestrian in a blue shirt is walking, 0], where the ID of the associated surveillance camera is 0, indicating that there is no associated surveillance camera.
[0052] Optionally, after receiving the video recording, the visual large model first divides the panoramic video recording and the surveillance video recording into corresponding panoramic video frames and surveillance video frames according to their frame rates; then identifies the static targets and dynamic targets in each panoramic video frame and each surveillance video frame, and obtains the corresponding static target attribute information and dynamic target attribute information; generates the corresponding static target video recording set and dynamic target video recording set according to the static target attribute information or the dynamic target attribute information; based on the static target attribute information, the dynamic target attribute information, the static target video recording set and the dynamic target video recording set, generates the corresponding static target relationship chain or dynamic target relationship chain, and stores the static target relationship chain and the dynamic target relationship chain in the form of text.
[0053] In the present invention, the visual large model can adopt the pre-trained visual large model in the prior art. Those skilled in the art can select it according to actual needs, and it is not limited here. For example, a visual large model based on the RAG system can be selected.
[0054] Among them, the static target attribute information includes: the name of the static target, the time when the static target appears in the recorded video, the position of the static target and the status of the static target; and / or, the dynamic target attribute information includes: the name of the dynamic target, the time when the dynamic target appears in the recorded video, the motion trajectory of the dynamic target, the motion status of the dynamic target and the ID of the associated surveillance camera.
[0055] It should be noted that the present invention not only allows for quick search of corresponding recorded videos using static and dynamic target relationship chains stored in text format, but also allows for the searched static and dynamic target relationship chains to be sent to users via text messages, etc. Users can directly play the corresponding recorded videos by clicking on the received static and dynamic target relationship chains. This method allows users to conveniently view recorded videos, and the static and dynamic target relationship chains stored in text format significantly conserve storage space.
[0056] Optionally, after identifying static targets and dynamic targets in each panoramic video frame and each surveillance video frame, the method of the present invention also includes: marking the identified static targets and dynamic targets in each panoramic video frame and each surveillance video frame for user viewing.
[0057] Optionally, the method of the present invention further comprises:
[0058] According to the motion trajectory of the dynamic target, determine whether the dynamic target is out of the target panoramic video;
[0059] If the dynamic target leaves the target panoramic video, the associated surveillance camera is determined according to the dynamic target relationship chain, and the target surveillance video after the dynamic target leaves the target panoramic video is found.
[0060] This method can effectively associate the target panoramic video and the target surveillance video and present them to the user, quickly and accurately achieving cross-lens tracking, saving a lot of manpower, material resources and other costs.
[0061] Step S13: Obtain the user's question information, filter out the corresponding feature information in the question information, and compare the feature information with the stored static target relationship chain and dynamic target relationship chain to filter out the target panoramic video and target surveillance video in the video;
[0062] For example, the user's question information is: From June 27, 2015 to June 30, 2015, where did a person wearing a red short-sleeved shirt and black trousers appear in the coal mine?
[0063] After receiving the question information, the visual big model screens the feature information therein, such as June 27, 2015 to June 30, 2015, red, short sleeves, black, trousers, and coal mine; after extracting the feature information, since it is a dynamic target pedestrian, the text similarity between the feature information and the dynamic target relationship chain is calculated, and the corresponding target panoramic video is screened out; according to the motion trajectory of the dynamic target, it is judged whether the dynamic target is out of the target panoramic video; if the dynamic target is out of the target panoramic video, the associated surveillance camera is determined according to the dynamic target relationship chain, and the target surveillance video after the dynamic target is out of the target panoramic video is found.
[0064] Step S14: Displaying the target panoramic video and the target monitoring video with a preset effect on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area, and a monitoring display area.
[0065] Alternatively, as Figure 1b As shown, step S14 specifically includes:
[0066] The enlarged display area 11 is provided at the upper portion of the display interface 10 and has a width equal to that of the display interface 10 , and is used to play the enlarged target panoramic video or target surveillance video;
[0067] A video play axis 110 may also be provided on the magnified display area 11 for displaying the play time and / or adjusting the play progress.
[0068] It should be noted that, when the user selects the corresponding target panoramic video or target surveillance video, the video played in the enlarged display area 11 will be replaced accordingly, and replaced with the target panoramic video or target surveillance video selected by the user.
[0069] The panoramic display area 12 is provided in the middle of the display interface 10, between the magnified display area 11 and the monitoring display area 13, and is used to play the target panoramic video;
[0070] The monitoring display area 13 is provided at the lower portion of the display interface 10 , below the panoramic display area 12 , and is used for playing and displaying target monitoring video.
[0071] In an optional embodiment of the present invention, Figure 1b As shown, the panoramic display area 12 further includes a plurality of window playback containers 121 and a video axis 122 of the same size;
[0072] The sum of the widths of the multiple window playback containers 121 is equal to the width of the panoramic display area 12. After the target panoramic video is screened out, the target panoramic video is loaded into the window playback container 121 for video playback.
[0073] The window playback container may be an mp4 container, an avi container, an mkv container, an flv container, a wmv container, etc., and is not limited here.
[0074] For example, in Figure 1b In the example, four window playback containers 121 are provided in the panoramic display area 12. The sum of the widths of the four window playback containers 121 is equal to the width of the panoramic display area 12. The four window playback containers 121 are loaded with target panoramic video clips a1, a2, a3, and a4, respectively. It should be noted that once the target panoramic video clips a1, a2, a3, and a4 are loaded into the corresponding window playback containers 121, they are automatically played, i.e., displayed to the user through video playback rather than as video images.
[0075] The video axis 122 is located below the multiple window playback containers 121 , and the video axis 122 correspondingly displays the playback time information of the target panoramic video in the multiple window playback containers 121 .
[0076] The playback time information includes the start playback time, the current playback time, the end playback time, etc. Those skilled in the art can flexibly set it according to actual needs, which is not limited here.
[0077] In another optional embodiment of the present invention, Figure 1cAs shown, when the number of target panoramic video recordings is greater than the number of window playback containers 121, a temporary video storage container (not shown in the figure) is established, and the target panoramic video recordings exceeding the number of window playback containers 121 are stored in the temporary video storage container as candidate target panoramic video recordings in chronological order.
[0078] Among them, the temporary video storage container can be Minio, and those skilled in the art can choose it according to actual needs, which is not limited here. Minio is a high-performance distributed object storage system based on open source technology that supports the management needs of massive unstructured data. This way of setting up a temporary video storage container can quickly and accurately find the corresponding candidate target panoramic video according to the user's event trigger operation, and quickly load it into the corresponding window playback container 121, avoiding problems such as slow loading and jamming.
[0079] Combine Figure 1b and Figure 1c When a candidate target panoramic video is stored in the temporary video storage container, it is detected whether there is a candidate target panoramic video in the temporary video storage container that is played earlier than the target panoramic video a1 played in the first window playback container 1211 in the panoramic display area;
[0080] If present, a first detection control 14 is set with the left window boundary line of the first window playback container 1211 as a reference point, wherein the width of the first detection control 14 is smaller than the width of the first window playback container 1211, and the height of the first detection control 14 is greater than or equal to the height of the first window playback container 1211, to ensure that the first detection control 14 can block the target panoramic video a1;
[0081] When an event trigger operation is received on the first detection control 14, a candidate target panoramic video that is adjacent in time to the target panoramic video played in the first window playback container 1211 and precedes the target panoramic video in the first window playback container 1211 is extracted from the temporary video storage container, and loaded into the first window playback container 1211, and the target panoramic video in the last window playback container 1212 is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container 121 is loaded into the next window playback container in sequence for playback.
[0082] Specifically, if Figure 1cAs shown, when an event trigger operation is received on the first detection control 14, a candidate target panoramic video a0 that is adjacent to the target panoramic video played in the first window playback container 1211 and precedes the target panoramic video in the first window playback container 1211 is extracted from the temporary video storage container, and loaded into the first window playback container 1211, and the target panoramic video a4 in the last window playback container 1212 is stored in the temporary video storage container in chronological order, and the remaining target panoramic video a1, a2, and a3 in the window playback container 121 are sequentially loaded into the next window playback container for playback, that is, Figure 1c The target panoramic video a1 is loaded into Figure 1c The window playback container corresponding to the target panoramic video a2, Figure 1c The target panoramic video a2 is loaded into Figure 1c The window playback container corresponding to the target panoramic video a3, Figure 1c The target panoramic video a3 is loaded into Figure 1c The window playback container corresponding to the target panoramic video a4.
[0083] Among them, the event triggering operation for the first detection control 14 can be to keep the mouse on the first detection control 14 for a preset time, or to double-click the first detection control 14, etc. Those skilled in the art can set it according to actual needs, which is not limited here.
[0084] Alternatively, as Figure 1c As shown, it is detected whether there is a candidate target panoramic video in the temporary video storage container whose time is later than the target panoramic video played in the last window 1212 playback container in the panoramic display area;
[0085] If it exists, a second detection control 15 is set with the right window boundary line of the last window playback container 1212 as a reference point, wherein the width of the second detection control 15 is less than the width of the last window playback container 1212, and the height of the second detection control 15 is greater than or equal to the height of the last window playback container 1212;
[0086] When an event trigger operation is received on the second detection control 15, a candidate target panoramic video that is adjacent in time to the target panoramic video played in the last window playback container 1212 and later in time than the target panoramic video in the last window playback container 1212 is extracted from the temporary video storage container, and loaded into the last window playback container, and the target panoramic video in the first window playback container 1211 is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is loaded into the previous window playback container in sequence for playback.
[0087] The description of the second detection control 15 can be obtained by referring to the description of the first detection control 14 and making corresponding modifications, and will not be repeated here.
[0088] Alternatively, as Figure 1b and Figure 1c As shown, the monitoring display area 13 further includes a monitoring window playback container 131 corresponding to the window playback container 121; when a single target panoramic video is associated with multiple target monitoring video videos, the first target monitoring video video among the multiple target monitoring video videos is loaded into the monitoring window playback container 131 for playback, and the remaining target monitoring video videos are displayed in the form of a time list 132 below the corresponding monitoring window playback container.
[0089] Specifically, if Figure 1c As shown, a monitoring window playback container 1311 corresponds to the first window playback container 1211; the target panoramic video a1 is associated with six target monitoring videos. The first target monitoring video b1 of the six target monitoring videos is loaded into the monitoring window playback container 1311 for playback. The remaining five target monitoring videos are displayed in the form of a time list 132 below the corresponding monitoring window playback container 1311. The rest is similar and will not be repeated here.
[0090] In the present invention, the first detection control 14 and the second detection control 15 can not only keep the target panoramic video in the corresponding window playback container in the playback state, but also serve as trigger controls for the user to update the target panoramic video in the window playback container, thereby allowing the user to view the target panoramic video more intuitively and quickly.
[0091] The video tracing and display method based on a visual macro model provided by the present invention extracts static and dynamic targets from each frame of the video after acquiring the video, generates corresponding static and dynamic target relationship chains, and stores the static and dynamic target relationship chains in text form. When a user enters a question, the method first filters out the corresponding feature information in the question, then compares the feature information with the stored static and dynamic target relationship chains to filter out the target panoramic video and target surveillance video in the video. Finally, the target panoramic video and target surveillance video are displayed with a preset effect on the client's display interface. The method of the present invention can quickly and accurately extract the corresponding video based on the user's question information and clearly present it to the user for playback.
[0092] Based on the same inventive concept, embodiments of the present invention further provide a video tracing and display device based on a visual macromodel for implementing the aforementioned video tracing and display method based on a visual macromodel. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the video tracing and display device based on a visual macromodel provided below can be found in the above-mentioned limitations of the video tracing and display method based on a visual macromodel, and will not be further elaborated here.
[0093] like Figure 2 As shown, the present invention provides a video tracing and display device based on a visual large model, which includes: an acquisition module 21, a visual large model 22 and a calculation and display module 23; wherein,
[0094] An acquisition module 21 is used to acquire video recordings, wherein the video recordings include panoramic video recordings and surveillance video recordings;
[0095] The visual macro model 22 is connected to the acquisition module 21 and is used to extract static and dynamic targets from each frame of the video, generate corresponding static and dynamic target relationship chains, and store the static and dynamic target relationship chains in text form; and obtain user question information, filter out corresponding feature information from the question information, and compare the feature information with the stored static and dynamic target relationship chains to filter out the target panoramic video and target surveillance video in the video.
[0096] The computing and display module 23 is connected to the visual large model 22 and the client 24 respectively, and is used to display the target panoramic video and the target monitoring video in a preset effect on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area and a monitoring display area.
[0097] Optionally, the visual big model 22 extracts static targets and dynamic targets in each frame of video recording, generates corresponding static target relationship chains and dynamic target relationship chains, and stores the static target relationship chains and dynamic target relationship chains in the form of text, specifically: respectively dividing the panoramic video recording and the surveillance video recording into corresponding panoramic video frames and surveillance video frames according to their frame rates; identifying static targets and dynamic targets in each panoramic video frame and each surveillance video frame, and obtaining corresponding static target attribute information and dynamic target attribute information; generating corresponding static target video recording sets and dynamic target video recording sets according to the static target attribute information or the dynamic target attribute information; generating corresponding static target relationship chains or dynamic target relationship chains based on the static target attribute information, the dynamic target attribute information, the static target video recording sets and the dynamic target video recording sets, and storing the static target relationship chains and the dynamic target relationship chains in the form of text.
[0098] Optionally, the static target relationship chain includes: the name of the static target, the time when the static target appears in the recorded video, the position of the static target and the status of the static target; and / or, the dynamic target relationship chain includes: the name of the dynamic target, the time when the dynamic target appears in the recorded video, the motion trajectory of the dynamic target, the motion status of the dynamic target and the ID of the associated surveillance camera.
[0099] Optionally, the visual large model 22 is also used to: determine whether the dynamic target is out of the target panoramic video according to the motion trajectory of the dynamic target; if the dynamic target is out of the target panoramic video, determine the associated surveillance camera according to the dynamic target relationship chain, and find the target surveillance video after the dynamic target is out of the target panoramic video.
[0100] Optionally, the computing and display module 23 is specifically used for: the magnified display area is set at the upper part of the display interface, and its width is equal to the width of the display interface, for playing the magnified target panoramic video or target monitoring video; the panoramic display area is set in the middle part of the display interface, between the magnified display area and the monitoring display area, for playing the target panoramic video; the monitoring display area is set at the lower part of the display interface, below the panoramic display area, for playing and displaying the target monitoring video.
[0101] Optionally, the panoramic display area further includes multiple window playback containers and video axes of the same size; the sum of the widths of the multiple window playback containers is equal to the width of the panoramic display area. After the target panoramic video is screened out, the target panoramic video is loaded into the window playback container for video playback; the video axis is located below the multiple window playback containers, and the video axis corresponds to displaying the playback time information of the target panoramic video in the multiple window playback containers.
[0102] Optionally, the calculation and display module 23 is also used to: when the number of target panoramic video recordings is greater than the number of window playback containers, establish a temporary video storage container, and store the target panoramic video recordings that exceed the number of window playback containers as candidate target panoramic video recordings in chronological order in the temporary video storage container.
[0103] Optionally, the calculation and display module 23 is also used to: detect whether there is a candidate target panoramic video in the temporary video storage container that is earlier than the target panoramic video played in the first window playback container in the panoramic display area; if so, set a first detection control with the left window boundary line of the first window playback container as the reference point, wherein the width of the first detection control is less than the width of the first window playback container, and the height of the first detection control is greater than or equal to the height of the first window playback container; when an event trigger operation is received on the first detection control, extract from the temporary video storage container a candidate target panoramic video that is adjacent in time to the target panoramic video played in the first window playback container and earlier in time than the target panoramic video in the first window playback container, load it into the first window playback container, and store the target panoramic video in the last window playback container into the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is sequentially loaded into the next window playback container for playback.
[0104] Optionally, the calculation and display module 23 is also used to: detect whether there is a candidate target panoramic video in the temporary video storage container that is later than the target panoramic video played in the last window playback container in the panoramic display area; if so, set a second detection control with the right window boundary line of the last window playback container as the reference point, wherein the width of the second detection control is less than the width of the last window playback container, and the height of the second detection control is greater than or equal to the height of the last window playback container; when an event trigger operation is received on the second detection control, extract the candidate target panoramic video that is adjacent to the target panoramic video played in the last window playback container and later than the target panoramic video in the last window playback container from the temporary video storage container, load it into the last window playback container, and store the target panoramic video in the first window playback container in chronological order into the temporary video storage container, and the remaining target panoramic video in the window playback container is sequentially loaded into the previous window playback container for playback.
[0105] Optionally, the monitoring display area further includes a monitoring window playback container corresponding to the window playback container; when a single target panoramic video is associated with multiple target monitoring video videos, the first target monitoring video video among the multiple target monitoring video videos is loaded into the monitoring window playback container for playback, and the remaining target monitoring video videos are displayed in the form of a time list below the corresponding monitoring window playback container.
[0106] The present invention provides a video tracing and display device based on a visual macro model. After acquiring the video, the acquisition module extracts the static and dynamic targets in each frame of the video, generates corresponding static and dynamic target relationship chains, and stores the static and dynamic target relationship chains in text form. When the visual macro model receives the question information input by the user, it first filters out the corresponding feature information in the question information, then compares the feature information with the stored static and dynamic target relationship chains, filters out the target panoramic video and target surveillance video in the video, and finally, the calculation and display module displays the target panoramic video and target surveillance video in a preset effect on the display interface of the client. The device of the present invention can quickly and accurately extract the corresponding video based on the user's question information and clearly present it to the user for playback.
[0107] It should be noted that, in the present invention, a plurality includes two or more.
[0108] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0109] Each module in each device of the present invention may be implemented in whole or in part by software, hardware, or a combination thereof. Each of the modules may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0110] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 3As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data required or generated for executing the above-mentioned method for tracing and displaying recorded videos based on a visual large model. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for tracing and displaying recorded videos based on a visual large model is implemented.
[0111] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, memory, a communication interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal via wired or wireless communication. Wireless communication can be achieved via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a method for retrospective display of recorded video based on a visual large model. The display screen of the computer device can be a liquid crystal display or an electronic ink display. The input device of the computer device can be a touch layer covering the display screen, or keys, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse.
[0112] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0113] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0114] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0115] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties.
[0117] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, database, or other media used in the embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, and the like.
[0118] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0119] The above-described embodiments merely represent several implementation methods of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A video tracing and display method based on a visual macro model, characterized in that: The method comprises: Obtaining video footage, wherein the video footage includes panoramic video footage and surveillance video footage; Extract static targets and dynamic targets in each frame of video, generate corresponding static target relationship chains and dynamic target relationship chains, and store the static target relationship chains and dynamic target relationship chains in text form; Obtain the user's question information, filter out the corresponding feature information in the question information, and compare the feature information with the stored static target relationship chain and dynamic target relationship chain to filter out the target panoramic video and target surveillance video in the video; The target panoramic video and the target monitoring video are displayed with preset effects on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area and a monitoring display area.
2. The method according to claim 1, characterized in that The step of extracting static objects and dynamic objects from each frame of video, generating corresponding static object relationship chains and dynamic object relationship chains, and storing the static object relationship chains and dynamic object relationship chains in text form includes: Split the panoramic video and the surveillance video into corresponding panoramic video frames and surveillance video frames according to their frame rates; Identify static targets and dynamic targets in each panoramic video frame and each surveillance video frame, and obtain corresponding static target attribute information and dynamic target attribute information; Generate corresponding static target video recording sets and dynamic target video recording sets according to static target attribute information or dynamic target attribute information; Based on the static target attribute information, the dynamic target attribute information, the static target video set and the dynamic target video set, a corresponding static target relationship chain or a dynamic target relationship chain is generated, and the static target relationship chain and the dynamic target relationship chain are stored in text form.
3. The method according to claim 1, characterized in that A static target relationship chain includes: the name of the static target, the time when the static target appears in the recorded video, the position of the static target, and the status of the static target; and / or a dynamic target relationship chain includes: the name of the dynamic target, the time when the dynamic target appears in the recorded video, the motion trajectory of the dynamic target, the motion status of the dynamic target, and the ID of the associated surveillance camera.
4. The method according to claim 3, characterized in that The method further comprises: According to the motion trajectory of the dynamic target, determine whether the dynamic target is out of the target panoramic video; If the dynamic target leaves the target panoramic video, the associated surveillance camera is determined according to the dynamic target relationship chain, and the target surveillance video after the dynamic target leaves the target panoramic video is found.
5. The method according to claim 1, wherein The step of displaying the target panoramic video and the target surveillance video on the display interface of the client with a preset effect includes: The enlarged display area is set at the upper part of the display interface, and its width is equal to the width of the display interface, and is used to play the enlarged target panoramic video or target monitoring video; The panoramic display area is set in the middle of the display interface, between the magnified display area and the monitoring display area, and is used to play the target panoramic video; The monitoring display area is set at the lower part of the display interface, below the panoramic display area, and is used to play and display the target monitoring video.
6. The method according to claim 1, characterized in that The panoramic display area further includes a plurality of window playback containers and video axes of the same size; The sum of the widths of the multiple window playback containers is equal to the width of the panoramic display area. After the target panoramic video is screened out, the target panoramic video is loaded into the window playback container for video playback; The video axis is located below the multiple window playback containers, and the video axis correspondingly displays the playback time information of the target panoramic video in the multiple window playback containers.
7. The method according to claim 6, characterized in that When the number of target panoramic video recordings is greater than the number of window playback containers, a temporary video storage container is established, and the target panoramic video recordings exceeding the number of window playback containers are taken as candidate target panoramic video recordings and stored in the temporary video storage container in chronological order.
8. The method according to claim 7, characterized in that Detecting whether there is a candidate target panoramic video in the temporary video storage container that is played earlier than the target panoramic video played in the first window playback container in the panoramic display area; If it exists, a first detection control is set with the left window boundary of the first window playback container as a reference point, wherein the width of the first detection control is less than the width of the first window playback container, and the height of the first detection control is greater than or equal to the height of the first window playback container; When an event trigger operation is received on the first detection control, a candidate target panoramic video that is adjacent to the target panoramic video played in the first window playback container and precedes the target panoramic video in the first window playback container is extracted from the temporary video storage container, and loaded into the first window playback container. The target panoramic video in the last window playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is sequentially loaded into the next window playback container for playback; and / or detecting whether there exists a candidate target panoramic video in the temporary video storage container whose playing time is later than the target panoramic video played in the last window playback container in the panoramic display area; If it exists, a second detection control is set with the right window boundary line of the last window playback container as the reference point, wherein the width of the second detection control is less than the width of the last window playback container, and the height of the second detection control is greater than or equal to the height of the last window playback container; When an event trigger operation is received for the second detection control, a candidate target panoramic video that is adjacent to and later than the target panoramic video played in the last window playback container is extracted from the temporary video storage container and loaded into the last window playback container, and the target panoramic video in the first window playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is loaded into the previous window playback container in sequence for playback.
9. The method according to claim 6, characterized in that The monitoring display area further includes a monitoring window playback container corresponding to the window playback container; when a single target panoramic video is associated with multiple target monitoring video videos, the first target monitoring video video among the multiple target monitoring video videos is loaded into the monitoring window playback container for playback, and the remaining target monitoring video videos are displayed in the form of a time list below the corresponding monitoring window playback container.
10. A video tracing display device based on a visual macro model, characterized in that: The device comprises: An acquisition module is used to acquire video recordings, wherein the video recordings include panoramic video recordings and surveillance video recordings; The visual large model is connected to the acquisition module and is used to extract static and dynamic targets in each frame of video, generate corresponding static target relationship chains and dynamic target relationship chains, and store the static target relationship chains and dynamic target relationship chains in text form; and obtain user question information, filter out corresponding feature information in the question information, and compare the feature information with the stored static target relationship chains and dynamic target relationship chains to filter out target panoramic video and target monitoring video in the video; The computing and display module is connected to the visual large model and the client respectively, and is used to display the target panoramic video and the target monitoring video with preset effects on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area and a monitoring display area.
Citation Information
Patent Citations
Processing method and device for digital video monitoring system
CN116866534A
Panorama video intelligent monitoring method and system
CN101123722A
Full space-time three-dimensional visualization method
CN103795976A
Intelligent public security video retrieval system
CN110188238A
Fixed scene monitoring video backtracking system
CN114817638A