Video recording playback display method and device based on visual large model
Through the video tracing display method based on the visual big model, the static and dynamic target relationship chains in the video are extracted and stored, and multi-source video is screened according to user queries and displayed collaboratively on the client side. This solves the problems of low retrieval efficiency and single information presentation of video surveillance systems in the existing technology, and realizes fast and accurate video presentation and user-friendly experience.
Patent Information
- Application Number
- CN202511152134.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing video surveillance systems have low timeliness in event tracing and single information presentation, and are unable to effectively associate panoramic and close-up videos, resulting in low retrieval efficiency and poor user experience.
A method based on a large visual model is used to extract static and dynamic targets from recorded videos, generate and store relationship chains, filter target videos based on user query information, and display preset effects on the client, including zoom, panorama, and collaborative display of monitoring display areas.
It achieves the rapid and accurate extraction and presentation of relevant video recordings, improves the timeliness of event tracing and the diversity of information presentation, simplifies user operation processes, and saves manpower and material costs.
Smart Images

Figure CN120640026B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video monitoring, in particular to a video recording tracing display method and device based on a visual large model. BACKGROUND
[0002] The current video monitoring system generally adopts a video recording retrieval mechanism based on a time period. For example, a processing method and device for a digital video monitoring system are disclosed in Chinese patent CN116866534A, in which a user must pre-configure a monitoring configuration time period, and on this basis, the corresponding monitoring video can be retrieved and displayed. This approach has the following problems: first, the retrieval efficiency is low. When the user cannot determine the specific time of an event (such as only knowing the behavior characteristics), the user needs to manually segment and screen a large amount of video recording, which is time-consuming and prone to missing key information, seriously affecting the timeliness of event tracing. Second, the information presentation is single. The existing technology only linearly splices and plays the extracted video clips in time sequence, lacks intelligent association and fusion display of multi-source videos, and especially cannot present the macroscopic scene of a wide-angle panoramic video and the detailed behavior of a close-up monitoring video in time and space coordination, which makes the user need to repeatedly switch and compare, and it is difficult to quickly build a complete event chain.
[0003] Therefore, there is an urgent need for a visual large model-based approach that can directly analyze semantic query requests, automatically locate and extract target video clips, and realize the collaborative visualization of panoramic and close-up videos, thereby fundamentally solving the dual bottlenecks of retrieval efficiency and presentation effect. SUMMARY
[0004] Therefore, it is necessary to provide a visual large model-based video recording tracing display method and device that can quickly and accurately extract corresponding video recording and clearly present it to the user for playing according to the user's question information.
[0005] In a first aspect, the present application provides a visual large model-based video recording tracing display method, which comprises:
[0006] obtaining video recording, wherein the video recording comprises panoramic video recording and monitoring video recording;
[0007] extracting static targets and dynamic targets in each frame of video recording, generating corresponding static target relationship chains and dynamic target relationship chains, and storing the static target relationship chains and dynamic target relationship chains in the form of text;
[0008] obtaining the user's question information, screening the corresponding feature information in the question information, and comparing the feature information with the stored static target relationship chains and dynamic target relationship chains to screen out target panoramic video recording and target monitoring video recording in the video recording;
[0009] The target panoramic video and the target monitoring video are displayed on a preset display interface of a client, wherein the display interface comprises an enlarged display area, a panoramic display area and a monitoring display area.
[0010] Optionally, static targets and dynamic targets in each frame of the video are extracted, corresponding static target relationship chains and dynamic target relationship chains are generated, and the static target relationship chains and the dynamic target relationship chains are stored in the form of text, comprising:
[0011] The panoramic video and the monitoring video are respectively divided into corresponding panoramic video frames and monitoring video frames according to their frame rates;
[0012] Static targets and dynamic targets in each panoramic video frame and each monitoring video frame are identified to obtain corresponding static target attribute information and dynamic target attribute information;
[0013] According to the static target attribute information or the dynamic target attribute information, corresponding static target video sets and dynamic target video sets are generated;
[0014] Based on the static target attribute information, the dynamic target attribute information, the static target video sets and the dynamic target video sets, corresponding static target relationship chains or dynamic target relationship chains are generated, and the static target relationship chains and the dynamic target relationship chains are stored in the form of text.
[0015] Optionally, the static target relationship chain comprises a name of a static target, a time when the static target appears in the video, a position of the static target and a state of the static target; and / or the dynamic target relationship chain comprises a name of a dynamic target, a time when the dynamic target appears in the video, a motion trajectory of the dynamic target, a motion state of the dynamic target and an ID of an associated monitoring camera.
[0016] Optionally, the method further comprises:
[0017] According to the motion trajectory of the dynamic target, it is determined whether the dynamic target has left the target panoramic video;
[0018] If the dynamic target has left the target panoramic video, the associated monitoring camera is determined according to the dynamic target relationship chain, and the target monitoring video after the dynamic target has left the target panoramic video is found.
[0019] Optionally, the target panoramic video and the target monitoring video are displayed on a preset display interface of a client, comprising:
[0020] The enlarged display area is arranged at the upper part of the display interface, has a width equal to that of the display interface, and is used for playing the target panoramic video or the target monitoring video after enlargement.
[0021] The panoramic display area is arranged in the middle of the display interface, between the zoomed display area and the monitoring display area, and is used to play the target panoramic video;
[0022] The monitoring display area is arranged in the lower part of the display interface, below the panoramic display area, and is used to play and display the target monitoring video.
[0023] Optionally, the panoramic display area further comprises a plurality of windowed playback containers of the same size and a video axis;
[0024] The sum of the widths of the plurality of windowed playback containers is equal to the width of the panoramic display area, and when the target panoramic video is filtered out, the target panoramic video is loaded into the windowed playback container for video playback;
[0025] The video axis is located below the plurality of windowed playback containers, and the video axis displays the play time information of the target panoramic video in the plurality of windowed playback containers.
[0026] Optionally, when the number of target panoramic videos is greater than the number of windowed playback containers, a temporary video storage container is established, and the target panoramic videos exceeding the number of windowed playback containers are stored in the temporary video storage container as candidate target panoramic videos in chronological order.
[0027] Optionally, it is detected whether there is a candidate target panoramic video in the temporary video storage container that is earlier in time than the target panoramic video played in the first windowed playback container in the panoramic display area;
[0028] If so, a first detection control is set with the left window boundary line of the first windowed playback container as the reference point, wherein the width of the first detection control is less than the width of the first windowed playback container, and the height of the first detection control is greater than or equal to the height of the first windowed playback container;
[0029] When an event triggering operation on the first detection control is received, a candidate target panoramic video adjacent in time to the target panoramic video played in the first windowed playback container and earlier in time than the target panoramic video in the first windowed playback container is extracted from the temporary video storage container, loaded into the first windowed playback container, and the target panoramic video in the last windowed playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic videos in the windowed playback containers are loaded into the next windowed playback container in sequence for playback;
[0030] And / or, it is detected whether there is a candidate target panoramic video in the temporary video storage container that is later in time than the target panoramic video played in the last windowed playback container in the panoramic display area;
[0031] If there is, a second detection control is set based on the right window boundary line of the last window playback container, wherein the width of the second detection control is less than the width of the last window playback container, and the height of the second detection control is greater than or equal to the height of the last window playback container;
[0032] When an event triggering operation on the second detection control is received, a candidate target panoramic video adjacent in time to the target panoramic video played in the last window playback container and later in time than the target panoramic video is extracted from the temporary video storage container, loaded into the last window playback container, and the target panoramic video in the first window playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic videos in the window playback container are loaded into the previous window playback container in turn for playing.
[0033] Optionally, the monitoring display area further comprises a monitoring window playback container corresponding to the window playback container; when a single target panoramic video is associated with multiple target monitoring videos, the first target monitoring video in the multiple target monitoring videos is loaded into the monitoring window playback container for playing, and the remaining target monitoring videos are displayed in the form of a time list below the corresponding monitoring window playback container.
[0034] In a second aspect, the present application provides a video tracking display device based on a visual large model, which comprises:
[0035] An acquisition module is configured to acquire a video, wherein the video comprises a panoramic video and a monitoring video;
[0036] A visual large model is connected to the acquisition module and configured to extract static targets and dynamic targets in each frame of the video, generate corresponding static target relationship chains and dynamic target relationship chains, and store the static target relationship chains and the dynamic target relationship chains in the form of text; and acquire user query information, filter out corresponding feature information in the query information, and compare the feature information with the stored static target relationship chains and dynamic target relationship chains to filter out target panoramic videos and target monitoring videos in the video.
[0037] A calculation and display module is connected to the visual large model and the client respectively, and configured to display the target panoramic videos and the target monitoring videos on a display interface of the client with a preset effect, wherein the display interface comprises an enlarged display area, a panoramic display area, and a monitoring display area.
[0038] The application provides a video recording backtracking display method and device based on a visual large model. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1a A flowchart of a video recording backtracking display method based on a visual large model is provided for an embodiment of the application.
[0040] Figure 1b A display diagram of preset effects of a display interface of the video recording backtracking display method based on the visual large model is provided for an embodiment of the application.
[0041] Figure 1c Another display diagram of preset effects of the display interface of the video recording backtracking display method based on the visual large model is provided for an embodiment of the application.
[0042] Figure 2 A circuit module structure diagram of the video recording backtracking display device based on the visual large model is provided for an embodiment of the application.
[0043] Figure 3 An internal structure diagram of a computer device in an embodiment of the application. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions and advantages of the application clearer, the application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.
[0045] As shown in Figure 1a The application provides a video recording backtracking display method based on a visual large model, which comprises the following steps:
[0046] Step S11: acquiring a video recording, wherein the video recording comprises a panoramic video recording and a monitoring video recording;
[0047] Optionally, the panoramic video is obtained by storing a panoramic video corresponding to the scene after the panoramic camera shoots the panoramic video, and the monitoring video is obtained by storing a monitoring video corresponding to the scene after the monitoring camera shoots the monitoring video. The panoramic camera can be an LC150 panoramic camera with a billion-pixel, an NV180E panoramic camera with a billion-pixel, or the like, and the monitoring camera can be a DS-2CD5A4XYZ-LS full-color intelligent warning camera, a DH-IPC-HFW3449DM-A-LED camera, or the like.
[0048] Step S12: Extracting static targets and dynamic targets in each frame of the video, generating corresponding static target relationship chains and dynamic target relationship chains, and storing the static target relationship chains and the dynamic target relationship chains in the form of text.
[0049] The static target relationship chain includes the name of the static target, the time when the static target appears in the video, the position of the static target, and the state of the static target. The dynamic target relationship chain includes the name of the dynamic target, the time when the dynamic target appears in the video, the motion trajectory of the dynamic target, the motion state of the dynamic target, and the ID of the associated monitoring camera. It should be noted that the static target can be a street lamp, a building, or the like, i.e., a static target is an object that cannot move under normal circumstances; the dynamic target can be a pedestrian, a pet, a bicycle, a car, or the like.
[0050] The static target relationship chain stored in the form of text is: [name of static target, time when static target appears in video, position of static target, state of static target] = [street lamp, panoramic-2012.5.1 12:30:30-2013.9.1013:00:29, left middle of the picture - beside the bench, solar wind street lamp].
[0051] The dynamic target relationship chain stored in the form of text is: [name of dynamic target, time when dynamic target appears in video, motion trajectory of dynamic target, motion state of dynamic target, ID of associated monitoring camera] = [pedestrian, panoramic-2012.5.1 12:30:30-2013.9.10 13:00:29, coordinates (54, 67) - (54, 90) - (54, 120), a pedestrian wearing a red shirt jogging, 10256], or [name of dynamic target, time when dynamic target appears in video, motion trajectory of dynamic target, motion state of dynamic target, ID of associated monitoring camera] = [pedestrian, monitoring-2015.5.1 12:30:30-2017.9.10 13:00:29, coordinates (54, 67) - (54, 90) - (54, 120), a pedestrian wearing a blue shirt walking, 0], where the ID of the associated monitoring camera is 0, indicating that there is no associated monitoring camera.
[0052] Optionally, after receiving the video recording, the visual large model first divides the panoramic video recording and the monitoring video recording into corresponding panoramic video recording frames and monitoring video recording frames according to their frame rates respectively; then identifies static targets and dynamic targets in each panoramic video recording frame and each monitoring video recording frame to obtain corresponding static target attribute information and dynamic target attribute information; generates corresponding static target video recording set and dynamic target video recording set according to the static target attribute information or the dynamic target attribute information; and generates corresponding static target relationship chain or dynamic target relationship chain based on the static target attribute information, the dynamic target attribute information, the static target video recording set and the dynamic target video recording set, and stores the static target relationship chain and the dynamic target relationship chain in the form of text.
[0053] In the present application, the visual large model can use a pre-trained visual large model in the prior art, which can be selected by those skilled in the art according to actual needs, which is not limited here. For example, a visual large model based on RAG system can be selected.
[0054] The static target attribute information includes the name of the static target, the time when the static target appears in the video recording, the position of the static target and the state of the static target; and / or the dynamic target attribute information includes the name of the dynamic target, the time when the dynamic target appears in the video recording, the motion trajectory of the dynamic target, the motion state of the dynamic target and the ID of the associated monitoring camera.
[0055] It should be noted that the present application can not only quickly find the corresponding video recording through the static target relationship chain and the dynamic target relationship chain stored in the form of text, but also can send the found corresponding static target relationship chain and dynamic target relationship chain to the user in the form of short message, etc. The user can directly play the corresponding video recording by clicking the received static target relationship chain and dynamic target relationship chain. This way can make the user conveniently view the video recording, and the static target relationship chain and the dynamic target relationship chain stored in the form of text can fully save storage space.
[0056] Optionally, after identifying the static targets and dynamic targets in each panoramic video recording frame and each monitoring video recording frame, the method of the present application further comprises marking the identified static targets and dynamic targets in each panoramic video recording frame and each monitoring video recording frame to facilitate user viewing.
[0057] Optionally, the method of the present application further comprises:
[0058] According to the motion trajectory of the dynamic target, it is judged whether the dynamic target is out of the target panoramic video recording;
[0059] If the dynamic target leaves the target panoramic video, the associated surveillance camera is determined according to the dynamic target relationship chain, and the target surveillance video after the dynamic target leaves the target panoramic video is found.
[0060] This method can effectively associate the target panoramic video and the target surveillance video and present them to the user, quickly and accurately achieving cross-lens tracking, saving a lot of manpower, material resources and other costs.
[0061] Step S13: Obtain the user's question information, filter out the corresponding feature information in the question information, and compare the feature information with the stored static target relationship chain and dynamic target relationship chain to filter out the target panoramic video and target surveillance video in the video;
[0062] For example, the user's question information is: From June 27, 2015 to June 30, 2015, where did a person wearing a red short-sleeved shirt and black trousers appear in the coal mine?
[0063] After receiving the question information, the visual big model screens the feature information therein, such as June 27, 2015 to June 30, 2015, red, short sleeves, black, trousers, and coal mine; after extracting the feature information, since it is a dynamic target pedestrian, the text similarity between the feature information and the dynamic target relationship chain is calculated, and the corresponding target panoramic video is screened out; according to the motion trajectory of the dynamic target, it is judged whether the dynamic target is out of the target panoramic video; if the dynamic target is out of the target panoramic video, the associated surveillance camera is determined according to the dynamic target relationship chain, and the target surveillance video after the dynamic target is out of the target panoramic video is found.
[0064] Step S14: Displaying the target panoramic video and the target monitoring video with a preset effect on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area, and a monitoring display area.
[0065] Alternatively, as Figure 1b As shown, step S14 specifically includes:
[0066] The enlarged display area 11 is provided at the upper portion of the display interface 10 and has a width equal to that of the display interface 10 , and is used to play the enlarged target panoramic video or target surveillance video;
[0067] A video play axis 110 may also be provided on the magnified display area 11 for displaying the play time and / or adjusting the play progress.
[0068] It should be noted that when the user selects the corresponding target panoramic video or target monitoring video, the video played in the enlarged display area 11 will be replaced accordingly, and the target panoramic video or target monitoring video selected by the user is played.
[0069] The panoramic display area 12 is arranged in the middle of the display interface 10, between the enlarged display area 11 and the monitoring display area 13, and is used to play the target panoramic video.
[0070] The monitoring display area 13 is arranged at the lower part of the display interface 10, below the panoramic display area 12, and is used to play and display the target monitoring video.
[0071] In an optional embodiment of the present application, as shown in Figure 1b The panoramic display area 12 further includes a plurality of window play containers 121 of the same size and a video axis 122.
[0072] The sum of the widths of the plurality of window play containers 121 is equal to the width of the panoramic display area 12, and when the target panoramic video is filtered out, the target panoramic video is loaded into the window play container 121 for video playing.
[0073] The window play container can be an mp4 container, an avi container, an mkv container, an flv container, a wmv container, etc., which is not limited here.
[0074] For example, in Figure 1b The panoramic display area 12 is provided with four window play containers 121, the sum of the widths of the four window play containers 121 is equal to the width of the panoramic display area 12, and the target panoramic videos a1, a2, a3 and a4 are loaded into the four window play containers 121 respectively. It should be noted that after the target panoramic videos a1, a2, a3 and a4 are loaded into the corresponding window play containers 121, they are automatically played, i.e., they are displayed to the user in the form of video playing, rather than video pictures.
[0075] The video axis 122 is arranged below the plurality of window play containers 121, and the video axis 122 displays the playing time information of the target panoramic video in the plurality of window play containers 121.
[0076] The playing time information includes the starting playing time, the current playing time and the ending playing time, etc., which can be flexibly set by those skilled in the art according to actual needs, which is not limited here.
[0077] In another optional embodiment of the present application, as shown in Figure 1cAs shown, when the number of target panoramic video is greater than the number of window playback containers 121, a temporary video storage container (not shown in the figure) is established, and the target panoramic video exceeding the number of window playback containers 121 is stored in the temporary video storage container as a candidate target panoramic video in chronological order.
[0078] The temporary video storage container can be Minio, which can be selected by those skilled in the art according to actual needs, and is not limited here. Minio is a high-performance distributed object storage system based on open source technology, which supports the management needs of massive unstructured data. In this way, the corresponding candidate target panoramic video can be quickly and accurately found according to the user's event triggering operation, and loaded into the corresponding window playback container 121 quickly, avoiding slow loading, stuttering and other problems.
[0079] In combination with Figure 1b And Figure 1c When the temporary video storage container stores the candidate target panoramic video, it is detected whether there is a candidate target panoramic video in the temporary video storage container that is earlier in time than the target panoramic video a1 played in the first window playback container 1211 in the panoramic display area.
[0080] If so, a first detection control 14 is set with the first window playback container 1211 left window boundary line as the reference point, wherein the width of the first detection control 14 is less than the width of the first window playback container 1211, and the height of the first detection control 14 is greater than or equal to the height of the first window playback container 1211, so as to ensure that the first detection control 14 can block the target panoramic video a1.
[0081] When receiving the event triggering operation on the first detection control 14, the candidate target panoramic video adjacent in time to the target panoramic video played in the first window playback container 1211 and earlier in time than the target panoramic video in the first window playback container 1211 is extracted from the temporary video storage container, loaded into the first window playback container 1211, and the target panoramic video in the last window playback container 1212 is stored in the temporary video storage container in chronological order. The remaining target panoramic video in the window playback container 121 is loaded into the next window playback container in turn for playing.
[0082] Specifically, as Figure 1cAs shown, when an event trigger operation is received on the first detection control 14, a candidate target panoramic video a0 that is adjacent to the target panoramic video played in the first window playback container 1211 and precedes the target panoramic video in the first window playback container 1211 is extracted from the temporary video storage container, and loaded into the first window playback container 1211, and the target panoramic video a4 in the last window playback container 1212 is stored in the temporary video storage container in chronological order, and the remaining target panoramic video a1, a2, and a3 in the window playback container 121 are sequentially loaded into the next window playback container for playback, that is, Figure 1c The target panoramic video a1 is loaded into Figure 1c The window playback container corresponding to the target panoramic video a2, Figure 1c The target panoramic video a2 is loaded into Figure 1c The window playback container corresponding to the target panoramic video a3, Figure 1c The target panoramic video a3 is loaded into Figure 1c The window playback container corresponding to the target panoramic video a4.
[0083] Among them, the event triggering operation for the first detection control 14 can be to keep the mouse on the first detection control 14 for a preset time, or to double-click the first detection control 14, etc. Those skilled in the art can set it according to actual needs, which is not limited here.
[0084] Alternatively, as Figure 1c As shown, it is detected whether there is a candidate target panoramic video in the temporary video storage container whose time is later than the target panoramic video played in the last window 1212 playback container in the panoramic display area;
[0085] If it exists, a second detection control 15 is set with the right window boundary line of the last window playback container 1212 as a reference point, wherein the width of the second detection control 15 is less than the width of the last window playback container 1212, and the height of the second detection control 15 is greater than or equal to the height of the last window playback container 1212;
[0086] When an event trigger operation is received on the second detection control 15, a candidate target panoramic video that is adjacent in time to the target panoramic video played in the last window playback container 1212 and later in time than the target panoramic video in the last window playback container 1212 is extracted from the temporary video storage container, and loaded into the last window playback container, and the target panoramic video in the first window playback container 1211 is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is loaded into the previous window playback container in sequence for playback.
[0087] For the description of the second detection control 15, corresponding modifications can be made to the description of the first detection control 14, which will not be repeated here.
[0088] Optionally, as shown in Figure 1c and Figure 1b , the monitoring display area 13 further comprises a monitoring window playback container 131 corresponding to the window playback container 121; when a single target panoramic video is associated with multiple target monitoring video, the first target monitoring video in the multiple target monitoring video is loaded into the monitoring window playback container 131 for playing, and the remaining target monitoring video is displayed in the form of a time list 132 below the corresponding monitoring window playback container.
[0089] Specifically, as shown in Figure 1c , the monitoring window playback container 1311 corresponding to the first window playback container 1211; the target panoramic video a1 is associated with 6 target monitoring videos, the first target monitoring video b1 in the 6 target monitoring videos is loaded into the monitoring window playback container 1311 for playing, and the remaining 5 target monitoring videos are displayed in the form of a time list 132 below the corresponding monitoring window playback container 1311. The others are the same, which will not be repeated here.
[0090] In the present application, the first detection control 14 and the second detection control 15 not only keep the target panoramic video in the corresponding window playback container playing, but also serve as the trigger control for the user to update the target panoramic video in the window playback container, so that the user can more intuitively and quickly view the target panoramic video.
[0091] The video tracking display method based on the visual large model provided by the present application extracts static targets and dynamic targets in each frame of video after obtaining the video, generates corresponding static target relationship chain and dynamic target relationship chain, and stores the static target relationship chain and the dynamic target relationship chain in the form of text; when the user inputs the question information, first, the corresponding feature information in the question information is screened out, then the feature information is compared with the stored static target relationship chain and dynamic target relationship chain, the target panoramic video and the target monitoring video in the video are screened out, and finally the target panoramic video and the target monitoring video are displayed on the display interface of the client with a preset effect. The method of the present application can quickly and accurately extract the corresponding video according to the user's question information and clearly present it to the user for playing.
[0092] Based on the same inventive concept, the embodiments of the present application also provide a visual large model-based video playback display device for implementing the visual large model-based video playback display method as described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more visual large model-based video playback display device embodiments provided below can refer to the limitations of the visual large model-based video playback display method described above, and will not be repeated here.
[0093] As shown in Figure 1c The present application provides a visual large model-based video playback display device, which comprises an acquisition module 21, a visual large model 22 and a calculation and display module 23; wherein,
[0094] The acquisition module 21 is used for acquiring video, wherein the video comprises panoramic video and monitoring video;
[0095] The visual large model 22 is connected with the acquisition module 21 and is used for extracting static targets and dynamic targets in each frame of video, generating corresponding static target relationship chains and dynamic target relationship chains, and storing the static target relationship chains and the dynamic target relationship chains in the form of text; and acquiring user's question information, screening out corresponding feature information in the question information, and comparing the feature information with the stored static target relationship chains and dynamic target relationship chains to screen out target panoramic video and target monitoring video in the video;
[0096] The calculation and display module 23 is connected with the visual large model 22 and a client 24 respectively, and is used for displaying the target panoramic video and the target monitoring video in a preset effect on a display interface of the client, wherein the display interface comprises an enlarged display area, a panoramic display area and a monitoring display area.
[0097] Optionally, the visual large model 22 extracts static targets and dynamic targets in each frame of the video, generates corresponding static target relationship chains and dynamic target relationship chains, and stores the static target relationship chains and the dynamic target relationship chains in the form of text, specifically: the panoramic video and the monitoring video are respectively divided into corresponding panoramic video frames and monitoring video frames according to their frame rates; static targets and dynamic targets in each panoramic video frame and each monitoring video frame are identified to obtain corresponding static target attribute information and dynamic target attribute information; corresponding static target video sets and dynamic target video sets are generated according to the static target attribute information or the dynamic target attribute information; and corresponding static target relationship chains or dynamic target relationship chains are generated based on the static target attribute information, the dynamic target attribute information, the static target video sets and the dynamic target video sets, and the static target relationship chains and the dynamic target relationship chains are stored in the form of text.
[0098] Optionally, the static target relationship chain includes: a name of the static target, a time when the static target appears in the video, a position of the static target, and a state of the static target; and / or, the dynamic target relationship chain includes: a name of the dynamic target, a time when the dynamic target appears in the video, a motion trajectory of the dynamic target, a motion state of the dynamic target, and an ID of an associated monitoring camera.
[0099] Optionally, the visual large model 22 is further configured to: determine whether the dynamic target has left the target panoramic video according to the motion trajectory of the dynamic target; if the dynamic target has left the target panoramic video, determine the associated monitoring camera according to the dynamic target relationship chain, and find the target monitoring video after the dynamic target has left the target panoramic video.
[0100] Optionally, the computing and displaying module 23 is specifically configured to: set the enlarged display area at the upper part of the display interface, the width of the enlarged display area being equal to the width of the display interface, for playing the target panoramic video or the target monitoring video after being enlarged; set the panoramic display area at the middle part of the display interface, between the enlarged display area and the monitoring display area, for playing the target panoramic video; and set the monitoring display area at the lower part of the display interface, below the panoramic display area, for playing and displaying the target monitoring video.
[0101] Optionally, the panoramic display area further includes a plurality of window playing containers with the same size and a video axis; the sum of the widths of the plurality of window playing containers is equal to the width of the panoramic display area, and when the target panoramic video is filtered out, the target panoramic video is loaded into the window playing container for video playing; the video axis is located below the plurality of window playing containers, and the video axis displays the playing time information of the target panoramic video in the plurality of window playing containers.
[0102] Optionally, the computing and displaying module 23 is further configured to: when the number of target panoramic video is greater than the number of window playing containers, establish a temporary video storage container, and store the target panoramic video exceeding the number of window playing containers as candidate target panoramic video in the temporary video storage container in chronological order.
[0103] Optionally, the computing and displaying module 23 is further configured to: detect whether there is candidate target panoramic video in the temporary video storage container which is earlier in time than the target panoramic video played in the first window playing container in the panoramic display area; if so, set a first detection control based on the left window boundary line of the first window playing container as a reference point, wherein the width of the first detection control is less than the width of the first window playing container, and the height of the first detection control is greater than or equal to the height of the first window playing container; when receiving an event triggering operation on the first detection control, extract the candidate target panoramic video adjacent in time to the target panoramic video played in the first window playing container and earlier in time than the target panoramic video in the first window playing container from the temporary video storage container, load it into the first window playing container, and store the target panoramic video in the last window playing container in chronological order into the temporary video storage container, and the remaining target panoramic video in the window playing container is loaded into the next window playing container in turn for playing.
[0104] Optionally, the computing and displaying module 23 is further configured to: detect whether there is candidate target panoramic video in the temporary video storage container which is later in time than the target panoramic video played in the last window playing container in the panoramic display area; if so, set a second detection control based on the right window boundary line of the last window playing container as a reference point, wherein the width of the second detection control is less than the width of the last window playing container, and the height of the second detection control is greater than or equal to the height of the last window playing container; when receiving an event triggering operation on the second detection control, extract the candidate target panoramic video adjacent in time to the target panoramic video played in the last window playing container and later in time than the target panoramic video in the last window playing container from the temporary video storage container, load it into the last window playing container, and store the target panoramic video in the first window playing container in chronological order into the temporary video storage container, and the remaining target panoramic video in the window playing container is loaded into the previous window playing container in turn for playing.
[0105] Optionally, the monitoring display area further comprises a monitoring window playback container corresponding to the window playback container; when a single target panoramic video is associated with multiple target monitoring video, a first target monitoring video of the multiple target monitoring video is loaded into the monitoring window playback container for playing, and the remaining target monitoring video is displayed in the form of a time list below the corresponding monitoring window playback container.
[0106] The video tracking display device based on the visual large model provided by the application extracts static targets and dynamic targets in each frame of the video after the acquisition module acquires the video, generates corresponding static target relationship chains and dynamic target relationship chains, and stores the static target relationship chains and the dynamic target relationship chains in the form of text; when the visual large model receives the question information input by the user, the corresponding feature information in the question information is first screened out, and then the feature information is compared with the stored static target relationship chains and dynamic target relationship chains, the target panoramic video and the target monitoring video in the video are screened out, and finally the target panoramic video and the target monitoring video are displayed on the display interface of the client in a preset effect by the calculation and display module. The device can quickly and accurately extract the corresponding video according to the question information of the user and clearly present the video to the user.
[0107] It should be noted that the plurality in the present application includes two or more.
[0108] It should be understood that, although each step in the flowchart involved in each embodiment described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0109] Each module in each device in the present application can be realized by software, hardware and their combination in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to the above modules by the processor.
[0110] In one embodiment, a computer device, which can be a server, is provided, and the internal structure diagram thereof can be as shown in Figure 2As shown in the figure. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data required or generated in the execution of the video tracking display method based on the visual large model. The network interface of the computer device is used to communicate with the terminal outside through the network connection. The computer program is executed by the processor to implement a video tracking display method based on a visual large model.
[0111] In one embodiment, a computer device is provided, which can be a terminal, and its internal structure diagram can be as shown in the figure. Figure 3 As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with the terminal outside in a wired or wireless manner. The wireless manner can be realized through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement a video tracking display method based on a visual large model. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0112] Those skilled in the art can understand that, Figure 3 Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0113] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in each of the method embodiments.
[0114] In an embodiment, a computer readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, implements the steps of any of the above method embodiments.
[0115] In an embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of any of the above method embodiments.
[0116] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.
[0117] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to a memory, database or other medium used in the embodiments provided by the present application can include at least one of a non-volatile and volatile memory. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory, an optical storage, a high-density embedded non-volatile memory, a resistive memory (ReRAM), a magnetoresistive random access memory (MRAM), a ferroelectric memory (FRAM), a phase change memory (PCM), a graphene memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0118] Any combination of the technical features in the above embodiments can be made, and for the sake of brevity, not all possible combinations are described, however, as long as there is no conflict, any combination of the technical features should be considered within the scope of the present disclosure.
[0119] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that for ordinary skilled persons in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A video tracing and display method based on a visual macro model, characterized in that: The method comprises: Obtaining video footage, wherein the video footage includes panoramic video footage and surveillance video footage; Extract static and dynamic targets from each frame of video, generate corresponding static and dynamic target relationship chains, and store the static and dynamic target relationship chains in text form; the dynamic target relationship chain includes the name of the dynamic target, the time when the dynamic target appears in the video, the motion trajectory of the dynamic target, the motion state of the dynamic target, and the ID of the associated surveillance camera; Obtain the user's question information, filter out the corresponding feature information in the question information, and compare the feature information with the stored static target relationship chain and dynamic target relationship chain to filter out the target panoramic video and target surveillance video in the video; Display the target panoramic video and the target surveillance video in a preset effect on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area, and a surveillance display area; The method further comprises: According to the motion trajectory of the dynamic target, determine whether the dynamic target is out of the target panoramic video; If the dynamic target leaves the target panoramic video, the associated surveillance camera is determined according to the dynamic target relationship chain, and the target surveillance video after the dynamic target leaves the target panoramic video is found.
2. The method according to claim 1, characterized in that The step of extracting static objects and dynamic objects from each frame of video, generating corresponding static object relationship chains and dynamic object relationship chains, and storing the static object relationship chains and dynamic object relationship chains in text form includes: Split the panoramic video and the surveillance video into corresponding panoramic video frames and surveillance video frames according to their frame rates; Identify static targets and dynamic targets in each panoramic video frame and each surveillance video frame, and obtain corresponding static target attribute information and dynamic target attribute information; Generate corresponding static target video recording sets and dynamic target video recording sets according to static target attribute information or dynamic target attribute information; Based on the static target attribute information, the dynamic target attribute information, the static target video set and the dynamic target video set, a corresponding static target relationship chain or a dynamic target relationship chain is generated, and the static target relationship chain and the dynamic target relationship chain are stored in text form.
3. The method according to claim 1, characterized in that The static target relationship chain includes: the name of the static target, the time when the static target appears in the recorded video, the location of the static target, and the status of the static target.
4. The method according to claim 1, wherein The step of displaying the target panoramic video and the target surveillance video on the display interface of the client with a preset effect includes: The enlarged display area is set at the upper part of the display interface, and its width is equal to the width of the display interface, and is used to play the enlarged target panoramic video or target monitoring video; The panoramic display area is set in the middle of the display interface, between the magnified display area and the monitoring display area, and is used to play the target panoramic video; The monitoring display area is set at the lower part of the display interface, below the panoramic display area, and is used to play and display the target monitoring video.
5. The method according to claim 1, wherein The panoramic display area further includes a plurality of window playback containers and video axes of the same size; The sum of the widths of the multiple window playback containers is equal to the width of the panoramic display area. After the target panoramic video is screened out, the target panoramic video is loaded into the window playback container for video playback; The video axis is located below the multiple window playback containers, and the video axis correspondingly displays the playback time information of the target panoramic video in the multiple window playback containers.
6. The method according to claim 5, characterized in that When the number of target panoramic video recordings is greater than the number of window playback containers, a temporary video storage container is established, and the target panoramic video recordings exceeding the number of window playback containers are taken as candidate target panoramic video recordings and stored in the temporary video storage container in chronological order.
7. The method according to claim 6, characterized in that Detecting whether there is a candidate target panoramic video in the temporary video storage container that is played earlier than the target panoramic video played in the first window playback container in the panoramic display area; If it exists, a first detection control is set with the left window boundary of the first window playback container as a reference point, wherein the width of the first detection control is less than the width of the first window playback container, and the height of the first detection control is greater than or equal to the height of the first window playback container; When an event trigger operation is received on the first detection control, a candidate target panoramic video that is adjacent to the target panoramic video played in the first window playback container and precedes the target panoramic video in the first window playback container is extracted from the temporary video storage container, and loaded into the first window playback container. The target panoramic video in the last window playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is sequentially loaded into the next window playback container for playback; and / or detecting whether there exists a candidate target panoramic video in the temporary video storage container whose playing time is later than the target panoramic video played in the last window playback container in the panoramic display area; If it exists, a second detection control is set with the right window boundary line of the last window playback container as the reference point, wherein the width of the second detection control is less than the width of the last window playback container, and the height of the second detection control is greater than or equal to the height of the last window playback container; When an event trigger operation is received for the second detection control, a candidate target panoramic video that is adjacent to and later than the target panoramic video played in the last window playback container is extracted from the temporary video storage container and loaded into the last window playback container, and the target panoramic video in the first window playback container is stored in the temporary video storage container in chronological order, and the remaining target panoramic video in the window playback container is loaded into the previous window playback container in sequence for playback.
8. The method according to claim 5, characterized in that The monitoring display area further includes a monitoring window playback container corresponding to the window playback container; when a single target panoramic video is associated with multiple target monitoring video videos, the first target monitoring video video among the multiple target monitoring video videos is loaded into the monitoring window playback container for playback, and the remaining target monitoring video videos are displayed in the form of a time list below the corresponding monitoring window playback container.
9. A video tracing and display device based on a visual macromodel, using the video tracing and display method based on a visual macromodel according to any one of claims 1 to 8, characterized in that: The device comprises: An acquisition module is used to acquire video recordings, wherein the video recordings include panoramic video recordings and surveillance video recordings; The visual large model is connected to the acquisition module and is used to extract static and dynamic targets in each frame of video, generate corresponding static target relationship chains and dynamic target relationship chains, and store the static target relationship chains and dynamic target relationship chains in text form; and obtain user question information, filter out corresponding feature information in the question information, and compare the feature information with the stored static target relationship chains and dynamic target relationship chains to filter out target panoramic video and target monitoring video in the video; The computing and display module is connected to the visual large model and the client respectively, and is used to display the target panoramic video and the target monitoring video with preset effects on the display interface of the client, wherein the display interface includes an enlarged display area, a panoramic display area and a monitoring display area.
Citation Information
Patent Citations
Processing method and device for digital video monitoring system
CN116866534A
Panorama video intelligent monitoring method and system
CN101123722A
Full space-time three-dimensional visualization method
CN103795976A