Video recording method and device, computer device and medium

CN122802643APending Publication Date: 2026-09-22SHENZHEN XIAOPAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611029243.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,当用户收到报警通知并试图回溯查看现场情况时,传统检索方式依赖时间戳或简单标签,使得用户在海量视频录像中查找特定目标时效率低下

Benefits of technology

[0014]本发明提供一种视频录像方案,响应于触发的当前报警事件,确定对应于当前报警事件的图像组对齐信息,图像组对齐信息包括当前报警事件对应的当前报警视频片段在存储介质中的起始物理偏移量和结束物理偏移量;确定当前报警事件指示的当前报警区域相较于报警摄像头的空间方位角,并根据空间方位角确定至少一个协助摄像头,以及控制协助摄像头朝向当前报警区域采集得到协助视频片段;调用视觉语言模型对当前报警视频片段与协助视频片段进行联合语义分析,得到当前报警事件的当前语义描述;根据图像组对齐信息和当前语义描述生成对应于当前报警事件的报警通知,并将报警通知推送至用户终端。以此,通过多视角视频数据的协同采集与智能语义理解,能够将报警视频片段与语义描述精准关联,结合图像组对齐信息,使得用户能够通过语义快速定位并回溯关键视频片段,同时,在报警通知中携带当前报警事件的语义描述,还能够提升报警通知的可读性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802643A_ABST
    Figure CN122802643A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of monitoring, and in particular to a video recording method and device, computer equipment and medium, which determines image group alignment information corresponding to the current alarm event in response to the triggered current alarm event, the image group alignment information including the starting physical offset and the ending physical offset of the current alarm video segment in the storage medium; determines the spatial azimuth angle of the current alarm region indicated by the current alarm event relative to the alarm camera, and determines at least one assisting camera according to the spatial azimuth angle, and controls the assisting camera to collect the assisting video segment towards the current alarm region; calls the visual language model to jointly analyze the semantics of the current alarm video segment and the assisting video segment, and obtains the current semantic description of the current alarm event; generates an alarm notification according to the image group alignment information and the current semantic description, and pushes the alarm notification to the user terminal, which can improve the video retrieval efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surveillance technology, and in particular to a video recording method, a video recording device, a computer device, and a computer-readable storage medium. Background Technology

[0002] Currently, surveillance equipment, including network video recorders, typically operates in a linear "event trigger → recording → notification" mode. In this mode, when a predefined alarm event is detected, such as a moving object or unusual sound, the monitoring equipment records video and generates an alarm notification, which is then sent to the user's terminal. However, when a user receives an alarm notification and attempts to review the scene, traditional retrieval methods rely on timestamps or simple tags, making it inefficient for users to find specific targets within massive amounts of video recordings. Summary of the Invention

[0003] This invention provides a video recording method, a video recording device, a computer device, and a computer-readable storage medium, enabling users to directly locate targets through semantics and improving video retrieval efficiency.

[0004] In a first aspect, the video recording method provided by the present invention is applicable to surveillance equipment, wherein the surveillance equipment is connected to at least two cameras, including: In response to the triggered current alarm event, determine the image group alignment information corresponding to the current alarm event. The image group alignment information includes the start physical offset and end physical offset of the current alarm video segment corresponding to the current alarm event in the storage medium. Determine the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, determine at least one auxiliary camera based on the spatial azimuth angle, and control the auxiliary camera to face the current alarm area to acquire auxiliary video clips. The visual language model is invoked to perform joint semantic analysis on the current alarm video clip and the assistance video clip to obtain the current semantic description of the current alarm event; An alarm notification corresponding to the current alarm event is generated based on the image group alignment information and the current semantic description, and then pushed to the user terminal.

[0005] Secondly, the video recording device provided by the present invention is suitable for monitoring equipment, which is connected to at least two cameras, including: The information determination module is used to determine the image group alignment information corresponding to the current alarm event in response to the triggered current alarm event. The image group alignment information includes the start physical offset and end physical offset of the current alarm video segment corresponding to the current alarm event in the storage medium. The assistance determination module is used to determine the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, and to determine at least one assistance camera based on the spatial azimuth angle, and to control the assistance camera to face the current alarm area to acquire assistance video clips. The semantic analysis module is used to call the visual language model to perform joint semantic analysis on the current alarm video clip and the assistance video clip to obtain the current semantic description of the current alarm event. The alarm notification module is used to generate an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description, and push the alarm notification to the user terminal.

[0006] Optionally, in one embodiment, the assistance determination module is used to determine the required rotation angle of other cameras besides the alarm camera toward the current alarm area based on the spatial azimuth angle; acquire image acquisition capability information and busy / idle status information of each other camera; calculate the coverage utility value of each other camera based on the rotation angle, image acquisition capability information and busy / idle status information corresponding to each other camera; and select the other camera with the highest coverage utility value as the assistance camera.

[0007] Optionally, in one embodiment, the alarm notification module is further configured to calculate the semantic similarity between the current semantic description and the historical semantic description of the historical alarm event marked as a false alarm; calculate the event confidence of the current alarm event based on the semantic similarity; and when the event confidence reaches a confidence threshold, generate an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description.

[0008] Optionally, in one embodiment, the video recording device provided by the present invention further includes a video playback module, which is used to respond to a recording viewing request returned by a user terminal, locate the current physical storage address of the current alarm video segment according to the image group alignment information, and push the current alarm video segment to the user terminal for playback according to the current physical storage address.

[0009] Optionally, in one embodiment, the video recording device provided by the present invention further includes a report generation module, which is used to receive current event feedback information returned by the user terminal for the current alarm event; and when a preset report generation period is reached, generate a situation analysis report based on the semantic description and event feedback information of the alarm events accumulated within the report generation period and push it to the user terminal.

[0010] Optionally, in one embodiment, the video recording device provided by the present invention further includes a live video module, which, in response to a live viewing request returned by a user terminal, combines a first real-time video stream from an assisting camera and a second real-time video stream from an alarm camera into a composite video stream with an expanded field of view, and pushes the composite video stream to the user terminal for playback.

[0011] Optionally, in one embodiment, the video recording device provided by the present invention further includes a video retrieval module, configured to receive a video retrieval request sent by a user terminal, the video retrieval request carrying a reference semantic description of the target to be retrieved; retrieve a target alarm video segment whose semantic similarity to the reference semantic description reaches a similarity threshold; locate the target physical storage address according to the image group alignment information corresponding to the target alarm video segment, and push the target alarm video segment to the user terminal for playback according to the target physical storage address.

[0012] Thirdly, the computer device provided by the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the video recording method provided by the present invention.

[0013] Fourthly, the computer-readable storage medium provided by the present invention stores a computer program, which, when executed by a processor, implements the video recording method provided by the present invention.

[0014] This invention provides a video recording scheme. In response to a triggered current alarm event, it determines image group alignment information corresponding to the current alarm event. The image group alignment information includes the start and end physical offsets of the current alarm video segment corresponding to the current alarm event in the storage medium. It determines the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, and determines at least one auxiliary camera based on the spatial azimuth angle. It also controls the auxiliary camera to face the current alarm area to acquire auxiliary video segments. A visual language model is invoked to perform joint semantic analysis on the current alarm video segment and the auxiliary video segments to obtain a current semantic description of the current alarm event. Based on the image group alignment information and the current semantic description, an alarm notification corresponding to the current alarm event is generated and pushed to the user terminal. Thus, through the collaborative acquisition of multi-view video data and intelligent semantic understanding, alarm video segments can be accurately associated with semantic descriptions. Combined with image group alignment information, users can quickly locate and recall key video segments through semantics. Furthermore, carrying the semantic description of the current alarm event in the alarm notification improves the readability of the alarm notification. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1This is a schematic diagram illustrating an application scenario of the video recording method provided in this embodiment of the invention; Figure 2 This is a flowchart illustrating the video recording method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the video recording device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0017] To make the technical problems solved, the technical solutions, and the beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0018] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0019] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0020] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0021] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0022] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0023] Please refer to Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment of the video recording method provided by the present invention. As one embodiment, the video recording method provided by the present invention can be applied to a monitoring device, where the monitoring device and the user terminal are connected via a network. The network serves as a medium for providing a communication link between the monitoring device and the user terminal, and can include various connection types, such as wired communication links, wireless communication links, etc., which are not limited in this embodiment of the present invention.

[0024] It should be noted that, Figure 1 The monitoring equipment, network, and user terminals shown are merely illustrative. Depending on actual needs, there can be any number of user terminals. For example, monitoring equipment can be a network video recorder, a digital video recorder, or a smart camera with video capture and storage capabilities, while user terminals can be smartphones, tablets, desktop computers, or laptops—terminal devices with video processing capabilities.

[0025] In some embodiments, the monitoring device may, in response to a triggered current alarm event, determine image group alignment information corresponding to the current alarm event, the image group alignment information including the start and end physical offsets of the current alarm video segment corresponding to the current alarm event in the storage medium; determine the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, and determine at least one auxiliary camera based on the spatial azimuth angle, and control the auxiliary camera to capture auxiliary video segments facing the current alarm area; invoke a visual language model to perform joint semantic analysis on the current alarm video segments and auxiliary video segments to obtain the current semantic description of the current alarm event; generate an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description, and push the alarm notification to the user terminal.

[0026] Please refer to Figure 2 This is a flowchart illustrating a video recording method disclosed in an embodiment of the present invention, as shown below. Figure 2As shown, the process of this video recording method can be as follows: In S110, in response to the triggered current alarm event, image group alignment information corresponding to the current alarm event is determined. The image group alignment information includes the start physical offset and end physical offset of the current alarm video segment corresponding to the current alarm event in the storage medium.

[0027] It should be noted that in the following embodiments, a monitoring device is used as the execution subject for description, but those skilled in the art should understand that the method can also be executed by other computer devices that can provide equivalent functions.

[0028] In this embodiment of the invention, the monitoring system is connected to at least two cameras, continuously receiving real-time video streams from each camera and caching them in local storage. When a camera or monitoring device detects abnormal movement, sudden sound changes, or intrusion of a specific target in the video frame, an alarm event is generated. The current alarm event refers to the alarm event triggered at the current moment, which can be any moment, depending on when the triggering condition occurs.

[0029] In response to a triggered alarm event, the monitoring device first locates the video segment in the storage medium that matches the timestamp of the alarm event, marking it as the current alarm video segment. It then extracts the start and end physical offsets of this current alarm video segment from the storage medium to construct image group alignment information. The physical offset refers to the relative byte position of the video segment in the storage medium. Using the start and end physical offsets, the monitoring device can directly and quickly locate and read the corresponding video segment in the storage medium without traversing the entire file index, thus significantly reducing I / O overhead.

[0030] For example, assuming the current alarm event occurs between frames 1500 and 1800 of the video stream, the monitoring device finds the nearest key frame starting from frame 1500 and uses the physical offset of that key frame as the starting physical offset. It then finds the nearest key frame starting from frame 1800 and uses the physical offset of that key frame as the ending physical offset.

[0031] In S120, the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera is determined, and at least one auxiliary camera is determined based on the spatial azimuth angle, and the auxiliary camera is controlled to face the current alarm area to acquire an auxiliary video segment.

[0032] Furthermore, the monitoring equipment determines the pixel coordinates of the current alarm area indicated by the current alarm event in the video frame. Combining this with the pre-calibrated homography mapping matrix and intrinsic parameter matrix of the alarm camera, the pixel coordinates are converted into a three-dimensional spatial position in the world coordinate system. Then, based on this three-dimensional spatial position and the installation pose of the alarm camera, the spatial azimuth angle of the current alarm area relative to the alarm camera is calculated. Here, the alarm camera refers to the camera that triggered the current alarm event, and this spatial azimuth angle reflects the specific direction of the current alarm area in three-dimensional space.

[0033] Subsequently, based on the calculated spatial azimuth angle, the monitoring equipment identifies at least one auxiliary camera from the other connected cameras. For example, it can identify a camera whose field of view covers the current alarm area as an auxiliary camera based on the spatial azimuth angle.

[0034] Once the assisting camera is identified, the monitoring equipment sends control commands to it, instructing it to adjust its pan-tilt-zoom (PTZ) orientation to capture video of the current alarm area. This results in an assisting video clip that is spatiotemporally related to the current alarm video segment. This assisting video clip not only fills in blind spots from a single viewpoint but also provides a reference for spatiotemporal correlation.

[0035] It should be noted that the length of the collected assistive video clips is not limited in the embodiments of the present invention, and their duration can be dynamically adjusted according to the duration and complexity of the actual alarm event. For example, for an intrusion alarm that is triggered instantaneously, only a few seconds of key video clips can be collected as assistive video clips, while for continuous abnormal behaviors, such as loitering or gathering, the collection time can be appropriately extended.

[0036] In S130, the visual language model is invoked to perform joint semantic analysis on the current alarm video clip and the assistance video clip to obtain the current semantic description of the current alarm event.

[0037] A Vision-Language Model (VLM) is a multimodal artificial intelligence model that can simultaneously understand images / videos and natural language. It is like giving a large language model "eyes," enabling it not only to understand text but also to "understand" the content of images and describe it in words.

[0038] As shown above, after acquiring the assistance video clips, the monitoring equipment inputs the current alarm video clip and the assistance video clip into the visual language model. The visual language model performs joint semantic analysis on the current alarm video clip and the assistance video clip, deeply mining the spatiotemporal correlation features of the current alarm video clip and the assistance video clip, thereby generating an accurate and contextually coherent current semantic description.

[0039] For example, in the case of a nighttime intrusion, the alarm camera only captures a blurry black figure, while the assist camera clearly records the intruder's facial features and the tools he is holding from the side. The visual language model can then integrate multi-perspective information to reconstruct fragmented visual cues into a complete semantic description, such as "a man wearing a black mask is holding a crowbar and is trying to break the west-facing window."

[0040] In S140, an alarm notification corresponding to the current alarm event is generated based on the image group alignment information and the current semantic description, and the alarm notification is pushed to the user terminal.

[0041] In this embodiment of the invention, after obtaining the current semantic description of the current alarm event, the monitoring device, on the one hand, associates and stores the current semantic description with the current alarm video segment, enabling rapid retrieval and backtracking of the alarm video segment based on semantics. On the other hand, the monitoring device also generates a structured alarm notification based on the current semantic description and image group alignment information, and pushes it to the user terminal. The specific form of the alarm notification is not strictly limited here; it can be a text message containing the current semantic description, a multimedia card integrating the current semantic description and key video frames, or even a voice broadcast, etc.

[0042] On the other hand, after receiving an alarm notification, the user terminal can visually present the notification through a graphical interface, allowing users to quickly determine the authenticity of the alarm based on a clear semantic description and thus provide accurate feedback. For example, users can select options such as "false alarm" or "handled" based on the actual situation. Correspondingly, the user terminal transmits the user feedback back to the monitoring equipment. The monitoring equipment then associates and stores the user feedback with the corresponding semantic description and alarm video clip, forming a closed-loop data chain.

[0043] As described above, the video recording solution provided by this invention, in response to a triggered current alarm event, determines image group alignment information corresponding to the current alarm event. This image group alignment information includes the start and end physical offsets of the current alarm video segment corresponding to the current alarm event in the storage medium. It also determines the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, determines at least one auxiliary camera based on the spatial azimuth angle, and controls the auxiliary camera to capture auxiliary video segments facing the current alarm area. Furthermore, it invokes a visual language model to perform joint semantic analysis on the current alarm video segment and the auxiliary video segments to obtain a current semantic description of the current alarm event. Finally, it generates an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description, and pushes the alarm notification to the user terminal. Thus, through the collaborative acquisition of multi-view video data and intelligent semantic understanding, alarm video segments can be accurately associated with semantic descriptions. Combined with image group alignment information, this allows users to quickly locate and trace key video segments through semantics. Simultaneously, carrying the semantic description of the current alarm event in the alarm notification also improves the readability of the alarm notification.

[0044] Optionally, in one embodiment, determining at least one assisting camera based on a spatial azimuth angle includes: Based on the spatial azimuth angle, determine the required rotation angle of other cameras besides the alarm camera towards the current alarm area; Acquire image acquisition capability information and busy / idle status information for each of the other cameras; The coverage utility value of each other camera is calculated based on its rotation angle, image acquisition capability, and busy / idle status. Select the other camera with the highest coverage utility value as the auxiliary camera.

[0045] Optionally, in one embodiment, an optional strategy for assisting in camera selection is provided.

[0046] The monitoring camera calculates the rotation angle required for each other camera to turn towards the current alarm area based on the spatial azimuth angle of the current alarm area relative to the alarm camera, combined with the installation posture and current orientation of each other camera. This rotation angle represents the mechanical adjustment range required for the camera to turn towards the current alarm area, and is directly related to response delay and energy consumption cost.

[0047] In addition, the monitoring equipment also obtains image acquisition capability information from other cameras, such as resolution, focal length and low-light performance, as well as their current busy / idle status, such as whether they are performing other tasks or in a sleep state.

[0048] After acquiring the rotation angle, image acquisition capability, and busy / idle status of each camera, the monitoring equipment evaluates the coverage utility value of each candidate camera according to preset utility evaluation rules. This coverage utility value comprehensively reflects the camera's overall performance in response speed, image quality, and resource availability. There are no specific restrictions on the setting of the utility evaluation rules, as long as they can reasonably quantify the aforementioned multi-dimensional indicators. For example, they can be flexibly configured according to the dynamic needs of the actual monitoring scenario. For instance, in low-light nighttime scenarios, higher weight can be given to low-light performance, while in bright daylight or complex dynamic scenarios, emphasis can be placed on resolution and frame rate, and so on. Through this dynamic weight adjustment mechanism, the optimal assisting camera can be selected under different environmental conditions, thereby maximizing the effectiveness of collaborative monitoring.

[0049] After evaluating the coverage utility value of other cameras besides the alarm camera, the monitoring equipment can sort the other cameras according to their coverage utility value and prioritize the other cameras with the highest coverage utility value as auxiliary cameras.

[0050] Optionally, in one embodiment, before generating an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description, the method further includes: Calculate the semantic similarity between the current semantic description and the historical semantic description of the historical alarm events marked as false alarms; Calculate the event confidence score of the current alarm event based on semantic similarity; If the event confidence level reaches the confidence threshold, an alarm notification corresponding to the current alarm event is generated based on the image group alignment information and the current semantic description.

[0051] In this embodiment of the invention, upon receiving the semantic description of the current alarm event, the monitoring device does not directly generate an alarm notification. Instead, it first identifies historical alarm events marked as false alarms (i.e., alarm events reported as false alarms by the user before the current alarm event), records them as historical false alarm events, and calculates the semantic similarity between the current semantic description and the historical semantic descriptions corresponding to these historical false alarm events. For example, the monitoring device can map the current semantic description and the historical false alarm semantic descriptions to the same high-dimensional semantic space, and quantify the semantic similarity between the current semantic description and the historical semantic description by calculating the cosine similarity or Euclidean distance between vectors.

[0052] Subsequently, the monitoring equipment further derives and calculates the event confidence of the current alarm event based on the calculated maximum semantic similarity. The higher the semantic similarity, the smaller the difference between the current alarm event and the false alarm historical alarm events, and the lower the event confidence of the current alarm event. Conversely, the lower the semantic similarity, the greater the difference between the current alarm event and the false alarm historical alarm events, and the higher the event confidence of the current alarm event.

[0053] After calculating the event confidence score of the current alarm event, the monitoring device compares the event confidence score with a preset confidence score threshold. If the event confidence score of the current alarm event is higher than or equal to the confidence score threshold, the monitoring device determines that the current alarm event is a valid alarm event, and generates an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description. If the event confidence score of the current alarm event is lower than the confidence score threshold, the monitoring device determines that the current alarm event is an invalid alarm event, filters out this trigger, and does not generate an alarm notification, thereby effectively reducing the false alarm rate.

[0054] Optionally, in one embodiment, after pushing the alarm notification to the user terminal, the video recording method provided by the present invention further includes: In response to the video recording viewing request returned by the user terminal, the current physical storage address of the current alarm video segment is located based on the image group alignment information, and the current alarm video segment is pushed to the user terminal for playback based on the current physical storage address.

[0055] Understandably, if a user has doubts about the authenticity of the alarm or wishes to review the details of the scene after viewing the alarm notification of the current alarm event displayed on the user terminal, they can initiate a video recording viewing request through the user terminal.

[0056] Accordingly, after receiving a video viewing request from the user terminal for the current alarm event, the monitoring equipment directly locates the current physical storage address of the current alarm video segment in the storage medium based on the image group alignment information of the current alarm video segment, achieving millisecond-level accurate retrieval without traversing the entire data. Subsequently, based on the current physical storage address, the monitoring equipment directly reads the current alarm video segment from the storage medium and pushes it to the user terminal for real-time playback.

[0057] In practice, monitoring equipment can attach the image group alignment information of the current alarm event to the alarm notification. This allows the user terminal to directly include the image group alignment information in the recording viewing request when the user initiates a recording viewing request based on the alarm notification, and send it to the monitoring equipment along with the request. In this way, the monitoring equipment can quickly parse the image group alignment information in the recording viewing request, thereby rapidly locating the physical storage address of the current alarm video segment and efficiently pushing the current alarm video segment to the user terminal for playback.

[0058] Optionally, in one embodiment, after pushing the current alarm video clip to the user terminal for playback based on the current physical storage address, the method further includes: Receive current user feedback from the user terminal regarding the current alarm event; When the preset report generation cycle is reached, a situation analysis report is generated and pushed to the user terminal based on the semantic description of the accumulated alarm events within the report generation cycle and user feedback.

[0059] In this embodiment of the invention, the monitoring device can also receive user feedback from the user terminal regarding the pushed alarm events, including but not limited to user confirmation of the authenticity of the alarm, false alarm marking, or supplementary explanations.

[0060] It should be noted that the embodiments of the present invention also include a pre-configured report generation cycle for triggering periodic situation analysis report generation tasks. The specific duration of the report generation cycle is not limited here and can be flexibly set according to actual business needs, such as daily, weekly, or monthly cycles.

[0061] When the preset report generation cycle is reached, the monitoring equipment aggregates the accumulated semantic descriptions and corresponding user feedback within that cycle, and inputs these aggregated semantic descriptions and user feedback into the big language model. Through the deep semantic analysis and logical reasoning capabilities of the big language model, it performs multi-dimensional correlation mining and trend analysis on these semantic descriptions and user feedback, generating a summary situation analysis conclusion, such as, "A total of 3 incidents occurred today, of which 1 was a false alarm, and the overall security situation is stable."

[0062] Subsequently, based on the situation analysis conclusions and relevant information such as semantic descriptions of alarm events within the report generation period, occurrence times, alarm cameras, and assisting cameras, the monitoring equipment generates a structured situation analysis report. The specific format of this situation analysis report is not limited here; it can be flexibly selected by those skilled in the art according to actual needs. For example, the following shows a sample content of a situation analysis report: [Daily Report, June 25, 2026] Overview: A total of 3 incidents occurred today, including 1 false alarm.

[0063] Event list: 09:15: Personnel loitering detected at the front door (CH1 triggered, CH5 assists) Semantic: A man in a blue coat, paused for 10 seconds.

[0064] Action: [Click to view video] 14:30: Abnormal movement in the garage area (triggered by CH2, assisted by CH1) Semantic meaning: Pet cat activities.

[0065] Action: [Click to view video] 18:45: Package delivery at the door (CH1 triggered, CH5 assists) Semantic meaning: A courier places a package.

[0066] Action: [Click to view video] The "[Click to view video]" option allows users to directly access and play the corresponding alarm video clip by clicking the link, enabling rapid retrieval and accurate verification of alarm events.

[0067] The above example is only for illustrating the report structure. In actual applications, fields can be adjusted or visualization charts can be added according to needs. Through this structured presentation, users can not only quickly grasp the overall security dynamics, but also delve into the details and handling status of specific events, thereby significantly improving monitoring efficiency and decision-making accuracy.

[0068] Optionally, in one embodiment, after pushing the alarm notification to the user terminal, the method further includes: In response to a live streaming viewing request from the user terminal, the first real-time video stream from the assisting camera and the second real-time video stream from the alarm camera are combined into a composite video stream with an expanded field of view, and the composite video stream is pushed to the user terminal for playback.

[0069] This invention also provides a live streaming capability that expands the field of view for monitoring devices to user terminals.

[0070] If a user needs to further verify the on-site situation after reviewing the alarm notification of the current alarm event displayed on the user terminal, they can initiate a live broadcast viewing request through the user terminal.

[0071] Accordingly, after receiving a live viewing request for the current alarm event from the user terminal, the monitoring equipment combines the first real-time video stream from the assisting camera and the second real-time video stream from the alarm camera into a composite video stream with an expanded field of view, and pushes the composite video stream to the user terminal for real-time playback.

[0072] In practice, the monitoring equipment first synchronizes and spatially calibrates the first and second real-time video streams to ensure seamless connection in terms of timing and viewing angle. Then, using image stitching or fusion algorithms, the two video streams are integrated into a composite image with a wider viewing angle. The resulting composite video stream eliminates blind spots from single perspectives, providing users with a panoramic view of the scene. For example, the second real-time video stream from the alarm camera focuses on the area in front of the door that triggered the alarm, while the first real-time video stream from the assisting camera covers the side of the porch and the steps. After merging, the composite video stream will fully present the entire view of the area in front of the door and its surroundings.

[0073] Optionally, in one embodiment, the video recording method provided by the present invention further includes: Receive video retrieval requests sent by user terminals. The video retrieval requests carry a reference semantic description of the target to be retrieved. Retrieve target alarm video clips whose semantic similarity to the reference semantic description reaches a similarity threshold; The target alarm video clip is located based on the image group alignment information corresponding to the target alarm video clip, and the target alarm video clip is pushed to the user terminal for playback based on the target physical storage address.

[0074] In this embodiment of the invention, the user can input a natural language description, such as "a person wearing red clothes," on the user terminal as needed. The user terminal then uses the natural language description input by the user as a reference semantic description to generate a video retrieval request and send it to the monitoring equipment.

[0075] Accordingly, after receiving a video retrieval request from a user terminal, the monitoring equipment first extracts the reference semantic description carried in the video retrieval request and calculates the semantic similarity between the reference semantic description and the semantic descriptions of each alarm video segment in the storage medium. For example, the monitoring equipment can map the reference semantic description and the semantic descriptions of the alarm video segments to the same high-dimensional semantic space, and quantify the semantic similarity between the reference semantic description and the semantic descriptions of the alarm video segments by calculating the cosine similarity or Euclidean distance between vectors.

[0076] Subsequently, the monitoring equipment filters out alarm video clips whose semantic similarity to the reference semantic description reaches a similarity threshold, and records them as target alarm video clips. Next, the monitoring equipment locates the target physical storage address based on the image group alignment information corresponding to the target alarm video clip, and then quickly reads the corresponding target alarm video clip from the storage medium and pushes the target alarm video clip to the user terminal for playback.

[0077] The above-mentioned video retrieval method based on semantic understanding has completely changed the inefficient mode of relying on timelines or manual frame-by-frame inspection. It allows users to locate targets in seconds by describing key semantics without having to remember precise time points, thus achieving efficient video retrieval.

[0078] To facilitate better implementation of the above video recording method, this embodiment of the invention also provides a corresponding video recording device. The meanings of the terms used are the same as in the above video recording method; for specific implementation details, please refer to the descriptions in the above method embodiments.

[0079] Please refer to Figure 3 The video recording device may include an information determination module 210, an assistance determination module 220, a semantic analysis module 230, and an alarm notification module 240. Detailed descriptions of each functional module are as follows: The information determination module 210 is used to determine the image group alignment information corresponding to the current alarm event in response to the triggered current alarm event. The image group alignment information includes the start physical offset and end physical offset of the current alarm video segment corresponding to the current alarm event in the storage medium. The assistance determination module 220 is used to determine the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, and to determine at least one assistance camera based on the spatial azimuth angle, and to control the assistance camera to face the current alarm area to acquire assistance video clips. The semantic analysis module 230 is used to call the visual language model to perform joint semantic analysis on the current alarm video segment and the assistance video segment to obtain the current semantic description of the current alarm event. The alarm notification module 240 is used to generate an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description, and push the alarm notification to the user terminal.

[0080] Optionally, in one embodiment, the assistance determination module 220 is used to determine the required rotation angle of other cameras besides the alarm camera toward the current alarm area based on the spatial azimuth angle; acquire image acquisition capability information and busy / idle status information of each other camera; calculate the coverage utility value of each other camera based on the rotation angle, image acquisition capability information and busy / idle status information of each other camera; and select the other camera with the highest coverage utility value as the assistance camera.

[0081] Optionally, in one embodiment, the alarm notification module 240 is further configured to calculate the semantic similarity between the current semantic description and the historical semantic description of the historical alarm event marked as a false alarm; calculate the event confidence of the current alarm event based on the semantic similarity; and generate an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description when the event confidence reaches a confidence threshold.

[0082] Optionally, in one embodiment, the video recording device provided by the present invention further includes a video playback module, which is used to respond to a recording viewing request returned by a user terminal, locate the current physical storage address of the current alarm video segment according to the image group alignment information, and push the current alarm video segment to the user terminal for playback according to the current physical storage address.

[0083] Optionally, in one embodiment, the video recording device provided by the present invention further includes a report generation module, which is used to receive current event feedback information returned by the user terminal for the current alarm event; and when a preset report generation period is reached, generate a situation analysis report based on the semantic description and event feedback information of the alarm events accumulated within the report generation period and push it to the user terminal.

[0084] Optionally, in one embodiment, the video recording device provided by the present invention further includes a live video module, which, in response to a live viewing request returned by a user terminal, combines a first real-time video stream from an assisting camera and a second real-time video stream from an alarm camera into a composite video stream with an expanded field of view, and pushes the composite video stream to the user terminal for playback.

[0085] Optionally, in one embodiment, the video recording device provided by the present invention further includes a video retrieval module, configured to receive a video retrieval request sent by a user terminal, the video retrieval request carrying a reference semantic description of the target to be retrieved; retrieve a target alarm video segment whose semantic similarity to the reference semantic description reaches a similarity threshold; locate the target physical storage address according to the image group alignment information corresponding to the target alarm video segment, and push the target alarm video segment to the user terminal for playback according to the target physical storage address.

[0086] For specific limitations regarding the video recording device, please refer to the limitations on the video recording method above, which will not be repeated here. Each module in the aforementioned video recording device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.

[0087] In one embodiment, a computer device is provided, which may be a monitoring device, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface connects to external wireless clients, providing wireless network access services to the connected clients. When the computer program is executed by the processor, it implements the video recording method provided by this invention.

[0088] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the video recording method described in the above embodiment.

[0089] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the video recording method described above.

[0090] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0092] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A video recording method, applicable to surveillance equipment, wherein the surveillance equipment is connected to at least two cameras, characterized in that, The video recording method includes: In response to a triggered current alarm event, image group alignment information corresponding to the current alarm event is determined. The image group alignment information includes the start physical offset and end physical offset of the current alarm video segment corresponding to the current alarm event in the storage medium. Determine the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, determine at least one auxiliary camera based on the spatial azimuth angle, and control the auxiliary camera to face the current alarm area to acquire auxiliary video clips. The visual language model is invoked to perform joint semantic analysis on the current alarm video segment and the assistance video segment to obtain the current semantic description of the current alarm event; An alarm notification corresponding to the current alarm event is generated based on the image group alignment information and the current semantic description, and the alarm notification is pushed to the user terminal.

2. The video recording method according to claim 1, characterized in that, The step of determining at least one assisting camera based on the spatial azimuth angle includes: Based on the spatial azimuth angle, determine the required rotation angles of other cameras besides the alarm camera toward the current alarm area; Obtain image acquisition capability information and busy / idle status information for each of the other cameras; The coverage utility value of each of the other cameras is calculated based on the rotation angle, image acquisition capability information, and busy / idle status information corresponding to each of the other cameras. Select the other camera with the highest coverage utility value as the auxiliary camera.

3. The video recording method according to claim 1, characterized in that, Before generating an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description, the method further includes: Calculate the semantic similarity between the current semantic description and the historical semantic descriptions of events marked as false alarms. The event confidence score of the current alarm event is calculated based on the semantic similarity. If the confidence level of the event reaches the confidence level threshold, an alarm notification corresponding to the current alarm event is generated based on the image group alignment information and the current semantic description.

4. The video recording method according to claim 1, characterized in that, After pushing the alarm notification to the user terminal, the method further includes: In response to the video recording viewing request returned by the user terminal, the current physical storage address of the current alarm video segment is located according to the image group alignment information, and the current alarm video segment is pushed to the user terminal for playback according to the current physical storage address.

5. The video recording method according to claim 4, characterized in that, After pushing the current alarm video clip to the user terminal for playback based on the current physical storage address, the method further includes: Receive current event feedback information returned by the user terminal regarding the current alarm event; When the preset report generation cycle is reached, a situation analysis report is generated and pushed to the user terminal based on the semantic description and event feedback information of the accumulated alarm events within the report generation cycle.

6. The video recording method according to claim 1, characterized in that, After pushing the alarm notification to the user terminal, the method further includes: In response to the live viewing request returned by the user terminal, the first real-time video stream from the assisting camera and the second real-time video stream from the alarm camera are combined into a composite video stream with an expanded field of view, and the composite video stream is pushed to the user terminal for playback.

7. The video recording method according to claim 1, characterized in that, Also includes: Receive a video retrieval request sent by the user terminal, wherein the video retrieval request carries a reference semantic description of the target to be retrieved; Retrieve target alarm video clips whose semantic similarity to the reference semantic description reaches a similarity threshold; The target physical storage address is located based on the image group alignment information corresponding to the target alarm video clip, and the target alarm video clip is pushed to the user terminal for playback based on the target physical storage address.

8. A video recording device, suitable for surveillance equipment, wherein the surveillance equipment is connected to at least two cameras, characterized in that, include: The information determination module is used to determine the image group alignment information corresponding to the current alarm event in response to the triggered current alarm event. The image group alignment information includes the start physical offset and end physical offset of the current alarm video segment corresponding to the current alarm event in the storage medium. The assistance determination module is used to determine the spatial azimuth angle of the current alarm area indicated by the current alarm event relative to the alarm camera, and to determine at least one assistance camera based on the spatial azimuth angle, and to control the assistance camera to face the current alarm area to acquire assistance video clips. The semantic analysis module is used to call the visual language model to perform joint semantic analysis on the current alarm video segment and the assistance video segment to obtain the current semantic description of the current alarm event; The alarm notification module is used to generate an alarm notification corresponding to the current alarm event based on the image group alignment information and the current semantic description, and push the alarm notification to the user terminal.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video recording method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the video recording method according to any one of claims 1 to 7.