Visitor processing method and device, storage medium and electronic equipment
By collecting multiple monitoring terminal videos in the intelligent security monitoring system and using a large security model to identify and generate visitor objects across the terminal, the problem of unintelligent visitor processing is solved, and efficient and accurate data integration and user-friendly guest event playback are achieved.
Patent Information
- Application Number
- CN202510571177.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-29
AI Technical Summary
In the existing intelligent security monitoring scenario, the visitor processing is not intelligent enough, and the space-time continuity and intelligence level of multi-camera data processing are insufficient, making it difficult for users to quickly locate and play back the complete visiting events of specific visitors.
By collecting videos of multiple security monitoring terminals in security scenarios, using security processing models to identify visiting objects across the terminals, determining the relationship tags between visiting objects and user objects, and generating visiting object events, displaying them in the security monitoring interface.
It realizes efficient, accurate and story-based security monitoring data integration, improves the space-time continuity and intelligence level of multi-camera data processing, and users can quickly locate and play back complete visits from specific visitors, improving the real-time and user experience of the security monitoring system.
Smart Images

Figure CN120388436A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and particularly to a visitor processing method, apparatus, storage medium, and electronic device. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and deep learning technologies, the intelligent security monitoring scenario is undergoing a brand-new transformation. The security video data analysis technology based on multi-camera collaborative monitoring in the intelligent security monitoring scenario has been widely applied, providing an efficient and real-time monitoring solution. At the same time, the robustness and accuracy of various video analysis algorithms in complex scenarios have been continuously improved, promoting the further development of video content automated processing and event detection technologies, thus laying a solid technical foundation for future applications such as intelligent security and behavior recognition. Summary of the Invention
[0003] Embodiments of this specification provide a visitor processing method, apparatus, storage medium, and electronic device. The technical solutions are as follows:
[0004] In a first aspect, embodiments of this specification provide a visitor processing method, which includes:
[0005] Collect multiple security monitoring terminal videos in the security scenario;
[0006] Based on each of the security monitoring terminal videos, use a security processing large model to perform cross-terminal identification of visiting objects to obtain an object visiting scenario video set of at least one visiting object;
[0007] Determine the visiting object relationship label between the visiting object and the user object, and generate a visiting object event for each of the visiting objects based on the visiting object relationship label and the object visiting scenario video set;
[0008] Display each of the visiting objects and the corresponding visiting object events on the security monitoring interface.
[0009] In a feasible implementation, the step of using a security processing large model to perform cross-terminal identification of visiting objects based on each of the security monitoring terminal videos to obtain an object visiting scenario video set of at least one visiting object includes:
[0010] Perform visiting target detection processing on each of the security monitoring terminal videos to obtain visiting target detection meta-information;
[0011] Use the security processing large model to perform cross-terminal identification of visiting objects based on the multiple security monitoring terminal videos and the visiting target parsing meta-information to determine at least one visiting object, and extract the object visiting scenario video set corresponding to the visiting object from the multiple security monitoring terminal videos.
[0012] In a feasible implementation manner, the security processing large model determines at least one visiting object through cross-terminal identification of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information, and extracts an object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos, including:
[0013] The security processing large model determines at least one visiting object through cross-terminal identification of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information;
[0014] The security processing large model determines a target motion description text and target motion description key frames for the visiting object, and performs multi-modal fusion based on the target motion description text and the target motion description key frames to obtain a potential visiting object and obtain target motion description event information;
[0015] The security processing large model generates a visiting video set editing script for the multiple security monitoring end videos based on the target motion description event information;
[0016] Based on the visiting video set editing script, an object visiting scenario video set corresponding to the visiting object is extracted from the multiple security monitoring end videos.
[0017] In a feasible implementation manner, the security processing large model generates a visiting video set editing script for the multiple security monitoring end videos based on the target motion description event information, including:
[0018] The security processing large model performs key clip segment parsing processing on the multiple security monitoring end videos based on the target motion description event information to obtain the visiting video set editing time corresponding to the visiting object;
[0019] Determine the visiting object transition mode and visiting story description subtitles for multiple key clip segments of the visiting object;
[0020] Based on the visiting video set editing time, the visiting object transition mode, and the visiting story description subtitles, a visiting video set editing script for the multiple security monitoring end videos is generated.
[0021] In a feasible implementation manner, the extracting an object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the visiting video set editing script includes:
[0022] Based on the visiting video set editing time in the visiting video set editing script, multiple key clip segments for the visiting object are determined from the multiple security monitoring end videos;
[0023] Determine the transition display effect configuration between each of the key clip segments according to the visitor transition mode in the clip script of the visit video set, and generate a video segment splicing sequence based on the transition display effect configuration and multiple key clip segments;
[0024] Add story description subtitles to the video segment splicing sequence based on the visit story description subtitles in the clip script of the visit video set to generate an object visit scenario video set corresponding to the visitor object.
[0025] In a feasible implementation manner, the obtaining of the visit target detection meta-information by performing visit target detection processing on each of the security monitoring end videos includes:
[0026] Input each of the security monitoring end videos into a visit target detection model, and determine visit target detection information, visit event information, cross-end tracking information of the visit target, visit target feature information, and visit event behavior information through the visit target detection model;
[0027] Generate visit target detection meta-information based on the visit target detection information, visit event information, cross-end tracking information of the visit target, visit target feature information, and visit event behavior information.
[0028] In a feasible implementation manner, after displaying each of the visitor objects and the visitor object events corresponding to the visitor objects on the security monitoring interface, it further includes:
[0029] In response to an event selection operation on the target visit object event of the target visit object, output the object visit scenario video set.
[0030] In a second aspect, an embodiment of the present specification provides a visitor processing device, and the device includes:
[0031] A video acquisition module, configured to acquire multiple security monitoring end videos in a security scenario;
[0032] A video recognition module, configured to perform cross-end recognition of visitor objects on each of the security monitoring end videos by using a security processing large model to obtain an object visit scenario video set of at least one visitor object;
[0033] An event processing module, configured to determine a visit object relationship label between the visit object and the user object, and generate a visit object event of each of the visit objects based on the visit object relationship label and the object visit scenario video set;
[0034] The event processing module is configured to display each of the visit objects and the visit object events corresponding to the visit objects on the security monitoring interface.
[0035] In a feasible implementation manner, the object visit scenario video set of at least one visiting object obtained by cross-terminal recognition of the visiting object by using a security processing large model based on each of the security monitoring end videos includes:
[0036] Performing visiting target detection processing on each of the security monitoring end videos to obtain visiting target detection meta-information;
[0037] Using a security processing large model to perform cross-terminal recognition of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object, and extracting the object visit scenario video set corresponding to the visiting object from the multiple security monitoring end videos.
[0038] In a feasible implementation manner, the using a security processing large model to perform cross-terminal recognition of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object, and extracting the object visit scenario video set corresponding to the visiting object from the multiple security monitoring end videos includes:
[0039] Using a security processing large model to perform cross-terminal recognition of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object;
[0040] Using a security processing large model to determine the target motion description text and the target motion description key frames for the visiting object, and performing multi-modal fusion based on the target motion description text and the target motion description key frames to obtain a potential visiting object to obtain target motion description event information;
[0041] Using a security processing large model to generate a visiting video set clip script for the multiple security monitoring end videos based on the target motion description event information;
[0042] Extracting the object visit scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the visiting video set clip script.
[0043] In a feasible implementation manner, the using a security processing large model to generate a visiting video set clip script for the multiple security monitoring end videos based on the target motion description event information includes:
[0044] Using a security processing large model to perform key clip segment parsing processing on the multiple security monitoring end videos based on the target motion description event information to obtain the visiting video set clip time corresponding to the visiting object;
[0045] Determining the visiting object transition mode and the visiting story description subtitles for the multiple key clip segments of the visiting object;
[0046] Generate a visiting video set editing script for the multiple security monitoring terminal videos based on the editing time of the visiting video set, the transition mode of the visiting object, and the caption of the visiting story description.
[0047] In a feasible implementation manner, the extracting, from the multiple security monitoring terminal videos, an object visiting scenario video set corresponding to the visiting object based on the visiting video set editing script includes:
[0048] Determine multiple key editing segments for the visiting object from the multiple security monitoring terminal videos based on the editing time of the visiting video set in the visiting video set editing script;
[0049] Determine the transition switching display effect configuration between each of the key editing segments according to the transition mode of the visiting object in the visiting video set editing script, and generate a video segment splicing sequence based on the transition switching display effect configuration and the multiple key editing segments;
[0050] Add the caption of the story description to the video segment splicing sequence based on the caption of the visiting story description in the visiting video set editing script to generate an object visiting scenario video set corresponding to the visiting object.
[0051] In a feasible implementation manner, the obtaining, by performing visiting target detection processing on each of the security monitoring terminal videos, visiting target detection meta-information includes:
[0052] Input each of the security monitoring terminal videos into a visiting target detection model, and determine visiting target detection information, visiting event information, visiting target cross-terminal tracking information, visiting target feature information, and visiting event behavior information through the visiting target detection model;
[0053] Generate visiting target detection meta-information based on the visiting target detection information, visiting event information, visiting target cross-terminal tracking information, visiting target feature information, and visiting event behavior information.
[0054] In a feasible implementation manner, after displaying each of the visiting objects and the visiting object events corresponding to the visiting objects on the security monitoring interface, it further includes:
[0055] In response to an event selection operation on the target visiting object event of the target visiting object, output the object visiting scenario video set.
[0056] In a third aspect, an embodiment of the present specification provides a computer storage medium, which stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the above method steps.
[0057] Fourthly, an embodiment of this specification provides an electronic device, which may include: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the above method steps.
[0058] The beneficial effects brought by the technical solutions provided by some embodiments of this specification at least include:
[0059] In one or more embodiments of this specification, by collecting multiple security monitoring end videos in a security scenario, based on each security monitoring end video, using a security processing large model to perform cross-end identification of visiting objects to obtain an object visiting scenario video set of at least one visiting object, determining a visiting object relationship label between the visiting object and the user object, generating a visiting object event for each visiting object based on the visiting object relationship label and the object visiting scenario video set, and displaying each visiting object and the corresponding visiting object event on the security monitoring interface, efficient, accurate and story-based security monitoring data integration is achieved, avoiding the limitation of non-intelligent visitor processing in the related art of security monitoring scenarios. It not only improves the spatio-temporal continuity and intelligent level of multi-camera data processing in the smart home security scenario, but also enables users to quickly locate and playback the complete visiting events of specific visitors, greatly enhancing the real-time performance, accuracy and user experience of the security monitoring system. Description of the Drawings
[0060] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 is a schematic flowchart of a visitor processing method provided by an embodiment of this specification;
[0062] Figure 2 is a schematic diagram of a security monitoring interface provided by an embodiment of this specification;
[0063] Figure 3 is a schematic flowchart of a cross-end identification processing of visiting objects provided by an embodiment of this specification;
[0064] Figure 4 is a schematic flowchart of a processing based on a security processing large model provided by an embodiment of this specification;
[0065] Figure 5 is a schematic flowchart of a script generation processing provided by an embodiment of this specification;
[0066] Figure 6 It is a schematic flowchart for obtaining a video set provided by an embodiment of this specification;
[0067] Figure 7 It is a schematic structural diagram of a visitor processing device provided by an embodiment of this specification;
[0068] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of this specification;
[0069] Figure 9 It is a schematic structural diagram of an operating system and a user space provided by an embodiment of this specification;
[0070] Figure 10 is Figure 9 the architecture diagram of the Android operating system in
[0071] Figure 11 is Figure 9 the architecture diagram of the IOS operating system in Detailed implementation manners
[0072] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0073] In the description of this specification, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In the description of this specification, it should be noted that unless otherwise clearly specified and limited, "including" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood in specific situations. In addition, in the description of this specification, unless otherwise stated, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0074] The following will describe this specification in detail with specific embodiments.
[0075] In one embodiment, as Figure 1 shown, a visitor processing method is proposed. This method can be implemented depending on a computer program and can run on a visitor processing device based on the von Neumann architecture. The computer program can be integrated into an application or run as an independent tool-type application. The visitor processing device can be an electronic device with a visitor processing system, and this electronic device can be used in cooperation with multiple devices in the smart home field, including but not limited to: smart cameras, servers, smart doorbells, smart door locks, personal computers, tablets, handheld devices, in-vehicle devices, wearable devices, computing devices, or other processing devices connected to a wireless modem, etc.
[0076] Specifically, the visitor processing method includes:
[0077] S102: Collect multiple security monitoring end videos in a security scenario;
[0078] Security monitoring end: Refers to cameras or monitoring devices deployed in the visitor processing system. These devices are distributed at different locations (such as at the door, in the corridor, in the living room, etc.) and independently collect video data.
[0079] The electronic device records videos in real time through the security monitoring end or periodically captures video segments to form continuous or discrete security monitoring end videos, thereby obtaining multiple security monitoring end videos in the security scenario.
[0080] Exemplarily, in a home security system, the door camera, the indoor corridor camera, and the living room camera simultaneously collect security monitoring end videos. Each security monitoring end transmits these security monitoring end videos to an electronic device such as a local server, and through preprocessing, the videos are denoised and timestamp synchronized to ensure that when performing cross-camera analysis later, each video segment can be aligned on a unified timeline.
[0081] S104: Based on each of the security monitoring end videos, use a security processing large model to perform cross-end identification of the visiting object to obtain an object visiting scenario video set of at least one visiting object;
[0082] Security processing large model: Refers to a security processing large model obtained by adapting a basic large language model (LLM) to the security processing scenario using sample data in the security processing scenario. After scenario adaptation, the security processing large model can not only extract fine-grained features from videos but also perform semantic understanding and logical reasoning on metadata. The basic large language model (LLM) can be, for example, GPT series large models, Wenxin Yiyan large models, DeepSeek series large models;
[0083] Cross - end identification of visiting objects: It can be understood as, in multiple surveillance - end videos, through feature matching and spatio - temporal association, determining and associating the same visiting objects detected across cameras and extracting the object visiting scenario video set of the visiting objects.
[0084] Object visiting scenario video set: A complete video set generated by integrating the activity segments of the same visiting object at different surveillance points, which shows the entire itinerary and behavior trajectory of the visiting object.
[0085] Schematically, in each security surveillance - end video, a target detection mechanism is used to identify visiting objects, and a multi - object tracking (MOT) mechanism is used to assign local tracking IDs to them and extract the information corresponding to each security surveillance - end video, including but not limited to one or more fittings of visiting target detection information, visiting event information, visiting target cross - end tracking information, visiting target feature information, and visiting event behavior information. Then, this information will be used as visiting target detection meta - information. Subsequently, a security processing large model is used to perform cross - end identification of visiting objects based on the visiting target detection meta - information and the security surveillance - end videos. Then, the various visiting segments of the identified same visiting object are spliced together in spatio - temporal order to form a complete video set, so as to obtain the object visiting scenario video set of at least one visiting object.
[0086] S106: Determine the visiting object relationship label between the visiting object and the user object, and generate the visiting object events of each visiting object based on the visiting object relationship label and the object visiting scenario video set;
[0087] Visiting object relationship label: A classification label that identifies the relationship between a visiting object and a user (family member, property owner, etc.), such as "relative", "acquaintance", "neighbor", "stranger", etc.
[0088] Visiting object event: An event record that describes a specific visiting object in a security scenario. For example, a visiting object event can contain information such as a visiting object relationship label, time, location, behavior description, and the corresponding video segment set, etc.
[0089] Schematically, the visiting object can be classified by comparing face recognition results, historical records, user - preset information, and behavior patterns to determine the visiting object relationship label. Based on the generated object visiting scenario video set and relationship label, visiting object events are automatically constructed, recording detailed information including visiting time, activity trajectory, behavior description, and security level, etc., and associating these visiting object events with the corresponding visiting objects.
[0090] S108: Display each visiting object and the visiting object events corresponding to the visiting object on the security surveillance interface.
[0091] Security monitoring interface: It refers to the display terminal or APP interface of the user in the intelligent security system, which is used to display real-time or historical monitoring information, event records and related data. The security monitoring interface is displayed in the form of pictures and texts, video thumbnails, event lists, etc., presenting the visiting objects and their events to the user intuitively, so as to facilitate the user to quickly understand the security situation.
[0092] Exemplarily, such as Figure 2 shown, Figure 2 is a schematic diagram of a security monitoring interface. In Figure 2 , the guiding copywriting for displaying visitors is "What's new today? Visitors", which generally displays the visitor situation of the day to the user, and combines the events of all visiting objects to mark identity statistics such as "4 family members, 1 stranger". The visiting objects detected on the day are listed in sequence in the form of circular avatars (showing "xinxin", "baobao", "mama", etc. in the example), and are distinguished according to the recognition results or the tags preset by the user, such as "family members", "strangers", etc. In Figure 2 , the event list of the visiting objects is also displayed in the form of "event stream of visiting objects". Each event of a visiting object is displayed with an event thumbnail, time and event description. The user can click on the avatar of a visiting object or an event item of a visiting object to further view the video details or perform management operations.
[0093] In a feasible implementation manner, after the security monitoring interface displays each of the visiting objects and the visiting object events corresponding to the visiting objects, in response to an event selection operation for a target visiting object event of a target visiting object, an object visiting scenario video set is output.
[0094] Schematically, in the security monitoring interface, the user selects a target visiting object event corresponding to a certain target visiting object through an input event selection operation (such as a click or touch operation), and the generated object visiting scenario video set is output to the user terminal for playback and display.
[0095] Example illustration: Taking Figure 2 as an example, assume that the user sees the event that the visitor "xinxin" enters the house at 10:05 in the security monitoring interface. After clicking on this event, the system automatically searches for the corresponding object visiting scenario video set, and finally outputs a complete object visiting scenario video set for the user to view.
[0096] In the embodiments of this specification, by collecting multiple security monitoring end videos in a security scenario, based on each security monitoring end video, using a large security processing model to perform cross-end identification of visiting objects to obtain an object visiting scenario video set of at least one visiting object, determining a visiting object relationship label between the visiting object and the user object, generating a visiting object event for each visiting object based on the visiting object relationship label and the object visiting scenario video set, and displaying each visiting object and the corresponding visiting object event on the security monitoring interface, efficient, accurate, and story-based security monitoring data integration is achieved, avoiding the limitation of non-intelligent visitor processing in the related art security monitoring scenario. It not only improves the spatio-temporal continuity and intelligent level of multi-camera data processing, but also enables users to quickly locate and playback the complete visiting events of specific visitors, greatly enhancing the real-time performance, accuracy, and user experience of the security monitoring system.
[0097] Please refer to Figure 3 , Figure 3 FIG. is a schematic flow diagram of a cross-end identification process for visiting objects proposed in this specification. Specifically, to perform the step of using a large security processing model to perform cross-end identification of visiting objects based on each security monitoring end video to obtain an object visiting scenario video set of at least one visiting object, the following method can be referred to:
[0098] S1002: Perform visiting target detection processing on each security monitoring end video to obtain visiting target detection meta-information;
[0099] First, perform target detection processing on the security monitoring end videos collected by each security monitoring end. Specifically, by applying a deep learning target detection mechanism (such as YOLO, RetinaNet, or MTCNN, etc.), automatically identify the visiting targets appearing in the security monitoring end videos, and extract the detection meta-information of each target. The detection meta-information includes but is not limited to:
[0100] Detection box information: indicating the position and size of the visiting target in the video frame (such as the coordinates of the upper left corner and the lower right corner);
[0101] Detection confidence: describing the reliability of the detection result;
[0102] Timestamp: marking the specific time when the target is detected;
[0103] Camera identification and location information: indicating which security monitoring end the target comes from and its position in the scenario;
[0104] Key points or feature vectors: used for subsequent cross-camera target comparison and re-identification;
[0105] For example, when a person entering is detected in the video of the door camera, the system generates detection metadata containing the above information, laying a data foundation for subsequent cross-end identification.
[0106] In a feasible implementation manner, the detection and processing of the visiting target based on the videos of each security monitoring end to obtain the detection meta-information of the visiting target includes:
[0107] Input the videos of each security monitoring end into the visiting target detection model, and determine the visiting target detection information, visiting event information, cross-end tracking information of the visiting target, visiting target feature information, and visiting event behavior information through the visiting target detection model; generate the detection meta-information of the visiting target based on the visiting target detection information, visiting event information, cross-end tracking information of the visiting target, visiting target feature information, and visiting event behavior information.
[0108] The visiting target detection information may include a detection box / bounding box: the position and size of the target in the video frame (usually represented in coordinate form, such as the coordinates of the upper left corner and the lower right corner); detection confidence: the confidence or probability score of the model for this detection result, used to judge the reliability of the recognition; target category label: such as "pedestrian", "vehicle", "face", etc., indicating the category of the detected object.
[0109] The visiting event information may include event visiting event information, frame sequence numbers corresponding to the visiting event in the video stream, etc.;
[0110] The cross-end tracking information of the visiting target may include a tracking ID: a unique identifier assigned to each target in multi-object tracking (MOT) to connect the detection results of the same object in consecutive frames; trajectory data: the movement path of the target in the video, including coordinate points, movement directions, and speed information in consecutive frames.
[0111] The visiting target feature information may include: an embedding vector / feature vector: a high-dimensional feature representation extracted from the target (such as a face or a pedestrian), used for cross-camera comparison and re-identification. Key point information: such as face key points (the positions of eyes, nose, mouth) or human body pose key points, assisting in alignment, recognition, and behavior analysis.
[0112] The visiting event behavior information may include: a behavior label: a description of the target's behavior, such as "enter", "leave", "linger", etc.; event duration: records the start and end times of the event, which is helpful for event analysis and video editing; anomaly detection metrics: such as the target staying for a long time, abnormal movement, etc., which can be used as the basis for security warnings.
[0113] Schematically, the video of each security monitoring end is first input into the visiting target detection model. This model uses deep learning algorithms to comprehensively analyze the video data to obtain visiting target detection information, visiting event information, visiting target cross-end tracking information, visiting target feature information, and visiting event behavior information. Based on the visiting target detection information, visiting event information, visiting target cross-end tracking information, visiting target feature information, and visiting event behavior information, visiting target detection meta-information is generated.
[0114] The target detection model is a target detection visual model trained based on a machine learning model, which is used to automatically identify and locate target objects in images or video frames. It can not only determine whether a specific object exists in the image, but also determine the position of these objects in the image (usually represented in the form of a bounding box) and their class labels. The target detection model learns to extract features from images by training a large amount of data with annotations (such as object categories and bounding boxes). When processing a new image, the target detection model performs feature extraction and multi-scale detection on the image to locate the target and output the bounding box, confidence score, and class information.
[0115] In security monitoring, the target detection model is used to automatically identify visiting targets in videos, such as detecting people entering the monitoring area. It can quickly analyze each frame of the video and output the position information and behavior annotations of the target, providing basic data for subsequent target tracking, cross-end identification, and behavior analysis. In the above way, the target detection model realizes the automatic detection and positioning of visiting objects in security videos, enabling the entire monitoring system to have the ability to capture abnormal or important behaviors in real time, efficiently, and accurately, thus providing solid data support for further security management and event handling.
[0116] S1004: Use the security processing large model to perform cross-end identification of visiting objects based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object, and extract the object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos.
[0117] Schematically, a cross-end identification prompt word for the security monitoring end video and the visiting target parsing meta-information is pre-constructed. The cross-end identification prompt word, the security monitoring end video, and the visiting target parsing meta-information are jointly input into the security processing large model. Based on the security processing large model, combined with the detection meta-information generated in the previous step, cross-camera visiting object identification is performed. The security processing large model performs step-by-step processing:
[0118] Metadata parsing and matching step: The security processing large model first semantically parses the visiting target metadata detected by different monitoring ends, and converts information such as timestamps, camera positions, detection frames, and feature vectors into natural language descriptions;
[0119] Cross - end target association step: Based on the parsed natural language description, the security processing large model compares the target features and spatio - temporal information through the visited target metadata, matches the detection data from different security monitoring ends, and determines whether it is the same visiting object. For example, if the targets detected by the door camera and the living room camera are highly similar in feature vectors and are closely connected in time, the security processing large model will determine that they belong to the same object;
[0120] Video clip extraction and splicing step: After the security processing large model determines the visiting object, it will extract the activity clips corresponding to the same visiting object from the videos of each monitoring end. According to the time sequence and scene information of the identified target (object) appearing in different monitoring areas, the video clips are spliced into a complete "video set of the object's visiting scene", showing the complete trajectory of the visiting object from entry to departure.
[0121] For example, if the same visitor is detected by the door camera at 10:05 and appears in the corridor camera at 10:07, the system will automatically extract and splice the video clips of these two time periods to form a coherent video set, intuitively showing the visiting process of the visitor.
[0122] In this specification, by extracting detailed detection metadata through S1002 and then combining the cross - end association and video clip extraction processing of the large model in S1004, the entire solution can achieve efficient identification and integration of visiting targets in multi - camera video data, and then output a complete and coherent video set of the object's visiting scene, providing a more intuitive and comprehensive display of visitor behavior for security monitoring.
[0123] Optionally, please refer to Figure 4 , Figure 4 is a schematic flowchart of a process based on the security processing large model. Specifically, to execute the cross - end identification of visiting objects by using the security processing large model based on the videos of the multiple security monitoring ends and the parsed meta - information of the visiting target to determine at least one visiting object, and extract the video set of the object's visiting scene corresponding to the visiting object from the videos of the multiple security monitoring ends, the following method can be referred to:
[0124] S202: Use the security processing large model to perform cross - end identification of visiting objects based on the videos of the multiple security monitoring ends and the parsed meta - information of the visiting target to determine at least one visiting object;
[0125] Schematically, the video data of multiple monitoring terminals and the pre-generated parsed meta-information of visiting targets are input into the security processing large model, and then cross-terminal matching is performed. The security processing large model uses meta-information of visiting targets such as target features (such as face or human feature vectors), timestamps, and spatial location information to perform multi-modal visiting object detection and inference on videos from multiple security monitoring terminals to determine whether the detection record data from different monitoring terminals belongs to the same visiting object. Then, based on the detection and inference results, the security processing large model assigns a unique identifier to each visiting object target to achieve cross-camera association, ensuring that the data of the same visiting object in different video sources can be merged together.
[0126] S204: Use the security processing large model to determine the target motion description text and target motion description key frames for the visiting object, and perform multi-modal fusion based on the target motion description text and the target motion description key frames to obtain potential visiting object and target motion description event information;
[0127] Target motion description text: Use the security processing large model to convert the spatio-temporal trajectories and behavior data generated during the detection and tracking process into natural language description text, expressing the motion path, speed, transition situation, etc. of the visiting object.
[0128] Target motion description key frames: Representative key frames extracted from the video, and these frame images can intuitively display the posture and scene information of the visiting object at key time points.
[0129] Multi-modal fusion: Combine the text description and image information to generate richer and more accurate motion description event information.
[0130] Schematically, the security processing large model parses the cross-terminal tracking information, converts data such as the motion path of the detection target, the time nodes of entering and leaving each monitoring area, and the motion speed into natural language description to obtain the target motion description text. Then, representative image frames are automatically extracted from the video data at each key time point as the target motion description key frames, and then the target motion description text and the target motion description key frames are subjected to multi-modal fusion to obtain potential visiting object and target motion description event information;
[0131] S206: Use the security processing large model to generate a visiting video set editing script for the videos of the multiple security monitoring terminals based on the target motion description event information;
[0132] Visiting video set editing script: Includes video editing time (start and end times of each video segment), transition modes of visiting objects (such as fade-in and fade-out, switching effects), and target motion description subtitles, etc., to guide the subsequent automatic extraction and splicing of video segments.
[0133] Target motion description event information: The comprehensive information integrating text and images obtained in the previous stage provides basic data and logical basis for the generation of the editing script.
[0134] Schematically, use the security processing large model to parse the target motion description event information, identify the appearance time, position and motion trajectory of the visiting object in different cameras, and then automatically determine the start and end times of each video segment for editing based on the time nodes of the target appearance and disappearance to ensure natural connection of video paragraphs. And according to the motion of the visiting object target between different monitoring terminals, configure the corresponding transition mode, such as using a smooth transition effect when the target crosses the camera switch. Embed the motion description text as subtitle information into the editing script for displaying the key descriptions of the target behavior after the video splicing. Finally, perform script output: Finally, generate a complete editing script, which details the time period, transition effect and corresponding subtitle description of each video segment in the script, providing operation instructions for subsequent video extraction and splicing.
[0135] In this specification, by executing S202, use the security processing large model and parsing meta-information to perform cross-terminal identification on the video data from multiple monitoring terminals to achieve unified identification of the same visiting object. Then execute S204 to generate the target motion description text and key frames, and after multi-modal fusion, form detailed target motion description event information, providing semantic and visual basis for subsequent video processing. Then execute S206 to generate a visiting video set editing script containing editing time, transition mode and subtitle description based on the target motion description event information, thereby guiding the system to automatically extract and splice the video segments of each monitoring terminal, output a complete visiting scene video set, making the integration and display of cross-camera video data automated and intelligent, and providing users with continuous and coherent playback of visiting behaviors.
[0136] In a feasible implementation manner, please refer to Figure 5 , Figure 5 is a schematic flowchart of a script generation process. To execute the generation of a visiting video set editing script for the multiple security monitoring terminal videos based on the target motion description event information using the security processing large model, the following method can be adopted:
[0137] S302: Use the security processing large model to perform key editing segment parsing processing on the multiple security monitoring terminal videos based on the target motion description event information to obtain the visiting video set editing time corresponding to the visiting object;
[0138] Schematically, through the security processing large model, using the previously generated target motion description event information, the input multiple security monitoring end videos are analyzed to automatically identify the key video segments related to the target visiting object. The model determines the key entry and transition moments of the visiting object in each monitoring video by analyzing the time stamps, motion trajectories, and behavior patterns, so as to obtain the start and end times of each key clip segment, that is, the "clip time of the visiting video set".
[0139] For example, if the visitor appears in the door camera from 10:05 to 10:07, in the corridor camera from 10:07 to 10:09, and in the living room camera from 10:09 to 10:12, then S302 generates the corresponding time period information as the key clip segment.
[0140] S304: Determine the transition mode of the visiting object and the caption description of the visiting story for the multiple key clip segments of the visiting object;
[0141] Schematically, after obtaining the key clip segment time information, the large model further performs semantic parsing on these segments to determine the transition mode between different video segments of the visiting object and generate the corresponding story description captions. The transition mode of the visiting object can be that the model selects a suitable visual transition effect (such as fade-in / fade-out, slide switch, etc.) according to the motion changes and scene switching of the target among the monitoring end videos to ensure natural connection of the video splicing. The caption description of the visiting story can be that the large model automatically generates the caption description of each segment according to the target motion description text, such as "Enter the door", "Pass through the corridor", "Arrive at the living room", providing a text description for the subsequent video display to help users understand the overall visiting plot.
[0142] S306: Generate a clip script for the visiting video set of the multiple security monitoring end videos based on the clip time of the visiting video set, the transition mode of the visiting object, and the caption description of the visiting story.
[0143] According to the clip time information obtained in S302 and the transition mode and story description captions determined in S304, the large model synthesizes various information to generate a complete clip script for the visiting video set. This clip script details the start and end times of each key video segment, the corresponding transition effects, and the caption descriptions to be superimposed, forming an operation instruction for guiding the subsequent video segment extraction and splicing process.
[0144] For example, the clip script may contain instructions: "From 10:05 to 10:07 (door camera), fade-in transition, caption 'Enter the door'; From 10:07 to 10:09 (corridor camera), slide switch, caption 'Pass through the corridor'; From 10:09 to 10:12 (living room camera), fade-out transition, caption 'Arrive at the living room'."
[0145] S308: Extract the object visit scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the editing script of the visiting video set.
[0146] Finally, based on the generated editing script of the visiting video set, automatically extract the corresponding key segments from each security monitoring end video, and splice the videos according to the preset time, transition, and subtitle requirements in the editing script to form a complete target visit scenario video set. This process ensures that the video segments obtained from different cameras are seamlessly connected in time sequence and naturally transition visually, while embedding descriptive subtitles, enabling the entire video to present a coherent storyline and intuitively showing the full behavior trajectory of the target visiting object.
[0147] In the embodiments of this specification, through the above steps, not only the automatic parsing of the target visiting behavior and the extraction of key segments in the multi-camera video data are realized, but also a detailed and semantic-rich editing script is generated using the security processing large model, providing accurate guidance for subsequent video splicing, thereby greatly improving the intelligence and user experience of the security monitoring system in event playback and security analysis.
[0148] In a feasible implementation manner, please refer to Figure 6 , Figure 6 is a schematic flow diagram of video set acquisition. To specifically execute the extraction of the object visit scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the editing script of the visiting video set, the following method can be referred to:
[0149] S402: Determine multiple key editing segments for the visiting object from the multiple security monitoring end videos based on the editing time of the visiting video set in the editing script of the visiting video set;
[0150] Schematically, the electronic device scans each security monitoring end video according to the "editing time" information (i.e., the start and end times of each key video segment) predefined in the editing script, and uses these timestamp information to accurately locate the key editing segments related to the target visiting object in each monitoring video, thereby ensuring that the extracted segments cover the entire process of the target visit.
[0151] For example, it is stipulated in the editing script that "10:05 - 10:07" is the key segment of the door camera, "10:07 - 10:09" is the key segment of the corridor camera, and "10:09 - 10:12" is the key segment of the living room camera. By sequentially reading the recordings of each camera and extracting the corresponding video segments according to the time periods defined in the editing script, multiple key editing segments for the same visiting object are formed.
[0152] S404: Determine the transition display effect configuration between each of the key clip segments according to the visitor transition mode in the visiting video set editing script, and generate a video segment splicing sequence based on the transition display effect configuration and the multiple key clip segments;
[0153] The main purpose of step S404 is to ensure a visually natural and smooth transition when splicing each key clip segment. The "transition mode" is preset in the editing script and describes how to switch between different monitoring end video segments, such as using fade-in / fade-out, sliding switch, or other dynamic effects.
[0154] Schematically, analyze the transition mode configuration in the editing script and generate transition effect instructions between every two consecutive key clip segments. For example, use a "fade-out - fade-in" effect between the door video segment and the corridor video segment, and a "sliding switch" between the corridor video segment and the living room video segment. Based on these configurations, the system constructs a video segment splicing sequence, which includes not only the order of each key segment but also the corresponding transition effect parameters, ensuring that the final output video is visually continuous and natural.
[0155] S406: Add story description subtitles to the video segment splicing sequence based on the visiting story description subtitles in the visiting video set editing script to generate an object visiting scenario video set corresponding to the visiting object.
[0156] In the generated spliced video sequence, the editing script also contains "visiting story description subtitles", which are used to describe the key behaviors or plots during the target visit process and play a role in explaining and guiding viewers to understand the video content. This step of S406 instructs to add text descriptions to the video splicing sequence to form the final story-based display effect.
[0157] Schematically, first analyze the subtitle information recorded in the editing script, such as "enter the door", "pass through the corridor", "arrive at the living room", etc. Then, in the generated video splicing sequence, overlay the corresponding subtitle information on each key segment or transition node, and set appropriate display positions, fonts, and durations. Finally, the spliced video sequence after subtitle overlay constitutes a complete object visiting scenario video set with a story description effect, and users can intuitively understand the whole process of the visiting object from entry to departure and the key plots of each scene transition when watching.
[0158] In this specification, through the implementation steps of S402 to S406, the system can automatically extract and accurately locate the key video segments of the target visit from multiple security monitoring terminal videos according to the pre-generated visit video set editing script; realize the natural switching between video segments according to the preset transition mode; and finally generate a complete video set that storytells and coherently displays the whole process of the target visit behavior through subtitle overlay, so as to provide users with an intuitive, complete and easy-to-understand replay of the visit scenario.
[0159] The following will be combined with Figure 7 to introduce the visitor processing device provided in the embodiments of this specification in detail. It should be noted that Figure 7 The visitor processing device shown is used to execute the method of the embodiments of this specification Figures 1 to 6 shown. For the sake of convenience of description, only the parts related to the embodiments of this specification are shown. For the specific technical details not disclosed, please refer to the embodiments Figures 1 to 6 shown in this specification.
[0160] Please refer to Figure 7 , which shows the structural schematic diagram of the visitor processing device of the embodiments of this specification. The visitor processing device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the visitor processing device 1 includes a video acquisition module 11, a video recognition module 12, and an event processing module 13, which are specifically used for:
[0161] The video acquisition module 11 is used to acquire multiple security monitoring terminal videos in a security scenario;
[0162] The video recognition module 12 is used to perform cross-terminal recognition of the visiting object based on each of the security monitoring terminal videos by using a security processing large model to obtain an object visiting scenario video set of at least one visiting object;
[0163] The event processing module 13 is used to determine the visiting object relationship label between the visiting object and the user object, and generate a visiting object event for each of the visiting objects based on the visiting object relationship label and the object visiting scenario video set;
[0164] The event processing module 13 is used to display each of the visiting objects and the visiting object events corresponding to the visiting objects on the security monitoring interface.
[0165] In a feasible implementation manner, the performing cross-terminal recognition of the visiting object based on each of the security monitoring terminal videos by using a security processing large model to obtain an object visiting scenario video set of at least one visiting object includes:
[0166] Performing visiting target detection processing on each of the security monitoring terminal videos to obtain visiting target detection meta-information;
[0167] Using a security processing large model, based on the multiple security monitoring end videos and the visitor target parsing meta-information, perform cross-end identification of the visitor object to determine at least one visitor object, and extract from the multiple security monitoring end videos an object visit scenario video set corresponding to the visitor object.
[0168] In a feasible implementation manner, the using a security processing large model to perform cross-end identification of the visitor object based on the multiple security monitoring end videos and the visitor target parsing meta-information to determine at least one visitor object, and extract from the multiple security monitoring end videos an object visit scenario video set corresponding to the visitor object includes:
[0169] Using a security processing large model to perform cross-end identification of the visitor object based on the multiple security monitoring end videos and the visitor target parsing meta-information to determine at least one visitor object;
[0170] Using a security processing large model to determine a target motion description text and target motion description key frames for the visitor object, and perform multi-modal fusion based on the target motion description text and the target motion description key frames to obtain a potential visitor object to obtain target motion description event information;
[0171] Using a security processing large model to generate a visit video set editing script for the multiple security monitoring end videos based on the target motion description event information;
[0172] Based on the visit video set editing script, extract from the multiple security monitoring end videos an object visit scenario video set corresponding to the visitor object.
[0173] In a feasible implementation manner, the using a security processing large model to generate a visit video set editing script for the multiple security monitoring end videos based on the target motion description event information includes:
[0174] Using a security processing large model to perform key clip segment parsing processing on the multiple security monitoring end videos based on the target motion description event information to obtain the visit video set editing time corresponding to the visitor object;
[0175] Determine a visitor object transition mode and a visitor story description subtitle for multiple key clip segments of the visitor object;
[0176] Based on the visit video set editing time, the visitor object transition mode, and the visitor story description subtitle, generate a visit video set editing script for the multiple security monitoring end videos.
[0177] In a feasible implementation manner, extracting the object visit scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the visiting video set editing script includes:
[0178] Determining multiple key editing segments for the visiting object from the multiple security monitoring end videos based on the visiting video set editing time in the visiting video set editing script;
[0179] Determining the transition switching display effect configuration between each of the key editing segments according to the visiting object transition mode in the visiting video set editing script, and generating a video segment splicing sequence based on the transition switching display effect configuration and the multiple key editing segments;
[0180] Adding story description subtitles to the video segment splicing sequence based on the visiting story description subtitles in the visiting video set editing script to generate the object visit scenario video set corresponding to the visiting object.
[0181] In a feasible implementation manner, the obtaining the visiting target detection meta information by performing visiting target detection processing based on each of the security monitoring end videos includes:
[0182] Inputting each of the security monitoring end videos into a visiting target detection model, and determining visiting target detection information, visiting event information, visiting target cross-end tracking information, visiting target feature information, and visiting event behavior information through the visiting target detection model;
[0183] Generating visiting target detection meta information based on the visiting target detection information, visiting event information, visiting target cross-end tracking information, visiting target feature information, and visiting event behavior information.
[0184] In a feasible implementation manner, after displaying each of the visiting objects and the visiting object events corresponding to the visiting objects on the security monitoring interface, it further includes:
[0185] In response to an event selection operation for the target visiting object event of the target visiting object, outputting the object visit scenario video set.
[0186] It should be noted that when the visitor processing device provided in the above embodiment executes the visitor processing method, only the above division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the visitor processing device provided in the above embodiment and the visitor processing method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.
[0187] The serial numbers of the embodiments in the present specification are only for description and do not represent the superiority or inferiority of the embodiments.
[0188] The embodiments of the present specification further provide a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the visitor processing method in the embodiments as described above. Figures 1 to 6 The specific execution process can be referred to Figures 1 to 6 the specific description of the embodiments shown, and will not be elaborated here.
[0189] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to perform the visitor processing method in the embodiments as described above. Figures 1 to 6 The specific execution process can be referred to Figures 1 to 6 the specific description of the embodiments shown, and will not be elaborated here.
[0190] Please refer to Figure 8 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of the present specification. The electronic device in the present specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected through the bus 150.
[0191] The processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect various parts within the entire electronic device, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120, it performs various functions of the electronic device and processes data. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used for processing wireless communication. It can be understood that the above modem may not be integrated into the processor 110 and may be implemented separately through a communication chip.
[0192] The memory 120 may include a random access memory (RAM) and may also include a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system developed based on the Android system in depth, an IOS system developed by Apple Inc., including a system developed based on the IOS system in depth, or other systems. The data storage area may also store data created during the use of the electronic device, such as a phone book, audio and video data, chat record data, etc.
[0193] See Figure 9 As shown, the memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party application programs run in the user space. To ensure that different third-party application programs can all achieve good running effects, the operating system allocates corresponding system resources for different third-party application programs. However, there are also differences in the system resource requirements of different application scenarios in the same third-party application program. For example, in the local resource loading scenario, the third-party application program has a higher requirement for the disk read speed; in the animation rendering scenario, the third-party application program has a higher requirement for the GPU performance. The operating system and the third-party application program are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application program, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application program.
[0194] In order to enable the operating system to distinguish the specific application scenarios of third-party application programs, it is necessary to establish data communication between the third-party application programs and the operating system, so that the operating system can obtain the current scenario information of the third-party application programs at any time, and then perform targeted system resource adaptation based on the current scenario.
[0195] Taking the Android system as an example of the operating system, the programs and data stored in the memory 120 are as Figure 10As shown, the memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380. The Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space. The Linux kernel layer 320 provides underlying drivers for various hardware components of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 340 provides major feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library (Android runtime), which mainly provides some core libraries that allow developers to write Android applications using the Java language. The application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider management, package management, call management, resource management, and location management. The application layer 380 runs at least one application. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera applications, or third-party applications developed by third-party developers, such as games, instant messaging programs, and photo enhancement programs.
[0196] Taking the operating system as the IOS system as an example, the programs and data stored in the memory 120 are as follows: Figure 11As shown in the figure, the IOS system includes: the Core OS layer 420, the Core Services layer 440, the Media layer 460, and the Cocoa Touch Layer 480. The Core OS layer 420 includes the operating system kernel, drivers, and underlying program frameworks, which provide functions closer to the hardware for use by the program frameworks in the Core Services layer 440. The Core Services layer 440 provides the system services and / or program frameworks required by applications, such as the Foundation framework, the Accounts framework, the Advertising framework, the Data Storage framework, the Network Connectivity framework, the Location framework, the Motion framework, and so on. The Media layer 460 provides interfaces related to audio and video for applications, such as interfaces related to graphics and images, audio technology, video technology, and the AirPlay interface for audio and video transmission technology. The Cocoa Touch Layer 480 provides various common interface-related frameworks for application development and is responsible for the touch interaction operations of users on electronic devices. For example, the Local Notification service, the Remote Push service, the Advertising framework, the Game Tools framework, the Message User Interface (UI) framework, the UIKit framework for the user interface, the Map framework, and so on.
[0197] In Figure 11 Among the frameworks shown, the frameworks related to most applications include, but are not limited to: the Foundation framework in the Core Services layer 440 and the UIKit framework in the Cocoa Touch Layer 480. The Foundation framework provides many basic object classes and data types, provides the most basic system services for all applications, and is independent of the UI. The classes provided by the UIKit framework are the basic UI class libraries, used to create touch-based user interfaces. iOS applications can provide the UI based on the UIKit framework, so it provides the infrastructure for applications to build user interfaces, draw, process, and handle user interaction events, respond to gestures, and so on.
[0198] Among them, the methods and principles for implementing data communication between third-party applications and the operating system in the IOS system can refer to the Android system, and will not be elaborated in this specification.
[0199] Among them, the input device 130 is used to receive input instructions or data. The input device 130 includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is used to output instructions or data. The output device 140 includes, but is not limited to, a display device, a speaker, etc. In one example, the input device 130 and the output device 140 can be integrated. The input device 130 and the output device 140 are a touch display screen, which is used to receive touch operations of a user using any suitable object such as a finger or a stylus on or near it, and to display the user interfaces of various applications. The touch display screen is usually set on the front panel of the electronic device. The touch display screen can be designed as a full-screen, a curved screen, or a special-shaped screen. The touch display screen can also be designed as a combination of a full-screen and a curved screen, or a combination of a special-shaped screen and a curved screen. The embodiments of this specification do not limit this.
[0200] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above drawings does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the drawings, or combine some components, or have different component arrangements. For example, the electronic device also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WiFi) module, a power supply, a Bluetooth module, etc., which will not be elaborated here.
[0201] In the embodiments of this specification, the execution subject of each step can be the electronic device introduced above. Optionally, the execution subject of each step is the operating system of the electronic device. The operating system can be the Android system, the IOS system, or other operating systems. The embodiments of this specification do not limit this.
[0202] The electronic device according to the embodiment of this specification may also be equipped with a display device, which can be various devices capable of realizing the display function, such as: cathode ray tube display (abbreviated as CR), light-emitting diode display (abbreviated as LED), electronic ink screen, liquid crystal display (abbreviated as LCD), plasma display panel (abbreviated as PDP), etc. Users can use the display device on the electronic device to view information such as text, images, and videos displayed. The electronic device may be a smart phone, a tablet computer, a game device, an AR (Augmented Reality) device, an automobile, a data storage device, an audio playback device, a video playback device, a notebook, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, an electronic clothing, etc.
[0203] In Figure 8 the electronic device shown, the processor 110 may be used to call the application programs stored in the memory 120 and specifically perform the following operations:
[0204] Collect multiple security monitoring end videos in a security scenario;
[0205] Based on each of the security monitoring end videos, use a security processing large model to perform cross-end identification of visiting objects to obtain an object visiting scenario video set of at least one visiting object;
[0206] Determine the visiting object relationship label between the visiting object and the user object, and generate a visiting object event for each of the visiting objects based on the visiting object relationship label and the object visiting scenario video set;
[0207] Display each of the visiting objects and the visiting object events corresponding to the visiting objects on the security monitoring interface.
[0208] In a feasible implementation manner, the step of using a security processing large model to perform cross-end identification of visiting objects based on each of the security monitoring end videos to obtain an object visiting scenario video set of at least one visiting object includes:
[0209] Perform visiting target detection processing on each of the security monitoring end videos to obtain visiting target detection meta-information;
[0210] Use the security processing large model to perform cross-terminal identification of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object, and extract the object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos.
[0211] In a feasible implementation manner, the step of using the security processing large model to perform cross-terminal identification of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object, and extracting the object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos includes:
[0212] Use the security processing large model to perform cross-terminal identification of the visiting object based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object;
[0213] Use the security processing large model to determine the target motion description text and the target motion description key frames for the visiting object, and perform multi-modal fusion based on the target motion description text and the target motion description key frames to obtain the potential visiting object and obtain the target motion description event information;
[0214] Use the security processing large model to generate a visiting video set editing script for the multiple security monitoring end videos based on the target motion description event information;
[0215] Extract the object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the visiting video set editing script.
[0216] In a feasible implementation manner, the step of using the security processing large model to generate a visiting video set editing script for the multiple security monitoring end videos based on the target motion description event information includes:
[0217] Use the security processing large model to perform key clip segment parsing processing on the multiple security monitoring end videos based on the target motion description event information to obtain the visiting video set editing time corresponding to the visiting object;
[0218] Determine the visiting object transition mode and the visiting story description subtitles for multiple key clip segments of the visiting object;
[0219] Generate a visiting video set editing script for the multiple security monitoring end videos based on the visiting video set editing time, the visiting object transition mode, and the visiting story description subtitles.
[0220] In a feasible implementation manner, the step of extracting the object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the visiting video set editing script includes:
[0221] Determine a plurality of key clip segments for the visiting object from the plurality of security monitoring terminal videos based on the visiting video set editing time in the visiting video set editing script of the visiting video set;
[0222] Determine the transition switching display effect configuration between each of the key clip segments according to the visiting object transition mode in the visiting video set editing script, and generate a video segment splicing sequence based on the transition switching display effect configuration and the plurality of key clip segments;
[0223] Add story description subtitles to the video segment splicing sequence based on the visiting story description subtitles in the visiting video set editing script to generate an object visiting scenario video set corresponding to the visiting object.
[0224] In a feasible implementation manner, the obtaining of the visiting target detection meta-information by performing visiting target detection processing based on each of the security monitoring terminal videos includes:
[0225] Input each of the security monitoring terminal videos into a visiting target detection model, and determine visiting target detection information, visiting event information, visiting target cross-terminal tracking information, visiting target feature information, and visiting event behavior information through the visiting target detection model;
[0226] Generate visiting target detection meta-information based on the visiting target detection information, visiting event information, visiting target cross-terminal tracking information, visiting target feature information, and visiting event behavior information.
[0227] In a feasible implementation manner, after displaying each of the visiting objects and the visiting object events corresponding to the visiting objects on the security monitoring interface, it further includes:
[0228] In response to an event selection operation on the target visiting object event of the target visiting object, output an object visiting scenario video set.
[0229] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory, or a random access memory, etc.
[0230] The above-disclosed are only the preferred embodiments of this specification. Of course, the scope of the rights of this specification cannot be limited by this. Therefore, equivalent changes made according to the claims of this specification still fall within the scope covered by this specification.
Claims
1. A visitor processing method, characterized in that, The method includes: Collecting multiple security monitoring end videos in a security scenario; Based on each of the security monitoring end videos, using a security processing large model to perform cross-end identification of visiting objects to obtain an object visiting scenario video set of at least one visiting object; Determining a visiting object relationship label between the visiting object and the user object, and generating a visiting object event for each of the visiting objects based on the visiting object relationship label and the object visiting scenario video set; Displaying each of the visiting objects and the visiting object events corresponding to the visiting objects on a security monitoring interface.
2. The method according to claim 1, wherein The step of using a security processing large model to perform cross-end identification of visiting objects based on each of the security monitoring end videos to obtain an object visiting scenario video set of at least one visiting object includes: Performing visiting target detection processing on each of the security monitoring end videos to obtain visiting target detection meta-information; Using a security processing large model to perform cross-end identification of visiting objects based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object, and extracting an object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos.
3. The method according to claim 2, wherein The step of using a security processing large model to perform cross-end identification of visiting objects based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object, and extracting an object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos includes: Using a security processing large model to perform cross-end identification of visiting objects based on the multiple security monitoring end videos and the visiting target parsing meta-information to determine at least one visiting object; Using a security processing large model to determine a target motion description text and target motion description key frames for the visiting object, and performing multi-modal fusion based on the target motion description text and the target motion description key frames to obtain a potential visiting object and obtain target motion description event information; Using a security processing large model to generate a visiting video set clip script for the multiple security monitoring end videos based on the target motion description event information; Extracting an object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the visiting video set clip script.
4. The method according to claim 3, wherein The step of using a security processing large model to generate a visiting video set clip script for the multiple security monitoring end videos based on the target motion description event information includes: Using a security processing large model to perform key clip segment parsing processing on the multiple security monitoring end videos based on the target motion description event information to obtain the visiting video set clip time corresponding to the visiting object; Determining a visiting object transition mode and a visiting story description subtitle for multiple key clip segments of the visiting object; Generating a visiting video set clip script for the multiple security monitoring end videos based on the visiting video set clip time, the visiting object transition mode, and the visiting story description subtitle.
5. The method according to claim 3, characterized in that, The step of extracting an object visiting scenario video set corresponding to the visiting object from the multiple security monitoring end videos based on the visiting video set clip script includes: Determine a plurality of key clip segments for the visiting object from the plurality of security monitoring end videos based on the visiting video set clip time in the visiting video set clip script; Determine the transition switching display effect configuration between each of the key clip segments according to the visiting object transition mode in the visiting video set clip script, and generate a video segment splicing sequence based on the transition switching display effect configuration and the plurality of key clip segments; Add story description subtitles to the video segment splicing sequence based on the visiting story description subtitles in the visiting video set clip script to generate an object visiting scenario video set corresponding to the visiting object.
6. The method according to claim 2, wherein The visiting target detection processing based on each of the security monitoring end videos to obtain visiting target detection meta-information includes: Input each of the security monitoring end videos into a visiting target detection model, and determine visiting target detection information, visiting event information, visiting target cross-end tracking information, visiting target feature information, and visiting event behavior information through the visiting target detection model; Generate visiting target detection meta-information based on the visiting target detection information, visiting event information, visiting target cross-end tracking information, visiting target feature information, and visiting event behavior information.
7. The method according to claim 1, wherein After displaying each of the visiting objects and the visiting object events corresponding to the visiting objects on the security monitoring interface, it further includes: In response to an event selection operation for a target visiting object event of a target visiting object, output an object visiting scenario video set.
8. A visitor processing device, characterized in that, The device includes: A video acquisition module for acquiring a plurality of security monitoring end videos in a security scenario; A video recognition module for performing cross-end recognition of visiting objects using a security processing large model based on each of the security monitoring end videos to obtain an object visiting scenario video set of at least one visiting object; An event processing module for determining a visiting object relationship label between the visiting object and the user object, and generating a visiting object event for each of the visiting objects based on the visiting object relationship label and the object visiting scenario video set; The event processing module for displaying each of the visiting objects and the visiting object events corresponding to the visiting objects on the security monitoring interface.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded and executed by a processor to perform the method steps of any one of claims 1 to 7.
10. An electronic device, characterized in that, Including: A processor and a memory; wherein, the memory stores a computer program, and the computer program is suitable for being loaded and executed by the processor to perform the method steps of any one of claims 1 to 7.