Method and device for determining early warning strategy of target area, equipment and medium

By using deep learning models to generate semantically labeled images in the physical space of financial institutions, and identifying and segmenting sub-target regions, the accuracy and reliability issues of traditional security systems are solved, achieving a more intelligent and efficient security effect.

CN120977069APending Publication Date: 2025-11-18AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511066422.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing technologies, the physical security systems of financial institutions rely on traditional monitoring equipment and manual patrols, which suffer from problems such as low security accuracy, poor reliability, high cost, poor real-time performance, and subjective judgment errors.

Method used

By acquiring video frames of the target area and using a deep learning model to generate semantically labeled images, stationary objects, potentially moving objects, and movable objects are identified, sub-target areas are divided, and corresponding early warning strategies are determined.

Benefits of technology

It has achieved smarter and more efficient security, improved the accuracy and reliability of regional early warning, reduced costs and improved real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977069A_ABST
    Figure CN120977069A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for determining a target area early warning strategy, equipment and a medium. The method comprises the following steps: acquiring a plurality of video frames associated with a target area; for any two continuous video frames, inputting the video frames into the first model, and outputting a first semantic tag image corresponding to the first video frame and a second semantic tag image corresponding to the second video frame; according to at least one first feature point in the first video frame and the first semantic tag image, determining a first semantic category tag corresponding to the first feature point, and further determining a second semantic category tag corresponding to the second feature point; and determining at least one sub-target area of the target area according to the first semantic category label and the second semantic category label, and determining an early warning strategy of the sub-target area according to the grade of the sub-target area. According to the technical scheme provided by the embodiment of the invention, the plurality of sub-target areas and the early warning strategy are determined, so that security and protection are more intelligent and more efficient, and the accuracy and reliability of area early warning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of deep learning, and in particular to a method and device for determining a target area early warning strategy, equipment and a medium. BACKGROUND

[0002] With the development of financial technology, financial institutions not only face information risks such as network attacks, but also have to deal with security risks in physical space. Security risks in physical space not only can cause property loss, but also can pose a threat to the personal safety of financial institution staff and customers.

[0003] Currently, when dealing with security risks in physical space, the security system of a financial institution mainly relies on traditional monitoring equipment and manual patrol to manage security. However, traditional video monitoring equipment mainly relies on fixed cameras to obtain pictures, and has the problem of low accuracy and reliability of security. The manual patrol security mode has the problems of high cost, poor real-time performance, subjective judgment errors, and the like, and is difficult to meet the growing security needs of financial institutions. SUMMARY

[0004] Embodiments of the present disclosure provide a method, device, equipment and medium for determining a target area early warning strategy, to determine a plurality of sub-target areas and an early warning strategy, so that security is more intelligent and efficient, and the accuracy and reliability of the target area early warning effect are obtained.

[0005] In a first aspect, embodiments of the present disclosure provide a method for determining a target area early warning strategy, the method comprising:

[0006] obtaining a plurality of video frames associated with a target area; wherein the video frames are obtained by a device configured with a visual sensor when the device moves in the target area; the visual sensor collects the video frames at a preset frame rate;

[0007] For any two consecutive video frames, input the consecutive video frames into a first model to output a first semantic label image corresponding to the first video frame and a second semantic label image corresponding to the second video frame; wherein the pixel value of each pixel in the first semantic label image and the second semantic label image represents the semantic class label of the pixel; the semantic class label is divided into three categories: static object class, potential moving object class and movable object class;

[0008] According to at least one first feature point in the first video frame and the first semantic label image, determine the first semantic class label corresponding to the first feature point, and according to at least one second feature point in the second video frame and the second semantic label image, determine the second semantic class label corresponding to the second feature point;

[0009] determine at least one sub-target region of the target region according to the first semantic category label and the second semantic category label, and determine a pre-warning strategy of the sub-target region according to a level of the sub-target region.

[0010] In a second aspect, an apparatus for determining a pre-warning strategy of a target region is provided. The apparatus includes:

[0011] a video frame acquisition module configured to acquire a plurality of video frames associated with the target region, wherein the video frames are acquired by a device configured with a visual sensor when the device moves in the target region, and the visual sensor acquires the video frames at a preset frame rate;

[0012] a semantic label image output module configured to, for any two consecutive video frames, input the consecutive video frames into a first model, and output a first semantic label image corresponding to a first video frame and a second semantic label image corresponding to a second video frame, wherein a pixel value of each pixel in the first semantic label image and the second semantic label image represents a semantic category label of the pixel, and the semantic category label is divided into three categories, i.e., a static object category, a potential moving object category, and a movable object category;

[0013] a semantic category label determination module configured to determine a first semantic category label corresponding to at least one first feature point in the first video frame according to the first semantic label image and the at least one first feature point, and determine a second semantic category label corresponding to at least one second feature point in the second video frame according to the second semantic label image and the at least one second feature point;

[0014] a pre-warning strategy determination module configured to determine at least one sub-target region of the target region according to the first semantic category label and the second semantic category label, and determine a pre-warning strategy of the sub-target region according to a level of the sub-target region.

[0015] In a third aspect, an electronic device is provided. The electronic device includes:

[0016] one or more processors;

[0017] a memory configured to store one or more programs,

[0018] when the one or more programs are executed by the one or more processors, the one or more processors implement a method for determining a pre-warning strategy of a target region according to any of the embodiments of the present application.

[0019] In a fourth aspect, the embodiments of the present application further provide a storage medium containing computer executable instructions for executing the method for determining a target area early warning strategy according to any of the embodiments of the present application when executed by a computer processor.

[0020] In a fifth aspect, the embodiments of the present application further provide a computer program product comprising a computer program, characterized in that the computer program, when executed by a processor, implements the method for determining a target area early warning strategy according to any of the embodiments of the present application.

[0021] The technical scheme of the embodiments of the present application acquires a plurality of video frames associated with a target area, wherein the video frames are acquired by a device configured with a visual sensor when the device moves in the target area, and the visual sensor acquires the video frames at a preset frame rate. Then, for any two consecutive video frames, the consecutive video frames are input into a first model to output a first semantic label image corresponding to the first video frame and a second semantic label image corresponding to the second video frame, wherein the pixel value of each pixel in the first semantic label image and the second semantic label image represents a semantic class label of the pixel, and the semantic class label is divided into three categories: a stationary object class, a potential moving object class, and a movable object class. Further, according to at least one first feature point in the first video frame and the first semantic label image, a first semantic class label corresponding to the first feature point is determined, and according to at least one second feature point in the second video frame and the second semantic label image, a second semantic class label corresponding to the second feature point is determined. Finally, according to the first semantic class label and the second semantic class label, at least one sub-target area of the target area is determined, and an early warning strategy of the sub-target area is determined according to the level of the sub-target area. The present application solves the problems of low accuracy and reliability of security in the prior art when dealing with security risks in a physical space, and the problems of high cost, poor real-time performance, subjective judgment errors, and the like in the security mode of manual patrol. The embodiments of the present application determine a plurality of sub-target areas and early warning strategies, making security more intelligent and efficient, and achieving the effect of improving the accuracy and reliability of area early warning. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the example embodiments of the present application, the drawings needed in the description of the embodiments are briefly introduced as follows. Obviously, the drawings introduced are only a part of the drawings of the embodiments to be described by the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0023] Figure 1 is a flowchart of a method for determining a target area early warning strategy provided by the embodiments of the present application;

[0024] Figure 2 is a flowchart of another method for determining a target area early warning strategy provided by an embodiment of the present disclosure;

[0025] Figure 3 is a schematic diagram of a semantic map construction thread provided by an embodiment of the present disclosure;

[0026] Figure 4 is a flowchart of a device for determining a target area early warning strategy provided by an embodiment of the present disclosure;

[0027] Figure 5 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The present application will be further described below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings, but not all the structures.

[0029] Before introducing the technical solutions provided by the embodiments of the present disclosure, the application scenarios can be exemplarily described. The technical solutions provided by the embodiments of the present disclosure can be applied in scenarios of determining a target area early warning strategy and constructing a semantic map of the target area. For example, in a physical space of a financial institution, the physical space of the financial institution can be divided to obtain a plurality of sub-regions, and the early warning strategies of different sub-target regions are determined, and the semantic map of the physical space of the financial institution is constructed. Based on the technical solutions of the embodiments of the present disclosure, the plurality of sub-target regions and the early warning strategies are determined, so that the security is more intelligent and efficient, and the accuracy and reliability of the target area early warning are achieved.

[0030] Embodiment one

[0031] Figure 1 is a flowchart of a method for determining a target area early warning strategy provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the scenarios of determining a target area early warning strategy and constructing a semantic map of the target area. The method can be executed by a device for determining a target area early warning strategy. The device can be realized in the form of software and / or hardware. The hardware can be an electronic device of a mobile terminal. The electronic device can execute the method for determining a target area early warning strategy provided by the present technical solution.

[0032] As shown in Figure 1 , the method comprises:

[0033] S110, acquiring a plurality of video frames associated with a target area.

[0034] The video frames are acquired by a device configured with a visual sensor when the device moves in the target area.

[0035] It should be noted that the target area refers to a specific geographical area that needs to be monitored and analyzed to determine a plurality of sub-areas and construct a semantic map. In the embodiments of the present application, the target area can be a financial institution. The video frames, as single static images, each composed of a series of pixels, capture the scene information of the target area at a specific time point. The device refers to the hardware tool used to shoot the video frames. The device can be a mobile robot, a camera, or a drone, etc. The device is configured with a visual sensor, which is a sensor used to capture video frames and can convert optical signals into electronic signals. The quality and clarity of the video frames are usually related to the resolution and settings of the visual sensor.

[0036] It should also be noted that the device can move autonomously in the target area, usually relying on navigation systems such as GPS, laser radar, or visual SLAM to determine the position and path; the device can also move in the target area according to a preset path or trajectory, such as moving in a specific track or route. The device should be able to cover the entire target area, ensuring that all key areas are monitored through a reasonable movement strategy. The preset frame rate refers to the number of frames captured per second when the visual sensor is collecting video frames. The frame rate directly affects the detail capture capability of the video frames.

[0037] Specifically, during the autonomous movement or movement according to the preset path or trajectory of the device in the target area, the visual sensor configured is used to collect a plurality of video frames at a preset frame rate.

[0038] S120, for any two consecutive video frames, input the consecutive video frames into the first model, output the first semantic label image corresponding to the first video frame and the second semantic label image corresponding to the second video frame.

[0039] Among them, the pixel value of each pixel in the first semantic label image and the second semantic label image represents the semantic class label of the pixel. The semantic class label is divided into three categories: static object class, potential moving object class, and movable object class.

[0040] It should be noted that for any two consecutive video frames, the first video frame refers to the former one of the two consecutive video frames, and the second video frame refers to the latter one of the two consecutive video frames. The first model refers to a deep learning model used to process video frames and generate semantic label images. The first model is based on a convolutional neural network or other deep learning architecture, which can accept image input and classify each pixel, thereby generating a semantic label image corresponding to the input image.

[0041] It is also to be noted that the first semantic label image refers to an output image corresponding to the first video frame, in which the value of each pixel represents the semantic class label of the object or region represented by the pixel. The first semantic label image usually has the same size as the original video frame, and the pixel value can be an integer representing different semantic classes. For example, 0 represents the background, 1 represents a stationary object, 2 represents a potential moving object, 3 represents a movable object, and so on. The second semantic label image refers to an output image corresponding to the second video frame, which has the same features as the first semantic label image. The value of each pixel in the second semantic label image also represents the semantic class label of the object or region represented by the pixel. By comparing the first semantic label image and the second semantic label image, the behavior and state change of the object in the target region can be analyzed. The semantic class label refers to a label used to identify the class of the object or region represented by each pixel in the semantic label image. The stationary object class refers to an object whose position does not change in two consecutive video frames. For example, in a financial institution, the stationary object class can include counters, ATMs, chairs, and fixed decorations within the financial institution, etc. The potential moving object class refers to an object that may move at a certain time. For example, in a financial institution, the potential moving object class can include customers waiting in line and objects placed on the counter, etc. The movable object class refers to an object whose position is likely to change. For example, in a financial institution, the movable object class can include customers conducting business, financial institution staff, customers entering and leaving the financial institution, and customers operating the automatic teller machine, etc.

[0042] Specifically, for any two consecutive video frames associated with the target region obtained, the first video frame and the second video frame are input into the first model. After processing by the first model, a first semantic label image corresponding to the first video frame is output, and a second semantic label image corresponding to the second video frame is output.

[0043] When training the first model, a semantic label image corresponding to the video frame is determined according to at least one video frame, and a first training sample in the training sample set is obtained based on the at least one video frame and the corresponding semantic label image. In order to improve the accuracy of the first model obtained by training, as many and as rich first training samples as possible can be obtained.

[0044] For each first training sample, the video frame in the current first training sample is input into the first model to obtain an actual semantic label image corresponding to the current first training sample.

[0045] It is to be noted that each first training sample can be trained in this way to obtain the required first model.

[0046] The first model is a model with initial parameters or default parameters. The actual semantic label image is obtained by inputting the video frame in the first training sample into the first model.

[0047] It should be noted that the model parameters in the first model do not meet the expected requirements, and therefore there is a certain difference between the actual semantic label image output based on the model parameters at this time and the theoretical semantic label image. Therefore, based on the actual semantic label image and the theoretical semantic label image corresponding to each video frame, the corresponding error loss value can be determined.

[0048] In this embodiment, the first model can be a YOLOv7 network model. It should be noted that only the semantic label image corresponding to the video frame can be obtained, i.e., the specific model type is not limited.

[0049] The actual semantic label image and the theoretical semantic label image of the current first training sample are processed based on the first preset loss function in the first model to correct the model parameters in the first model according to the obtained loss value.

[0050] It should be noted that the training parameters can be set to default values before training the first model. When training the first model, the training parameters in the model can be corrected based on the output results of the first model, i.e., the final applied first model can be obtained by correcting the loss function in the first model. Each video frame has a loss value corresponding to it, which is determined based on the actual semantic label image and the theoretical semantic label image of each video frame.

[0051] Specifically, after inputting the video frame in the first training sample into the first model, the first model can obtain the actual semantic label image corresponding to the video frame. According to the actual semantic label image and the theoretical semantic label image, the loss value corresponding to the video frame can be determined, and the model parameters in the first model can be corrected using the backpropagation method.

[0052] The convergence of the first preset loss function is taken as the training target to obtain the final first model that can be applied.

[0053] Specifically, the training error of the loss function, i.e., a loss parameter, can be used as a condition for detecting whether the loss function currently reaches convergence, such as whether the training error is less than a preset error or whether the error change trend is stable, or whether the current iteration number is equal to a preset number. If the convergence condition is detected, such as the training error of the loss function reaching less than the preset error or the error change tending to be stable, it indicates that the first model training is completed, at which time the iterative training can be stopped. If it is detected that the current convergence condition is not reached, the first training sample can be further obtained to train the first model until the training error of the loss function is within a preset range. When the training error of the loss function reaches convergence, an applicable first model can be obtained.

[0054] S130, determining a first semantic category label corresponding to the first feature point according to the at least one first feature point in the first video frame and the first semantic label image, and determining a second semantic category label corresponding to the second feature point according to the at least one second feature point in the second video frame and the second semantic label image.

[0055] The first feature point refers to one or more points with saliency and recognizability identified in the first video frame. The first feature point is usually a point with unique texture, color or shape in the first video frame, which can be used as a feature point in subsequent image processing or analysis. The second feature point refers to one or more points with saliency and recognizability identified in the second video frame. Similar to the first feature point, the second feature point is used to analyze the object in the second video frame and can be compared with the first feature point to check the motion or change of the object. The first semantic category label refers to a semantic category label associated with the first feature point in the first video frame, indicating the category of the object represented by the first feature point. The first semantic category label can be one of a stationary object category, a potential moving object category or a movable object category, reflecting the actual meaning of the first feature point in the first video frame. The second semantic category label refers to a semantic category label associated with the second feature point in the second video frame, indicating the category of the object represented by the second feature point. Like the first semantic category label, the second semantic category label can also be one of a stationary object category, a potential moving object category or a movable object category.

[0056] It should be noted that an algorithm for feature extraction can be used to obtain at least one first feature point from the first video frame and at least one second feature point from the second video frame. For example, a scale-invariant feature transform or a speeded up robust features can be used to detect a plurality of feature points from the first video frame, and finally at least one feature point is selected from the plurality of feature points as the first feature point.

[0057] Optionally, for the at least one first feature point, a first semantic category label corresponding to a first coordinate of the first feature point in the first semantic label image is determined according to the first coordinate of the first feature point in the first video frame; and for the at least one second feature point, a second semantic category label corresponding to a second coordinate of the second feature point in the second semantic label image is determined according to the second coordinate of the second feature point in the second video frame.

[0058] The first coordinate refers to the specific position of the first feature point in the first video frame, which is usually represented by a two-dimensional coordinate. In the feature extraction process, the algorithm can generate a corresponding first coordinate value for each first feature point. For example, if the coordinate of a first feature point in the first video frame is (150, 200), the first coordinate of the first feature point is (150, 200). The second coordinate refers to the specific position of the second feature point in the second video frame, which is usually represented by a two-dimensional coordinate.

[0059] It should be noted that for each first feature point in the first video frame, a coordinate of the first feature point is recorded. According to the first coordinate, a pixel point corresponding to the first coordinate in the first semantic label image is found. Finally, the first semantic category label of the pixel point corresponding to the first coordinate in the first semantic label image is read. The process of determining the second semantic category label corresponding to the second feature point is similar to the process of determining the first semantic category label corresponding to the first feature point, which will not be described in detail in the embodiment of the present application.

[0060] Specifically, according to the feature extraction algorithm, at least one first feature point in the first video frame and at least one second feature point in the second video frame can be determined. For the at least one first feature point, a first semantic category label corresponding to the first feature point can be determined according to the first semantic label image corresponding to the first video frame. For the at least one second feature point, a second semantic category label corresponding to the second feature point can be determined according to the second semantic label image corresponding to the second video frame.

[0061] S140, at least one sub-target region of the target region is determined according to the first semantic category label and the second semantic category label, and a warning strategy of the sub-target region is determined according to the level of the sub-target region.

[0062] The sub-target region refers to a smaller region divided based on the first semantic category label and the second semantic category label within the target region. The level of the sub-target region is a classification of the importance, risk level or priority of different sub-target regions. The level of the sub-target region includes special level, first level, second level and third level. The warning strategy is a response plan and measures formulated for the level of different sub-target regions.

[0063] It should be noted that when the level of the sub-target area is the special level, the early warning strategy can be strict control, rapid response, and multiple detection; when the level of the sub-target area is the first level, the early warning strategy can be real-time monitoring, rapid support, and key protection; when the level of the sub-target area is the second level, the early warning strategy can be balancing safety and experience, and focusing on environmental monitoring; and when the level of the sub-target area is the third level, the early warning strategy can be daily monitoring, flexible handling, and focusing on comfort.

[0064] Optionally, according to at least one pair of feature points, the target area is divided into at least one sub-target area; according to the semantic category label corresponding to the sub-target area, the level of the sub-target area is evaluated; and according to the level of the sub-target area, the early warning strategy of the sub-target area is determined.

[0065] It should be noted that for the extracted at least one first feature point and at least one second feature point, feature point matching can be performed. When the first feature point and the second feature point are extracted, the descriptors of the feature points are calculated, and the matching between the first feature point and the second feature point is performed according to the descriptors. The feature point descriptors in the first feature point and the second feature point can be matched using a matching algorithm such as brute force matching or FLANN matching to find a matched feature point pair. For each matched feature point pair, it is determined whether the semantic category labels corresponding to the feature points are consistent. When the first semantic category label corresponding to the first feature point and the second semantic category label corresponding to the second feature point are consistent, the feature point pair is regarded as a valid feature point pair.

[0066] It should be further noted that the feature point pair refers to a pair of first feature points and second feature points whose descriptor matches successfully and whose semantic category labels are consistent. According to the coordinate information of the feature point pair, their actual positions in the target region can be determined. By analyzing the distribution of these feature point pairs, the target region can be divided into multiple sub-target regions. And when dividing the target region, it is ensured that each sub-target region has the same semantic category label, so as to guarantee the rationality and effectiveness of the division. Each sub-target region is associated with its corresponding semantic category label, and according to the pre-defined evaluation standard, each sub-target region can be evaluated. The pre-defined evaluation standard can be: according to the semantic category label, the sub-target region is divided into special level, first level, second level and third level. For example, when the object corresponding to the semantic category label is a vault or a data center machine room, the level of the corresponding sub-target region is evaluated as special level; when the object corresponding to the semantic category label is a cash counter, an ATM refilling room or a bill archive, the level of the corresponding sub-target region is evaluated as first level; when the object corresponding to the semantic category label is a business hall or a self-service area, the level of the corresponding sub-target region is evaluated as second level; when the object corresponding to the semantic category label is an office corridor or an employee lounge, the level of the corresponding sub-target region is evaluated as third level.

[0067] Specifically, after determining the first semantic category label corresponding to the first feature point and the second semantic category label corresponding to the second feature point, at least one pair of feature point pairs is determined according to the first feature point, the first semantic category label corresponding to the first feature point, the second feature point and the second semantic category label corresponding to the second feature point. According to the feature point pair, the region consistent with the feature point pair can be divided into a sub-target region. Further, according to the semantic category label corresponding to the sub-target region, it is determined whether the sub-target region is special level, first level, second level or third level. For sub-target regions of different levels, the warning strategy of each sub-target region is determined.

[0068] The technical scheme of the embodiment of the present disclosure is to acquire a plurality of video frames associated with a target area, wherein the video frames are acquired by a device configured with a visual sensor when the device moves in the target area, and the visual sensor collects the video frames at a preset frame rate. Then, for any two consecutive video frames, the consecutive video frames are input into a first model to output a first semantic label image corresponding to the first video frame and a second semantic label image corresponding to the second video frame, wherein the pixel value of each pixel in the first semantic label image and the second semantic label image represents a semantic class label of the pixel, and the semantic class label is divided into three categories: a static object class, a potential moving object class, and a movable object class. Further, according to at least one first feature point in the first video frame and the first semantic label image, a first semantic class label corresponding to the first feature point is determined, and according to at least one second feature point in the second video frame and the second semantic label image, a second semantic class label corresponding to the second feature point is determined. Finally, according to the first semantic class label and the second semantic class label, at least one sub-target area of the target area is determined, and a warning strategy of the sub-target area is determined according to the level of the sub-target area. The problem of low precision and reliability of security in the prior art when dealing with safety hazards in a physical space is solved, and the problems of high cost, poor real-time performance, subjective judgment errors, and the like of the artificial patrol security mode are solved. The embodiment of the present disclosure determines a plurality of sub-target areas and a warning strategy, so that the security is more intelligent and efficient, and the precision and reliability of the area warning are achieved.

[0069] Embodiment two

[0070] Figure 2 The flowchart of the method for determining a target area warning strategy provided by the embodiment of the present disclosure is based on the foregoing embodiment and further refines the construction of the semantic map of the target area. The specific implementation can be referred to the technical scheme of the embodiment. The same or corresponding technical terms as the above embodiments are not repeated here.

[0071] As Figure 2 shown, the method specifically includes the following steps:

[0072] S210, acquiring a plurality of video frames associated with a target area.

[0073] S220, for any two consecutive video frames, inputting the consecutive video frames into a first model to output a first semantic label image corresponding to the first video frame and a second semantic label image corresponding to the second video frame.

[0074] S230, determining a first semantic category label corresponding to the first feature point according to the at least one first feature point and the first semantic label image in the first video frame, and determining a second semantic category label corresponding to the second feature point according to the at least one second feature point and the second semantic label image in the second video frame.

[0075] S240, determining at least one pair of feature points according to the at least one first feature point, the first semantic category label corresponding to the first feature point, the at least one second feature point, and the second semantic category label corresponding to the second feature point.

[0076] Optionally, at least one third feature point corresponding to the at least one first feature point when the category of the first semantic category label corresponding to the first feature point is a stationary object category is screened out from the at least one first feature point according to the first semantic category label corresponding to the first feature point; at least one fourth feature point corresponding to the at least one second feature point when the category of the second semantic category label corresponding to the second feature point is a stationary object category is screened out from the at least one second feature point according to the second semantic category label corresponding to the second feature point; and at least one pair of feature points is determined according to the at least one third feature point and the at least one fourth feature point.

[0077] The third feature point refers to a feature point corresponding to a stationary object category when the category of the corresponding semantic category label is screened out from the first feature point. The fourth feature point refers to a feature point corresponding to a stationary object category when the category of the corresponding semantic category label is screened out from the second feature point. In the embodiment of the application, the pair of feature points refers to a pair of third feature point and fourth feature point whose descriptors match successfully and whose semantic category labels are consistent.

[0078] Specifically, a pair of feature points whose category of the semantic category label is a stationary object category is screened out according to the first feature point, the first semantic category label corresponding to the first feature point, the second feature point, and the second semantic category label corresponding to the second feature point. By screening out the pair of feature points of the stationary object category, it can be ensured that the selected pair of feature points has consistency in semantics, which helps to reduce the possibility of false matching and improve the accuracy of feature point matching. The stationary object is usually an important part of the scene, and by focusing on these feature points, the structure and content of the scene can be better understood. By focusing on stationary objects, noise introduced by dynamic objects can be reduced, thereby optimizing the efficiency of subsequent processing. And by focusing on stationary objects, the system can better adapt to dynamic changes in the environment such as changes in lighting or moving objects, thereby improving the overall robustness and stability.

[0079] S250, determining the pose information of the device according to the at least one pair of feature points.

[0080] The pose information of the device refers to the translation vector and the rotation matrix of the second video frame relative to the first video frame.

[0081] Optionally, according to the calibration process of the visual sensor, an intrinsic matrix corresponding to the visual sensor is obtained; for the at least one pair of feature points, according to the intrinsic matrix and the first coordinate corresponding to the first feature point, a first normalized coordinate corresponding to the first feature point is determined; according to the intrinsic matrix and the second coordinate corresponding to the second feature point, a second normalized coordinate corresponding to the second feature point is determined; and according to the at least one first normalized coordinate and the at least one second normalized coordinate, the pose information of the device is determined.

[0082] The intrinsic matrix at least includes focal length and principal point position. The intrinsic matrix is a matrix describing the internal characteristics of the visual sensor, and is usually represented as a 3*3 matrix. The focal length is a basic parameter of the visual sensor, representing the distance from the lens to the imaging sensor. The focal length determines the magnification of the image, and is usually measured in millimeters. In the intrinsic matrix, the focal length is represented in pixels. The principal point is the origin of the image coordinate system, and is usually the center point of the image. According to the focal length of the visual sensor in the x and y directions and the horizontal and vertical coordinates of the principal point position, the intrinsic matrix corresponding to the visual sensor is determined.

[0083] It should be noted that the calibration process of the visual sensor refers to the process of obtaining the intrinsic matrix of the visual sensor through a series of known standard scenes or objects, such as a chessboard, a calibration board, etc. Specifically, images of known objects are taken at different angles and positions, and these video frames are recorded. Then, feature points are extracted from the video frames. Further, computer vision algorithms such as Zhang Zhengyou calibration method are used to estimate the intrinsic matrix corresponding to the visual sensor. For example, the intrinsic matrix can be represented as:

[0084]

[0085] where f x and f y represent the focal length of the visual sensor in the x and y directions; c x and c y represent the horizontal and vertical coordinates of the principal point position.

[0086] It should be noted that the first coordinate refers to the coordinate corresponding to the first feature point. The second coordinate refers to the coordinate corresponding to the second feature point. The first normalized coordinate refers to the result of converting the first coordinate to a normalized coordinate by using the intrinsic matrix. The second normalized coordinate refers to the result of converting the second coordinate to a normalized coordinate by using the intrinsic matrix. The first normalized coordinate and the second normalized coordinate are coordinates after removing the influence of the internal parameters of the visual sensor. It should be noted that for the first coordinate corresponding to the first feature point, the focal length and the principal point position in the intrinsic matrix are used to convert the first coordinate to the first normalized coordinate. For example, the normalized coordinate can be represented as:

[0087]

[0088] Here, x and y represent the first or second coordinate.

[0089] It should also be noted that after determining the first normalized coordinates corresponding to the first feature point and the second normalized coordinates corresponding to the second feature point, the first and second normalized coordinates can be used as feature point pairs for analysis. The PnP algorithm can be used to determine the device's pose information, outputting a translation vector representing the device's position change in space, and a rotation matrix representing the device's orientation change.

[0090] Specifically, based on multiple feature point pairs, the translation vector and rotation matrix of the second video frame relative to the first video frame are determined.

[0091] S260. Based on the pose information, determine the target coordinates of the second feature point in at least one pair of feature points in three-dimensional space, and construct a semantic map based on the target coordinates and the second semantic category label corresponding to the second feature point.

[0092] Here, the target coordinates refer to the exact location of the second feature point in three-dimensional space, usually represented in a three-dimensional Cartesian coordinate system, in the form of ( x (y, z). Collect multiple target coordinates and their corresponding second semantic category labels in the scene, and use appropriate data structures, such as point clouds, grids, or raster maps, to store the target coordinates and corresponding semantic category labels to construct a semantic map.

[0093] Optionally, the target coordinates of the second feature point in three-dimensional space are determined based on the translation vector, rotation matrix, and second normalized coordinates; the target coordinates and the second semantic category label corresponding to the second feature point are added to the semantic map to construct the semantic map.

[0094] It should be noted that the depth value corresponding to the second feature point needs to be determined. The depth value z can typically be obtained using a depth sensor or stereo vision. After determining the second normalized coordinates based on the translation vector and rotation matrix, the target coordinates of the second feature point in 3D space can be calculated using the second normalized coordinates and the depth value z. For example, X norm =z*x norm ;Y norm =z*y norm Z norm =z. Appropriate data structures such as dictionaries, arrays, or databases can be used to store the target coordinates of the second feature point in three-dimensional space and its corresponding second semantic category label, thus constructing a semantic map.

[0095] Specifically, after determining the pose information of the device, the target coordinates of the second feature points in the plurality of feature point pairs in the three-dimensional space are determined according to the pose information, and a semantic map is constructed according to the target coordinates and the second semantic category labels corresponding to the second feature points. For the constructed semantic map, a visualization tool such as point cloud visualization software is used to display the semantic map, helping users understand the scene structure. And save the semantic map to a file or a database for subsequent analysis and use.

[0096] The technical scheme of the embodiment of the present disclosure determines the first semantic category label corresponding to the first feature point and the second semantic category label corresponding to the second feature point, and then determines at least one pair of feature point pairs according to at least one first feature point, the first semantic category label corresponding to the first feature point, at least one second feature point, and the second semantic category label corresponding to the second feature point. Then, the pose information of the device is determined according to the at least one pair of feature point pairs. Finally, the target coordinates of the second feature points in the at least one pair of feature point pairs in the three-dimensional space are determined according to the pose information, and a semantic map is constructed according to the target coordinates and the second semantic category labels corresponding to the second feature points. According to the matching of the semantic category labels of the feature points, it is ensured that the selected feature point pairs have consistency in semantics, thereby improving the accuracy and reliability of feature point matching. Using multiple feature point pairs for pose estimation can improve the robustness of the system and reduce the positioning inaccuracy caused by single feature point error. Through the pose information and the normalized coordinates, the target coordinates of the second feature points in the three-dimensional space can be accurately determined. This positioning accuracy is crucial for subsequent environment modeling and analysis. Combining the target coordinates with the corresponding semantic category labels to construct a semantic map can provide more rich and useful environment information. In a dynamic environment, the update and maintenance of the semantic map can provide support for real-time feedback, helping the system to quickly adapt to environmental changes.

[0097] Embodiment three

[0098] As an optional embodiment of the present application, the invention is further illustrated with an example.

[0099] It should be noted that when constructing the semantic map, the semantic information in the scene is obtained, thereby enhancing the semantic understanding ability of the objects, structures and scenes in the environment. It can help the security system of the financial institution to more accurately perceive and understand the environment, and provide a reliable basis for subsequent dynamic target detection and protection area division. For dynamic targets in the scene, the system can effectively identify whether there is an abnormal motion trajectory through motion consistency detection based on optical flow calculation. Specifically, this method can distinguish between potential moving objects and abnormal moving objects, thereby filtering out false positives in the scene caused by factors such as light changes, personnel misoperation, and further accurately positioning illegal intruders. In addition, combined with the semantic map, the system can divide the area of the financial institution into different levels of protection zones. And according to the safety needs of different areas, take corresponding early warning strategies, so as to realize intelligent and differentiated security protection.

[0100] It should be noted that, see Figure 3In order to estimate the position and pose of the visual sensor during motion, first, the visual sensor continuously acquires image frames. By analyzing the consecutive video frames, the motion trajectory of the sensor is calculated using feature points or edge information in the images. In the video frames, the feature points usually refer to the pixels with unique and stable characteristics, such as corners or edges, etc. These feature points have high matching in different video frames, and the feature points of the current frame are matched with the feature points of the previous frame to identify the same feature points. Through the matched feature points, the motion parameters of the visual sensor can be calculated, such as translation and rotation. Finally, according to the motion parameters, the current pose of the visual sensor is updated. Further, through target detection technology, semantic information in the financial institution scene is extracted to enhance the scene understanding ability. In the financial institution scene, target detection is used to identify and locate key objects or areas, such as ATMs, counters, or safes, etc. According to the semantic information of the identified key objects, the bank physical space is divided, such as the vault area, the cash counter area, or the self-service area, etc. Loop closure detection is an important part of preventing positioning drift. When the sensor returns to a position that has been visited before, loop closure detection can identify this and reduce the cumulative positioning error through loop correction. During operation, the system continuously constructs a semantic map and stores key video frames. The current video frame is matched with the stored key video frames to find similar scenes. If the current video frame matches a key video frame successfully, it is considered that the visual sensor has returned to the previous position. Through the matching result, the current pose estimation is adjusted to reduce the cumulative error. The results of target detection are combined with traditional geometric maps to generate semantic maps containing semantic information. The semantic map not only contains geometric structures, but also contains information such as the category and purpose of the object, greatly improving the semantic understanding ability of the map. The process includes geometric map construction, semantic information integration, semantic labeling, and map updating, i.e. through feature detection, the geometric structure of the environment is constructed, such as point cloud map, grid map, etc. The results of target detection, such as object category, location, etc. are integrated into the geometric map. The objects in the map are semantically labeled, such as counters, ATMs, safes, etc. According to the new perception information, the semantic map is dynamically updated to maintain the real-time and accuracy of the map.

[0101] It is also necessary to point out that object detection aims to identify and locate specific objects or objects in video frames. It combines the tasks of object classification and object localization, and can automatically detect objects of interest in video frames and mark the location of the object with a bounding box or similar. When introducing object detection technology into SLAM, the model needs to balance between speed and accuracy, as SLAM usually has high real-time requirements. The YOLOv7 network model has the advantages of fast processing speed, high accuracy, lightweight design, strong small target detection capability, and support for multi-target detection. In the embodiment of the present application, the semantic information of the target of interest in the bank scene is extracted by object detection. For example, in the cash counter, ATM refilling room, bill archive library or business hall of a financial institution, a lot of important semantic information can be obtained. These information can help the model to better understand the structure and function of the financial institution scene. The important semantic information that can be obtained in different scenes and the key objects required for training the model are different. For example, the cash counter semantic information needs objects such as cash, currency counting machine, currency checking machine, printer, teller operation table and display for model training, while the semantic information of the business hall scene needs objects such as chairs, sofas, floor queuing lines, signs and call machines in the customer rest area for model training. For important places of the financial institution, such as the vault and the bill archive library, which are non-staff access areas, the face information of the staff needs to be trained for the model to further accurately locate the illegal intruders.

[0102] In the embodiment of the present application, the optical flow method analyzes the position changes of pixels in the sequence of video frames to infer the motion information of objects in the scene. The optical flow describes the motion vector of each pixel in the video frame from the current frame to the next frame. The core assumption of the optical flow method is that the pixels of the same object in adjacent frames have similar gray values or color features in space. Motion consistency detection refers to whether the motion trajectory of an object in the scene conforms to the expected physical law or the motion pattern of other objects in the scene. By analyzing the distribution of motion vectors in the optical flow field, it can be detected whether there is an inconsistent motion trajectory, such as sudden acceleration or direction mutation, etc.

[0103] Assuming that the pixels in the video frame collected by the visual sensor move over time, the video frame is regarded as a function I(t) with respect to time. Given the coordinates (x, y) of a pixel point in the video frame, the gray value of the pixel is described as I(x, y, t). Over time, the pixel point coordinates change. When t+dt, the pixel is located at (x+dx, y+dy), and based on the gray value invariance assumption, I(x, y, t) = I(x+dx, y+dy, t+dt).

[0104] Taking Taylor expansion on the above formula, we can get

[0105] Solving the two formulas together, we get arranged

[0106] Finally, the motion speed of the pixel points is solved to track the pixel points in multiple frames. By analyzing whether the motion trajectory of the object is smooth, whether there are abrupt broken lines or curves, whether the motion trajectory of the object is consistent with the motion trajectory of other objects, and whether there is deviation, the motion of the object is determined whether it is consistent with the normal motion mode in the background or scene, so as to divide the potential moving object and the abnormal motion object, detect the personnel target in the scene, and analyze whether the motion trajectory of the personnel target is consistent with the normal behavior mode. A light flow algorithm robust to light changes is selected to reduce the influence of light changes on the calculation of motion vectors and suppress the noise caused by personnel misoperation and light changes. A semantic map is constructed in real time by SLAM in the bank scene, and the financial institution environment is divided into different levels of protection zones. For example, the financial institution area is divided into a special protection zone, including a vault, a data center machine room, etc.; a first protection zone, including a cash counter, an ATM cash adding room, a bill archive, etc.; a second protection zone, including a business hall, a self-service area, etc.; and a third protection zone, including an office corridor, an employee rest room, etc. Different levels of protection zones are taken corresponding early warning strategies. The special protection zone: strict control, rapid response, multiple detection. The first protection zone: real-time monitoring, rapid support, key protection. The second protection zone: balance safety and experience, focus on environmental monitoring. The third protection zone: daily monitoring, flexible handling, focus on comfort. For the special protection zone, once unauthorized personnel enter or abnormal behavior such as equipment damage or illegal operation is detected, the highest level of alarm is triggered. A linkage mechanism is adopted, and alarm information should be sent to the security center, relevant responsible persons and emergency teams at the same time to ensure rapid response. The robot should record the whole process of the event to provide evidence for subsequent investigation. The system automatically blocks the area to limit personnel access. If an abnormal device is found, the robot should automatically cut off the power or start the standby system. For the first protection zone, the robot should monitor the behavior of personnel in the area in real time and identify abnormal actions such as theft tendencies. When an anomaly is found, the robot can issue a warning sound to remind the staff to pay attention. Alarm information should be immediately sent to security personnel and relevant responsible persons to ensure rapid arrival at the scene. If cash is stolen or abnormal operation is found, the robot should immediately notify the relevant departments and start the emergency plan. For the second protection zone, the robot should monitor the flow density in the area in real time and identify abnormal gathering or queuing. Identify abnormal behaviors such as fighting and theft, and monitor environmental abnormalities such as fire and water leakage, and alarm in time. When an abnormal behavior is found, the robot should remind the staff to pay attention and guide the customers to the safe area. If a fire or other emergency is found, the robot should guide the customers to evacuate in an orderly manner through the broadcast system. For the third protection zone, the robot should monitor the activities of personnel in the area in real time and identify abnormal behaviors such as long-time lingering. When an anomaly is found, a lower level of alarm is issued to remind the staff to pay attention and avoid excessive interference with normal office work.

[0107] The technical solutions of the embodiments of the present disclosure realize accurate scene understanding: the semantic SLAM combines the space mapping capability of SLAM and the target recognition capability of deep learning, so that the system can not only construct a three-dimensional map of a financial institution, but also semantically label objects, personnel, and signs in the environment, realize higher-level environmental understanding, and the system can not only construct a spatial map of a financial institution, but also label the functions of different regions, such as counters, access control, ATM machines, and the like, to provide more abundant information for security decision-making. In combination with motion consistency detection, the accuracy of abnormal behavior recognition is improved. Based on the semantic map, real-time early warning of the financial institution region is realized, and once abnormal behavior is detected, the security personnel or the security system can be automatically triggered to provide early warning, thereby improving the response speed.

[0108] Embodiment four

[0109] Figure 4 is a structural schematic diagram of a device for determining a target region early warning strategy provided by an embodiment of the present disclosure, as shown in Figure 4 The device includes a video frame acquisition module 310, a semantic label image output module 320, a semantic category label determination module 330, and an early warning strategy determination module 340.

[0110] The video frame acquisition module is configured to acquire a plurality of video frames associated with a target region; wherein the video frames are acquired by a device configured with a visual sensor when the device moves in the target region; the visual sensor acquires the video frames at a preset frame rate; the semantic label image output module is configured to input any two consecutive video frames to a first model, and output a first semantic label image corresponding to a first video frame and a second semantic label image corresponding to a second video frame; wherein the pixel value of each pixel in the first semantic label image and the second semantic label image represents the semantic category label of the pixel; the semantic category label is divided into three categories: static object category, potential moving object category, and movable object category; the semantic category label determination module is configured to determine a first semantic category label corresponding to at least one first feature point in the first video frame according to the first feature point and the first semantic label image, and determine a second semantic category label corresponding to at least one second feature point in the second video frame according to the second feature point and the second semantic label image; the early warning strategy determination module is configured to determine at least one sub-target region of the target region according to the first semantic category label and the second semantic category label, and determine an early warning strategy of the sub-target region according to the level of the sub-target region.

[0111] The technical scheme of the embodiment of the present disclosure is as follows: a plurality of video frames associated with a target region are acquired, wherein the video frames are acquired by a device configured with a visual sensor when the device moves in the target region, and the visual sensor collects the video frames at a preset frame rate. Then, for any two continuous video frames, the continuous video frames are input into a first model, and a first semantic label image corresponding to the first video frame and a second semantic label image corresponding to the second video frame are output, wherein the pixel value of each pixel in the first semantic label image and the second semantic label image represents a semantic class label of the pixel, and the semantic class label is divided into three categories: a static object class, a potential moving object class, and a movable object class. Further, a first semantic class label corresponding to at least one first feature point in the first video frame is determined according to the first feature point and the first semantic label image, and a second semantic class label corresponding to at least one second feature point in the second video frame is determined according to the second feature point and the second semantic label image. Finally, at least one sub-target region of the target region is determined according to the first semantic class label and the second semantic class label, and a warning strategy of the sub-target region is determined according to the level of the sub-target region. The problems of low accuracy and reliability of security in the prior art when dealing with security risks in a physical space, and the problems of high cost, poor real-time performance, subjective judgment errors, and the like of the artificial patrol security mode are solved. The embodiment of the present disclosure determines a plurality of sub-target regions and a warning strategy, so that the security is more intelligent and efficient, and the accuracy and reliability of the regional warning are achieved.

[0112] On the basis of the above technical methods, the device further includes a feature point pair determination module, a pose information determination module, and a semantic map construction module.

[0113] The feature point pair determination module is configured to determine at least one feature point pair according to the at least one first feature point, the first semantic class label corresponding to the first feature point, the at least one second feature point, and the second semantic class label corresponding to the second feature point.

[0114] The pose information determination module is configured to determine the pose information of the device according to the at least one feature point pair.

[0115] The semantic map construction module is configured to determine a target coordinate of the second feature point in the at least one feature point pair in a three-dimensional space according to the pose information, and construct a semantic map according to the target coordinate and the second semantic class label corresponding to the second feature point.

[0116] On the basis of the above technical methods, the semantic class label determination module 330 includes a first semantic class label determination submodule and a second semantic class label determination submodule.

[0117] The first semantic category label determination submodule is configured to determine, for at least one first feature point, a first semantic category label corresponding to a first coordinate of the first feature point in the first video frame in the first semantic label image.

[0118] The second semantic category label determination submodule is configured to determine, for at least one second feature point, a second semantic category label corresponding to a second coordinate of the second feature point in the second video frame in the second semantic label image.

[0119] On the basis of the above technical methods, the early warning strategy determination module 340 comprises a sub-target region division submodule, a level evaluation submodule, and an early warning strategy determination submodule.

[0120] The sub-target region division submodule is configured to divide a target region into at least one sub-target region according to at least one pair of feature points.

[0121] The level evaluation submodule is configured to evaluate the level of the sub-target region according to a semantic category label corresponding to the sub-target region.

[0122] The early warning strategy determination submodule is configured to determine an early warning strategy of the sub-target region according to the level of the sub-target region.

[0123] On the basis of the above technical methods, the feature point pair determination module comprises a third feature point screening submodule, a fourth feature point screening submodule, and a feature point pair matching submodule.

[0124] The third feature point screening submodule is configured to screen, from the at least one first feature point, at least one third feature point corresponding to a static object category when a first semantic category label corresponding to the first feature point is a static object category.

[0125] The fourth feature point screening submodule is configured to screen, from the at least one second feature point, at least one fourth feature point corresponding to a static object category when a second semantic category label corresponding to the second feature point is a static object category.

[0126] The feature point pair matching submodule is configured to determine at least one pair of feature points according to the at least one third feature point and the at least one fourth feature point.

[0127] On the basis of the above technical methods, the pose information determination module comprises an intrinsic matrix acquisition submodule, a normalized coordinate determination submodule, and a device pose information determination submodule.

[0128] An intrinsic parameter matrix obtaining submodule is configured to obtain an intrinsic parameter matrix corresponding to the visual sensor according to a calibration process of the visual sensor, wherein the intrinsic parameter matrix at least includes a focal length and a principal point position.

[0129] A normalized coordinate determining submodule is configured to determine a first normalized coordinate corresponding to the first feature point pair according to the intrinsic parameter matrix and a first coordinate corresponding to the first feature point pair, and determine a second normalized coordinate corresponding to the second feature point pair according to the intrinsic parameter matrix and a second coordinate corresponding to the second feature point pair.

[0130] A device pose information determining submodule is configured to determine the pose information of the device according to the at least one first normalized coordinate and the at least one second normalized coordinate, wherein the pose information is a translation vector and a rotation matrix of the second video frame relative to the first video frame.

[0131] On the basis of the above technical methods, the semantic map construction module includes a target coordinate determining submodule and a second semantic category label adding submodule.

[0132] The target coordinate determining submodule is configured to determine a target coordinate of the second feature point in a three-dimensional space according to the translation vector, the rotation matrix, and the second normalized coordinate.

[0133] The second semantic category label adding submodule is configured to add the target coordinate and a second semantic category label corresponding to the second feature point to the semantic map to construct the semantic map.

[0134] The device for determining a target region early warning strategy provided in the embodiments of the present disclosure can execute the method for determining a target region early warning strategy provided in any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of executing the method.

[0135] It should be noted that each unit and module included in the above device is only divided according to a function logic, but is not limited to the above division, as long as the corresponding functions can be implemented; in addition, the specific names of each functional unit are only for convenient mutual distinction, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0136] Embodiment Five

[0137] Figure 5 is a structural schematic diagram of an electronic device provided in the embodiments of the present disclosure. The following refers to Figure 5 which shows an electronic device (for example, a mobile phone) suitable for being used to implement the embodiments of the present disclosure. Figure 5The diagram below shows the structure of the terminal device (or server) 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals). Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0138] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.

[0139] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0140] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0141] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.

[0142] The electronic device provided by the embodiments of the present disclosure and the method for determining a target area warning strategy provided by the above embodiments belong to the same inventive concept, and technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.

[0143] Embodiment six

[0144] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the method for determining a target area warning strategy provided by the above embodiments.

[0145] It should be noted that the computer readable medium of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0146] In some embodiments, the server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., communications networks) of any form or medium, such as a local area network ("LAN"), a wide area network ("WAN"), the Internet, and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0147] The computer readable medium described above can be included in the electronic device described above; or can exist separately, without being assembled into the electronic device.

[0148] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to:

[0149] Obtain a plurality of video frames associated with a target area; wherein the video frames are obtained when a device configured with a visual sensor moves within the target area; the visual sensor collects the video frames at a preset frame rate;

[0150] For any two consecutive video frames, input the consecutive video frames into a first model, and output a first semantic label image corresponding to the first video frame and a second semantic label image corresponding to the second video frame; wherein the pixel value of each pixel in the first semantic label image and the second semantic label image represents the semantic class label of the pixel; the semantic class label is divided into three categories: static object class, potential moving object class, and movable object class;

[0151] According to at least one first feature point in the first video frame and the first semantic label image, determine the first semantic class label corresponding to the first feature point, and according to at least one second feature point in the second video frame and the second semantic label image, determine the second semantic class label corresponding to the second feature point;

[0152] According to the first semantic class label and the second semantic class label, determine at least one sub-target area of the target area, and according to the level of the sub-target area, determine the early warning strategy of the sub-target area.

[0153] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0154] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0155] The units described in the embodiments of the present disclosure can be implemented by hardware, software, or a combination of hardware and software. In some cases, the names of the units do not constitute a limitation on the units themselves.

[0156] The functions described in this specification can be implemented in part or in whole through one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0157] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0158] The above description is only the preferred embodiment of the present disclosure and the explanation of the principles of the applied technology. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0159] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0160] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A method for determining an early warning strategy for a target area, characterized in that, The method includes: Multiple video frames associated with a target area are acquired; wherein the video frames are acquired by a device equipped with a visual sensor as it moves within the target area; the visual sensor acquires the video frames at a preset frame rate; For any two consecutive video frames, the consecutive video frames are input into the first model, and the first semantic label image corresponding to the first video frame and the second semantic label image corresponding to the second video frame are output; wherein, the pixel value of each pixel in the first semantic label image and the second semantic label image represents the semantic category label of the pixel; the semantic category labels are divided into three categories: stationary object class, potential moving object class, and movable object class. Based on at least one first feature point in the first video frame and the first semantic label image, a first semantic category label corresponding to the first feature point is determined; based on at least one second feature point in the second video frame and the second semantic label image, a second semantic category label corresponding to the second feature point is determined. Based on the first semantic category label and the second semantic category label, at least one sub-target region of the target region is determined, and a warning strategy for the sub-target region is determined based on the level of the sub-target region.

2. The method according to claim 1, characterized in that, After determining the first semantic category label corresponding to the first feature point based on at least one first feature point in the first video frame and the first semantic label image, and determining the second semantic category label corresponding to the second feature point based on at least one second feature point in the second video frame and the second semantic label image, the method further includes: Based on the at least one first feature point, the first semantic category label corresponding to the first feature point, at least one second feature point, and the second semantic category label corresponding to the second feature point, at least one pair of feature points are determined; Based on the at least one pair of feature points, the pose information of the device is determined; Based on the pose information, the target coordinates of the second feature point in the at least one pair of feature points in the three-dimensional space are determined, and a semantic map is constructed based on the target coordinates and the second semantic category label corresponding to the second feature point.

3. The method according to claim 1, characterized in that, The step of determining a first semantic category label corresponding to the first feature point based on at least one first feature point in the first video frame and the first semantic label image, and determining a second semantic category label corresponding to the second feature point based on at least one second feature point in the second video frame and the second semantic label image, includes: For at least one first feature point, a first semantic category label corresponding to the first coordinate in the first video frame is determined based on the first coordinate of the first feature point. For at least one second feature point, based on the second coordinates corresponding to the second feature point in the second video frame, the second semantic category label corresponding to the second coordinates in the second semantic label image is determined.

4. The method according to claim 1, characterized in that, The step of determining at least one sub-target region of the target region based on the first semantic category label and the second semantic category label, and determining the early warning strategy of the sub-target region based on the level of the sub-target region, includes: The target region is divided into at least one sub-target region based on at least one pair of feature points; The level of the sub-target region is evaluated based on the semantic category label corresponding to the sub-target region; Based on the level of the sub-target area, determine the early warning strategy for the sub-target area.

5. The method according to claim 2, characterized in that, The step of determining at least one pair of feature points based on the at least one first feature point, the first semantic category label corresponding to the first feature point, at least one second feature point, and the second semantic category label corresponding to the second feature point includes: Based on the first semantic category label corresponding to the first feature point, at least one third feature point is selected from the at least one first feature point when the category of the first semantic category label is a static object; Based on the second semantic category label corresponding to the second feature point, at least one fourth feature point is selected from the at least one second feature point when the category of the second semantic category label is a static object; Based on the at least one third feature point and the at least one fourth feature point, at least one pair of feature points is determined.

6. The method according to claim 2, characterized in that, Determining the pose information of the device based on the at least one pair of feature points includes: Based on the calibration process of the vision sensor, the intrinsic parameter matrix corresponding to the vision sensor is obtained; wherein, the intrinsic parameter matrix includes at least the focal length and the principal point position; For the at least one pair of feature points, the first normalized coordinates corresponding to the first feature point are determined based on the intrinsic parameter matrix and the first coordinates corresponding to the first feature point; the second normalized coordinates corresponding to the second feature point are determined based on the intrinsic parameter matrix and the second coordinates corresponding to the second feature point. The pose information of the device is determined based on at least one first normalized coordinate and at least one second normalized coordinate; wherein the pose information is the translation vector and rotation matrix of the second video frame relative to the first video frame.

7. The method according to claim 2, characterized in that, The step of determining the target coordinates of the second feature point in the at least one pair of feature points in three-dimensional space based on the pose information, and constructing a semantic map based on the target coordinates and the second semantic category label corresponding to the second feature point, includes: Based on the translation vector, rotation matrix, and second normalized coordinates, determine the target coordinates of the second feature point in three-dimensional space; The semantic map is constructed by adding the target coordinates and the second semantic category label corresponding to the second feature point to the semantic map.

8. An apparatus for determining a target area early warning strategy, characterized in that, include: A video frame acquisition module is used to acquire multiple video frames associated with a target area; wherein, the video frames are acquired by a device equipped with a visual sensor moving within the target area; the visual sensor acquires the video frames at a preset frame rate; The semantic label image output module is used to input any two consecutive video frames into a first model and output a first semantic label image corresponding to the first video frame and a second semantic label image corresponding to the second video frame; wherein, the pixel value of each pixel in the first semantic label image and the second semantic label image represents the semantic category label of the pixel; the semantic category labels are divided into three categories: static object class, potential moving object class, and movable object class. The semantic category label determination module is used to determine the first semantic category label corresponding to the first feature point based on at least one first feature point in the first video frame and the first semantic label image, and to determine the second semantic category label corresponding to the second feature point based on at least one second feature point in the second video frame and the second semantic label image. The early warning strategy determination module is used to determine at least one sub-target region of the target region based on the first semantic category label and the second semantic category label, and to determine the early warning strategy of the sub-target region based on the level of the sub-target region.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the method for determining a target area early warning strategy as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the method for determining a target area early warning strategy as described in any one of claims 1-7.