Remote monitoring early warning processing method and device

By integrating the features of monitoring images, depth images and semantic segmented images in the remote monitoring system, identifying and determining the type of early warning object of the target object, the accuracy of target recognition and early warning processing in complex traffic scenarios in the prior art is solved, and a more efficient and reliable remote monitoring and early warning is achieved.

CN119649614BActive Publication Date: 2025-05-13SHENZHEN KEAN DIGITAL CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510186294.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-13
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

When dealing with complex traffic scenarios, existing remote monitoring systems face challenges in target identification, early warning processing and real-time response, especially under interference factors such as lighting changes, occlusion and complex backgrounds, resulting in a reduced accuracy of target identification, which in turn affects the accuracy of remote monitoring and early warning.

Method used

By obtaining the monitoring images, depth images and semantic segmented images collected by the remote camera, visual features, depth features and semantic features are extracted respectively, and feature fusion processing is performed, and similarity is found in the preset object feature set, thereby identifying and determining the warning object type of the target object and triggering the corresponding monitoring and warning processing operation.

Benefits of technology

By integrating multimodal data, this solution improves the accuracy of early warning object recognition of surveillance images, thereby improving the accuracy and reliability of remote monitoring and early warnings, and reducing false alarms and missed reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649614B_ABST
    Figure CN119649614B_ABST
Patent Text Reader

Abstract

The present application provides a remote monitoring early warning processing method and device, the method can obtain monitoring images and depth images, semantic segmentation images collected by a remote camera; respectively extract visual features, depth features and semantic features of the target object and perform feature fusion processing to obtain fused features; search for object type features corresponding to multiple early warning object types in a preset object feature set, and calculate the similarity between the fused features and each early warning object category feature; screen out at least one target type feature based on the similarity; identify the target early warning object type to which at least one target type feature belongs, and count the number of target type features in the target early warning object type whose similarity with the fused features is greater than a preset similarity threshold; determine the early warning object type of the target object according to the number of features corresponding to each target early warning object type; and trigger the corresponding monitoring early warning processing operation according to the early warning object type of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of information technology and artificial intelligence technology, and in particular to a method and device for processing early warnings for remote monitoring. Background Art

[0002] With the rapid development of computer vision and image processing technology, remote monitoring systems have been widely used in the field of traffic management. These systems provide important support for traffic flow optimization, traffic violation detection and road safety monitoring by real-time monitoring and analysis of traffic image data. However, existing technologies still face many challenges in dealing with complex traffic scenarios, especially in target recognition, warning processing and real-time response. In the process of target object recognition analysis and warning, traditional remote monitoring systems are prone to reduce the accuracy of target recognition when faced with interference factors such as lighting changes, occlusion, and complex background, which in turn affects the accuracy of remote monitoring warnings. Summary of the invention

[0003] The embodiments of the present application provide a remote monitoring warning processing method and device, which can improve the accuracy of remote monitoring warnings.

[0004] In one aspect, an embodiment of the present application provides a remote monitoring early warning processing method, which is applicable to a remote monitoring platform, and includes:

[0005] Acquire a monitoring image captured by a remote camera and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object; perform feature extraction on the monitoring image, the depth image and the semantic segmentation image respectively to obtain visual features, depth features and semantic features of the target object;

[0006] Performing feature fusion processing on the visual features, depth features and semantic features to obtain fused features;

[0007] Searching for object type features corresponding to multiple warning object types in a preset object feature set, and calculating the similarity between the fusion feature and each warning object category feature;

[0008] Based on the similarity, at least one target type feature is selected from the object type features corresponding to the plurality of warning object types;

[0009] Identify the target warning object type to which the at least one target type feature belongs, and count the number of target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold, to obtain the number of features corresponding to each target warning object type;

[0010] Determining the warning object type of the target object from the target warning object types according to the number of features corresponding to each target warning object type;

[0011] The corresponding monitoring warning processing operation is triggered according to the warning object type of the target object.

[0012] In one aspect, an embodiment of the present application provides a remote monitoring early warning processing device, which is applicable to a remote monitoring platform, and includes:

[0013] An acquisition unit is used to acquire a monitoring image collected by a remote camera and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object; and perform feature extraction on the monitoring image, the depth image and the semantic segmentation image to obtain visual features, depth features and semantic features of the target object;

[0014] A fusion unit, used for performing feature fusion processing on the visual features, depth features and semantic features to obtain fusion features;

[0015] A similarity unit, used to search for object type features corresponding to multiple warning object types in a preset object feature set, and calculate the similarity between the fusion feature and each warning object category feature;

[0016] A selection unit, configured to select at least one target type feature from the object type features corresponding to the plurality of warning object types based on the similarity;

[0017] an identification unit, configured to identify the target warning object type to which the at least one target type feature belongs, and to count the number of target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold, to obtain the number of features corresponding to each target warning object type;

[0018] An object determination unit, configured to determine the warning object type of the target object from the target warning object types according to the number of features corresponding to each target warning object type;

[0019] The warning unit is used to trigger a corresponding monitoring warning processing operation according to the warning object type of the target object.

[0020] The embodiment of the present application can obtain a monitoring image collected by a remote camera and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object; feature extraction is performed on the monitoring image, depth image and semantic segmentation image respectively to obtain visual features, depth features and semantic features of the target object; feature fusion processing is performed on the visual features, depth features and semantic features to obtain fusion features; object type features corresponding to multiple warning object types are searched in a preset object feature set, and the similarity between the fusion feature and each warning object category feature is calculated; based on the similarity, at least one target type feature is screened out from the object type features corresponding to the multiple warning object types; the target warning object type to which the at least one target type feature belongs is identified, and the number of features of the target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold is counted to obtain the number of features corresponding to each target warning object type; the warning object type of the target object is determined from the target warning object type according to the number of features corresponding to each target warning object type; and the corresponding monitoring warning processing operation is triggered according to the warning object type of the target object. This solution realizes warning object recognition in surveillance images by fusing multimodal data (fusion of surveillance images, depth images and semantic segmentation images), which can improve the accuracy of warning object recognition and thus improve the accuracy and reliability of remote monitoring warnings.

[0021] Furthermore, the embodiments of the present application can comprehensively identify warning objects based on the fusion feature similarity and feature quantity, which can greatly improve the accuracy of warning object identification, and further improve the accuracy and reliability of remote monitoring warnings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the implementation examples or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 This is a schematic diagram of the architecture of a remote monitoring and early warning system provided in an embodiment of the present application;

[0024] Figure 2 It is a flowchart of a remote monitoring early warning processing method provided in an embodiment of the present application;

[0025] Figure 3 It is a schematic diagram of a feature adjustment process provided in an embodiment of the present application;

[0026] Figure 4It is a structural diagram of a remote monitoring early warning processing device provided in an embodiment of the present application;

[0027] Figure 5 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. The terms that may be involved in this application are explained below.

[0029] Remote monitoring: Remote monitoring refers to the technology of real-time monitoring and control of equipment, places or systems far away from the monitoring center through the network or other communication means. It uses technical means such as sensors, cameras, communication networks and computer software to realize real-time data collection, transmission, analysis and management of the target area. The remote monitoring system is usually composed of monitoring equipment (such as cameras, sensors), transmission equipment (such as wired or wireless networks) and management platforms (such as computers, mobile phone apps), which can realize functions such as status monitoring, data collection, analysis and processing and remote control of remote equipment or systems. This technology is widely used in public safety, traffic management, industrial automation, health care, agriculture, environmental monitoring and other fields, and plays an important role in improving management efficiency, ensuring safety and optimizing resource allocation.

[0030] Early warning processing of remote monitoring refers to the process of timely discovering abnormal situations or potential risks in a remote monitoring system through real-time analysis and processing of monitoring data, and taking corresponding measures for early warning and intervention. It uses advanced computer vision, image processing, data analysis, and machine learning technologies to intelligently analyze images, videos, or other sensor data collected by remote cameras, identify possible safety hazards, equipment failures, environmental changes, and other abnormal events, and generate early warning information based on preset rules and thresholds, notify relevant personnel, or automatically trigger corresponding processing mechanisms, so that timely measures can be taken to avoid or reduce losses. Early warning processing plays a vital role in remote monitoring systems, and can effectively improve the active defense capabilities and emergency response speed of the monitoring system, and ensure the safety and stable operation of the monitoring area.

[0031] Warning object recognition refers to the process of intelligently analyzing and processing monitoring data in a remote monitoring system using computer vision, image processing, machine learning and other technologies to automatically detect and identify target objects that need attention and determine whether they trigger warning conditions. This process usually involves feature extraction, classification and behavior analysis of specific objects (such as people, vehicles, objects, etc.) in monitoring images or video streams to determine whether they meet the preset warning rules. For example, in security monitoring, the system can identify blacklisted personnel or strangers through face recognition technology and issue an alarm; in traffic monitoring, it can identify speeding vehicles or violations and issue warnings. The application of warning object recognition technology can significantly improve the intelligence level of the monitoring system and the accuracy of warnings, reduce false alarms and missed alarms, and provide more reliable protection for public safety, traffic management and other fields.

[0032] In the field of remote monitoring, with the development of technology, the importance of early warning processing technology has become increasingly prominent. However, the existing remote monitoring early warning processing technology still has some shortcomings. Traditional remote monitoring systems usually rely on single-modal data (such as two-dimensional images) for analysis. In complex scenarios, such as lighting changes, occlusions, and complex backgrounds, the accuracy of target recognition and the timeliness of early warning are often insufficient. The existence of these problems limits the application effect of remote monitoring systems in public safety, traffic management, industrial automation and other fields. Therefore, there is an urgent need for a new technical solution that can effectively solve the above problems and improve the accuracy and reliability of remote monitoring early warning processing.

[0033] In order to solve the above technical problems, the embodiment of the present application provides a remote monitoring early warning processing method, which overcomes the limitations of single modality data in complex scenes by fusing the features of monitoring images, depth images and semantic segmentation images. For example, in the case of lighting changes, occlusion or complex background, a single image may not be able to accurately identify the target, but after combining depth information and semantic information, the target object can be more accurately located and identified, thereby improving the accuracy and stability of remote monitoring early warnings.

[0034] The solution provided in the embodiment of the present application can be implemented by using computer vision technology (Computer Vision, CV) and machine learning (Machine Learning, ML) in the field of artificial intelligence. Specifically:

[0035] Artificial intelligence (AI) is a technology that simulates and extends human intelligence through digital computers or computer-controlled machines. It enables machines to perceive the environment, acquire knowledge, and use this knowledge to achieve optimal results. As an important branch of computer science, artificial intelligence aims to reveal the nature of intelligence and develop intelligent machines that can respond in a human-like manner. This technology covers a wide range of hardware and software fields, including basic technologies such as sensor technology, dedicated chips, cloud computing, big data processing, interactive systems, and mechatronics, as well as software technologies such as computer vision, speech processing, natural language processing, and machine learning.

[0036] Computer vision (CV) is a science dedicated to giving machines visual perception capabilities. It uses cameras and computers to replace human eyes to identify, track and measure targets, and process images to make them more suitable for human observation or instrument detection. This field focuses on developing systems that can extract information from images or multidimensional data, covering technologies such as image processing, target recognition, semantic understanding, image retrieval, optical character recognition (OCR), video processing, behavior recognition, 3D reconstruction, virtual reality, augmented reality, and simultaneous positioning and mapping, as well as biometric recognition technologies such as face recognition and fingerprint recognition.

[0037] Machine learning (ML) is an interdisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It focuses on studying how to enable computers to simulate human learning behavior to acquire new knowledge or skills and continuously optimize their own performance. As the core of artificial intelligence, machine learning enables computers to learn and improve from data through various algorithms and techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and demonstration learning.

[0038] refer to Figure 1 , Figure 1 1 is a schematic diagram of the architecture of a remote monitoring and early warning system provided in an embodiment of the present application, and the remote monitoring and early warning system may include multiple remote cameras 100, early warning terminal devices 101, and a remote monitoring platform 102. The remote cameras 100 and the terminal devices 101 are connected to the remote monitoring platform 102 through a network.

[0039] Among them, the remote monitoring platform 102 can be an independent cloud server or a cloud server cluster or distributed system composed of multiple cloud servers. The cloud server can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0040] The remote camera 100 can collect image data of the monitored object in the monitoring area and upload it to the remote monitoring platform. The monitoring area can be divided according to different remote monitoring scenarios, such as indoor monitoring area, parking lot monitoring area, etc.

[0041] The remote cloud monitoring platform 102 can obtain the monitoring image collected by the remote camera 100 and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object; feature extraction is performed on the monitoring image, depth image and semantic segmentation image respectively to obtain visual features, depth features and semantic features of the target object; feature fusion processing is performed on the visual features, depth features and semantic features to obtain fusion features; object type features corresponding to multiple warning object types are searched in a preset object feature set, and the similarity between the fusion feature and each warning object category feature is calculated; based on the similarity, at least one target type feature is selected from the object type features corresponding to the multiple warning object types; the target warning object type to which the at least one target type feature belongs is identified, and the number of target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold is counted to obtain the number of features corresponding to each target warning object type; the warning object type of the target object is determined from the target warning object type according to the number of features corresponding to each target warning object type; and the corresponding monitoring warning processing operation is triggered according to the warning object type of the target object. For example, in one embodiment, after determining the warning object type of the target object, warning notification information or the like may be sent to the warning terminal device 101 to prompt the risk.

[0042] The early warning terminal device 101 may be one or more of a smart phone, a camera, a desktop computer, a tablet computer, an MP4 player and a laptop computer.

[0043] The early warning terminal device 101 receives the early warning notification information of remote monitoring and can display the information, for example, display the processing results in the form of a document. The content of the early warning display includes but is not limited to information such as the target object, description information associated with the target object, risk warnings, etc.

[0044] refer to Figure 2 , Figure 2 This is a flowchart of a remote monitoring early warning process provided by an embodiment of the present application. The execution subject of the method may be a remote monitoring platform, which may be a computer device or a cluster of multiple computer devices. The computer device may be a terminal device or a server, etc. The early warning process provided by the embodiment of the present application specifically includes:

[0045] S201, obtaining a monitoring image captured by a remote camera and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object.

[0046] In one embodiment, the remote camera can collect monitoring images in real time and upload them to the remote monitoring platform via the network for early warning analysis and processing.

[0047] In one embodiment, remote camera devices can be deployed in multiple monitoring areas respectively, with one remote camera device deployed in each monitoring area to collect monitoring image data of objects in the monitoring area and upload them to the remote cloud monitoring platform in real time, so that the remote cloud monitoring platform can generate and output remote monitoring videos based on the monitoring image data.

[0048] The monitoring image may specifically include an image of at least one target object, which may be a living object such as a person or an animal, or a static object, etc., and is not specifically limited here.

[0049] Among them, the depth image includes the depth value information of each pixel in the monitoring image, and each pixel value in the depth image represents the depth value of the corresponding pixel, usually in millimeters or centimeters. The depth image can provide three-dimensional structural information of the scene, so that the computer can more accurately understand the shape, position and spatial relationship of the object. In one embodiment, the monitoring image can be processed by pixel depth calculation to obtain the depth value information corresponding to the monitoring image, thereby obtaining the depth image. In one embodiment, the depth image is an image obtained by depth estimation of the monitoring image, and there is a one-to-one correspondence between the pixels in the monitoring image and the pixels in the depth image. In one embodiment, the depth calculation can form a depth image by calculating the distance between the object and the shooting point, which can also be called an optical flow map, wherein the depth image contains the parallax of the object between different images. The parallax is used to obtain the depth information of the target object, that is, the distance between the target object and the shooting point.

[0050] Among them, semantic segmentation image technology is a technology in the field of computer vision, whose goal is to assign each pixel in the image to a predefined semantic category. For example, the semantic segmentation model assigns a category label to each pixel in the input image, thereby forming a semantic map (also known as a semantic segmentation image) with the same resolution as the input image. This process not only divides the image into several regions, but also assigns a semantic label to each region to indicate the object or category represented by the region. For example, in an image containing people, vehicles, and buildings, each pixel will be labeled as "person", "vehicle", or "building". In this way, we can clearly understand the position and shape of each object in the image.

[0051] In an embodiment of the present application, the semantic segmentation image is a semantic segmentation image obtained by performing semantic segmentation processing on the surveillance image, and the area where the target object is located in the surveillance image is segmented out to form a semantic segmentation image. The semantic segmentation image may include the location information and shape information of the target object. In one embodiment, it may also include a preliminary category label of the target object, etc.

[0052] There are many ways to obtain semantic segmentation images. In one embodiment, it can be implemented based on machine learning technology. For example, a semantic segmentation model based on machine learning can be used to perform semantic segmentation processing on the surveillance image to obtain a semantic segmentation image. The semantic segmentation model can include a fully convolutional network (FCN), DeepLab or U-Net, etc. The following introduces a method for obtaining a semantic segmentation image:

[0053] First, collect a large amount of surveillance image data to ensure that these images cover a variety of scenes, lighting conditions, and target objects to improve the generalization ability of the model. Then, annotate these images at the pixel level and assign a semantic category label to each pixel, such as "person", "car", "road", "building", etc. The annotation work is usually done manually by professionals or assisted by semi-automated tools. Finally, preprocess the images, including normalization, cropping, resizing, etc., to make them meet the model input requirements and ensure that the format and size of the input images are consistent, which is convenient for model training and reasoning.

[0054] Secondly, select a deep learning model suitable for semantic segmentation tasks, such as FCN, DeepLab, or U-Net. Then, use the labeled dataset to train the model, and adjust the model parameters through back propagation and optimization algorithms to minimize the difference between the predicted label and the actual label. During the training process, you can use cross-validation and other techniques to evaluate the performance of the model, and adjust the model's hyperparameters as needed to ensure that the model can achieve good performance on both the training set and the validation set.

[0055] Finally, the monitoring image is input into the trained model, and the model will output the semantic category prediction result of each pixel to generate a preliminary semantic segmentation image. Then, the preliminary semantic segmentation image is post-processed, such as removing small area noise, smoothing boundaries, etc., to improve the accuracy and visual effect of the segmentation result. The purpose of post-processing is to eliminate errors and irregular areas that may be generated during the segmentation process, so that the segmentation results are more accurate and beautiful, so as to better serve the subsequent monitoring tasks.

[0056] S202 , performing feature extraction on the monitoring image, the depth image, and the semantic segmentation image respectively to obtain visual features, depth features, and semantic features of the target object.

[0057] Among them, visual features refer to features extracted from monitoring images that can describe the appearance and shape of the target object. These features generally include color features, texture features, and shape features. In one embodiment, visual features can be extracted by ordinary algorithms, such as color features can be obtained by calculating the color histogram or color moment of the image, and these features can describe the distribution of colors in the image. Texture features can be extracted by methods such as gray level co-occurrence matrix (GLCM) or local binary pattern (LBP), which can capture the texture information of pixels or regions in the image. Shape features are usually extracted by edge detection algorithms (such as Canny, Sobel, etc.), which can identify edges and contours in the image to describe the shape of the object. In addition, in one embodiment, deep learning models (such as convolutional neural networks) can also be used to automatically extract high-level visual features of the image, which can more comprehensively represent the content of the image.

[0058] For example, it can be extracted through a convolutional neural network. Specifically, the surveillance image is input into the convolutional neural network. The convolutional neural network consists of multiple convolutional layers and pooling layers, which extract local features in the image through a series of convolution kernel operations. Each convolutional layer is responsible for extracting features at different levels, from the bottom-level edge and texture features to the high-level shape and object part features. For example, the first convolutional layer may extract edge information in the image, while deeper convolutional layers can recognize more complex shapes and patterns.

[0059] Among them, depth features refer to features extracted from depth images that can describe the three-dimensional structure of the target object. The depth image records the depth information of each pixel in the scene, that is, the distance between the object surface and the camera. Depth features can be obtained by calculating the depth gradient and depth change rate of the depth image. These features can describe the depth change of the object surface. For example, the depth gradient can reflect the degree of inclination of the object surface, and the depth change rate can reflect the degree of concavity and convexity of the object surface.

[0060] In one embodiment, a deep learning model (such as a 3D convolutional neural network) can also be used to extract high-level depth features of the depth image, which can more comprehensively represent the three-dimensional structural information of the object. Depth features play an important role in tasks such as object recognition and scene reconstruction. For example, a depth image is input into a 3D convolutional neural network. The 3D convolutional neural network can simultaneously consider information in the spatial and depth dimensions by performing convolution operations on the input data using a three-dimensional convolution kernel. Each convolutional layer is responsible for extracting features at different levels, from local features at the bottom layer to global features at the top layer. For example, the first convolutional layer may extract local depth changes in the depth image, while deeper convolutional layers can identify more complex three-dimensional structural features.

[0061] Among them, semantic features refer to features extracted from semantic segmentation images that can describe the semantic information of the target object. The semantic segmentation image assigns each pixel in the image to a predefined semantic category, thereby forming a semantic map with the same resolution as the input image. Semantic features usually include semantic category labels for each pixel, which can indicate the object or scene category to which the pixel belongs. For example, in an image containing people, vehicles, and buildings, each pixel will be labeled as "person", "vehicle", or "building". In addition, regional features in semantic segmentation images can also be extracted, such as the size, shape, position, etc. of the region, which can further describe the semantic information of the object.

[0062] In the embodiment of the present application, the extraction of semantic features is mainly achieved through a deep learning model. First, the image is processed using a pre-trained semantic segmentation model (such as FCN, DeepLab, U-Net, etc.) to generate a semantic segmentation image. In the semantic segmentation process, the model assigns each pixel in the image to a semantic category, such as "person", "car", "building", etc., by performing operations such as convolution, pooling, and upsampling on the image. Then, semantic features are extracted from the semantic segmentation image, which includes the semantic category label of each pixel and the regional features composed of these labels. Regional features may include the size, shape, position of the region, and the spatial relationship between regions. In addition, the semantic features can be further enriched by calculating the statistical features of the region (such as area, perimeter, aspect ratio, etc.) and semantic context associations (such as adjacency and overlapping relationships between different regions).

[0063] S203: performing feature fusion processing on the visual features, depth features and semantic features to obtain fused features.

[0064] After obtaining the visual features, depth features and semantic features, the remote monitoring platform can fuse the three features. There are many ways to fuse features. For example, in one embodiment, the visual features, depth features and semantic features can be preprocessed to have the same dimension and format. The preprocessed features are then concatenated or added to form fused features.

[0065] In practical applications, you can also choose a suitable fusion method based on task requirements and model performance. For example, you can use a simple concatenation operation to connect visual features, depth features, and semantic features to form a high-dimensional fused feature vector. In addition, you can use a weighted fusion method to assign different weights to each feature to highlight important features. Finally, the fused features are input into the deep learning model for training and prediction to improve the performance and generalization ability of the model.

[0066] S204: Search for warning object type features corresponding to multiple warning object types in a preset object feature set, and calculate the similarity between the fusion feature and each warning object category feature.

[0067] Among them, the object feature set includes warning object type features corresponding to multiple warning object types. The warning object type is the type of object that needs warning processing. For example, in one embodiment, it can be divided into high-risk warning objects, low-risk warning objects, etc. For example, in some remote monitoring scenarios, the warning objects are people and objects that generate risks, such as flammables, dangerous goods, illegal intruders, etc. The object feature set can be expressed in the form of a database, which is pre-set in the remote monitoring platform or other storage devices for the remote monitoring platform to obtain.

[0068] The object feature set also includes at least one warning object type feature for each warning object type. Multiple warning object type features can be pre-set for each warning object type, which can be the warning object type features of each warning object type in different scenes or environments (such as different lighting conditions).

[0069] Among them, the warning object type features refer to the features used to characterize the warning object type, which can reflect the key attributes and behavior patterns of the warning object. In the remote monitoring and early warning system, the warning object type features usually include visual features, depth features, semantic features and other aspects. These features are integrated through feature fusion technology to form a comprehensive feature representation for identifying and classifying the warning object type. For example, in a monitoring image, the warning object type features can include the shape, color, texture, depth information and semantic category of the target. By extracting and fusing these features, the early warning system can more accurately identify potential dangerous objects or abnormal behaviors, so as to issue early warning signals in a timely manner.

[0070] In one embodiment, each warning object type feature may correspond to one warning object type, and each warning object type may also correspond to one or more warning object type features. Different warning object type features may correspond to the same warning object type, and the warning object type features corresponding to different warning object types may also be similar.

[0071] For example, there are 8 warning object types in the object feature set database, and the number of warning object type features is 40. The number of warning object type features corresponding to each warning object type is {3, 4, 5, 7, 4, 6, 4, 7}, among which the 4th and 5th warning object types are relatively close due to similar categories. The warning object type features of the 1st and 2nd warning object types are also similar. In addition, the 3 object type features of the 3rd warning object type are different due to different collection methods or environments, resulting in different corresponding object type features. The object type features of other different warning types are different.

[0072] In one embodiment, the object feature set may be an offline object feature database or an online object feature database. In the embodiment of the present application, a large number of sample monitoring images of known target object types (including monitoring images of different target object types in different environments or monitoring images of the same target object type in different environments) may be pre-collected or obtained from an image database, and then the sample monitoring images are subjected to depth calculation and semantic segmentation processing to obtain sample depth images and sample semantic segmentation images (the specific processing method may refer to the above-mentioned image processing), and then the sample monitoring images, sample depth images, and sample semantic segmentation images are subjected to feature extraction and fusion to obtain at least one sample fusion feature (i.e., object type feature) of the known target object type.

[0073] In one embodiment, it is assumed that the number of warning object types corresponding to all object type features in the object feature database is defined as K. For example, there are 150,000 warning object types in the object feature database, and there are 210,000 total object type features. These warning object types can be collected from open source databases or customized from professional organizations.

[0074] It should be noted that in the embodiment of the present application, the warning object type is the smallest granularity description of the object, such as the category of the warning object is accurate to the "species". The object feature database can include a large number of different warning objects, such as animals, microorganisms, plants, non-living things such as cars, liquids, etc.

[0075] In an embodiment of the present application, after finding the object type feature, a similarity calculation method can be used to calculate the similarity between the fused feature and each warning object category feature. Calculate the similarity between the fused feature and each warning object category feature. The similarity calculation method can use distance-based similarity calculation, direction-based similarity calculation, or a combination of the two. For example, Euclidean distance, Manhattan distance, cosine similarity, Pearson correlation coefficient, etc.

[0076] In the embodiment of the present application, in order to improve the calculation accuracy of the similarity, the following method can be used for calculation:

[0077]

[0078] Among them, sim(X,Y) represents the similarity between the fusion feature X and the warning object category feature Y, ωi represents the weight of the i-th feature, and sim(Xi,Yi) represents the local similarity of the i-th feature.

[0079] S205: Based on the similarity, at least one target type feature is selected from the object type features corresponding to the multiple warning object types.

[0080] Since the similarity between features represents the degree of similarity between two features, based on these similarities, L object type features can be screened out from the object type features corresponding to the K warning object types to obtain the target type features, where L can be equal to or less than K.

[0081] In one embodiment, an object type feature whose similarity is greater than a preset threshold may be selected as a target type feature, wherein the preset threshold may be set according to demand.

[0082] In one embodiment, in order to reduce the amount of calculation and improve the warning efficiency, the specific implementation method of determining L target type features can be as follows: sort the features corresponding to the K warning object types in order from large to small according to the similarity between the fusion features and the features of each object type, and then determine the object type features ranked in the top L as the target type features according to the sorting results.

[0083] S206. Identify the target warning object type to which the at least one target type feature belongs, and count the number of target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold, to obtain the number of features corresponding to each target warning object type.

[0084] In a possible embodiment, in one embodiment, after L target type features are screened, the number of warning object types corresponding to the L target type features may be L or less than L. Therefore, it is necessary to refine and select the similarity less than or equal to the set similarity threshold from the L similarities again, so as to obtain the target type features corresponding to these selected similarities and the warning object types corresponding to the target type features, and then the target type features that meet the set similarity threshold requirements in the L target type features can be obtained, and finally the number of features of the warning object types to which these target type features that meet the set similarity threshold requirements belong can be obtained, and the number of features can also be understood as the frequency of occurrence of the same warning object type in the set. For example, in one embodiment, assuming that L is 8, there are 4 types of warning object types to which the 8 target type features belong, assuming that they are warning object 1, warning object 2, warning object 3, and warning object 4. Among these 8 target type features, 7 target type features meet the corresponding set similarity threshold conditions, that is, they are greater than the preset similarity threshold, which are 4 target type features corresponding to warning object 1 and 2 features corresponding to warning object 3. Therefore, the number of target type features 4 can be used as the number of features corresponding to warning object 1, and the number of target type features 2 can be used as the number of features corresponding to warning object 3.

[0085] S207: Determine the warning object type of the target object from the target warning object types according to the number of features corresponding to each target warning object type.

[0086] After obtaining the number of features of each target warning object type that meets the set similarity threshold, the remote monitoring platform can determine the warning object type of the target object in the monitoring image based on the number.

[0087] For example, in one embodiment, the target warning object type with the largest number of features may be selected as the final warning object type of the target object.

[0088] In one embodiment, the remote monitoring warning system can better adapt to different monitoring environments and scene changes. For example, under different lighting conditions or when the target distance changes, by adjusting the number of features, the target warning object type can be more accurately screened, reducing false alarms and missed alarms. The number of features can be dynamically adjusted according to the image acquisition parameters of the remote camera, so that the system can better adapt to different monitoring environments and scene changes. Specifically, refer to Figure 3 , determining the warning object type of the target object from the target warning object types according to the number of features corresponding to each target warning object type, may include the following steps:

[0089] S301, obtaining image acquisition parameters of a remote camera.

[0090] Among them, the image acquisition parameters are shooting parameters when the remote camera shoots the monitoring image or picture, etc., and may include: optical parameters, environmental parameters, image processing parameters, etc. Among them, the optical parameters may include: resolution, frame rate, aperture, focal length, etc. Environmental parameters include shooting angle (referring to the angle of the camera relative to the object being photographed, including horizontal angle and vertical angle. Different shooting angles will affect the position and shape of the object in the image), light conditions (referring to the light intensity and light type during shooting, such as natural light, artificial light, strong light, weak light, etc. Light conditions directly affect the brightness, contrast and color saturation of the image), shooting distance (referring to the distance between the camera and the object being photographed. Different shooting distances will affect the size and clarity of the object in the image). Image processing parameters may include exposure, noise reduction, etc. Specific image acquisition parameters can be selected according to the actual scene. For example, in one embodiment, shooting angle, shooting distance, etc. can be selected.

[0091] S302: adjusting the number of features corresponding to each target warning object type according to the image acquisition parameters to obtain the adjusted number of features for each target warning object type.

[0092] In one embodiment, based on the acquired image acquisition parameters, the number of features corresponding to each target warning object type can be adjusted. For example, a preset weight relationship set can be set, which contains the corresponding relationship between different image acquisition parameters and feature quantity adjustment weights. According to the current image acquisition parameters, the corresponding adjustment weights are obtained from the weight relationship set, and the number of features corresponding to each target warning object type is weighted to obtain the adjusted number of features. Finally, based on the adjusted number of features, the warning object type of the target object is determined from the target warning object types, and the target warning object type with the highest adjusted number of features is selected as the final warning object type.

[0093] In one embodiment, in order to improve the accuracy and stability of object recognition and early warning, the influence of image parameters such as shooting angle and distance on the depth image can be taken into account. Therefore, the depth image can be corrected according to the image acquisition parameters, and the number of features can be adjusted based on the corrected depth image. Specific methods may include:

[0094] Correcting the depth image according to the image acquisition parameters to obtain a corrected depth image;

[0095] Calculate the average depth value information corresponding to the corrected depth image;

[0096] According to the average depth value information, the number of features corresponding to each target warning object type is adjusted to obtain the adjusted number of features for each target warning object type.

[0097] In one embodiment, the adjustment weight corresponding to each target warning object type is obtained from a preset weight relationship set according to the average depth value; the number of features corresponding to each target warning object type is weighted according to the adjustment weight corresponding to each target warning object type to obtain the adjusted number of features for each target warning object type.

[0098] In one embodiment, this can be achieved by performing geometric transformation, grayscale transformation and other operations on the depth image based on image acquisition parameters such as shooting angle, distance, etc., so as to eliminate the influence of image acquisition parameters on the depth image. Then, the average depth value information corresponding to the corrected depth image is calculated, that is, the depth values ​​of all pixels in the corrected depth image are averaged to obtain the average depth value. Then, according to the average depth value information, the number of features corresponding to each target warning object type is adjusted to obtain the adjusted number of features for each target warning object type. Specifically, the adjustment weight of the number of features of each target warning object type can be determined based on the correspondence between the average depth value and the preset depth value range, and then the number of features can be weighted adjusted.

[0099] Taking the shooting angle and lighting conditions, and the scenes of day and night as examples, the process of determining the number of features and object types is introduced:

[0100] First, obtain the current image acquisition parameters, including shooting angle and light conditions. Assume that during the day, the light is sufficient and the shooting angle is 0 degrees (normal angle); at night, the light is dim and the shooting angle is adjusted to 30 degrees.

[0101] According to the acquired image acquisition parameters, the depth image is corrected. In the daytime, when the light is sufficient, the depth image does not need special correction; at night, when the light is dim, the depth image may need to be enhanced to improve the visibility of the depth information. Assume that after correction, the average depth value of the depth image is 1.5 meters.

[0102] Calculate the average depth value information of the corrected depth image. Assume that during the day, the average depth value is 1.2 meters; at night, the average depth value is 1.8 meters.

[0103] According to the average depth value information, the number of features corresponding to each target warning object type is adjusted. Assume that the initial number of features for vehicles is 100 and the initial number of features for pedestrians is 80. During the day, the average depth value is 1.2 meters. According to the preset depth value range, the feature number adjustment weight for vehicles is 1.2, and the feature number adjustment weight for pedestrians is 1.0. Therefore, the adjusted feature number of vehicles is 100×1.2=120, and the adjusted feature number of pedestrians is 80×1.0=80. At night, the average depth value is 1.8 meters, the feature number adjustment weight for vehicles is 1.5, and the feature number adjustment weight for pedestrians is 1.2. Therefore, the adjusted feature number of vehicles is 100×1.5=150, and the adjusted feature number of pedestrians is 80×1.2=96.

[0104] Based on the adjusted feature number, the warning object type of the target object is determined from the target warning object type. During the day, the adjusted feature number of vehicles is 120, and the adjusted feature number of pedestrians is 80, so the warning object type of the target object is vehicle. At night, the adjusted feature number of vehicles is 150, and the adjusted feature number of pedestrians is 96, so the warning object type of the target object is still vehicle.

[0105] The embodiment of the present application can eliminate the influence of image acquisition parameters (such as shooting angle, light, etc.) on depth information by correcting the depth image, so that the depth image more accurately reflects the real depth information of the target object. This helps to improve the accuracy of subsequent feature extraction and reduce feature extraction errors caused by image acquisition parameters. In addition, by calculating the average depth value information of the corrected depth image, the overall distribution of the target object in the depth dimension can be obtained, which provides a basis for the subsequent adjustment of the number of features according to the average depth value, making the adjustment of the number of features more scientific and reasonable, such as making the number of features of different target warning object types more consistent with their performance in actual scenarios. This helps to improve the accuracy of warning processing, reduce false alarms and missed alarms, and enable the remote monitoring system to more accurately identify and warn of potential security threats.

[0106] S303: Determine the warning object type of the target object from the target warning object types based on the adjusted feature quantity of each target warning object type.

[0107] In one embodiment, the highest number of adjusted feature quantities in each warning object type is identified; and the warning object type of the target object is determined from the target warning object types corresponding to the highest number. For example, when the highest number exists at 1 and is greater than a preset number value, the warning object type corresponding to the highest number can be directly determined as the final warning object type of the target object.

[0108] In the embodiment of the present application, if the maximum number is less than the preset number value, then the maximum number cannot meet the set conditions. In this case, a preset method or the like can be used for processing, or a direct feedback of recognition failure can be given.

[0109] Since the highest target number in the number of features corresponding to the target object type may be one or more, that is, the highest target number has the same situation, then the recognition result can be one or more. For example, there are 4 target warning object types that meet the similarity threshold, and the number of features corresponding to each target warning object type is 3, 4, 5, and 5 respectively. At this time, the highest number of features is 5, and it is greater than the set number threshold. At this time, there are two warning object types corresponding to the highest number. Since the target object in the middle usually needs to return a unique result, it is necessary to determine a final warning object type as the target object. At this time, in order to improve the accuracy of warning object type recognition, the final warning object type of the target object is determined by determining the average similarity corresponding to the target warning object object category to which each highest number belongs, such as taking the object type corresponding to the maximum level value from the average similarity value corresponding to the highest number as the final warning object type of the target object.

[0110] In one embodiment, if there are multiple maximum numbers, and the maximum number is not greater than the predetermined number value, in this case, it means that there is no object type that meets the requirements, and feedback of recognition failure can be given. Alternatively, in order to improve the stability and accuracy of the warning, in one embodiment, the warning object type can be identified based on a separate semantic feature. Specifically:

[0111] If there are multiple maximum numbers, and the maximum number is not greater than a predetermined value, similarity calculation is performed based on the semantic feature and the sample semantic feature of the target warning object corresponding to the maximum number to obtain the semantic similarity between the semantic feature and the sample semantic feature;

[0112] The warning object type of the target object is determined from the target warning object types corresponding to the highest number according to the semantic similarity.

[0113] For example, when the highest number is 3, the corresponding warning object types are warning object 1, warning object 2, and warning object 3. At this time, if these 3 highest numbers are all less than the preset number threshold, then the semantic similarity between the extracted semantic features and the sample semantic features in the object type features can be calculated based on the semantic features extracted above. The specific calculation method can refer to the similarity calculation method mentioned above. Then, the warning object corresponding to the highest semantic similarity is selected, such as warning object 1 as the warning object type of the target object, thereby improving the success rate and stability of the warning object.

[0114] S208: Trigger a corresponding monitoring and warning processing operation according to the warning object type of the target object.

[0115] After determining the warning object type of the target object, the remote monitoring platform triggers the corresponding monitoring warning processing operations, including:

[0116] Real-time alarm: The system will immediately send an alarm signal to the monitoring center or relevant personnel to inform them of abnormal situations. The alarm signal can be sent through sound, light, text message, email and other methods to ensure that relevant personnel can receive early warning information in time.

[0117] Recording and storage: The system will record and store relevant surveillance images and data for subsequent analysis and processing. These records can be used as evidence for accident investigation, responsibility determination, etc.

[0118] Automatic tracking: The system can automatically track the movement trajectory of the target object and continuously monitor its behavior until it is confirmed that the danger has been eliminated or appropriate measures have been taken.

[0119] Linkage with other systems: The system can be linked with other security systems (such as access control systems, fire protection systems, etc.) to take further measures, such as closing the access control of relevant areas, activating fire protection equipment, etc.

[0120] In one embodiment, in order to improve the accuracy and effect of the warning, the warning level corresponding to the warning object type to which the target object belongs may also be determined; and monitoring and warning processing operations are performed according to the warning level.

[0121] For example, a corresponding warning level can be set for each warning object type according to preset rules and standards. Warning levels can be divided into multiple levels, such as low risk, medium risk, and high risk. For example, in a remote traffic monitoring scenario, high-risk warning object types may include "speeding vehicles" or "vehicles running red lights", while low-risk warning object types may include "normal driving vehicles".

[0122] From the above, it can be seen that the embodiment of the present application provides a warning processing method for remote monitoring. This solution realizes warning object recognition of monitoring images by fusing multimodal data (fusion of monitoring images, depth images and semantic segmentation images), which can improve the accuracy of warning object recognition, thereby improving the accuracy and reliability of remote monitoring warnings.

[0123] Furthermore, the embodiments of the present application can comprehensively identify warning objects based on the fusion feature similarity and feature quantity, which can greatly improve the accuracy of warning object identification, and further improve the accuracy and reliability of remote monitoring warnings.

[0124] In one embodiment, based on the above-mentioned remote monitoring early warning processing method, a remote monitoring early warning processing device is also provided. The early warning processing device can be integrated in the remote monitoring platform. Figure 4 , which is a warning processing device for remote monitoring, specifically including

[0125] The acquisition unit 401 is used to acquire a monitoring image collected by a remote camera and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object; and perform feature extraction on the monitoring image, the depth image and the semantic segmentation image to obtain visual features, depth features and semantic features of the target object;

[0126] A fusion unit 402 is used to perform feature fusion processing on the visual features, depth features and semantic features to obtain fusion features;

[0127] A similarity unit 403 is used to search for object type features corresponding to multiple warning object types in a preset object feature set, and calculate the similarity between the fusion feature and each warning object category feature;

[0128] A selection unit 404, configured to select at least one target type feature from the object type features corresponding to the plurality of warning object types based on the similarity;

[0129] The identification unit 405 is used to identify the target warning object type to which the at least one target type feature belongs, and count the number of target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold, to obtain the number of features corresponding to each target warning object type;

[0130] An object determination unit 406, configured to determine the warning object type of the target object from the target warning object types according to the number of features corresponding to each target warning object type;

[0131] The warning unit 407 is used to trigger a corresponding monitoring warning processing operation according to the warning object type of the target object.

[0132] In one embodiment, the object determination unit 406 is specifically configured to:

[0133] Get the image acquisition parameters of the remote camera;

[0134] Adjusting the number of features corresponding to each target warning object type according to the image acquisition parameters to obtain an adjusted number of features for each target warning object type;

[0135] Based on the adjusted number of features of each target warning object type, a warning object type of the target object is determined from the target warning object types.

[0136] In one embodiment, the object determination unit 406 is specifically used to: correct the depth image according to the image acquisition parameters to obtain a corrected depth image; calculate the average depth value information corresponding to the corrected depth image; and adjust the number of features corresponding to each target warning object type according to the average depth value information to obtain the adjusted number of features for each target warning object type.

[0137] In one embodiment, the object determination unit 406 is specifically used to: obtain the adjustment weight corresponding to each target warning object type from a preset weight relationship set according to the average depth value; perform weighted processing on the number of features corresponding to each target warning object type according to the adjustment weight corresponding to each target warning object type to obtain the adjusted number of features for each target warning object type.

[0138] In one embodiment, the object determination unit 406 is specifically configured to: identify the highest number of adjusted feature quantities in each warning object type; and determine the warning object type of the target object from the target warning object types corresponding to the highest number.

[0139] In one embodiment, the object determination unit 406 is specifically configured to: if the highest number exists and is greater than a preset number threshold, determine that the target object type corresponding to the highest number is the warning object type of the target object;

[0140] If there are multiple maximum numbers and they are greater than a preset number threshold, then the average similarity value corresponding to the maximum number is obtained; and the warning object type of the target object is determined according to the average similarity value.

[0141] In one embodiment, the object determination unit 406 is specifically used to: if there are multiple maximum quantities and the maximum quantity is not greater than a predetermined quantity value, perform similarity calculation processing based on the semantic feature and the sample semantic feature of the target warning object corresponding to the highest quantity to obtain the semantic similarity between the semantic feature and the sample semantic feature; determine the warning object type of the target object from the target warning object types corresponding to the highest quantity based on the semantic similarity.

[0142] In one embodiment, the object determination unit 406 is specifically configured to select the target warning object type corresponding to the highest semantic similarity as the warning object type of the target object.

[0143] In one embodiment, the warning unit 407 is specifically used to: determine the warning level corresponding to the warning object type to which the target object belongs; and perform monitoring and warning processing operations according to the warning level.

[0144] It is understandable that the functions of the functional modules of the image device described in the embodiment of the present application can be based on the above

[0145] The specific implementation of the method in the method embodiment can refer to the relevant description of the above method embodiment for its specific implementation process, which will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated.

[0146] Figure 5 : is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 500 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 521 (for example, one or more processors) and a memory 532, and one or more storage media for storing application programs or data (for example, one or more mass storage devices). Among them, the memory 532 and the storage medium may be temporary storage or permanent storage. The program stored in the storage medium may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device. Furthermore, the central processing unit 521 may be configured to communicate with the storage medium to execute a series of instruction operations in the storage medium on the computer device 500.

[0147] The steps performed by the computer device in the above embodiment can be based on the Figure 5 The computer device structure shown in the figure. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0148] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0149] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0151] On one hand, an embodiment of the present application provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method provided in one aspect of the embodiment of the present application.

[0152] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and any variations thereof are intended to cover

[0153] Non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules that are not listed, or may optionally include other steps and units that are inherent to these processes, methods, devices, products or equipment.

[0154] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0155] The method and related apparatus provided by the embodiment of the present application are described with reference to the method flow chart and / or structural diagram provided by the embodiment of the present application. Specifically, each process and / or box in the method flow chart and / or structural diagram, as well as the combination of the processes and / or boxes in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the process. Figure 1 A process or multiple processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A flow or multiple flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.

[0156] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A remote monitoring early warning processing method, characterized in that: The early warning processing method is applicable to a remote monitoring platform, and the method comprises: Acquire a monitoring image captured by a remote camera and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object; perform feature extraction on the monitoring image, the depth image and the semantic segmentation image respectively to obtain visual features, depth features and semantic features of the target object; Performing feature fusion processing on the visual features, depth features and semantic features to obtain fused features; Searching for object type features corresponding to multiple warning object types in a preset object feature set, and calculating the similarity between the fusion feature and each warning object category feature; Based on the similarity, at least one target type feature is selected from the object type features corresponding to the plurality of warning object types; Identify the target warning object type to which the at least one target type feature belongs, and count the number of target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold, to obtain the number of features corresponding to each target warning object type; Get the image acquisition parameters of the remote camera; Adjusting the number of features corresponding to each target warning object type according to the image acquisition parameters to obtain an adjusted number of features for each target warning object type; Determining the warning object type of the target object from the target warning object types based on the adjusted number of features of each target warning object type; The corresponding monitoring warning processing operation is triggered according to the warning object type of the target object.

2. The early warning processing method according to claim 1, characterized in that: The number of features corresponding to each target warning object type is adjusted according to the image acquisition parameters to obtain the adjusted number of features of each target warning object type, including: Correcting the depth image according to the image acquisition parameters to obtain a corrected depth image; Calculate the average depth value information corresponding to the corrected depth image; According to the average depth value information, the number of features corresponding to each target warning object type is adjusted to obtain the adjusted number of features for each target warning object type.

3. The early warning processing method according to claim 2, characterized in that: According to the average depth value information, the number of features corresponding to each target warning object type is adjusted to obtain the adjusted number of features of each target warning object type, including: Obtaining an adjustment weight corresponding to each target warning object type from a preset weight relationship set according to the average depth value; The number of features corresponding to each target warning object type is weighted according to the adjustment weight corresponding to each target warning object type to obtain the adjusted number of features for each target warning object type.

4. The early warning processing method according to any one of claims 2-3, characterized in that: Determining the warning object type of the target object from the target warning object types based on the adjusted number of features of each target warning object type includes: Identify the highest number of adjusted features in each alert object type; The warning object type of the target object is determined from the target warning object types corresponding to the highest number.

5. The early warning processing method according to claim 4, characterized in that: Determining the warning object type of the target object from the target warning object types corresponding to the highest number includes: If the highest number exists and is greater than a preset number threshold, determining that the target object type corresponding to the highest number is the warning object type of the target object; If there are multiple maximum numbers and they are greater than a preset number threshold, then the average similarity value corresponding to the maximum number is obtained; and the warning object type of the target object is determined according to the average similarity value.

6. The early warning processing method according to claim 4, characterized in that: The object type feature includes at least a sample semantic feature, and determining the warning object type of the target object from the target warning object types corresponding to the highest number, further comprising: If there are multiple maximum numbers, and the maximum number is not greater than a predetermined value, similarity calculation is performed based on the semantic feature and the sample semantic feature of the target warning object corresponding to the maximum number to obtain the semantic similarity between the semantic feature and the sample semantic feature; The warning object type of the target object is determined from the target warning object types corresponding to the highest number according to the semantic similarity.

7. The early warning processing method according to claim 6, characterized in that: Determining the warning object type of the target object from the target warning object types corresponding to the highest number according to the semantic similarity includes: The target warning object type corresponding to the highest semantic similarity is selected as the warning object type of the target object.

8. The early warning processing method according to claim 7, characterized in that: The corresponding monitoring warning processing operation is triggered according to the warning object type of the target object, including: Determine the warning level corresponding to the warning object type to which the target object belongs; A monitoring and warning processing operation is performed according to the warning level.

9. A remote monitoring early warning processing device, characterized in that: The early warning processing device is suitable for a remote monitoring platform, and the early warning processing device includes: An acquisition unit is used to acquire a monitoring image collected by a remote camera and its corresponding depth image and semantic segmentation image, wherein the monitoring image includes a target object; and perform feature extraction on the monitoring image, the depth image and the semantic segmentation image to obtain visual features, depth features and semantic features of the target object; A fusion unit, used for performing feature fusion processing on the visual features, depth features and semantic features to obtain fusion features; A similarity unit, used to search for object type features corresponding to multiple warning object types in a preset object feature set, and calculate the similarity between the fusion feature and each warning object category feature; A selection unit, configured to select at least one target type feature from the object type features corresponding to the plurality of warning object types based on the similarity; an identification unit, configured to identify the target warning object type to which the at least one target type feature belongs, and to count the number of target type features in the target warning object type whose similarity with the fusion feature is greater than a preset similarity threshold, to obtain the number of features corresponding to each target warning object type; An object determination unit is used to obtain image acquisition parameters of a remote camera; adjust the number of features corresponding to each target warning object type according to the image acquisition parameters to obtain the adjusted number of features of each target warning object type; and determine the warning object type of the target object from the target warning object types based on the adjusted number of features of each target warning object type; The warning unit is used to trigger a corresponding monitoring warning processing operation according to the warning object type of the target object.

Citation Information

Patent Citations

  • Target retrieval method and device

    CN111581423A

  • Target object recognition method and device and computer equipment

    CN113837174A

  • Method and device for identifying position relation of objects in image and storage medium

    CN114820785A