Method and system for detecting security device

By improving the architecture of the SparseInst network, the problem of deploying instance segmentation model in embedded systems is solved, and the security device detection is realized when computing power and development board support package software are limited, which improves the detection accuracy and efficiency.

CN119992520APending Publication Date: 2025-05-13VIA TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510060148.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deploy an instance segmentation model in embedded systems, especially under the limitations of computing power and development board support package software, and it is impossible to accurately and efficiently detect whether the driver and passengers are equipped with safety equipment.

Method used

By improving the architecture of the SparseInst network, including the use of lightweight backbone modules such as MobileNetV2, removing the pyramid pooling module and coordinate concatenation operations, and introducing heat map loss functions, the instance segmentation model is optimized to adapt to the resource limitations of embedded systems.

Benefits of technology

The possibility of deploying an instance segmentation model in an embedded system is realized, the accuracy and efficiency of security device detection is improved, and the demand for computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992520A_ABST
    Figure CN119992520A_ABST
Patent Text Reader

Abstract

The invention provides a method and system for detecting safety equipment, and the method comprises the steps: obtaining an input image containing a driver and passengers, inputting the input image into a pre-trained instance segmentation model to obtain an instance segmentation result outputted by the instance segmentation model, calculating attribute data related to the safety equipment based on the instance segmentation result, and judging whether the driver is equipped with safety equipment or not based on the attribute data. According to the invention, the accuracy requirement of security equipment detection can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image analysis technology, and in particular to a detection method and system for security equipment. Background Art

[0002] Passive safety equipment is applied to various vehicles (such as cars, motorcycles, trains, airplanes and yachts), engineering vehicles (such as forklifts, excavators and cranes) and amusement facilities (such as roller coasters, pirate ships and free fall facilities), such as seat belts, helmets, life jackets and goggles. Its main appeal is to reduce the damage to the drivers and passengers in the event of an accident. Taking the automotive industry as an example, the application of artificial intelligence technology in Advanced Driving Assistant System (ADAS) has made ADAS more and more popular. Among them, the detection of whether the drivers and passengers (especially the driver) are equipped with safety equipment is one of the focuses of ADAS. Its purpose is to quickly and accurately detect the behavior of the drivers and passengers who are not equipped with safety equipment, so as to assist in issuing alarms in time. Facts have proved that the damage caused by the behavior of the drivers and passengers who are not equipped with safety equipment in various accidents cannot be underestimated. In various accidents, the failure to equip safety equipment will cause the drivers and passengers to suffer more serious or even fatal injuries. Therefore, the research and development of safety equipment detection is of great practical significance.

[0003] At present, object detection technology is usually used to detect unsafe behaviors of drivers and passengers. For example, in the application of detecting phone calls, the position of the mobile phone in the image is detected. However, the known object detection technology cannot meet the practical needs of safety equipment detection, because the object detection output is a rectangular bounding box that surrounds the object, but the properties of some safety equipment are not suitable for description only by a rectangular bounding box. Take the seat belt as an example. Because the seat belt occupies too small a proportion of the rectangular bounding box and there is too much background information, the false detection rate is too high, which makes it difficult to meet the accuracy requirements in practice.

[0004] On the other hand, instance segmentation technology involves segmenting the regions in the image at the pixel level and identifying the instance categories of each region to identify the location and range of specific objects in the image. Existing instance segmentation models, such as the SparseInst network, may enjoy real-time computing efficiency and fairly high accuracy on personal computers or server computers. However, due to the limitation of computing resources, instance segmentation models are difficult to deploy in ADAS or other similar embedded systems without sacrificing computing efficiency and accuracy. In addition to computing power, the board support package (BSP) software used by some embedded systems is also limited in its support for neural networks, and it is difficult to support or not support certain special types of neural network layers at all, such as downsampling layers.

[0005] Therefore, a detection method and system for security equipment is needed to overcome the above technical challenges. Summary of the invention

[0006] An embodiment of the present invention provides a safety device detection method implemented by a computer system, which includes obtaining an input image containing a driver and an occupant, inputting the input image into a pre-trained instance segmentation model to obtain an instance segmentation result output by the instance segmentation model, calculating attribute data associated with the safety device based on the instance segmentation result, and determining whether the driver and the occupant are equipped with safety equipment based on the attribute data.

[0007] One embodiment of the present invention also provides a safety device detection system, which includes a processing device that executes the steps of obtaining an input image containing a driver and an occupant, inputting the input image into a pre-trained instance segmentation model to obtain an instance segmentation result output by the instance segmentation model, calculating attribute data associated with the safety device based on the instance segmentation result, and determining whether the driver and the occupant are equipped with safety equipment based on the attribute data. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present invention will be better understood through the following description of typical embodiments in conjunction with the accompanying drawings. In addition, it should be understood that in the flowchart of the present invention, the execution order of each block can be changed, and / or some blocks can be changed, deleted or merged.

[0009] Figure 1 is a flow chart of a detection method of a security device according to an embodiment of the present invention.

[0010] Figure 2 A typical SparseInst network architecture is shown.

[0011] Figure 3 4 is an architecture diagram of an improved SparseInst network according to an embodiment of the present invention.

[0012] Figure 4A is a flow chart of more detailed steps for obtaining an input image containing a driver and an occupant in one embodiment.

[0013] Figure 4B is a schematic diagram of more detailed steps of obtaining an input image containing a driver and an occupant in one embodiment.

[0014] Figure 5 As an example, a minimum circumscribed rectangle 500 and a pixel area A of a seat belt are shown, wherein the minimum circumscribed rectangle has a height h, a width w, and a rotation angle θ.

[0015] Figure 6 is a system block diagram of a system for implementing a detection method for a security device according to an embodiment of the present invention.

[0016] The reference numerals are described as follows:

[0017] 10: Method; S101-S104: Steps; 20: SparseInst network; 21: Backbone module; 22: Encoder module; 23: Decoder module; 201: Input image; 202: Backbone; 203: Multi-scale features; 204: Pyramid pooling module; 205-209: Convolution; 210: Feature map; 211: Convolution; 212: Feature map; 213: Coordinate information; 214: Concatenation; 215: Encoded features; 216: Mask branch; 217: Instance branch; 218: Mask features; 219: Instance information; 220: Core features; 221: Prediction category confidence; 222: Prediction object confidence; 30: Improved SparseInst network; 31: Backbone module; 32: Encoder module; 33: decoder module; 301: input image; 302: backbone; 303: multi-scale features; 305-309: convolution; 310: feature map; 311: convolution; 315: encoded features; 316: mask branch; 317: instance branch; 318: mask features; 319: instance information; 320: core features; 321: predicted category confidence; 322: predicted object confidence; S401-S403: steps; 410: driver and occupants; 411: original image; 412: bounding box; 413: local image; 414: input image; 500: minimum enclosing rectangle; h: height; w: width; θ: rotation angle; A: pixel area; 60: system; 601: photographic device; 602: processing device; 603: output device. DETAILED DESCRIPTION

[0018] The following description lists various embodiments of the present invention, but is not intended to limit the content of the present invention. The actual scope of the invention is defined by the claims.

[0019] In the various embodiments listed below, the same reference numerals will be used to represent the same or similar elements or components.

[0020] Serial numbers in this specification and the claims, such as "first", "second", etc., are only for the convenience of explanation and have no sequential relationship with each other.

[0021] The following description of the embodiments of the apparatus or system is also applicable to the embodiments of the method, and vice versa.

[0022] In order to better describe the embodiments of the present invention, some of the most critical terms involved therein are first defined as follows.

[0023] Safety equipment: any equipment that can be applied to various vehicles (such as cars, motorcycles, trains, airplanes and yachts), engineering vehicles (such as forklifts, excavators and cranes) and amusement facilities (such as roller coasters, pirate ships and free fall facilities) to reduce the harm to the drivers and passengers in the event of an accident, such as seat belts, helmets, life jackets and goggles.

[0024] Instance segmentation: A technique in the field of computer vision that involves segmenting regions in an image at the pixel level and identifying the instance category of each region to identify the location and range of a specific object in the image.

[0025] Attribute data: associated with the security device and used to describe the visible physical properties and / or characteristics of the security device, such as the location, size, shape, orientation, color, etc. of the security device in the image.

[0026] Figure 1 FIG. 1 is a flow chart of a detection method 10 for a security device according to an embodiment of the present invention. Figure 1 As shown, method 10 may include steps S101 - S104 .

[0027] In step S101, an input image containing a driver and passenger is obtained. For example, the input image may be an image obtained by photographing the driver's seat in a car and then undergoing certain preprocessing (such as geometric transformation, normalization, smoothing, image enhancement and / or cropping). However, the various embodiments of the present invention do not limit the application scenario to the detection of car driving, nor do they limit the preprocessing operations performed on the captured image. In some embodiments, the input image may be derived from the application scenarios of various vehicles (such as cars, motorcycles, trains, airplanes and yachts), engineering vehicles (such as forklifts, excavators and cranes) or amusement facilities (such as roller coasters, pirate ships and free fall facilities).

[0028] In step S102, the input image is input into a pre-trained instance segmentation model to obtain an instance segmentation result output by the instance segmentation model. The instance segmentation model involves segmenting the regions in the image at the pixel level and identifying the instance category of each region, thereby identifying the location and range of a specific object in the image. In the instance segmentation result, the security device is used as an instance object, and the pixels in the region where it is located are marked in the form of a category label, for example, "1" represents the region of the security device and "0" represents the background region, so as to further identify the security device in the input image.

[0029] In step S103, based on the instance segmentation result, attribute data associated with the safety device is calculated. For example, the attribute data may include relevant information such as the position, size, shape, direction, color, etc. of the safety device in the image, and the present invention is not limited to this. The safety device can be applied to various vehicles (such as cars, motorcycles, trains, airplanes and yachts), engineering vehicles (such as forklifts, excavators and cranes) and amusement facilities (such as roller coasters, pirate ships and free fall facilities), and any equipment that reduces the harm to the driver and passengers in the event of an accident, such as seat belts, helmets, life jackets and goggles.

[0030] In step S104, based on the attribute data, it is determined whether the driver or passenger is equipped with safety equipment. Conceptually, if the driver or passenger is indeed equipped with safety equipment, the safety equipment will have a reasonable position and range configuration in the input image. Therefore, it is possible to determine whether the driver or passenger is equipped with safety equipment by analyzing the degree of consistency between the relevant information in the attribute data and the expected position and range.

[0031] In one embodiment, the instance segmentation model is implemented by an improved SparseInst network. The typical SparseInst network has the characteristics of large scale, complex structure, numerous parameters, special network type, etc., which is not conducive to deployment on computer systems (such as embedded systems) with limited computing resources and supported network types. Due to the above problems, some further embodiments of the present invention involve improvements to the SparseInst network to expand its application scope. Figure 2 and Figure 3 Various possible improvements to the SparseInst network proposed by the present invention are described.

[0032] Figure 2 A typical SparseInst network 20 architecture is shown. Figure 2 As shown, the SparseInst network 20 includes a backbone module 21, an encoder module 22 and a decoder module 23, and these modules will be introduced in sequence below.

[0033] The main function of the backbone module 21 is to extract image features. Figure 2 As shown, the backbone module 21 takes an input image 201 with a batch number B, a channel number 3 (or "3-dimensional"), a width W, and a height H (marked as (B, 3, W, H) in the figure), and uses a backbone 202 to extract a multi-scale feature 203 of the image 201. The multi-scale feature 203 includes feature maps of different levels such as C5, C4, and C3. These levels correspond to feature extraction layers of different depths. The C3 feature map contains a lower level of feature representation, which usually has higher spatial information and less semantic information. The C4 feature map contains a medium-level feature representation, which is lower in spatial information than the C3 feature map, but contains richer semantic information. The C5 feature map contains a higher level of feature representation, which has the least spatial information, but has the richest semantic information. In Figure 2 In the example, the sizes (i.e., width and height) of the C5, C4, and C3 feature maps are 1 / 32, 1 / 16, and 1 / 8 of the input image 201, respectively, and the number of channels are 256, 1024, and 512, respectively. The typical backbone 202 is implemented as a residual neural network (ResNet), which has a large number of parameters and requires considerable storage and computing resources in both the training and inference stages.

[0034] The main function of the encoder module 22 is to enhance contextual information and fuse multi-scale features 203. Figure 2 As shown, the encoder module 22 takes the multi-scale features 203 output by the backbone module 21 as input, wherein the C5, C4 and C3 feature maps are respectively subjected to the operations of the pyramid pooling module (PPM) 204, convolution 205 and convolution 206. The main function of the pyramid pooling module 204 is to expand the receptive field and fuse features of different scales to enhance the ability of SparseInst to perceive global information. The result of C5 after the pyramid pooling module will be convolved 207 to output a feature map with a size enlarged by 4 times (i.e., from 1 / 32 to 1 / 8 of the input image 201), and will be enlarged by 2 times and added to the result of the C4 feature map after convolution 205 (indicated by ⊕ in the figure); the result of the addition will be convolved 208 to output a feature map with a size enlarged by 2 times (i.e., from 1 / 16 to 1 / 8 of the input image 201), and will be added to the result of the C3 feature map after convolution 206, and then convolved 209 to output a feature map with a size of 1 / 8 of the input image 201. Next, the outputs of convolution 207, convolution 208, and convolution 209 are added to fuse the features to obtain a feature map 210 with a batch number B, a channel number 256, a width W / 8, and a height H / 8. The feature map 210 will be convolved 211 again to obtain a feature map 212. The feature map 212 is concatenated 214 with the coordinate information of the same size (i.e., x-coordinate and y-coordinate) to obtain a feature map of batch number B, channel number 258, width W / 8, and height H / 8, i.e., encoded feature 215, as the output of the encoder module 22. It is worth mentioning that, unlike Figure 2 As shown in FIG. 2 , some literature lists the feature map 212 as the output of the encoder module 22, and lists the concatenation 214 of the feature map 212 and the coordinate information 213 as the operation of the decoder module 23. However, it can be determined that the typical feature map 212 involves a coordinate concatenation operation, whether in the encoder module 22 or the decoder module 23. Since the positions of different instances are also different, the introduction of the coordinate information 213 helps the model distinguish between instances.

[0035] The decoder module 23 mainly utilizes a sparse set of instance activation maps to highlight distinct instance pixels and suppress useless pixels, so as to facilitate the classification and segmentation of foreground objects. Figure 2 As shown, the decoder module 23 includes a mask branch 216 and an instance branch 217. Both branches use the encoded features 215 output by the encoder module 22 as input and output mask features 218 and instance information 219 respectively. The mask branch 216 is mainly composed of a sequence of convolutions to derive the mask features 218 from the encoded features 215. The mask features 218 are abbreviated as "Mask" in some literatures, which indicates the probability of a certain instance belonging to a certain pixel, and can therefore be used to generate instance segmentation results, that is, to determine whether the pixel belongs to a certain instance. The instance branch 217 is also mainly composed of a sequence of convolutions to derive instance information 219 from the encoded features 215. The instance information 219 includes a core feature 220, a predicted category confidence 221, and a predicted object confidence 222. The core feature 220 is usually abbreviated as "Kernel", which includes a feature representation of the instance, that is, a 128-dimensional vector of num_mask potential instances. The predicted class confidence 221 is abbreviated as "pred_logits" in some literature, which refers to a tensor in which each element represents the model's predicted probability score for each pixel belonging to a certain instance, and thus contains the confidence scores of num_mask potential instances corresponding to num_class categories. The predicted object confidence 222 is abbreviated as "pred_objectness" in some literature to indicate the probability score of whether num_mask potential instances exist.

[0036] It should be noted that the mask feature 218 output by the mask branch 216 is not the final predicted mask of the model, and it is necessary to combine the instance information 219 provided by the instance branch 217 to obtain the instance segmentation result. Specifically, the instance segmentation result includes a predicted mask and its corresponding confidence, wherein the predicted mask is obtained by performing matrix multiplication and square root of the mask feature 218 and the core feature 220, and the confidence is obtained by multiplying the predicted category confidence 221 and the predicted object confidence 222.

[0037] The training phase of the SparseInst network 20 involves optimizing the SparseInst network 20 using a training dataset and a loss function. The training dataset includes multiple training data, each of which includes a training image and its corresponding labeled data. The training images can be collected by the developer and then obtained after some preprocessing, or they can be collected from open source datasets such as Pascal VOC or CoCo (Common Objects in Context). The labeled data can be, for example, the result of using labelme or other similar labeling tools to draw the edge contours of object instances in the training image and label their categories on the user interface provided by the tool, and then parsed and converted into a file in a specific format (such as JSON, XML or CSV) required for training. During training, the training image will be input into the SparseInst network 20 to obtain the predicted result of instance segmentation corresponding to the training image output by the SparseInst network 20. The loss function is used to evaluate the gap between the prediction results output by the SparseInst network 20 during training and the labeled data as the ground truth. This gap is used to determine the direction of the parameter update of the SparseInst network 20. The iterative update of the parameters is usually achieved through gradient descent or its variants, such as stochastic gradient descent (SGD), Momentum, AdaGrad or Adam, with the backpropagation algorithm. After the loss value converges, the SparseInst network 20 can be inferred using a test data set different from the training data set to evaluate its generalization ability and ensure that it can accurately analyze diverse data. The test results will help determine whether the model is overfitting and provide a basis for its further optimization.

[0038] The loss function used to evaluate and optimize the SparseInst network 20 involves focal loss, mask loss, and binary cross entropy loss, as shown in the following <Formula 1>:

[0039] L=λ c L cls +L mask +λ s L s <Formula 1>

[0040] Where L represents the loss function, L cls represents the focal loss for object classification, L mask represents the mask loss between the predicted mask and the ground truth mask, L s Represents the binary cross entropy loss of the object intersection over Union (IoU). L cls The coefficient λ c , and L s The coefficient λ s , are hyperparameters used to mitigate the imbalance between foreground and background. To further overcome the imbalance between foreground and background, the mask loss L mask A combination of Dice loss and pixel-by-pixel binary cross entropy loss is used, as shown in the following <Formula 2>:

[0041] L mask =λ dice ·L dice +λ pix ·L pix <Formula 2>

[0042] Where L dice and Lpix represents the Dice loss and pixel-wise binary cross entropy loss between the predicted mask and the ground truth mask, respectively, and λ dice and λ pix is the corresponding coefficient. More specifically, the Dice loss L dice The calculation is as follows <Formula 3>:

[0043]

[0044] Where m xy and t xy Represent the values ​​of the predicted mask and the ground truth mask at the (x, y) position respectively.

[0045] Figure 3is an architecture diagram of an improved SparseInst network 30 according to an embodiment of the present invention. Similar to the typical SparseInst network 20, the improved SparseInst network 30 also includes a backbone module 31, an encoder module 32 and a decoder module 33, and the main functions of each module are also similar to the aforementioned backbone module 21, encoder module 22 and decoder module 23. Specifically, the backbone module 31 involves extracting multi-scale features 303 of an input image 301 using a backbone 301. The encoder module 32 involves performing a series of feature fusion operations (including convolutions 305-309 and 311, and addition operations on the convolution results in the process) on the multi-scale features 303 to obtain encoded features 315. The decoder module 33 includes a mask branch 316 and an instance branch 317, wherein the mask branch 316 involves deriving a mask feature 318 from the encoded feature 315, and the instance branch 317 involves deriving instance information 319 from the encoded feature 315, wherein the instance information 319 includes a core feature 320, a predicted category confidence 321, and a predicted object confidence 322. The instance segmentation model also involves deriving an instance segmentation result from the mask feature 318 and the instance information 319, which has been described in detail previously, and the details will not be repeated here. The following will explain possible improvements of the improved SparseInst network 30 compared to the typical SparseInst network 20. It should be understood that although Figure 3 All possible improvements are shown, but the present invention does not limit all possible improvements to be applied. In various embodiments, some improvements can be selected for application according to practical requirements, such as the computing resources available to the system and / or the support for a specific network type.

[0046] In one embodiment, the backbone 302 used by the backbone module 31 is implemented by the MobileNetV2 network, replacing the Resnet used in the typical backbone 202. MobileNetV2 is a lightweight convolutional neural network with much fewer parameters than Resnet. Therefore, replacing Resnet with MobileNetV2 as the backbone 302 can significantly reduce the amount of calculation. In addition, the multi-scale features 303 extracted by the backbone 302 include C5, C4, and C3 feature maps, all of which have 128 dimensions, which is lower than the C5, C4, and C3 feature maps of the multi-scale features 203, which are 256, 1024, and 512 dimensions, respectively. Afterwards, the feature map 310 obtained by adding the results output by the convolutions 307, 308, and 309 also has a dimensionality of 128 dimensions, which is lower than Figure 2 The feature map 210 in is 256-dimensional. In this way, the amount of calculation of the encoder module 32 and the decoder module 33 can be reduced, and the inference speed can be improved.

[0047] In one embodiment, the encoder module 32 does not involve the pyramid pooling module. Specifically, the C5 feature Figure 1 On the one hand, it will directly pass through convolution 307, and on the other hand, it will be directly magnified by 2 times and added to the result of convolution 305 of the C4 feature map. This process is no longer involved Figure 2 The pyramid pooling module 204 is shown. The removal of the pyramid pooling module can further reduce the amount of calculation, and can also allow the improved SparseInst network 30 to be deployed in embedded systems with relatively monotonous neural network support categories.

[0048] In one embodiment, neither the encoder module 32 nor the decoder module 33 involves coordinate concatenation operations. Specifically, the feature map obtained after the feature map 310 is convolved 311 is directly used as the input of the encoded feature 315, that is, the mask branch 316 and the instance branch 317, and is no longer concatenated with the coordinate information. In this way, the improved SparseInst network 30 can be deployed in embedded systems that do not support coordinate concatenation operations. In addition, the training of the instance segmentation model also involves optimizing the improved SparseInst network 30 using an improved loss function. In order to compensate for the loss of coordinate information caused by the deletion of the coordinate concatenation operation, the improved loss function adds a heatmap loss on the basis of the original loss function (as previously described in <Formula 1>), as shown in the following <Formula 4>:

[0049] L′=λ c L cls +L mask +λ s L s +λ k L k <Formula 4>

[0050] Where L' represents the improved loss function, L k represents the heat map loss between the predicted category confidence and the ground truth category confidence, L k The coefficient λ k is a hyperparameter. More specifically, the heatmap loss L k The calculation is as follows <Formula 5>:

[0051]

[0052] Where α = 2, β = 4, N represents the total number of categories (if it is only used to identify security devices, it can be set to 1), c represents the category label, Represents the value of the predicted category confidence at the (x, y) position and category c, Y xyc Represents the value of the reference fact category confidence at position (x, y) and category c. The calculation of the reference fact category confidence is as follows <Formula 6>:

[0053]

[0054] where Y x,y represents the ground truth category confidence, m xy Represents the value of the ground truth mask at the (x,y) position. When (x,y) is the coordinate of the center point of the object, and is 1 if the value is true, otherwise it is 0.

[0055] In one embodiment, the instance branch 317 involves multiple convolution layers but does not involve any downsampling layers. In other words, the typical instance branch 217 involves a downsampling layer, while in the instance branch 317, the downsampling layer is replaced by a convolution layer. In this way, the improved SparseInst network 30 can be deployed in embedded systems that do not support downsampling layers. In addition, the core features 220, predicted category confidence 221, and predicted object confidence 222 derived by the typical instance branch 217 have a size fixed to num_mask, that is, the number of potential instance masks, for example, 20*20. As for the core features 320, predicted category confidence 321, and predicted object confidence 322 derived by the instance branch 317, their sizes depend on the size of the input image 301. Figure 3 In the example, the size of the core feature 320, the predicted category confidence 321, and the predicted object confidence 322 is 1 / 32 of the input image 301. For example, in the typical SparseInst network 20, for input images of 640*640 and 256*256, the size of the core feature 220, the predicted category confidence 221, and the predicted object confidence 222 derived by the instance branch 217 are all 20*20; while in the improved SparseInst network 30, the size of the core feature 320, the predicted category confidence 321, and the predicted object confidence 322 derived by the instance branch 317 will be 20*20 and 8*8, respectively. It can be seen that for the instance branch 317, as long as the input image is specified as a lower size, the required computation amount can be greatly reduced.

[0056] Figure 4A is a flowchart of more detailed steps S401-S403 of step S101 in one embodiment. Figure 4A , Figure 4B This is a schematic diagram of steps S401-S403. Please refer to Figure 4A and Figure 4B , to better understand this embodiment.

[0057] In step S401, a camera is used to capture an original image 411. The capture scene of the original image 411 may be, for example, various vehicles (such as cars, motorcycles, trains, airplanes, and yachts), engineering vehicles (such as forklifts, excavators, and cranes), or seats of amusement facilities (such as roller coasters, pirate ships, and free-fall facilities), but the present invention is not limited thereto.

[0058] In step S402, the driver and passenger 410 in the original image 411 are detected to crop a partial image 413 containing the driver and passenger 410 from the original image 411. The detection of the driver and passenger 410 can be achieved through an object detection model based on machine learning, or through a traditional object detection algorithm, such as a Haar cascade detector, template matching, and Hough transform, but the present invention is not limited thereto. The result of the detection is a rectangular bounding box 412 surrounding the driver and passenger 410. The original image 411 is cropped along the bounding box 412 to obtain a partial image 413 containing the driver and passenger 410.

[0059] In step S403, the size of the local image 413 is adjusted to a specified input size to obtain an input image 414. The specified input size may be a size required for inputting an instance segmentation model, such as 256*256.

[0060] During the execution of steps S401 - S403 , some additional processing may be performed on the original image 411 , the local image 413 and the input image 414 , such as geometric transformation, normalization, smoothing and / or image enhancement, etc., which is not limited in the present invention.

[0061] In one embodiment, the attribute data associated with the security device calculated in step S403 may include the height, width, and rotation angle of the minimum bounding rectangle (MBR) of the security device, and the pixel area of ​​the security device. Figure 5, which shows a minimum bounding rectangle 500 as an example and a pixel area A of a safety belt, wherein the minimum bounding rectangle has a height h, a width w and a rotation angle θ. The rotation angle θ can be defined as the angle between a certain dimension (height or width) of the minimum bounding rectangle and a horizontal line or a vertical line, and the present invention is not limited to this. The pixel area A can be defined as the number of pixels contained in the area where the safety device is located, or as the area ratio of the area where the safety device is located relative to the minimum bounding rectangle 500, and the present invention is not limited to this. In one implementation, the calculation of the minimum bounding rectangle and the pixel area can be implemented by functions provided by computer vision function libraries such as OpenCV, PyTorch, TensorFlow, scikit-image or MATLAB. For example, cv2.minAreaRect() provided by OpenCV accepts a contour as input and returns a rotated rectangle that just surrounds the entire contour. cv2.contourArea() provided by OpenCV accepts a contour as input and returns the pixel area of ​​the area surrounded by the contour.

[0062] In addition, step S401 also includes checking whether the height h, width w, rotation angle θ and pixel area A of the minimum circumscribed rectangle 500 are all within their respective specified ranges. The setting of the specified range can be determined according to the actual application scenario, and the present invention is not limited to this. For example, the judgment condition can be set as: the height h falls between 50 and 400 pixels in length, the width w falls between 100 and 300 pixels in length, the rotation angle θ falls between 30 degrees and 100 degrees, and the pixel area A occupies more than 70% of the area of ​​the minimum circumscribed rectangle 500. If all of the above conditions are met, it is judged that the driver and passengers are equipped with safety equipment; if any condition is not met, it is judged that the driver and passengers are not equipped with safety equipment.

[0063] Figure 6 1 is a system block diagram of a system 60 for implementing the above method 10 according to an embodiment of the present invention. The system 60 can be any computer system with computing power, such as a personal computer (such as a desktop computer or a notebook computer) or a server computer, or a mobile device such as a tablet computer or a smart phone, or an embedded system designed to process specific tasks with relatively limited computing resources, and the present invention is not limited thereto. In one embodiment, the system 60 is an embedded system using a development board support package software (BSP), and the types of neural networks supported by it are relatively monotonous.

[0064] The system 60 at least includes a processing device 602 to execute the above steps S101-S104, S402-S403 and various embodiments thereof. The processing device 602 may include any one or more general or special processors or combinations thereof for executing instructions, such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, a microcontroller, a single chip system (System on a Chip; SoC), an application specific integrated circuit (Application Specific Integrated Circuit; ASIC), a field programmable gate array (Field Programmable Gate Array; FPGA) or a combination thereof, and the present invention is not limited thereto.

[0065] In one embodiment, the system 60 further includes a photographing device 601 to perform the above step S401, i.e., to capture the original image. The photographing device 601 may include a lens and a conversion element. The lens may include one or more lenses, such as a zoom lens for magnifying or reducing the size of the target object, and a focus lens for adjusting the focal length of the target object. The conversion element may be, for example, a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS), which is used to receive an optical signal from the lens and then convert the optical signal into an electrical signal. In addition, the photographing device 601 may communicate with the processing device 602 through various wired or wireless communication interfaces, such as a universal serial bus (USB), a wireless network (Wi-Fi), an Ethernet network (Ethernet), Bluetooth (Bluetooth) or a near field communication (NFC), so as to transmit the captured image data to the processing device 602 for subsequent processing and analysis.

[0066] In one embodiment, the system 60 further includes an output device 603 to issue a warning message when it is detected that the driver is not equipped with safety equipment. The warning message can be in various forms, including but not limited to text, sound, vibration or light. The object of sending the warning message can be the driver himself, or the relevant supervisor. The output device 603 can be a device that transmits information to the driver or supervisor through visual, auditory or tactile perception, such as a display that outputs text information, a sound alarm that outputs sound information, a vibration alarm that outputs vibration information, or a warning light that outputs a light source signal. The output device 603 can communicate with the processing device through various communication interfaces, such as a high-definition multimedia interface (HDMI), a universal serial bus (USB), a wireless network (Wi-Fi) or Bluetooth. In response to determining that the driver is not equipped with safety equipment, the processing device 602 causes the output device 603 to issue a warning message.

[0067] The above paragraphs are described using a variety of situations. Obviously, the present invention can be implemented in a variety of ways, and any specific architecture or function disclosed in the examples is only a representative situation. According to the teachings of this article, those skilled in the art should understand that each aspect disclosed in this article can be implemented independently, or two or more aspects can be implemented in combination.

[0068] The above description is only a preferred embodiment of the present invention, but it is not intended to limit the scope of the present invention. Any person familiar with this technology can make further improvements and changes on this basis without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope defined by the claims of this application.

Claims

1. A method for detecting a security device, implemented by a computer system, comprising: Obtaining an input image containing a driver and an occupant; Inputting the input image into a pre-trained instance segmentation model to obtain an instance segmentation result output by the instance segmentation model; Based on the instance segmentation result, calculating attribute data associated with the security device; as well as Based on the attribute data, determine whether the driver or passenger is equipped with the safety equipment.

2. The method according to claim 1, wherein the instance segmentation model is implemented by an improved SparseInst network, the improved SparseInst network comprising: A backbone module, for extracting multi-scale features of the input image using a backbone; An encoder module, used for performing a series of feature fusion operations on the multi-scale features to obtain encoding features; as well as Decoder module, including: A mask branch, for deriving a mask feature from the encoded feature; and An instance branch, used to derive instance information from the encoded features, the instance information including core features, predicted category confidence, and predicted object confidence; The instance segmentation model is also used to derive the instance segmentation result from the mask feature and the instance information.

3. The method according to claim 2, wherein the backbone is implemented by a MobileNetV2 network, and the dimension of the multi-scale features extracted by the backbone is 128 dimensions.

4. The method of claim 2, wherein the encoder module does not include a pyramid pooling module.

5. The method according to claim 2, wherein the encoder module and the decoder module do not perform a coordinate concatenation operation; and The training of the instance segmentation model includes optimizing the improved SparseInst network using an improved loss function, wherein the improved loss function includes focal loss, mask loss, binary cross entropy loss and heat map loss.

6. The method according to claim 2, wherein the instance branch comprises a plurality of convolutional layers but does not include any downsampling layers, and the number of the core features, the predicted category confidence, and the predicted object confidence derived by the instance branch depends on the size of the input image.

7. The method according to claim 1, wherein obtaining the input image of the driver and passenger comprises: Use photographic devices to capture original images; Detecting the driver and occupant in the original image to crop a partial image containing the driver and occupant from the original image; and The size of the partial image is adjusted to a specified input size to obtain the input image.

8. The method according to claim 1, wherein the attribute data comprises a height, a width and a rotation angle of a minimum bounding rectangle, and a pixel area; and The step of judging whether the driver or passenger is equipped with the safety device based on the attribute data includes: Check whether the height, the width, the rotation angle and the pixel area of ​​the minimum bounding rectangle are all within their respective specified ranges.

9. The method according to claim 1, further comprising: In response to determining that the driver and passenger are not equipped with the safety equipment, a warning message is issued.

10. The method of claim 1, wherein the computer system is an embedded system using a development board support package software.

11. A detection system for a safety device, comprising a processing device, performing the following steps: Obtaining an input image containing a driver and an occupant; Inputting the input image into a pre-trained instance segmentation model to obtain an instance segmentation result output by the instance segmentation model; Based on the instance segmentation result, calculating attribute data associated with the security device; as well as Based on the attribute data, determine whether the driver or passenger is equipped with the safety equipment.

12. The system according to claim 11, wherein the instance segmentation model is implemented by an improved SparseInst network, the improved SparseInst network comprising: A backbone module, for extracting multi-scale features of the input image using a backbone; An encoder module, used for performing a series of feature fusion operations on the multi-scale features to obtain encoding features; as well as Decoder module, including: A mask branch, for deriving a mask feature from the encoded feature; and An instance branch, used to derive instance information from the encoded features, the instance information including core features, predicted category confidence, and predicted object confidence; The instance segmentation model is also used to derive the instance segmentation result from the mask feature and the instance information.

13. The system according to claim 12, wherein the backbone is implemented by a MobileNetV2 network, and the dimension of the multi-scale features extracted by the backbone is 128 dimensions.

14. The system of claim 12, wherein the encoder module does not include a pyramid pooling module.

15. The system of claim 12, wherein the encoder module and the decoder module do not perform coordinate concatenation operations; and The training of the instance segmentation model includes optimizing the improved SparseInst network using an improved loss function, wherein the improved loss function includes focal loss, mask loss, binary cross entropy loss and heat map loss.

16. The system of claim 12, wherein the instance branch comprises a plurality of convolutional layers but does not include any downsampling layers, and the number of the core features, the predicted class confidence, and the predicted object confidence derived by the instance branch depends on the size of the input image.

17. The system of claim 11, further comprising: A photographing device capable of communicating with the processing device to capture an original image; The processing device further detects the driver or passenger in the original image to crop a partial image containing the driver or passenger from the original image; and The processing device further adjusts the size of the local image to a specified input size to obtain the input image.

18. The system according to claim 11, wherein the attribute data comprises a height, a width, a rotation angle, and a pixel area of ​​a minimum bounding rectangle; and The processing device further determines whether the driver or passenger is equipped with the safety device by checking whether the height, the width, the rotation angle and the pixel area of ​​the minimum circumscribed rectangle are all within their respective specified ranges.

19. The system according to claim 11, further comprising an output device: In response to determining that the driver or passenger is not equipped with the safety device, the processing device causes the output device to issue a warning message.

20. The system of claim 11, wherein the system is an embedded system using a development board support package software.