A method, device, and storage medium for detecting three-dimensional properties.

By using sensors on the top of the elevator car to detect image data and perform self-calibration and target detection, the problems of low efficiency and poor accuracy in 3D attribute detection in elevators are solved, achieving efficient and accurate 3D attribute statistics.

CN119637660BActive Publication Date: 2025-10-31HITACHI BUILDING TECH GUANGZHOU CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510074779.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-10-31
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

In elevators, existing technologies suffer from low computing power, resulting in a large amount of 3D data computation and low efficiency. Furthermore, due to the complex relationships between target objects within the elevator car, 3D attribute detection is prone to errors.

Method used

By using sensors on the top of the elevator car to detect and perceive image data, self-calibration of pixel depth values ​​and target detection are performed, background is filtered out, and the three-dimensional attributes of the target object are gradually statistically analyzed.

Benefits of technology

It reduces the amount of data, improves computational efficiency, and enhances the accuracy of 3D attributes, making it suitable for elevator hardware with lower computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119637660B_ABST
    Figure CN119637660B_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, and storage medium for detecting three-dimensional attributes. The method includes: using a sensor located at the top of the elevator car to detect and perceive image data downwards; each pixel in the perceived image data has a depth value; performing self-correction on the depth values ​​of the pixels in the perceived image data; if self-correction is completed, detecting the region image data where the target object is located in the perceived image data based on the elevator car; and, after filtering out the elevator car background, statistically analyzing the three-dimensional attributes of the target object in the region image data. This embodiment significantly reduces the amount of data through target detection and background filtering, thereby effectively reducing the computational load and making it suitable for elevator hardware with low computing power, thus improving its computational efficiency. Furthermore, this embodiment gradually eliminates interference and converges to the target object through target detection and background filtering, and statistically analyzes the three-dimensional attributes of the target object, which can improve the accuracy of the three-dimensional attributes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of elevators, and more particularly to a method, device, and storage medium for detecting three-dimensional properties. Background Technology

[0002] In scenarios such as counting user density in elevator cars, three-dimensional data of the car is collected, the three-dimensional data is projected and transformed as a whole to restore it to the world coordinate system, and the three-dimensional attributes of the target object are calculated based on the world coordinate system.

[0003] However, this method involves a large amount of computation, while elevators often use hardware with lower computing power (such as embedded devices), resulting in low efficiency during computation.

[0004] At the same time, due to the many different relationships between different target objects (such as users) inside the elevator car, for example, two users are close to each other, or one user's head is higher and partially covers the other user's head in the image, etc., there are deviations in the three-dimensional attributes. Summary of the Invention

[0005] In view of this, the present invention provides a method, device and storage medium for detecting three-dimensional attributes, so as to improve the efficiency and accuracy of detecting three-dimensional attributes in elevators.

[0006] A first aspect of the present invention provides a method for detecting three-dimensional attributes, comprising:

[0007] Inside the elevator, a sensor located on the top of the car is used to detect and perceive image data downwards; the pixels in the perceived image data have depth values.

[0008] Self-correction is performed on the depth value of the pixel in the perceived image data;

[0009] If self-correction is completed, the car detects the image data of the area where the target object is located in the perceived image data;

[0010] Under the condition of filtering out the background of the car, the three-dimensional attributes of the target object are statistically analyzed in the regional image data.

[0011] A second aspect of the present invention provides a three-dimensional attribute detection device, comprising:

[0012] The elevator car detection module is used to call a sensor located on the top of the elevator car to detect and perceive image data downwards; the pixels in the perceived image data have depth values.

[0013] A depth correction module is used to self-correct the depth value of the pixel in the perceived image data;

[0014] The target detection module is used to detect the area image data where the target object is located in the perceived image data based on the car, if self-correction is completed.

[0015] The attribute statistics module is used to statistically analyze the three-dimensional attributes of the target object in the regional image data after filtering out the background of the car.

[0016] A third aspect of the present invention provides an electronic device, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the three-dimensional attribute detection method as described in the first aspect above.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the three-dimensional attribute detection method as described in the first aspect above.

[0021] A fifth aspect of the present invention provides a computer program product comprising a computer program that, when executed by a processor, implements the three-dimensional attribute detection method as described in the first aspect above.

[0022] In this embodiment, a sensor located at the top of the elevator car is used to detect and perceive image data downwards. Pixels in the perceived image data have depth values. The depth values ​​of these pixels are self-corrected within the perceived image data. If self-correction is complete, the image data of the area where the target object is located is detected within the perceived image data. After filtering out the elevator car background, the three-dimensional attributes of the target object are statistically analyzed in the area image data. This embodiment significantly reduces the amount of data through target detection and background filtering, thereby effectively reducing the computational load. It is suitable for elevators with low computing power, improving their computational efficiency. Furthermore, this embodiment gradually eliminates interference and converges to the target object through target detection and background filtering, and then statistically analyzes the three-dimensional attributes of the target object, improving the accuracy of the three-dimensional attributes.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a three-dimensional attribute detection method provided in Embodiment 1 of the present invention.

[0026] Figure 2 This is an example diagram of a type of perceived image data provided in Embodiment 1 of the present invention.

[0027] Figure 3 This is an example diagram of image data for a detection region provided in Embodiment 1 of the present invention.

[0028] Figure 4 This is an example diagram of a bucket provided in Embodiment 1 of the present invention.

[0029] Figure 5 This is an example diagram of a three-dimensional property provided in Embodiment 1 of the present invention.

[0030] Figure 6 This is a schematic diagram of the structure of a three-dimensional attribute detection device provided in Embodiment 2 of the present invention.

[0031] Figure 7 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate so that the embodiments of the invention described herein can cover implementations in sequences other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] Example 1

[0035] See Figure 1 The diagram illustrates a flowchart of a three-dimensional attribute detection method according to Embodiment 1 of the present invention. This method can be executed by a three-dimensional attribute detection device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0036] Step 101: In the elevator, the sensor located on the top of the car is used to detect and perceive image data downwards.

[0037] Different types of buildings, especially high-rise buildings, have different transportation needs for people, pets, and goods. Therefore, different types of elevators can be deployed in buildings according to different transportation needs, such as passenger elevators, freight elevators, sightseeing elevators, etc. The method in this embodiment can be applied to various types of elevators.

[0038] This embodiment can be applied to elevators. An elevator is a complex system, and the structure of an elevator varies in different types of elevators.

[0039] In one example, a certain type of elevator is configured with: a controller (also known as an elevator control system), call buttons distributed on each floor, a car (including car doors), a motor for pulling the car (also known as a traction machine), a control cabinet, a speed governor, a door operator, a car frame, car doors, counterweight guide rails, car guide rails, guide rail supports, traveling cables, a counterweight device, a compensating chain (cable), landing doors, a guide device for the compensating chain (cable), buffers, etc.

[0040] The car door is equipped with a motor, which drives the car door to open and close.

[0041] In some types of elevators, the traction machine, control cabinet, speed governor, traveling cable, etc., can be omitted.

[0042] These devices can be divided into different sets according to their functions, thus forming various subsystems that support the operation of the elevator. The controller is connected to multiple systems of the elevator via wired means such as serial port or serial clock line (SCL). The controller monitors each system and controls the operation of each subsystem, so that the car moves in the hoistway and reaches each floor of the building.

[0043] In one example, the controller includes a door system, a frequency conversion system, a call system, and a traction system. The door system controls the elevator doors. The car is equipped with a car door, and the elevator has hall doors on each floor. The elevator doors include the car door and the hall doors on each floor. The car door and the hall door are of the same type and open / close simultaneously. The frequency conversion system controls the frequency converter. The call system controls the logic of internal call (calling the elevator from inside the car) and external call (calling the elevator from the hall). The traction system controls the car to move vertically (vertically upward or vertically downward) in the hoistway.

[0044] The elevator manager or owner can choose whether to install an edge computing node on the elevator based on factors such as the elevator's load status. If no edge computing node is installed, the controller maintains the original control logic and does not affect the normal operation of the elevator. If an edge computing node is installed, a suitable computing device can be selected as the elevator's edge computing node according to the needs. The edge computing node is combined with the original controller to form a new controller and redefine the elevator's control logic.

[0045] Generally, edge computing nodes are computing devices with strong computing capabilities, such as computers, servers, or embedded devices. In addition, depending on the different intelligent services, edge computing nodes can be equipped with graphics processing units (GPUs) or embedded neural network processors (NPUs).

[0046] Edge computing nodes refer to new business platforms built at the network edge near elevators, providing storage, computing, and network resources. This allows some critical business applications to be offloaded to the edge of the access network, reducing bandwidth and latency losses caused by network transmission and multi-level forwarding. Located between the user and the cloud (server), edge computing nodes are closer to the user (data source) than traditional cloud computing, featuring miniaturization, distribution, and user-friendliness. Massive amounts of data (such as audio data) no longer need to be uploaded to the cloud for processing; data processing can be performed at the network edge, reducing request response time, reducing network bandwidth, and ensuring data security and privacy.

[0047] In addition, edge computing nodes can implement algorithm functions and model inference, communicate with the original controller, and provide the original controller with artificial intelligence (AI) and complex computing capabilities; edge computing can also communicate with the cloud to realize algorithm functions and model updates, and relay the original controller function calls, etc.

[0048] One or more sensors are installed on the top of the elevator car to detect events during elevator operation, such as... Figure 2 As shown, the sensor can be invoked to probe downwards and obtain perceived image data.

[0049] In one scenario, the sensor is a depth sensor, which can acquire depth values. In this case, the pixels in the perceived image data have depth values.

[0050] In another scenario, the sensor is an infrared sensor, which can collect infrared values. In this case, the pixels in the perceived image data also have infrared values.

[0051] Furthermore, the depth sensor and the infrared sensor can be configured in the same camera or in different cameras; this embodiment does not impose any restrictions on this.

[0052] The depth sensor and infrared sensor have been jointly calibrated. The transformation relationship (such as translation relationship, rotation relationship, etc.) between the coordinate system of the depth sensor and the coordinate system of the infrared sensor is calculated. Based on the transformation relationship, the depth value is projected from the coordinate system of the depth sensor to the coordinate system of the infrared sensor, or the infrared value is projected from the coordinate system of the infrared sensor to the coordinate system of the depth sensor, so that the pixels in the perceived image data have both depth value and infrared value.

[0053] Step 102: Perform self-correction on the depth values ​​of pixels in the perceived image data.

[0054] In practical applications, the depth values ​​of pixels in perceived image data can be self-corrected, for example, by filtering out abnormal depth values, correcting depth values ​​with low accuracy, etc., in order to improve the quality of the depth values ​​of pixels, which is beneficial to subsequent recognition and processing.

[0055] In one embodiment of the present invention, when the pixels in the perceived image data also have infrared values, considering that the accuracy of the distance between each pixel in the perceived image data is related to the strength of the returned signal received by the sensor, the depth value can be corrected based on the infrared value.

[0056] Specifically, due to the different surface coverings and postures (views relative to the sensor surface) of various objects or users inside the elevator car, and the complex reflection characteristics of the inner wall of the car, after the sensor actively emits light onto the surface of the object or user, the light returns to the sensor surface through multiple different paths and is superimposed (multipath reflection), which causes errors in the sensor's measurement of the distance to the object or user.

[0057] The accuracy of a sensor's distance measurement is related to the intensity of the received reflected light; higher light intensity results in more accurate distance measurement. Even in the presence of multipath reflection, higher light intensity makes the dominant peak in the reflected light (usually formed by the light from the shortest path) more significant, leading to more accurate distance measurement. Conversely, if there is no significant dominant peak, there may be significant interference between the distances measured from different paths, making accurate distance measurement difficult.

[0058] Since the sensor uses an active light source when detecting depth, and the infrared value can reflect the intensity of the reflected light from the corresponding pixel, the depth value can be compensated based on the infrared value to correct the depth value of each pixel and improve the accuracy of ranging.

[0059] In this embodiment, step 102 may include the following steps:

[0060] Step 1021: Determine the neighborhood of the current pixel in the perceived image data.

[0061] In this embodiment, each pixel in the perceived image data can be traversed in a certain order. During the traversal, the neighborhood of the current pixel can be divided.

[0062] For example, a square with a specified side length (such as 3×3 or 5×5) is determined as its neighborhood, centered on the current pixel.

[0063] Step 1022: Using a preset reference value as the correction standard for infrared values, calculate the standard value using the depth values ​​of pixels in the neighborhood.

[0064] In this embodiment, a reference value can be set for the infrared value by conducting experiments in the elevator car, etc., as a standard for correction. In this way, the depth values ​​of pixels in the neighborhood can be statistically analyzed to obtain a standard value.

[0065] In a practical implementation, pixels with infrared values ​​greater than a preset reference value can be selected from the neighborhood as target points, and the average depth value of the target points can be calculated as a standard value.

[0066] Pixels in the neighborhood whose infrared values ​​are less than or equal to a preset reference value can be ignored and no further compensation will be performed.

[0067] Step 1023: Subtract the depth value of the current pixel from the standard value to obtain the depth difference value.

[0068] For the current pixel, subtract the depth value of the current pixel from the standard value of other pixels in its neighborhood to obtain the depth difference value.

[0069] Step 1024: Generate a compensation coefficient based on the difference between the infrared value of the current pixel and the reference value.

[0070] In this embodiment, the difference between the infrared value of the current pixel and the reference value can be statistically analyzed, and a compensation coefficient can be generated based on the difference between the infrared value of the current pixel and the reference value.

[0071] Generally, the compensation coefficient is positively correlated with the difference. That is, the greater the difference between the infrared value of the current pixel and the reference value, the larger the compensation coefficient; conversely, the smaller the difference between the infrared value of the current pixel and the reference value, the smaller the compensation coefficient.

[0072] For example, the compensation coefficient is:

[0073]

[0074] Among them, ir (x,y) The infrared value of pixel (x,y), ir θ The value is a reference value, and max is the function to find the maximum value.

[0075] In this example, when the infrared value is greater than or equal to the reference value, the difference between the infrared value of the current pixel and the reference value is... If the value is less than or equal to 0, the depth value can be considered accurately measured, and no compensation is required; the compensation coefficient is 0. The difference between the infrared value of the current pixel and the reference value... As the value gradually increases, the compensation coefficient will increase exponentially, and the difference between the infrared value of the current pixel and the reference value will increase. When the value reaches its maximum (i.e., the infrared value is 0), the compensation coefficient is 1.

[0076] Step 1025: Multiply the depth difference value by the compensation coefficient to obtain the depth compensation value.

[0077] In this embodiment, the product of the depth difference value and the compensation coefficient can be calculated as the depth compensation value.

[0078] Step 1026: Add the depth compensation value to the depth value of the current pixel to complete the self-correction.

[0079] In this embodiment, a depth compensation value can be added to the current pixel's depth value to obtain a new depth value, thereby completing the self-correction of the current pixel's depth value.

[0080] For example, the self-correcting process is represented as follows:

[0081]

[0082] Where, depth (x,y) ir represents the depth value of the pixel (x, y). (x,y) The infrared value of pixel (x,y), ir θ As a reference value, σ is the radius of the neighborhood, f(ir) (x,y) ≥ir θ ) is an indicator function, in ir (x,y) ≥ir θ At that time, f(ir) (x,y) ≥ir θ ) = 1, in ir (x,y) <ir θ At that time, f(ir) (x,y) ≥ir θ ) = 0, and max is the function to find the maximum value.

[0083] Step 103: If self-calibration is completed, then the image data of the area where the target object is located is detected in the perceived image data of the car.

[0084] In this embodiment, target objects can be set according to different services in the elevator, such as the user's head and shoulders, robots, etc.

[0085] like Figure 3 As shown, when the perceived image data completes self-correction, the area image data where the target object is located can be detected in the perceived image data according to the characteristics of the car.

[0086] In a practical implementation, the target detection network can be trained in an offline environment using historical perceptual image data (especially perceptual image data after linear transformation of depth values) and the image data of the region where the labeled target object is located as samples. This enables the target detection network to perform target detection in perceptual image data with the target object as the target.

[0087] The target detection network is deployed in the elevator and can be loaded and run while the elevator is in operation.

[0088] Furthermore, the structure of the object detection network is not limited to manually designed neural networks. It can also be a neural network optimized by model quantization methods, a neural network searched by NAS (Neural Architecture Search) methods, and so on. This embodiment does not impose any restrictions on this.

[0089] Among them, object detection networks can be divided into one-stage and two-stage.

[0090] Two-stage refers to segment-to-segment object detection, which is completed in two steps. The first step is to use various convolutional neural networks as the backbone of the object detection network to extract features from the perceived image data, perform coarse classification (distinguishing between foreground and background) and coarse localization (anchor) based on the features, and obtain candidate regions. The second step is to classify the candidate regions (i.e., target objects) in the classification network of the object detection network.

[0091] For example, two-stage object detection operations may include R-CNN (Region-CNN), Fast R-CNN, Faster R-CNN, R-FCN (Region-based fully convolutional network), and so on.

[0092] One-stage refers to end-to-end object detection, which is completed in one step without searching for candidate regions separately. Instead, the perceived image data is input into a holistic network, and the generated detection result contains both the location and category information of the target object.

[0093] For example, one-stage object detection operations may include SSD (Single Shot Multibox Detector), YOLO (You Only Look Once), and so on.

[0094] Generally speaking, two-stage detection has higher accuracy but slightly slower detection speed, while one-stage detection is faster but slightly less accurate. One-stage or two-stage detection can be selected based on factors such as elevator resources and real-time detection requirements. This embodiment does not impose any restrictions on this.

[0095] For perceived image data, the upper limit of the sensor's detection range in the car can be calculated based on prior knowledge of the car (i.e., the distance from the pixel on the sensor surface to the edge of the car). If there is a non-zero pixel (i.e., effective detection distance) within the field of view and its depth value is greater than the upper limit, it is considered to be outside the detection range, and the depth value of the pixel is set to zero. This can filter out multipath interference or depth background formed by objects.

[0096] In practice, based on the elevator's specifications (such as the car door height, car interior width, car interior depth, etc.), sensor installation parameters (for example, when the sensor is installed in the middle of the car door light beam, the installation parameters include the sensor's left-right (x-axis) offset, front-back (y-axis) offset, tilt angle, etc. relative to the car door light's lateral center point), and sensor intrinsic parameters, a three-dimensional model of the car is constructed in the world coordinate system, and the three-dimensional model is inverted from the world coordinate system to the sensor's coordinate system.

[0097] The upper limit of the sensor's detection range in the car is obtained by measuring the point farthest from the sensor in the 3D model.

[0098] At this point, the depth value can be linearly varied based on the upper limit, thereby standardizing the depth value.

[0099] For example, a linear change is represented as:

[0100] val=(limit–dist)×G / limit;

[0101] Where val is the depth value after linear transformation, dist is the depth value before linear transformation, limit is the upper limit value, and G is the gray level, such as 255.

[0102] In this example, the perceived image data can be converted into 8-bit data through linear transformation, that is, single-channel 8-bit 2D data with width and height of x and y respectively.

[0103] If a linear transformation is completed, then for each frame of perceived image data, the target object is set as the target to be detected, with the depth value as the pixel value, and the perceived image data is input into the target detection network to detect the image data of the region where the target object is located.

[0104] Step 104: After filtering out the background of the car, statistically analyze the three-dimensional attributes of the target object in the regional image data.

[0105] In practical applications, the background in the elevator car can be adaptively filtered out from the regional image data based on prior knowledge of the elevator (such as the user's height, the robot's height, etc.) to obtain the foreground representing the target object. In this case, the three-dimensional attributes of the target object can be statistically analyzed using data in the regional image data (such as coordinates, depth values, etc.).

[0106] The background in the car may be existing facilities in the car (such as the interior walls) or other target objects that have entered the area image data.

[0107] In one embodiment of the present invention, step 104 may include the following steps:

[0108] Step 1041: In the regional image data, sort the pixels according to their depth values ​​to obtain a pixel sequence.

[0109] In regional image data, for pixels with non-zero values, the pixels can be sorted according to their depth values ​​to obtain a pixel sequence.

[0110] like Figure 4 As shown, if the pixels are sorted in ascending order of depth value, then in the pixel sequence, the closer to the sensor (i.e., the higher the actual height), the higher the pixel will be in the sequence.

[0111] Step 1042: Divide the pixel sequence into multiple buckets.

[0112] In this embodiment, the pixel sequence can be divided into multiple bins using an equal division method, with each bin containing multiple consecutively ordered pixels.

[0113] Step 1043: Calculate the gradient value of the change in overall depth for each bucket.

[0114] In this embodiment, the average depth value of the pixels in each bin can be calculated as the depth value of the bin as a whole, thereby calculating the gradient value of the change in the overall depth value for each bin.

[0115] Step 1044: Sort the gradient values ​​to obtain the gradient sequence.

[0116] In this embodiment, the gradient values ​​of each bin can be sorted from smallest to largest or from largest to smallest to obtain a gradient sequence.

[0117] Step 1045: Take the specified quantile value in the gradient sequence as the threshold.

[0118] In this embodiment, a specified quantile value (such as a quartile value) is taken from the gradient sequence as a threshold for dividing the foreground and background.

[0119] Step 1046: If the overall gradient value of the bucket is greater than or equal to the threshold, then the pixels in the bucket are determined to be the background of the car.

[0120] If the gradient value of a certain bin (such as the average depth value of the pixels in the bin) is greater than or equal to the threshold, then the pixels in that bin are determined to be the background of the car and are ignored.

[0121] Step 1047: If the overall gradient value of the bucket is less than the threshold, then use the pixels in the bucket to calculate the three-dimensional properties of the target object.

[0122] like Figure 5As shown, if the gradient value of a certain bucket (such as the average depth value of the pixels in the bucket) is less than the threshold, then the pixels in that bucket are determined to be the foreground of the car (i.e. the target object), and the three-dimensional attributes of the target object are calculated using the pixels in these buckets.

[0123] On one hand, the average pixel coordinates of the pixels in these bins are calculated and used as coordinate anchors. These anchors are then projected onto the world coordinate system to obtain the position of the target object.

[0124] On the other hand, the average depth value of the pixels in these bins is calculated and used as a distance anchor point. The distance anchor point is then projected onto the world coordinate system to obtain the height of the target object.

[0125] In this embodiment, a sensor located at the top of the elevator car is used to detect and perceive image data downwards. Pixels in the perceived image data have depth values. The depth values ​​of these pixels are self-corrected within the perceived image data. If self-correction is complete, the image data of the area where the target object is located is detected within the perceived image data. After filtering out the elevator car background, the three-dimensional attributes of the target object are statistically analyzed in the area image data. This embodiment significantly reduces the amount of data through target detection and background filtering, thereby effectively reducing the computational load. It is suitable for elevators with low computing power, improving their computational efficiency. Furthermore, this embodiment gradually eliminates interference and converges to the target object through target detection and background filtering, and then statistically analyzes the three-dimensional attributes of the target object, improving the accuracy of the three-dimensional attributes.

[0126] Example 2

[0127] See Figure 6 The diagram shows a structural schematic of a three-dimensional attribute detection device provided in Embodiment 2 of the present invention. Figure 6 As shown, the device includes:

[0128] The car detection module 601 is used to call a sensor located on the top of the car in the elevator to detect and perceive image data downwards; the pixels in the perceived image data have depth values;

[0129] A depth correction module 602 is used to self-correct the depth value of the pixel in the perceived image data;

[0130] The target detection module 603 is used to detect the area image data where the target object is located in the perceived image data based on the car in the self-correction if self-correction is completed.

[0131] The attribute statistics module 604 is used to count the three-dimensional attributes of the target object in the regional image data after filtering out the background of the car.

[0132] In one embodiment of the present invention, the pixels in the perceived image data also have infrared values;

[0133] The depth correction module 602 includes:

[0134] The neighborhood determination module is used to determine the neighborhood of the current pixel in the perceived image data;

[0135] A standard value calculation module is used to calculate a standard value using the depth value of the pixel in the neighborhood, with a preset reference value as the correction standard for the infrared value.

[0136] The depth difference value calculation module is used to subtract the current depth value of the pixel from the standard value to obtain the depth difference value.

[0137] The compensation coefficient generation module is used to generate a compensation coefficient based on the difference between the infrared value of the current pixel and the reference value.

[0138] The depth compensation value calculation module is used to multiply the depth difference value by the compensation coefficient to obtain the depth compensation value;

[0139] The depth compensation value addition module is used to add the depth compensation value to the current depth value of the pixel to complete the self-correction.

[0140] In one embodiment of the present invention, the standard value calculation module includes:

[0141] The target point filtering module is used to filter pixels in the neighborhood whose infrared values ​​are greater than a preset reference value as target points.

[0142] The depth calculation module is used to calculate the average value of the depth value of the target point as a standard value.

[0143] In one embodiment of the present invention, the compensation coefficient is:

[0144]

[0145] Among them, ir (x,y) The infrared value of a pixel, ir θ The reference value is given, and max is the function for finding the maximum value.

[0146] In one embodiment of the present invention, the target detection module 603 includes:

[0147] The object detection network loading module is used to load object detection networks.

[0148] The upper limit calculation module is used to calculate the upper limit value detected by the sensor in the car.

[0149] A linear transformation module is used to linearly transform the depth value based on the upper limit value;

[0150] An image detection module is used to input the perceived image data into the target detection network to detect the image data of the region where the target object is located, using the depth value as the pixel value if a linear change is completed.

[0151] In one embodiment of the present invention, the upper limit calculation module includes:

[0152] A 3D model building module is used to build a 3D model of the elevator car based on the elevator's specifications, the sensor's installation parameters, and the sensor's intrinsic parameters.

[0153] The upper limit measurement module is used to measure the point farthest from the sensor in the three-dimensional model to obtain the upper limit value that the sensor can detect in the car.

[0154] The linear change is expressed as:

[0155] val=(limit–dist)×G / limit;

[0156] Where val is the depth value after linear transformation, dist is the depth value before linear transformation, limit is the upper limit value, and G is the gray level.

[0157] In one embodiment of the present invention, the attribute statistics module 604 includes:

[0158] A depth sorting module is used to sort the pixels in the region image data according to the depth value to obtain a pixel sequence;

[0159] A bucket segmentation module is used to segment the pixel sequence into multiple buckets;

[0160] The gradient value calculation module is used to calculate the gradient value of each bucket as a function of the overall depth value.

[0161] The gradient sorting module is used to sort the gradient values ​​to obtain a gradient sequence;

[0162] A threshold determination module is used to select a specified quantile value in the gradient sequence as a threshold.

[0163] The background determination module is used to determine the pixel in the bucket as the background of the car if the gradient value of the bucket as a whole is greater than or equal to the threshold.

[0164] The foreground statistics module is used to calculate the three-dimensional attributes of the target object using the pixels in the bucket if the overall gradient value of the bucket is less than the threshold.

[0165] In one embodiment of the present invention, the foreground statistics module includes:

[0166] The coordinate anchor point determination module is used to calculate the average value of the pixel coordinates of the pixels in the bucket, and use it as the coordinate anchor point;

[0167] The position projection module is used to project the coordinate anchor point onto the world coordinate system to obtain the position of the target object;

[0168] The distance anchor point determination module is used to calculate the average value of the depth values ​​of the pixels in the bucket, and use it as the distance anchor point;

[0169] The height projection module is used to project the distance anchor point onto the world coordinate system to obtain the height of the target object.

[0170] The three-dimensional attribute detection device provided in the embodiments of the present invention can execute the three-dimensional attribute detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the three-dimensional attribute detection method.

[0171] Example 3

[0172] See Figure 7 This diagram illustrates a structural schematic of an electronic device according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0173] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0174] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0175] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for detecting three-dimensional attributes.

[0176] In some embodiments, the method for detecting three-dimensional attributes may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the three-dimensional attribute detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the three-dimensional attribute detection method by any other suitable means (e.g., by means of firmware).

[0177] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0178] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0179] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0181] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0182] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0183] Example 4

[0184] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the three-dimensional attribute detection method provided in any embodiment of this invention.

[0185] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0186] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0187] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for detecting three-dimensional attributes, characterized in that, include: The system uses sensors located on the top of the elevator car to detect and perceive image data downwards. The pixels in the perceived image data have depth values ​​and infrared values; Determine the neighborhood of the current pixel in the perceived image data; Pixels whose infrared values ​​are greater than a preset reference value are selected from the neighborhood and used as target points. Calculate the average depth value of the target point and use it as a standard value; Subtract the depth value of the current pixel from the standard value to obtain the depth difference value; A compensation coefficient is generated based on the difference between the infrared value of the current pixel and the reference value; Multiply the depth difference value by the compensation coefficient to obtain the depth compensation value; The depth compensation value is added to the current depth value of the pixel to complete the self-correction. If self-correction is completed, the car detects the image data of the area where the target object is located in the perceived image data; Under the condition of filtering out the background of the car, the three-dimensional attributes of the target object are statistically analyzed in the regional image data.

2. The method according to claim 1, characterized in that, The compensation coefficient is: ; Among them, ir (x,y) The infrared value of a pixel, ir θ The reference value is given, and max is the function for finding the maximum value.

3. The method according to claim 1, characterized in that, The step of detecting the region image data where the target object is located in the perceived image data based on the car includes: Load the object detection network; Calculate the upper limit value that the sensor can detect in the car; The depth value is linearly varied based on the upper limit value; If a linear transformation is completed, the depth value is used as the pixel value, and the perceived image data is input into the target detection network to detect the image data of the region where the target object is located.

4. The method according to claim 3, characterized in that The calculation of the upper limit value that the sensor can detect in the car includes: A three-dimensional model of the elevator car is constructed based on the elevator's specifications, the sensor's installation parameters, and the sensor's intrinsic parameters. The point furthest from the sensor in the 3D model is measured to obtain the upper limit of the sensor's detection range in the car; the linear change is expressed as: val=(limit–dist)×G / limit; Where val is the depth value after linear transformation, dist is the depth value before linear transformation, limit is the upper limit value, and G is the gray level.

5. The method according to any one of claims 1-4, characterized in that, The step of statistically analyzing the three-dimensional attributes of the target object in the regional image data after filtering out the car background includes: In the region image data, the pixels are sorted according to the depth value to obtain a pixel sequence; The pixel sequence is divided into multiple buckets; Calculate the gradient value of each bucket as a function of the overall depth value; The gradient values ​​are sorted to obtain a gradient sequence; A specified quantile value is taken from the gradient sequence as the threshold; If the gradient value of the bucket as a whole is greater than or equal to the threshold, then the pixel in the bucket is determined to be the background of the car. If the gradient value of the bucket as a whole is less than the threshold, then the three-dimensional properties of the target object are calculated using the pixels in the bucket.

6. The method according to claim 5, characterized in that, The step of calculating the three-dimensional attributes of the target object using the pixels in the bucket includes: Calculate the average pixel coordinates of the pixels in the bucket, and use it as the coordinate anchor point; The coordinate anchor point is projected onto the world coordinate system to obtain the position of the target object; Calculate the average depth value for the pixels in the bucket, and use it as the distance anchor point; The distance anchor point is projected onto the world coordinate system to obtain the height of the target object.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for detecting three-dimensional attributes as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for detecting three-dimensional attributes as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method, device and system for recognizing stretcher mode in elevator car

    CN108178031A

  • Bootlegging video identification method and system

    CN110197144A