State detection method, system and related device
By using the initial depth value to generate a distance threshold and transform the camera coordinate system during state detection, a ground coordinate system is established, which solves the problem of insufficient accuracy of traditional detection methods in different scenarios and achieves higher detection accuracy and alarm function for abnormal situations.
Patent Information
- Application Number
- CN202510710590.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional state detection methods are prone to accuracy issues in different scenarios, making it difficult to achieve high accuracy.
By acquiring the initial pixels of the target image, the target pixels of the target object are determined, and a distance threshold is generated using the initial depth value. A ground coordinate system is established, and combined with the camera coordinate system transformation relationship, the positioning information and status information of the target object are obtained.
It improves the accuracy of ground coordinate system determination and target object positioning in different scenarios, enhances the accuracy of status detection, and can generate alarm information when abnormal situations are detected.
Smart Images

Figure CN120876353A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a state detection method, system and related apparatus. Background Technology
[0002] To achieve state detection of target objects in a target space, traditional detection methods rely on acquisition devices to capture images of the target space and then analyze these images using real-time target detection technology based on deep learning to determine the corresponding state information. This state detection includes detecting the position of the target object in the target space or detecting the positional relationships between different target objects. The state information includes at least one of the following: the position information of the target object in the target space and the positional relationships between different target objects. However, these detection methods depend on the acquisition accuracy of the acquisition device, and their accuracy is easily affected by different scenarios.
[0003] Therefore, how to propose a state detection method with high accuracy has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide a state detection method, system, and related apparatus that can improve the accuracy of state detection of target objects.
[0005] To address the aforementioned technical problems, a first aspect of this application provides a state detection method, comprising: acquiring a target image; determining target pixels matching a target object from all initial pixels in the target image; wherein the initial pixels are matched with initial coordinates in a camera coordinate system, the initial coordinates including an initial depth value; generating a distance threshold matching the target image using the initial depth value; determining a ground coordinate system matching the target image based on the distance threshold; acquiring a transformation relationship between the ground coordinate system and the camera coordinate system; acquiring positioning information of the target object in a target space based on the transformation relationship and the target pixels; and acquiring state information of the target object based on the positioning information.
[0006] To address the aforementioned technical problems, a second aspect of this application provides a state detection system, comprising: an acquisition module, configured to acquire a target image and determine target pixels matching a target object from initial pixels in the target image; wherein the initial pixels are matched with initial coordinates in a camera coordinate system, the initial coordinates including an initial depth value; a first processing module, configured to generate a distance threshold matching the target image using the initial depth value, and determine a ground coordinate system matching the target image based on the distance threshold; a transformation module, configured to acquire a transformation relationship between the ground coordinate system and the camera coordinate system, and acquire positioning information of the target object in a target space based on the transformation relationship and the target pixels; and a second processing module, configured to acquire state information of the target object based on the positioning information.
[0007] To address the aforementioned technical problems, a third aspect of this application provides an electronic device comprising: a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor executes the program instructions to implement the method mentioned in the above technical solutions.
[0008] To address the aforementioned technical problems, a fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the method described in the above technical solutions.
[0009] The beneficial effects of this application are as follows: Unlike existing technologies, the state detection method proposed in this application, after acquiring the target image, adaptively generates a distance threshold matching the target image based on the initial depth values corresponding to each initial pixel in the target image. Using this distance threshold, the ground region in the target image is determined, thereby improving the accuracy of ground coordinate system determination in different scenarios. Based on the transformation relationship between the ground coordinate system and the camera coordinate system, the precise positioning information of the target object in the target space is determined. Based on the positioning information, the state information of the target object is obtained, improving the accuracy of state detection. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0011] Figure 1 This is a flowchart illustrating one embodiment of the status detection method of this application;
[0012] Figure 2 yes Figure 1The flowchart of step S102 corresponds to another embodiment;
[0013] Figure 3 yes Figure 2 The flowchart of step S202 corresponds to another embodiment;
[0014] Figure 4 yes Figure 1 A flowchart of an embodiment preceding step S103;
[0015] Figure 5 yes Figure 1 The flowchart of step S101 corresponds to another embodiment;
[0016] Figure 6 This is a schematic diagram of one embodiment of the status detection system of this application;
[0017] Figure 7 This is a schematic diagram of the structure of one embodiment of the electronic device of this application;
[0018] Figure 8 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments, and different implementation methods can be adaptively combined. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.
[0021] The state detection method provided in this application is used in a state detection device, and the corresponding execution subject is a processing unit capable of data processing. The processing unit integrates a data acquisition device or exists independently of the data acquisition device and interacts with the data acquisition device.
[0022] Please see Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the state detection method of this application. The method includes:
[0023] S101: Acquire the target image and determine the target pixel that matches the target object from all the initial pixels in the target image; wherein the initial pixel is matched with the initial coordinates in the camera coordinate system, and the initial coordinates include the initial depth value.
[0024] In one embodiment, a target image is acquired by the acquisition device at the current moment. The target image includes multiple initial pixels. The target image is converted into a depth image, and the initial coordinates of each initial pixel in the camera coordinate system are determined. The initial coordinates include an initial depth value, which represents the distance between the corresponding initial pixel and the acquisition device. Target detection is performed on the target image to identify at least one target object in the target image. The initial pixels in the region where the target object is located in the target image are used as target pixels for target object matching.
[0025] In one embodiment, after acquiring the target image, the target image is preprocessed. Based on the preprocessed target image, the initial coordinates of each initial pixel in the camera coordinate system are determined, and the target pixel matching the target object is determined. The preprocessing includes at least one of denoising, image enhancement, and distortion correction. Denoising removes noise from the target image using methods such as Gaussian filtering or median filtering. Image enhancement adjusts image enhancement parameters such as brightness and contrast to improve the sharpness of the target image. Distortion correction eliminates geometric distortions in the target image.
[0026] In some implementation scenarios, the data acquisition device is a camera, and the specific position of the camera can be set according to the actual scene. Furthermore, the camera's acquisition angle can be adjusted in real time to achieve image acquisition from different angles of the target space. The camera can be a monocular camera or a multi-view camera.
[0027] S102: Generate a distance threshold that matches the target image using the initial depth value, and determine the ground coordinate system that matches the target image based on the distance threshold.
[0028] In one embodiment, the standard deviation is calculated based on the initial depth value of each initial pixel in the target image, and this standard deviation is used as a distance threshold for matching the target image. Based on the distance threshold, a ground region is determined from the target image, and a ground coordinate system corresponding to the ground region is established.
[0029] In one embodiment, the variance is calculated based on the initial depth value of each initial pixel in the target image, and this variance is used as a distance threshold for matching the target image. Based on the distance threshold, a ground coordinate system for matching the target image is determined.
[0030] S103: Obtain the transformation relationship between the ground coordinate system and the camera coordinate system. Based on the transformation relationship and the target pixel, obtain the positioning information of the target object in the target space.
[0031] In one embodiment, after determining the ground coordinate system that matches the target image, the transformation relationship between the ground coordinate system and the camera coordinate system is obtained. Using this transformation relationship, the target coordinates of the target pixels corresponding to the target object in the camera coordinate system are transformed to the ground coordinate system. Based on the target pixels in the ground coordinate system, the positioning information of the target object in the target space is obtained.
[0032] In one implementation scenario, after determining the ground coordinate system, the rotation and translation matrices between the ground coordinate system and the camera coordinate system are determined. These rotation and translation matrices are used as the transformation relationship between the ground coordinate system and the camera coordinate system, so that the pixels in the camera coordinate system can be transformed to the ground coordinate system through the transformation relationship.
[0033] It should be noted that there can be one or more target objects in the target image. When there are multiple target objects, for each target object, the positioning information of the target object in the target space is obtained based on the transformation relationship and the corresponding target pixel.
[0034] S104: Obtain the status information of the target object based on the location information.
[0035] In one embodiment, the location area of the target object is determined based on the target object's location information, and the state information of the target object is obtained based on the location area. The state information is used to characterize whether the target object is located in a safe area, or to characterize the positional relationship between different target objects.
[0036] In some implementation scenarios, a safe zone matching the target object is pre-defined within the target space. Based on the target object's location, it is determined whether the target object is within the safe zone. If so, the target object's status is determined to be normal; otherwise, the target object's status is determined to be abnormal.
[0037] In some implementation scenarios, in response to an abnormal state information of the target object, an alarm is generated based on the target object's location and state information. This alarm is then fed back to the alarm terminal to prompt action on the target object's state information.
[0038] In one embodiment, in response to the presence of multiple target objects in the target image, state information between different target objects is obtained based on the positioning information corresponding to each target object.
[0039] Specifically, the target objects include a first object and a second object. After obtaining the first positioning information corresponding to the first object and the second positioning information corresponding to the second object through the corresponding implementation methods described above, the target distance between the first object and the second object is determined based on the first positioning information corresponding to the first object and the second positioning information corresponding to the second object. Based on the target distance, the state information between the first object and the second object is obtained.
[0040] In some implementation scenarios, in response to the target distance between the first object and the second object being less than a target threshold, the state information between the first object and the second object is determined to be abnormal, and an alarm message matching the aforementioned state information is generated and fed back. The aforementioned target threshold can be obtained through estimation, or it can be derived through multiple experiments conducted by relevant technical personnel.
[0041] In some implementation scenarios, the first positioning information includes the 3D point cloud of the first object in the target space. Based on the 3D point cloud corresponding to the first object, the first centroid of the first object in the target space is determined. Similarly, the second positioning information includes the 3D point cloud of the second object in the target space. Based on the 3D point cloud corresponding to the second object, the second centroid of the second object in the target space is determined. The distance between the first centroid and the second centroid in the target space is taken as the target distance between the first object and the second object.
[0042] In a specific application scenario, the first object is the worker, and the second object is the operating equipment. Based on the determined first and second positioning information, the target distance between the worker and the operating equipment is determined. When the target distance is greater than or equal to a target threshold, it indicates that the worker is within a safe range, and the corresponding status information is determined to be normal. When the target distance is less than the target threshold, it indicates that there is a safety hazard between the worker and the operating equipment, the corresponding status information is determined to be abnormal, and an alarm message is generated. The operating equipment can be a crane, industrial robot, or machine tool, etc.
[0043] The state detection method proposed in this application, after acquiring the target image, adaptively generates a distance threshold matching the target image based on the initial depth values corresponding to each initial pixel in the target image. Using this distance threshold, the ground region in the target image is determined, thereby improving the accuracy of ground coordinate system determination in different scenarios. Based on the transformation relationship between the ground coordinate system and the camera coordinate system, the precise positioning information of the target object in the target space is determined. Based on the positioning information, the state information of the target object is obtained, further improving the accuracy of state detection.
[0044] Please see Figure 2 , Figure 2 yes Figure 1The flowchart of step S102 corresponds to another embodiment. The specific implementation process of step S102 includes:
[0045] S201: Obtain the ground model corresponding to the target image in the current construction round, take the initial pixels that are less than the distance threshold from the ground model as ground pixels, and update the ground model using the ground pixels.
[0046] Specifically, initial pixels in the target image are randomly sampled to determine the ground model corresponding to the target image in the current construction round. Based on the constructed ground model and the generated distance threshold, multiple ground pixels are determined, and the ground model in the current construction round is updated using these ground pixels.
[0047] In some implementation scenarios, multiple sampled pixels matching the ground are randomly sampled from all initial pixels in the target image. A ground model is constructed based on these sampled pixels; this ground model is a plane equation matching the ground region in the target image obtained by sampling the pixels. For other initial pixels in the target image, the distance between each other initial pixel and the ground model is determined. Initial pixels whose distance to the ground model is less than a distance threshold are designated as ground pixels. The least squares method is then used to refit the plane equation between the ground model in the current construction round and all ground pixels, resulting in an updated ground model for the current construction round.
[0048] S202: Determine the ground coordinate system based on the updated ground model.
[0049] Specifically, the ground coordinate system is determined based on the number of ground pixels in the updated ground model in the current construction round.
[0050] In some implementation scenarios, a target ground model is acquired, which is determined based on the ground models from previous build rounds. If the number of ground pixels in the updated ground model for the current build round is greater than or equal to the number of ground pixels in the target ground model, the updated ground model for the current build round is used as the target ground model. Alternatively, if the number of ground pixels in the updated ground model for the current build round is less than the number of ground pixels in the target ground model, the target ground model remains unchanged.
[0051] Further, it is determined whether the construction round of the ground model has reached a preset iteration threshold. If it has, a point is selected from the target ground model as the origin to construct a ground coordinate system. If it has not, the current construction round is updated to a historical construction round, and the process returns to obtaining the ground model corresponding to the target image in the current construction round. Initial pixels with a distance less than the distance threshold from the ground model are used as ground pixels, and the ground model is updated using these ground pixels. This process returns to step S201. By performing multiple iterations, the accuracy of the target ground model is improved. The preset iteration threshold can be obtained through estimation or by relevant technical personnel through multiple experiments.
[0052] It should be noted that the process of obtaining the updated ground model under different construction rounds can refer to the model fitting algorithm (RANdom SAmple Consensus, RANSAC), and the specific process will not be described in detail here.
[0053] Please see Figure 3 , Figure 3 yes Figure 2 The flowchart of step S202 corresponds to another embodiment. The specific implementation process of step S202 includes:
[0054] S301: Determine whether the ground model in the current construction round meets the preset construction conditions; wherein, the preset construction conditions are related to the first number of ground pixels and the second number of initial pixels.
[0055] Specifically, after obtaining the updated ground model for the current construction round, it is determined whether the ground model meets the preset construction conditions. These preset construction conditions are related to a first number of ground pixels corresponding to the ground model in the current construction round and a second number of all initial pixels in the target image.
[0056] In some implementation scenarios, the historical inlier rate corresponding to the previous historical build round and the current inlier rate corresponding to the current build round are obtained, and the difference between the current inlier rate and the historical inlier rate is determined. If the difference is less than a preset threshold, the ground model in the current build round is determined to meet the preset build conditions. Alternatively, if the difference is greater than or equal to the preset threshold, the ground model in the current build round is determined not to meet the preset build conditions.
[0057] In some implementation scenarios, the aforementioned current inlier rate is the ratio of the first number of ground pixels corresponding to the ground model in the current construction round to the second number of all initial pixels in the target image; similarly, the aforementioned historical inlier rate is the ratio of the number of ground pixels corresponding to the ground model in the previous construction round to the second number of all initial pixels in the target image. Specifically, the specific formula for calculating the current inlier rate is as follows:
[0058]
[0059] Where p represents the current inlier rate, n represents the first number of ground pixels corresponding to the ground model in the current construction round, and N represents the second number of all initial pixels in the target image.
[0060] In some implementation scenarios, if the difference between the current inlier rate and the historical inlier rate is greater than or equal to a preset threshold, it is determined whether the construction round of the ground model has reached a preset iteration threshold. If it has, the ground model in the current construction round is determined to meet the preset construction conditions. If it has not yet reached the threshold, the ground model in the current construction round is determined to not meet the preset construction conditions.
[0061] S302: If so, determine the ground coordinate system based on the updated ground model in the current construction round.
[0062] Specifically, if the updated ground model in the current construction round meets the preset construction conditions, the iteration stops, and the target ground model is determined based on the ground model. A point in the target ground model is taken as the origin of the ground coordinate system, and the normal vector corresponding to the target ground model is taken as the Z-axis of the ground coordinate system. The X-axis and Y-axis of the ground coordinate system are determined within the target ground model, and the X-axis, Y-axis and Z-axis are perpendicular to each other.
[0063] In one implementation scenario, in response to meeting preset construction conditions, the updated ground model in the current construction round is used as the target ground model. Alternatively, in response to meeting preset construction conditions, the number of ground pixels in the target ground model determined based on historical construction rounds is compared with the number of ground pixels in the updated ground model in the current construction round. If the number of ground pixels in the target ground model is greater than the number of ground pixels in the ground model in the current construction round, the ground coordinate system is directly determined based on the target ground model. If the number of ground pixels in the target ground model is less than the number of ground pixels in the ground model in the current construction round, the ground model in the current construction round is updated to the target ground model, and the ground coordinate system is determined based on the target ground model.
[0064] S303: If not, update the current construction round to the historical construction round, return to the step of obtaining the ground model corresponding to the target image under the current construction round, take the initial pixel points that are less than the distance threshold from the ground model as ground pixels, and use the ground pixels to update the ground model.
[0065] Specifically, if the updated ground model in the current construction round does not meet the preset construction conditions, the iteration continues, the current construction round is updated to the historical construction round, and the process returns to the step of obtaining the ground model corresponding to the target image in the current construction round, taking the initial pixel points that are less than the distance threshold from the ground model as ground pixels, and using the ground pixels to update the ground model, that is, returning to the above step S201.
[0066] The above scheme improves the accuracy of ground coordinate system construction by dynamically controlling the number of iterations based on the interior point rate of the ground in adjacent construction rounds, while avoiding the waste of computing resources and improving the efficiency of ground coordinate system construction.
[0067] Please see Figure 4 , Figure 4 yes Figure 1 The flowchart preceding step S103 corresponds to one embodiment. The target image includes a reference object, and based on this, the process preceding step S103 includes:
[0068] S401: Obtain reference pixels that match the reference object from all initial pixels, and determine the predicted size information of the reference object based on the reference coordinates of the reference pixels in the camera coordinate system.
[0069] Specifically, a reference object is set in the target space for calibration, and the actual size information of the reference object is determined. After acquiring the target image, the reference object is detected from the target image, and the initial pixels of the region where the reference object is located in the target image are used as reference pixels to match the reference object. Based on the reference coordinates of the reference pixels in the camera coordinate system, the predicted size information of the reference object is determined.
[0070] In some implementation scenarios, a trained detection model is used to process the target image to detect reference objects within it. The detection model includes a size prediction network that determines the predicted size information of the reference object based on the reference coordinates of the matched reference pixels in the camera coordinate system.
[0071] In a specific application scenario, the aforementioned reference object can be a regularly shaped calibration object, such as a cylinder. Alternatively, the aforementioned reference object can also be a relatively regularly shaped operating device in the target space.
[0072] S402: Obtain the actual size information of the reference object in the target space, and determine the scale correction factor based on the actual size information and the predicted size information.
[0073] Specifically, the actual size information of the reference object in the target space is obtained. Based on the actual size information and the predicted size information, a scale correction factor is determined to adjust the initial depth value of each initial pixel in the target image. The actual size information of the reference object can be determined by relevant technicians through actual measurements of the reference object.
[0074] In some implementation scenarios, the ratio of the actual size information to the predicted size information of the reference object is used as a scale correction factor.
[0075] S403: Correct the target depth value of the target pixel by using the scale correction factor to obtain the corrected target coordinates of the target pixel in the camera coordinate system.
[0076] Specifically, for each target pixel corresponding to the target object, the target depth value corresponding to the target pixel is corrected using a scale correction factor to obtain the corrected target coordinates of the target pixel in the camera coordinate system.
[0077] In some implementation scenarios, the target depth value corresponding to the target pixel is multiplied by a scale correction factor to correct the target depth value in the target coordinates of the target pixel, thus obtaining the corrected target coordinates. Specifically, only the scale correction factor is used to correct the target pixels corresponding to the target object to improve correction efficiency.
[0078] Optionally, in other implementation scenarios, in order to improve the overall effect of the target image, the above-mentioned scale correction factor can also be used to correct the initial coordinates of all initial pixels in the target image.
[0079] The above scheme determines a scale correction factor that matches the target image by using the actual and predicted size information of the reference object in the target image. This scale correction factor is then used to correct the target coordinates of the target pixels in the camera coordinate system, improving the accuracy of the target coordinates and thus enhancing the accuracy of target object localization in complex scenes.
[0080] Please see Figure 5 , Figure 5 yes Figure 1 The flowchart of step S101 corresponds to another embodiment. The specific implementation process of step S101 includes:
[0081] S501: Acquire the target image acquired by the acquisition device, and use the depth estimation model to generate the initial depth value matching each initial pixel in the target image.
[0082] Specifically, the target image is acquired by the acquisition device. A depth estimation model is obtained, and for each initial pixel in the target image, the depth estimation model is used to predict and output the matching initial depth value.
[0083] In some implementation scenarios, the aforementioned acquisition device is a low-cost monocular camera. Since monocular cameras have low accuracy in depth perception, a depth estimation model is used to predict the initial depth value corresponding to each initial pixel in the target image to improve the accuracy of depth perception. The depth estimation model is the DepthAnythingV2 model.
[0084] In some implementation scenarios, the aforementioned depth estimation model can differ from the DepthAnythingV2 model. For example, the depth estimation model can be a pre-trained neural network model with image analysis and processing capabilities. The depth estimation model is trained to output the initial depth value corresponding to each initial pixel in the target image.
[0085] S502: Obtain the trained detection model, input the target image into the detection model, use the detection model to determine the target object in the target image, and take the initial pixel point corresponding to the target object as the target pixel point.
[0086] Specifically, the trained detection model is obtained, the target image is input into the detection model, the detection model is used to detect the target image, and the target region corresponding to the target object in the target image is determined. The initial pixels within the target region in the target image are taken as the target pixels.
[0087] In some implementation scenarios, the above detection model is an existing neural network model with target detection capabilities. This detection model is trained using multiple training samples. The specific training process can be referred to the training process of existing neural network models, and will not be elaborated here.
[0088] Please see Figure 6 , Figure 6 This is a schematic diagram of one embodiment of the state detection system of this application. The state detection system 60 includes an acquisition module 601, a first processing module 602, a conversion module 603, and a second processing module 604 that are coupled to each other. Specifically:
[0089] The acquisition module 601 is used to acquire the target image and determine the target pixel that matches the target object from the initial pixel in the target image; wherein the initial pixel is matched with the initial coordinate in the camera coordinate system, and the initial coordinate includes the initial depth value.
[0090] The first processing module 602 is used to generate a distance threshold that matches the target image using the initial depth value, and to determine the ground coordinate system that matches the target image based on the distance threshold.
[0091] The transformation module 603 is used to obtain the transformation relationship between the ground coordinate system and the camera coordinate system, and based on the transformation relationship and the target pixel, obtain the positioning information of the target object in the target space.
[0092] The second processing module 604 is used to obtain the status information of the target object based on the positioning information.
[0093] In one embodiment, the first processing module 602 determines the ground coordinate system matching the target image based on a distance threshold, including: obtaining the ground model corresponding to the target image in the current construction round, taking initial pixels that are less than the distance threshold from the ground model as ground pixels, updating the ground model using the ground pixels, and determining the ground coordinate system based on the updated ground model.
[0094] In one embodiment, the first processing module 602 determines the ground coordinate system based on the updated ground model, including: determining whether the ground model in the current construction round meets the preset construction conditions; wherein the preset construction conditions are related to the first number of ground pixels and the second number of initial pixels; if yes, determining the ground coordinate system based on the updated ground model in the current construction round; if no, updating the current construction round to a historical construction round, returning to the step of obtaining the ground model corresponding to the target image in the current construction round, taking the initial pixels that are less than a distance threshold from the ground model as ground pixels, and updating the ground model using the ground pixels.
[0095] In one embodiment, the first processing module 602 determines whether the ground model in the current construction round meets the preset construction conditions, including: obtaining the historical inlier rate corresponding to the previous historical construction round and the current inlier rate corresponding to the current construction round, and determining the difference between the current inlier rate and the historical inlier rate; wherein, the current inlier rate is the ratio of a first number to a second number; in response to the difference being less than a preset threshold, determining that the ground model in the current construction round meets the preset construction conditions; in response to the difference being greater than or equal to the preset threshold, determining that the ground model in the current construction round does not meet the preset construction conditions.
[0096] In one embodiment, the target image includes a reference object. Before the conversion module 603 obtains the positioning information of the target object in the target space based on the conversion relationship and the target pixels, it includes: obtaining reference pixels that match the reference object from all initial pixels; determining the predicted size information of the reference object based on the reference coordinates of the reference pixels in the camera coordinate system; obtaining the actual size information of the reference object in the target space; determining a scale correction factor based on the actual size information and the predicted size information; and correcting the target depth value matched by the target pixels using the scale correction factor to obtain the corrected target coordinates of the target pixels in the camera coordinate system.
[0097] In one embodiment, the acquisition module 601 acquires a target image and determines a target pixel that matches the target object from all initial pixels in the target image, including: acquiring the target image acquired by the acquisition device, generating an initial depth value matching each initial pixel in the target image using a depth estimation model; acquiring a trained detection model, inputting the target image into the detection model, using the detection model to determine the target object in the target image, and taking the initial pixel corresponding to the target object as the target pixel.
[0098] In one embodiment, the target object includes a first object and a second object. The second processing module 604 obtains the status information of the target object based on the positioning information, including: determining the target distance between the first object and the second object based on the first positioning information corresponding to the first object and the second object corresponding to the second object; and determining that the status information between the first object and the second object is abnormal in response to the target distance being less than the target threshold, and generating alarm information that matches the status information.
[0099] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 70 includes a memory 701 and a processor 702 coupled to each other. The memory 701 stores program data (not shown in the figure), and the processor 702 calls the program data to implement the method in any of the above embodiments. For the description of the relevant content, please refer to the detailed description of the above method embodiments, which will not be repeated here.
[0100] Please see Figure 8 , Figure 8 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 80 stores program data 801. When the program data 801 is executed by the processor, it implements the method in any of the above embodiments. For a description of the relevant content, please refer to the detailed description of the above method embodiments, which will not be repeated here.
[0101] It should be noted that the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0103] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0104] The above description is merely an embodiment of this application and does not limit the scope of protection of this application. Any equivalent structural or procedural transformations made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.
Claims
1. A state detection method, characterized in that, include: Acquire a target image, and determine the target pixel that matches the target object from all initial pixels in the target image; wherein, the initial pixel is matched with initial coordinates in the camera coordinate system, and the initial coordinates include an initial depth value; The initial depth value is used to generate a distance threshold that matches the target image, and based on the distance threshold, a ground coordinate system that matches the target image is determined; Obtain the transformation relationship between the ground coordinate system and the camera coordinate system, and based on the transformation relationship and the target pixel, obtain the positioning information of the target object in the target space; Based on the location information, the status information of the target object is obtained.
2. The method according to claim 1, characterized in that, Determining the ground coordinate system matching the target image based on the distance threshold includes: Obtain the ground model corresponding to the target image in the current construction round, take the initial pixel points that are less than the distance threshold from the ground model as ground pixels, and update the ground model using the ground pixels; Based on the updated ground model, the ground coordinate system is determined.
3. The method according to claim 2, characterized in that, Determining the ground coordinate system based on the updated ground model includes: Determine whether the ground model in the current construction round meets the preset construction conditions; wherein, the preset construction conditions are related to the first number of ground pixels and the second number of initial pixels; If so, determine the ground coordinate system based on the updated ground model in the current construction round; If not, update the current construction round to the historical construction round, return to the step of obtaining the ground model corresponding to the target image under the current construction round, take the initial pixel point that is less than the distance threshold from the ground model as the ground pixel point, and update the ground model using the ground pixel point.
4. The method according to claim 3, characterized in that, The determination of whether the ground model in the current construction round meets the preset construction conditions includes: Obtain the historical in-point rate corresponding to the previous historical build round and the current in-point rate corresponding to the current build round, and determine the difference between the current in-point rate and the historical in-point rate; wherein, the current in-point rate is the ratio of the first quantity to the second quantity; In response to the difference being less than a preset threshold, it is determined that the ground model in the current construction round meets the preset construction conditions; In response to the difference being greater than or equal to a preset threshold, it is determined that the ground model in the current construction round does not meet the preset construction conditions.
5. The method according to claim 1, characterized in that, The target image includes a reference object. Before obtaining the location information of the target object in the target space based on the transformation relationship and the target pixels, the process includes: Obtain reference pixels that match the reference object from all the initial pixels, and determine the predicted size information of the reference object based on the reference coordinates of the reference pixels in the camera coordinate system; Obtain the actual size information of the reference object in the target space, and determine the scale correction factor based on the actual size information and the predicted size information; The target depth value matched by the target pixel is corrected using the scale correction factor to obtain the corrected target coordinates of the target pixel in the camera coordinate system.
6. The method according to claim 1, characterized in that, The step of acquiring the target image, and determining the target pixel that matches the target object from all initial pixels of the target image, includes: The target image acquired by the acquisition device is obtained, and the initial depth value matching each initial pixel in the target image is generated using a depth estimation model; The trained detection model is obtained, the target image is input into the detection model, the target object in the target image is determined by the detection model, and the initial pixel corresponding to the target object is taken as the target pixel.
7. The method according to claim 1, characterized in that, The target object includes a first object and a second object. The step of obtaining the state information of the target object based on location information includes: Based on the first positioning information corresponding to the first object and the second positioning information corresponding to the second object, the target distance between the first object and the second object is determined. In response to the target distance being less than a target threshold, the state information between the first object and the second object is determined to be abnormal, and an alarm message matching the state information is generated.
8. A state detection system, characterized in that, include: An acquisition module is used to acquire a target image and determine target pixels that match the target object from initial pixels in the target image; wherein, the initial pixels are matched with initial coordinates in the camera coordinate system, and the initial coordinates include an initial depth value; The first processing module is used to generate a distance threshold that matches the target image using the initial depth value, and to determine a ground coordinate system that matches the target image based on the distance threshold; The transformation module is used to obtain the transformation relationship between the ground coordinate system and the camera coordinate system, and based on the transformation relationship and the target pixel, obtain the positioning information of the target object in the target space; The second processing module is used to obtain the status information of the target object based on the location information.
9. An electronic device, characterized in that, include: A memory and a processor are coupled to each other, the memory storing program instructions, and the processor executing the program instructions to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Automatic ground testing and relative camera pose estimation method in depth image
CN104361575A
Positioning method and device, equipment and storage medium
CN110728717A
Depth camera external parameter calibration method and device
CN112541950A
Tea garden disease and insect pest identification and positioning method and system based on binocular vision
CN117475373A
Positioning method, electronic device, and storage medium
US20220262039A1