Luggage compartment attitude recognition system, method, electronic device, and storage medium

The luggage compartment attitude recognition system uses binocular cameras and deep learning for accurate, cost-effective detection of truck cargo compartments, addressing high costs and complexity in existing systems, enabling large-scale outdoor unmanned autonomous driving.

JP7743642B2Active Publication Date: 2025-09-24GUANGXI LIUGONG MASCH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024549512
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-24
Filing Date
2023-09-21
Publication Date
2025-09-24
Estimated Expiration
2043-09-21

AI Technical Summary

Technical Problem

Existing unmanned autonomous driving systems for construction machinery, such as excavators, face high costs and complexity due to the need for RTK mobile stations, and lack intelligence for dynamic detection of truck cargo compartments, making them unsuitable for large-scale outdoor applications.

Method used

A luggage compartment attitude recognition system using a binocular camera and image recognition network model to determine the three-dimensional coordinates and posture of a truck's cargo compartment, utilizing stereoscopic geometric vision and deep convolutional neural networks for accurate, low-cost, single-end detection.

Benefits of technology

The system provides high-intelligence, low-cost detection of truck cargo compartment positions and postures, suitable for large-scale outdoor unmanned autonomous driving of construction machinery, reducing costs to a few thousand yuan per vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007743642000008
    Figure 0007743642000008
  • Figure 0007743642000009
    Figure 0007743642000009
  • Figure 0007743642000010
    Figure 0007743642000010
Patent Text Reader

Abstract

The present invention discloses a luggage compartment attitude recognition system, method, electronic device, and storage medium. The system includes a processor, a binocular camera, a first vehicle, a second vehicle, and a sign unit, the first vehicle includes a luggage compartment, the second vehicle is used to unload materials into the luggage compartment of the first vehicle, the sign unit is provided on the side of the luggage compartment, and the sign unit is provided with a number of reference points, and the binocular camera is provided on the second vehicle, the binocular camera collects an image including the sign unit and transmits the image to the processor, the processor recognizes and processes the image using an image recognition network model to obtain a target image including the recognition result of the number of reference points, determines three-dimensional coordinates of a target space for each reference point in the target image through a binocular camera stereoscopic geometric vision algorithm, and determines vehicle attitude information based on the three-dimensional coordinates of the target space of each reference point. The luggage compartment attitude recognition system has the characteristics of being simple in structure and suitable for large-scale outdoor unmanned driving scenes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of computer technology, and more particularly to a luggage compartment attitude recognition system, method, electronic device and storage medium. [Background technology]

[0002] Construction machinery and equipment such as earthmoving machinery, mining machinery, and road machinery require cooperative work among multiple machines as they are upgraded to unmanned autonomous driving. For example, when an excavator needs to unload cargo onto a truck, the excavator needs to be able to detect the real-time position and attitude of the truck's cargo compartment.

[0003] Typically, the truck's position in a map coordinate system can be detected using the phase-difference positioning principle of real-time kinematic (RTK) carrier phase difference technology. Furthermore, the location of the cargo compartment in the map coordinate system can be estimated using the size information of the truck itself. Furthermore, the excavator can obtain its own position in the map coordinate system using RTK positioning. By comparing the two, the relative position of the truck compartment and the excavator can be obtained, and then unloading can be completed.

[0004] While this method can achieve unmanned autonomous driving of excavators outdoors, the problem it poses is its high cost. To obtain the truck and excavator's location information on a map, a set of RTK mobile stations must be installed on each truck and excavator, as well as a public base station. This would require costs of tens of thousands to hundreds of thousands of yuan, making large-scale applications unfeasible. Therefore, this solution is not suitable for large-scale unmanned autonomous driving applications outdoors in construction machinery rooms. Summary of the Invention

[0005] The present invention aims to solve the technical problems that the systems used in conventional unmanned autonomous driving have complex structures, low levels of intelligence, high costs, and narrow ranges of application.

[0006] In order to solve the above problem, in one aspect, the present invention provides a luggage compartment attitude recognition system including a processor, a binocular camera, a first vehicle, a second vehicle, and a sign unit, the first vehicle includes a luggage compartment; the second vehicle is used to unload materials into the cargo compartment of the first vehicle; The sign is provided on a side surface of the luggage compartment, and a plurality of reference points are provided on the sign, the binocular camera is provided on the second vehicle, and the binocular camera is used to collect an image including the sign portion and transmit the image to the processor; The processor is communicatively connected to the binocular camera, and the processor recognizes and processes the image using an image recognition network model to obtain a target image including the recognition results of the plurality of reference points, determines three-dimensional coordinates in a target space for each of the plurality of reference points in the target image using a stereoscopic geometric vision algorithm of the binocular camera, and determines posture information of the luggage compartment based on the three-dimensional coordinates in the target space of each reference point.

[0007] Optionally, the second vehicle includes an unloading device and a body structure connected to each other; the unloading device is connected to a first side of the body structure; The binocular camera is provided on the first side surface.

[0008] Optionally, the first side includes a first attachment point and a second attachment point; the first attachment point and the second attachment point are located on an upper portion of the first side surface, a first preset distance exists between the first attachment point and the second attachment point; the binocular camera includes a first binocular camera and a second binocular camera; the first binocular camera is mounted at the first mounting point; The second binocular camera is mounted at the second mounting point.

[0009] Optionally, the marking portion includes at least two sub-marking portions; At least two of the sub-signal portions are respectively located at a first position point and a second position point on a second side surface of the luggage compartment, the second side surface is a surface facing the first side surface, There is a second preset distance between the first position point and the second position point, the second preset distance being longer than half the length of the second side in a first predetermined direction, the first predetermined direction being perpendicular to the second predetermined direction, and the second predetermined direction being an extension of the height direction of the first vehicle.

[0010] Optionally, the processor includes an image recognition module and a position determination module; the image recognition module is used to recognize and process the image using an image recognition network model to obtain a target image, and send the target image to a position determination module, the target image being an image including the recognition results of the plurality of reference points; The position determination module determines three-dimensional coordinates in the target space for each of a plurality of reference points in the target image using a stereoscopic geometric vision algorithm of the binocular camera, and determines the attitude information of the luggage compartment based on the three-dimensional coordinates in the target space of each reference point.

[0011] In another aspect, the present invention provides Collect images including the sign using a binocular camera, inputting the image into an image recognition network model to obtain a target image, the target image including a recognition result for each of the plurality of reference points; For each of the reference points, determine the coordinates of the reference point in a pixel coordinate system based on the target image and the recognition result of the reference point; acquiring camera parameters of the binocular camera; Determine the three-dimensional coordinates of the reference point in the target space based on the camera parameters of the binocular camera and the coordinates of the reference point in the pixel coordinate system using a binocular camera stereoscopic geometric vision algorithm; Determine three-dimensional coordinates of the target space of the luggage compartment based on three-dimensional coordinates of the target space of each reference point; determining attitude information of the luggage compartment based on three-dimensional coordinates of the target space of the luggage compartment; There is further provided a luggage compartment attitude recognition method, including:

[0012] Preferably, the camera parameters include a distance between a left eye camera and a right eye camera of a binocular camera, an attachment parameter, and an internal parameter.

[0013] Optionally, determining three-dimensional coordinates of the target space of the reference point based on camera parameters of the binocular camera and coordinates of the reference point in a pixel coordinate system, determining three-dimensional coordinates of the target space of the luggage compartment based on the target three-dimensional coordinates of each reference point, and determining attitude information of the luggage compartment based on the target three-dimensional coordinates of the luggage compartment determining three-dimensional coordinates of the reference point in a camera coordinate system based on a distance between the left eye camera and the right eye camera of the binocular camera, internal parameters of the binocular camera, and coordinates of the reference point in a pixel coordinate system; determining a first coordinate transformation matrix based on mounting parameters of the binocular camera; transforming the three-dimensional coordinates of the reference point in the camera coordinate system into three-dimensional coordinates in a target vehicle coordinate system according to the first coordinate transformation matrix, the target vehicle coordinate system being a coordinate system in which the second vehicle is located; determining three-dimensional coordinates of the luggage compartment in the target vehicle coordinate system based on three-dimensional coordinates of each reference point in the target vehicle coordinate system; determining attitude information of the luggage compartment based on three-dimensional coordinates of the luggage compartment in the target vehicle coordinate system; Includes:

[0014] In another aspect, the present invention provides an electronic device that includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, the at least one program, code set, or instruction set is loaded and executed by the processor, thereby realizing the luggage compartment posture recognition method.

[0015] In another aspect, the present invention provides a computer storage medium having at least one instruction or at least one program stored therein, the at least one instruction or at least one program being loaded and executed by a processor to realize the luggage compartment posture recognition method.

[0016] According to the above technical means, the luggage compartment attitude recognition method of the present invention has the following beneficial effects. The luggage compartment posture recognition system includes a processor, a binocular camera, a first vehicle, a second vehicle, and a sign unit, wherein the first vehicle is used by the second vehicle to unload materials into the luggage compartment of the first vehicle, the sign unit is mounted on a side of the luggage compartment, and the sign unit has a plurality of reference points. The binocular camera is mounted on the second vehicle, and the binocular camera collects images including the sign unit and transmits the images to the processor. The processor is connected to the binocular camera and performs recognition processing on the images using an image recognition network model to obtain a target image, which is an image including the recognition results of the plurality of reference points. A binocular camera stereoscopic geometric vision algorithm is used to determine target 3D coordinates for each of the plurality of reference points in the target image, and posture information of the vehicle is determined based on the target 3D coordinates of each reference point. The luggage compartment posture recognition system of the present invention has a simple structure and is applicable to large-scale outdoor unmanned driving scenes. [Brief explanation of the drawings]

[0017] In order to more clearly describe the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can also obtain other drawings based on these drawings without creative efforts. [Figure 1] 1 is a structural schematic diagram of a selectable luggage compartment posture recognition system of the present invention; [Figure 2] FIG. 2 is a structural schematic diagram of a selectable second vehicle of the present invention. [Figure 3] FIG. 2 is a structural schematic diagram of a selectable first vehicle of the present invention. [Figure 4] 1 is a structural schematic diagram of a selectable label portion of the present invention; [Figure 5] 3 is a flowchart of a selectable luggage compartment attitude recognition method of the present invention. [Figure 6] 4 is a flowchart of another selectable luggage compartment attitude recognition method of the present invention. [Figure 7] FIG. 10 is a diagram showing the relationship between a plurality of selectable coordinates of the present invention. [Figure 8] 1 is a schematic diagram of a selectable binocular stereoscopic vision camera model of the present invention; [Figure 9] 10 is a schematic diagram showing the relationship between a plurality of coordinates that can be selected in another embodiment of the present invention.

[0018] 1 - first vehicle; 101 - luggage compartment; 2 - second vehicle; 201 - unloading device; 202 - main body structure; 2021 - first side; 3 - sign unit; 301 - sub-sign unit; 4 - binocular camera; 401 - first binocular camera; 402 - second binocular camera; 5 - processor DETAILED DESCRIPTION OF THE INVENTION

[0019] Hereinafter, the technical means in the embodiments of the present invention will be described clearly and completely with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts are all included in the protection scope of the present invention.

[0020] The term "one example" or "example" herein refers to a particular feature, structure, or characteristic that may be included in at least one embodiment of the present invention. In describing the present invention, orientations or positional relationships indicated by terms such as "upper," "lower," "upper," "top," and "bottom" are based on orientations or positional relationships shown in the drawings. These terms are intended solely to illustrate and simplify the present invention and do not indicate or suggest that the described devices or components must have a particular orientation, be constructed, or operate in a particular orientation, and therefore should not be understood as limitations on the present invention. Furthermore, the terms "first" and "second" are used solely for descriptive purposes and should not be understood as indicating or suggesting relative importance or implicitly indicating the number of technical features described. Thus, a feature qualified by "first" or "second" can explicitly or implicitly include one or more of the aforementioned features. Furthermore, the terms "first," "second," and the like are not intended to describe a particular order, but rather to distinguish between similar objects. Data used in this manner are interchangeable where appropriate, and the embodiments of the present invention described herein may be implemented in an order other than that shown or described herein.

[0021] Also, the terms "comprise" and "have" and any variations thereof are intended to cover exclusive inclusions, for example, a process, method, system, product or server comprising a series of steps or units is not necessarily limited to the steps or units expressly recited, but may include other steps or units not expressly shown or inherent to those processes, methods, products or apparatus.

[0022] Generally, the luggage compartment position detection method includes:

[0023] First, the location of the cargo compartment is detected visually from a fixed position. For example, several key points on the cargo compartment that moves back and forth between multiple positions are detected visually to determine the location of the cargo compartment and then work operations are performed. This method is applied to detecting the location of cargo compartments that are used multiple times in a single scene, and is therefore a repetitive detection method with a low level of intelligence. It is not suitable for detecting the dynamic position and posture of a truck cargo compartment when an unmanned excavator is operating outdoors autonomously.

[0024] Second, relative position correction is performed using a laser spot. In the automated guided vehicle (AGV) industry, a laser transmitter attached to the AGV and a laser receiver on the loading platform are used to correct the relative position between the AGV and the loading platform, guiding the AGV's dump truck. This type of sensing of the AGV's shelf position can only be applied to precise terminal correction on a fixed route, and cannot be used to detect the dynamic position and posture of the truck loading platform for outdoor unmanned autonomous driving of excavators.

[0025] Third, truck detection during unmanned driving. Currently, unmanned driving also detects the truck in front and obtains the relative spatial position between the truck in front and the vehicle itself. However, unmanned driving does not focus on the exact position and posture of the truck's cargo compartment, but instead detects the truck as a whole and regards the rack as an obstacle. Therefore, the truck detection used in unmanned driving cannot be directly used to accurately detect the position and posture of the cargo compartment of outdoor unmanned autonomous driving trucks for construction machinery.

[0026] Fourth, there is the RTK truck position and attitude detection described in the background art above. RTK can detect the truck's position in a map coordinate system using differential positioning principles, and can estimate the location of the cargo compartment in a map coordinate system using the truck's own size information. Furthermore, the excavator can obtain its own position in a map coordinate system using RTK positioning. By comparing the two, the relative position of the truck compartment and the excavator can be determined, and unloading can be completed. While this method can complete unmanned outdoor autonomous driving of an excavator, it suffers from high costs. To obtain the map position information of the truck and excavator, a set of RTK mobile stations must be installed on the truck and the excavator, and a public base station must also be installed. This requires costs ranging from tens of thousands to hundreds of thousands of yuan, making it unsuitable for large-scale applications. Therefore, this solution is not suitable for large-scale outdoor unmanned autonomous driving of construction machinery rooms.

[0027] The above luggage compartment detection method has the following drawbacks. 1) Low level of intelligence: Multi-position visual cargo compartment detection and point-to-point laser relative position correction detection do not have the ability of intelligent detection, and due to their low level of intelligence, they cannot be used to sense the dynamic position and posture of the cargo compartment of a truck for outdoor unmanned autonomous driving of an excavator. 2) Complex structure and high cost: RTK positioning technology can be used for outdoor unmanned autonomous operation of construction machinery, but it is expensive, cannot be applied on a large scale, and is only suitable for early exploratory research. 3) Not suitable for outdoor construction machinery operations. None of the existing mass-produced schemes are suitable for outdoor unmanned autonomous driving of construction machinery. The track detection used for unmanned driving detects the truck, but it only detects the entire truck as an obstacle and does not pay attention to the precise position and posture of the cargo compartment. Other schemes, such as multi-position vision-based cargo compartment detection and point-to-point laser relative position correction detection, are not intelligent enough to be used for dynamic sensing of the position and posture of the cargo compartment of outdoor unmanned autonomous driving of construction machinery. RTK positioning technology is expensive and therefore not suitable for large-scale applications of outdoor unmanned autonomous driving of construction machinery.

[0028] For this purpose, please refer to Fig. 1, which is a structural schematic diagram of an optional luggage compartment attitude recognition system of the present invention. The present invention provides a luggage compartment attitude recognition system. The system includes a processor 5, a binocular camera 4, a first vehicle 1, a second vehicle 2, and a sign unit 3, wherein the first vehicle 1 includes a cargo compartment 101, the second vehicle 2 is used to unload materials into the cargo compartment 101 of the first vehicle 1, the sign unit 3 is provided on the side of the cargo compartment 101, and a plurality of reference points are provided on the sign unit 3, the binocular camera 4 is provided on the second vehicle 2, the binocular camera 4 collects images including the sign unit 3 and transmits the images to the processor 5, the processor 5 is connected in communication with the binocular camera 4, the processor 5 recognizes and processes the images using an image recognition network model to obtain a target image, the target image is an image including the recognition results of the plurality of reference points, and the stereoscopic geometric vision algorithm of the binocular camera 4 determines target three-dimensional coordinates for each of the plurality of reference points in the target image, and determines posture information of the cargo compartment 101 based on the target three-dimensional coordinates of each reference point.

[0029] Alternatively, the first vehicle 1 may be a vehicle having a luggage compartment 101, such as a truck, and the second vehicle 2 may be a vehicle having a robotic arm, such as a shovel.

[0030] Alternatively, the processor 5 may be located in the second vehicle 2 or in a terminal or server separate from this vehicle.

[0031] Alternatively, the server may be an independent physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud audio recognition model training, middleware services, domain name services, security services, CDNs (Content Delivery Networks), big data and artificial intelligence platforms, etc. The operating systems running on the servers may include, but are not limited to, Android systems, IOS systems, Linux, Windows, Unix, etc.

[0032] Optionally, the terminal may include, but is not limited to, a type of client such as a smartphone, a desktop computer, a tablet computer, a notebook computer, a smart speaker, a digital assistant, an augmented reality (AR) / virtual reality (VR) device, a smart wearable device, etc. The terminal may also be software executed on the client, such as an application program, an applet, etc. Optionally, the operating system executed on the client may include, but is not limited to, an Android system, an IOS system, Linux, Windows, Unix, etc.

[0033] In an optional example, refer to Fig. 2, which is a structural schematic diagram of an optional second vehicle of the present invention. The second vehicle 2 includes a loading / unloading device 201 and a main body structure 202 connected to each other, the loading / unloading device 201 is connected to a first side 2021 of the main body structure 202, and the binocular camera 4 is provided on the first side 2021.

[0034] To further ensure that the binocular camera can collect image information of all sign portions, in an alternative example, the first side 2021 includes a first mounting point and a second mounting point, the first mounting point and the second mounting point are located at the top of the first side 2021, a first preset distance exists between the first mounting point and the second mounting point, and the binocular camera 4 includes a first binocular camera 401 and a second binocular camera 402, the first binocular camera 401 is mounted at the first mounting point and the second binocular camera 402 is mounted at the second mounting point.

[0035] In order to improve the viewing angle of the binocular camera 4 and ensure the effectiveness of the collected images, optionally, the first preset distance is longer than half the width of the first side 2021, and preferably, the first mounting point is located at the leftmost side of the first side 2021 and the second mounting point is located at the rightmost side of the first side 2021 to further improve the image collection effect.

[0036] In an optional example, refer to FIG. 3, which is a structural schematic diagram of an optional first vehicle of the present invention. To improve the accuracy of the luggage compartment attitude to be determined later, the sign unit 3 includes at least two sub-sign units 301, each of which is located at a first position point and a second position point on a second side surface of the luggage compartment 101, the second side surface being a surface facing the first side surface 2021, a second preset distance being between the first position point and the second position point, the second preset distance being longer than half the length of the second side surface in a first predetermined direction, the first predetermined direction (x-axis direction in FIG. 3) being perpendicular to the second predetermined direction (y-axis direction in FIG. 3), and the second predetermined direction being an extension of the height direction of the first vehicle 1. As a result, the processor 5 can determine the spatial three-dimensional coordinates of the multiple sub-sign units 301 and position information of the luggage compartment 101, such as the distance between the luggage compartment 101 and the second vehicle 2 and the deflection angle relative to the second vehicle 2.

[0037] In this embodiment, the second side may refer to multiple sides. In an actual scene, the first vehicle 1 does not necessarily stop directly in front of the second vehicle 2, but may be at a certain angle with the second vehicle 2. In this case, the second side may be two sides corresponding to the first side 2021. In this case, the multiple sub-sign units 301 may be located on one of the two sides, or may be provided in numbers corresponding to the two sides. In this embodiment, an example will be described in which all the sub-sign units 301 are located on the same side.

[0038] Optionally, referring to Figure 4, Figure 4 is a structural schematic diagram of a selectable sign unit of the present invention. The sign unit 3 is provided with five reference points, but in practice, it may be five or more, for example, n, such as 6, 7, 8, 9, 10, etc., as needed. The more reference points there are, the more accurate the coordinate data of the sign unit 3 that is finally determined. However, if there are too many reference points, the calculation time will further increase and data processing will take too long, so the number of reference points can be set according to specific needs.

[0039] In an alternative example, the processor 5 includes an image recognition module and a position determination module, the image recognition module performs recognition processing on the image using an image recognition network model to obtain a target image including the recognition results of the plurality of reference points, and transmits the target image to a position determination module, the position determination module determines target three-dimensional coordinates for each of the plurality of reference points in the target image using a stereoscopic geometric vision algorithm of the binocular camera 4, and determines posture information of the luggage compartment 101 based on the target three-dimensional coordinates of each reference point. For details, see the description of the luggage compartment posture recognition method described below.

[0040] The luggage compartment 101 posture recognition system according to the present invention has the following advantages. 1) High level of intelligence: Based on deep convolutional neural network technology in the field of artificial intelligence, the binocular camera 4 can accurately detect the position and posture of the truck cargo compartment 101 relative to the truck under various weather conditions and in various working environments by utilizing the images and the stereoscopic geometric vision of the binocular camera 4. 2) Single-end detection. The binocular camera 4 performs single-end detection, recognizing and measuring objects in the external world like the human eye. Unlike RTK, there is no need to attach mobile stations to both the truck and the excavator. The RTK positioning results of the truck must be transmitted to the unmanned excavator via a communications terminal. After receiving the truck positioning data, the unmanned excavator can compare it with its own positioning data to determine the relative positional relationship between the two. 3) Low cost. The current vehicle-grade binocular camera 4 is waterproof and dustproof, and the cost can be reduced to a few thousand yuan in large-scale applications, making it low cost. 4) Satisfying the large-scale application of outdoor unmanned autonomous driving of construction machinery. Based on deep convolutional neural network technology in the field of artificial intelligence, and utilizing images and binocular stereoscopic geometric vision, the binocular camera 4 can detect the exact position and orientation of the cargo compartment 101 of the truck under various weather conditions and in various working environments. At the same time, it is low-cost and can realize the large-scale application of outdoor unmanned autonomous driving of construction machinery.

[0041] Specific embodiments of the method for recognizing the posture of a luggage compartment according to the present invention will be described below. FIG. 5 is a flowchart of a selectable method for recognizing the posture of a luggage compartment according to the present invention. This specification provides operational steps of the method as in the embodiments or flowcharts, but more or fewer operational steps may be included due to ordinary or non-inventive efforts. The order of steps listed in the embodiments is one of many possible execution orders and does not represent the only execution order. When executed by an actual system or server product, it may be executed according to or in parallel with the methods shown in the embodiments or drawings (in an environment such as a parallel processor 5 or multithread processing). Specifically, as shown in FIG. 5, the method may include the following steps:

[0042] S501: Images including the sign portion 3 are collected by the binocular camera 4.

[0043] In this embodiment, the binocular camera 4 is installed as shown in FIG. 2, and the system description above is specifically referred to.

[0044] Optionally, the binocular camera 4 includes a left-eye camera and a right-eye camera. In step S501, specifically, the left-eye camera collects a first image including the sign 3, the first image being a visible light image, and the right-eye camera collects a second image including the sign 3, the second image being a visible light image. In the following step S503, the input images include the first image and the second image, and later, the target 3D coordinates are calculated based on the recognition results of the sign in the two images.

[0045] S503: Input the image into an image recognition network model to obtain a target image, which includes a recognition result for each of the plurality of reference points.

[0046] Optionally, the image recognition network model includes a feature extraction network, a feature fusion network, and a predictive recognition network. Step S503 may include: performing a feature extraction operation on an image by the feature extraction network to obtain a feature image set (image collection); performing a feature fusion process on the feature image set by the feature fusion network to obtain a target feature image; and performing a prediction process on the target feature image by the predictive recognition network to obtain a target image.

[0047] Optionally, the feature extraction network includes an input layer and a sub-feature extraction network. Here, the input layer generally normalizes image data and inputs it into the neural network for inference. There are several normalization methods, one of which is to normalize image pixel values ​​from 0-255 to 0-1. The normalization formula for image pixel values ​​is as follows: X1=x / 255 Here, x represents a pixel value corresponding to a pixel point of the image, and X1 represents a value after normalization processing has been performed on the pixel value of the pixel point.

[0048] For example, if x=200, normalization results in X1=0.78.

[0049] Next, image features are continuously extracted from the normalized input layer image data through convolution and activation operations.

[0050] Optionally, in the feature extraction process, the convolution kernel may be set to 1*1 or 3*3, and is not limited here.

[0051] After the convolution operation, the output of the current convolution layer is obtained by the activation function. There are several types of activation functions, usually ReLU and Sigmoid.

[0052] Optionally, the deep convolutional neural network Yolo5 KeyPoints used in this embodiment has generalization capabilities and can effectively identify specific pattern markers 3 and their five control points in various weather conditions and working environments. Yolo5 KeyPoints uses CSP-Darknet53 as the backbone feature extraction network, SPPF and CSP-PAN as feature fusion networks, and then uses a class prediction subnetwork to predict the class of each network, a detection box regression subnetwork to regress detection boxes, and a control point regression subnetwork to regress control points, thereby obtaining the final network prediction results. The control point detection loss function Yolo5 KeyPoints uses is the Wing Loss function. When the loss is large, the parameter gradient is small and it is not sensitive to outliers. When the loss is small, the parameter gradient is large and the model converges better, significantly improving the control point detection accuracy.

[0053] As can be seen from the above description, the backbone feature extraction network is generally followed by a feature fusion network, which performs feature fusion on the different scale feature images extracted by the backbone feature extraction network, and then performs further convolutional feature extraction.

[0054] The feature fusion network may include a reinforced feature extraction network, which may be implemented by a pyramid stacking convolution operation of features. One typical reinforced feature extraction network is FPN. The FPN network fuses a lower-layer high-resolution feature image with a higher-layer high-semantic feature image, and then performs independent prediction on the fused feature image of each layer to obtain a prediction result.

[0055] After the feature fusion and enhanced feature extraction network, the network prediction results, i.e., detection box regression results, control point regression results and class confidence values, are generally obtained by 1x1 convolution.

[0056] Tests have shown that when the above-mentioned position and orientation recognition algorithm for the cargo compartment 101 is deployed on a CPU with a calculation speed of approximately 130 ms per cycle, it can meet the real-time requirements in low-speed scenarios such as unmanned autonomous driving of a shovel, while when deployed on AI acceleration hardware (e.g., GPU, AI chip), the real-time performance is further improved, reaching within 30 ms per cycle.

[0057] The deep convolutional neural network Yolo5 KeyPoints is only able to distinguish specific pattern markers 3 and their five reference points through prior data collection, data labeling, and model training.

[0058] Alternatively, the process of training the image recognition network model to recognize the object is as follows. 1) A training sample dataset is obtained, the training sample dataset including each sample image in a plurality of sample images and a corresponding labeled image, each of the sample images including a sign portion 3 and a plurality of reference points located on the sign portion 3, and the labeled image is an image corresponding to a bounding box mark of the sign portion 3 in the sample image and a mark of each reference point on the sign portion 3.

[0059] In this embodiment, the sample images may be images of the sign 3 captured using a binocular camera under various weather conditions and in various work environments. After image collection is complete, images that meet the requirements are manually selected and labeled. When selecting images, images that are meaningful under various weather conditions and in various work environments and that completely contain the pattern of the sign 3 are selected, and only one image that appears multiple times is retained.

[0060] In the process of generating the labeled image, data labeling is performed using the labelme tool to obtain the bounding box of each image label part 3 and the label file of the five reference point positions.

[0061] In the specific operation process, click the Create Target Rectangle Frame button in the annotation software Labelme, create a new target frame to frame the sign 3, enter the bounding box name "Signalboard", thereby obtaining the label for sign 3, then click the Create Point Target button, click the five reference points of the sign 3 pattern, enter the reference point names: point 1 (point1), point 2 (point2), point 3 (point3), point 4 (point4), point 5 (point5) to obtain the labels for the reference points, label all sign 3 and their five reference points in the complete image, and obtain the labeled image for this image.

[0062] 2) Construct a predetermined deep learning model and determine the predetermined deep learning model as the current deep learning model.

[0063] 3) Based on the current deep learning model, perform a label prediction operation on a sample image in a sample dataset to determine a label prediction result for the sample image.

[0064] 4) Determine a loss value based on the label of the sample image and the label prediction result.

[0065] 5) If the loss value is greater than a predetermined threshold, backpropagation is performed based on the loss value, and a model weight update is performed on the current deep learning model to obtain a deep learning model after the model weight update, and the deep learning model after the model weight update is newly determined as the current deep learning model, thereby completing one iteration of training. Based on the current deep learning model, a label prediction operation is performed on the sample image, and a loss value between the label of the sample image and the annotation prediction result is determined. If the loss value is greater than a threshold, backpropagation and the steps of updating the model weight are repeated.

[0066] 6) If the loss value is less than or equal to the predetermined threshold, or if a predetermined maximum number of iterations is reached, determine the current deep learning model as the target image recognition network model.

[0067] In an alternative embodiment, the model training process may be as follows: After loading pre-training weights, labeled data is input to perform model training. The image data is normalized by a pre-processing module, and the normalized image data is sent to the network model for forward propagation to obtain a prediction result. The prediction result and the target true value of the label file are used to calculate a loss value, which is the deviation between the model output and the target true value, through a damage function. The loss value includes three parts: target classification loss, detection box regression loss, and reference point regression loss. The loss is used to update the weights of each layer of the network through backpropagation, thereby completing one training session. The model is converged through continuous iterative training, and the entire model training is completed when the convergence target or maximum number of iterations is reached.

[0068] Next, the trained Yolo5 KeyPoints model is deployed to a CPU processor, GPU, or AI chip, and then the model performs inference and prediction to obtain the identification results of the sign 3 and its five reference points in the images captured in real time by the binocular camera 4.

[0069] The model training stage obtains a model with detection capabilities through field data training. The model can be deployed to a CPU processor via the OpenCV library or Libtorh library. It can also be deployed to a GPU or AI chip via the vendor-provided libraries and deployment requirements. Model format conversion is generally required.

[0070] The process of detecting the sign unit 3 and its five reference points includes three steps. When a real-time image is input, it first goes through a preprocessing step to complete the image normalization operation, then the normalized image is sent to the Yoo5 KeyPoints model to perform forward inference to obtain the prediction results for each grid point on the image, and finally the prediction results for each grid point are post-processed, and if non-maximum suppression is used, the final prediction results are obtained.

[0071] During the autonomous operation of the unmanned excavator, the binocular camera 4 collects image data in real time, and after completing recognition of the sign unit 3 and its five reference points, the following steps S505-S513 are carried out, that is, the image coordinates of the five reference points are calculated using the stereoscopic geometric vision of the binocular camera 4 to obtain the spatial three-dimensional coordinates of the five reference points on the sign unit 3. This method provides high detection robustness and high accuracy.

[0072] S505: For each reference point, the coordinates of the reference point in the pixel coordinate system are determined based on the target image and the recognition results of the reference point.

[0073] In this embodiment, the coordinates of the reference point in the pixel coordinate system may be (u, v).

[0074] S507: The camera parameters of the binocular camera 4 are acquired.

[0075] In an alternative example, the camera parameters include the distance between the left and right eye cameras of the binocular camera 4, mounting parameters and internal parameters.

[0076] Optionally, the binocular camera 4 includes a left eye camera and a right eye camera, the distance between the left eye and the right eye is T, the mounting parameters of the binocular camera include the translation distance of each of the left eye camera or the right eye camera from the origin of the vehicle coordinate system and the rotation angle relative to the origin of the vehicle coordinate system, and the internal parameters of the binocular camera include the internal parameters of the left eye camera and the internal parameters of the right eye camera.

[0077] S509: The stereoscopic geometric vision algorithm of the binocular camera 4 determines the target three-dimensional coordinates of the reference point based on the camera parameters of the binocular camera 4 and the coordinates of the reference point in the pixel coordinate system.

[0078] S511: The target three-dimensional coordinates of the luggage compartment 101 are determined based on the target three-dimensional coordinates of each reference point.

[0079] S513: The posture information of the luggage compartment 101 is determined based on the target three-dimensional coordinates of the luggage compartment 101.

[0080] Referring to Figure 6, Figure 6 is a flowchart of another selectable luggage compartment attitude recognition method of the present invention. Steps S509-S513 can be specifically detailed as follows:

[0081] S601: The coordinate of the reference point in the camera coordinate system is determined based on the distance between the left and right cameras of the binocular camera 4, the internal parameters of the binocular camera 4, and the coordinate of the reference point in the pixel coordinate system.

[0082] Optionally, the internal parameters of the binocular cameras include internal parameters of a left-eye camera and internal parameters of a right-eye camera.

[0083] Optionally, the following provides an example of determining the coordinates of the selectable reference points in the camera coordinate system. Referring to Figure 7, Figure 7 is a schematic diagram showing the relationship between multiple selectable coordinate systems of the present invention. World coordinate system: Xw, Yw, Zw; camera coordinate system: Xc, Yc, Zc; image coordinate system: x, y; pixel coordinate system: u, v (reflecting the pixel arrangement situation of the camera CCD chip).

[0084] As can be seen from the figure, (u0, v0) represents the coordinates in the uv coordinate system of O. Assuming that the length and width of one pixel are dx and dy, respectively, the relationship between the pixel coordinate system and the image coordinate system is as follows: JPEG0007743642000001.jpg30170If equations (1) and (2) are combined and written as a matrix, the result is as follows. JPEG0007743642000002.jpg42170The coordinates in the image coordinate system, that is, (x, y), of any pixel point on the image can be found using the above formula (3).

[0085] FIG. 8 is a schematic diagram of selectable camera models of the present invention. l , O r are the projection centers of the left and right eyes of the binocular camera 4, and the connecting line D between them is the binocular baseline, i.e., the distance between the center points of the two cameras. P is a point in space, and P l is the image point of point P in the left eye, and P r is the image point of point P in the right eye, and the disparity d=x l -x r is. According to the similar triangle theorem, ΔPO l O r is ΔPP l P r The depth Z can be calculated using the following formula: JPEG0007743642000003.jpg19170 where f is the focal length of the camera.

[0086] After obtaining the depth information Z, i.e., Zc, using the above equation (4), it is possible to further determine the Xc and Yc of the coordinates in the target binocular camera coordinate system of P (binocular camera coordinate system, i.e., a coordinate system with the center of the left eye camera as the origin, i.e., the camera coordinate system in the coordinate relationship diagram 7).

[0087] 9, which is a schematic diagram of the relationship between another selectable coordinates of the present invention. In the left-eye camera coordinate system of the binocular camera, the image point P in the left-eye image of the spatial point P is l From this, the horizontal coordinate X and vertical coordinate Y of spatial point P can be calculated using the similar triangle theorem. The coordinates of P(X, Y, Z) in the left camera coordinate system of the binocular camera are P(Xc, Yc, Zc), and P(x, y) are the image coordinates of the imaging point of spatial point P(Xc, Yc, Zc) in the left eye image coordinate system of the binocular camera.

[0088] ΔABO c ~ΔoCO c , ΔPBO c ~ΔpCO c Therefore, JPEG0007743642000004.jpg18170 can be derived. From the derivation of the above equation, Xc and Yc are It can be determined that the image is JPEG0007743642000005.jpg33170. At this point, all coordinates of spatial point P (Xc, Yc, Zc) have been determined. In this way, the three-dimensional coordinates of the five reference points on reference point 3 in the camera coordinate system can be determined sequentially, and the coordinates are coordinates in the left-eye camera coordinate system.

[0089] S603: Determine a first coordinate transformation matrix based on the mounting parameters of the binocular camera.

[0090] S605: According to the first coordinate transformation matrix, transform the three-dimensional coordinates of the reference point in the camera coordinate system into three-dimensional coordinates in a target vehicle coordinate system, where the target vehicle coordinate system is the coordinate system in which the second vehicle 2 is located.

[0091] In this embodiment, assuming that the origin of the vehicle coordinate system is the rotation center point of the second vehicle 2, the corresponding rotation matrix R and translation vector T can be determined by determining the translation distance and rotation angle of the binocular camera from the origin. When the second vehicle is a shovel, the second vehicle includes a robot arm and a base, a rotation end of the robot arm is rotatably connected to the base, and the rotation center point of the second vehicle is located on the rotation axis on which the rotation end is located.

[0092] The coordinate transformation matrix that transforms the three-dimensional coordinates of the reference point in the camera coordinate system into three-dimensional coordinates in the second vehicle 2 coordinate system is expressed as follows: JPEG0007743642000006.jpg20170According to equation (7), the relationship between the camera coordinate system and the coordinate system of the second vehicle 2 is as follows: Optionally, the above coordinate transformation matrix, i.e., equation (7), also differs depending on the definition of the vehicle coordinate system.

[0093] S607: The three-dimensional coordinates of the luggage compartment 101 in the target vehicle coordinate system are determined based on the three-dimensional coordinates of each reference point in the target vehicle coordinate system.

[0094] Alternatively, the sign unit 3 may include two sub-sign units 301. For example, if each sub-sign unit 301 includes five reference points, the coordinates of the five reference points for each sub-sign unit 301 may be sorted, and the position coordinate data of the reference points with large numerical deviations may be removed. The coordinate data of the remaining reference points that meet the requirements may be averaged to obtain the three-dimensional coordinate data of the sub-sign unit 301. In another alternative embodiment, weights may be assigned to the coordinates of reference points located at different positions on the sub-sign unit 301, and the three-dimensional coordinate data of the sub-sign unit 301 may be obtained by multiplying the coordinates of each reference point by the corresponding weight.

[0095] S609: The posture information of the luggage compartment 101 is determined based on the three-dimensional coordinates of the luggage compartment 101 in the target vehicle coordinate system.

[0096] The attitude information of the luggage compartment 101 can be determined based on the three-dimensional coordinates of the two sub-sign portions 301 in the second vehicle-2 coordinate system.

[0097] This attitude information is, for example, information on the relative distance and angle between the luggage compartment 101 and the second vehicle 2.

[0098] In a possible embodiment, the image recognition module includes a feature extraction sub-module, a feature fusion sub-module, and a predictive recognition module. The feature extraction sub-module performs a feature extraction operation on an image using a feature extraction network to obtain a set of feature images. The feature fusion sub-module performs a feature fusion process on the set of feature images using a feature fusion network to obtain a target feature image. The predictive recognition module performs a prediction process on the target feature image using a predictive recognition network to obtain a target image.

[0099] In a possible embodiment, the position determining module includes a pixel coordinate determining module, a camera parameter obtaining module, a target 3D coordinate determining module, and an attitude information determining module. A pixel coordinate determination module is used for determining, for each reference point, the coordinate of the reference point in a pixel coordinate system based on the target image and the recognition result of the reference point. The camera parameter acquisition module is used to acquire the camera parameters of the binocular camera 4 . The target 3D coordinate determination module is used to determine the target 3D coordinates of the reference points based on the camera parameters of the binocular camera 4 and the coordinates of the reference points in the pixel coordinate system using the stereoscopic geometric algorithm of the binocular camera 4, and to determine the target 3D coordinates of the luggage compartment 101 based on the target 3D coordinates of each reference point. The posture information determination module is used to determine posture information of the luggage compartment 101 based on the target three-dimensional coordinates of the luggage compartment 101 .

[0100] In a possible embodiment, the target three-dimensional coordinate determining module includes a first coordinate determining module, a first coordinate transformation matrix, and a second coordinate determining module. The first coordinate determination module is used to determine the three-dimensional coordinate of the reference point in the camera coordinate system based on the distance between the left eye camera and the right eye camera of the binocular camera 4, the internal parameters of the binocular camera 4, and the coordinate of the reference point in the pixel coordinate system. The first coordinate transformation matrix is ​​used to determine a first coordinate transformation matrix based on mounting parameters of the binocular camera. The second coordinate determination module transforms the three-dimensional coordinates of the reference points in the camera coordinate system into three-dimensional coordinates in the target vehicle coordinate system according to the first coordinate transformation matrix, and the target vehicle coordinate system is the coordinate system in which the second vehicle 2 is located, and is used to determine the three-dimensional coordinates of the luggage compartment 101 in the target vehicle coordinate system based on the three-dimensional coordinates of each reference point in the target vehicle coordinate system. The attitude information determination module is used to determine the attitude information of the luggage compartment 101 based on the three-dimensional coordinates of the luggage compartment 101 in the target vehicle coordinate system.

[0101] The specific implementation process of the above modules is the same as that of the above method, so the description will be omitted here. Note that the above modules may be actual sub-modules in a processor, or virtual modules configured by a program.

[0102] As described above, the luggage compartment posture recognition system provided by the present invention has the following selectable operation steps: The trained deep convolutional neural network model Yolo5 KeyPoints is deployed to the CPU. At this time, the time required for the neural network model Yolo5 KeyPoints to perform one detection and calculation to obtain the sign 3 and its five reference points is approximately 130 ms. If the calculation unit includes an image calculation unit GPU or AI accelerator chip, the neural network model Yolo5 KeyPoints can be deployed to the image calculation unit GPU or AI accelerator chip to improve real-time performance. In this case, the time required for the neural network model Yolo5 KeyPoints to perform one detection and calculation to obtain the sign 3 and its five reference points is approximately 30 ms. A five-reference-point space 3D coordinate algorithm, which calculates the five reference points using binocular stereoscopic geometric vision, is deployed to the CPU.

[0103] When power is supplied to the system, the binocular camera 4 confirms the position of the cargo compartment 101 of the truck, the attitude sensing system is activated, the unmanned excavator control system is activated, and the unmanned excavator enters an unmanned autonomous operation state.

[0104] The driver parks the first vehicle 1 (for example, a truck) in a parking space and prepares to load food.

[0105] After the truck has finished parking, the binocular camera 4 sends the captured left-eye and right-eye 1280x720 RGB images to the deep convolutional neural network Yolo5 KeyPoints model. The neural network Yolo5 KeyPoints model preprocesses the left-eye and right-eye 1280x720 RGB images before inference. This preprocessing includes image normalization and image scaling. In the Yolo5 KeyPoints model preprocessing, image normalization involves dividing the RGB values ​​of each pixel in the image by 127.5 and subtracting 1 to normalize all pixel values ​​to between -1 and 1. The image scaling operation scales the original 1280x720 RGB image to a 1280x736 RGB image to meet the model input image size requirements. After preprocessing, the image data is sent to the Yolo5 KeyPoints model for inference and model output. The model output includes object classification results for each anchor box at each grid point, detection box regression results, and five control point regression results. The Yolo5 KeyPoints model then performs post-processing such as non-maximum suppression to obtain the final sign 3 target detection box and its five reference point pixel coordinates.

[0106] After obtaining the pixel coordinates of the five reference points on the sign 3 in the left and right eye images of the binocular camera 4 through neural network inference, the pixel coordinates of the five reference points are calculated through the stereoscopic geometric vision of the binocular camera 4 to obtain the three-dimensional spatial coordinates of the five reference points.The luggage compartment posture recognition system has a simple structure and is applicable to large-scale outdoor unmanned driving scenes.

[0107] An embodiment of the present invention further provides an electronic device including a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, the at least one program, code set, or instruction set is loaded and executed by the processor to realize the above-described luggage compartment posture recognition method.

[0108] An embodiment of the present invention further provides a computer storage medium, which is provided in a server and implements at least one instruction, at least one program, code set, or instruction set related to the luggage compartment posture recognition method in a method embodiment, and the at least one instruction, the at least one program, code set, or instruction set is loaded and executed by the processor to implement the luggage compartment posture recognition method.

[0109] Optionally, in this embodiment, the storage medium may be located in at least one network server among a plurality of network servers of a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB disk, a read-only memory (ROM), a random access memory (RAM), a portable hard disk, a magnetic disk, or an optical disk.

[0110] It should be noted that the order of the above-described embodiments of the present invention is for illustrative purposes only and does not indicate superiority or inferiority of the embodiments. Certain embodiments of the present specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve desirable results. Also, the processes depicted in the figures do not necessarily achieve the desired results by following only the particular order or sequential order shown. In some embodiments, multitasking and parallel processing may also be possible or advantageous.

[0111] Each embodiment in this specification will be described step by step, and the same or similar parts between the embodiments can be referred to. The parts that are mainly described in each embodiment are the differences from other embodiments. In particular, the device embodiments are basically the same as the method embodiments, so the explanation is relatively simple. For the relevant parts, please refer to the explanation in the method embodiments.

[0112] It is understandable to those skilled in the art that all or part of the steps in the above embodiments may be realized by hardware, or may be completed by instructing relevant hardware by a program, and the program may be stored in a computer-readable storage medium, and the storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.

[0113] The above description is only a preferred embodiment of the present invention, and does not limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included within the protection scope of the present invention.

Claims

1. A luggage compartment attitude recognition system including a processor (5), a binocular camera (4), a first vehicle (1), a second vehicle (2), and a sign unit (3), The first vehicle (1) includes a luggage compartment (101), The second vehicle (2) is used to unload materials into the luggage compartment (101) of the first vehicle (1); The sign (3) is provided on a side surface of the luggage compartment (101), and a plurality of reference points are provided on the sign (3), the binocular camera (4) is provided on the second vehicle (2), and the binocular camera (4) is used to collect an image including the sign portion (3) and transmit the image to the processor (5); The processor (5) is connected in communication with the binocular camera (4), and the processor (5) recognizes and processes the image using an image recognition network model to obtain a target image including recognition results of the plurality of reference points, determines, for each of the reference points, coordinates of the reference points in a pixel coordinate system based on the target image and the recognition results of the reference points, and determines, using a stereoscopic geometric vision algorithm of the binocular camera (4), coordinates of the reference points in a camera coordinate system based on camera parameters of the binocular camera (4) and the coordinates of the reference points in the pixel coordinate system, and a first coordinate transformation matrix is ​​determined based on the mounting parameters of the reference point, and three-dimensional coordinates of the reference point in the camera coordinate system are transformed into three-dimensional coordinates of the reference point in a target vehicle coordinate system according to the first coordinate transformation matrix, the target vehicle coordinate system being a coordinate system in which the second vehicle is located; three-dimensional coordinates of the luggage compartment (101) in the target vehicle coordinate system are determined based on the three-dimensional coordinates of each of the reference points in the target vehicle coordinate system; and attitude information of the luggage compartment (101) is determined based on the three-dimensional coordinates of the luggage compartment (101) in the target vehicle coordinate system.

2. The second vehicle (2) includes an unloading device (201) and a body structure (202) connected to each other; The unloading device (201) is connected to a first side (2021) of the body structure (202); 2. The luggage compartment posture recognition system according to claim 1, wherein the binocular camera (4) is provided on the first side surface (2021).

3. The first side (2021) includes a first attachment point and a second attachment point; The first attachment point and the second attachment point are located on the upper part of the first side surface (2021), a first preset distance exists between the first attachment point and the second attachment point; The binocular camera (4) includes a first binocular camera (401) and a second binocular camera (402), The first binocular camera (401) is mounted at the first mounting point; 3. The luggage compartment posture recognition system according to claim 2, wherein the second binocular camera (402) is provided at the second mounting point.

4. The sign portion (3) includes at least two sub-sign portions (301), At least two of the sub-signal portions (301) are respectively located at a first position point and a second position point on a second side surface of the luggage compartment (101); The second side surface is a surface facing the first side surface (2021), 3. The luggage compartment posture recognition system of claim 2, wherein a second preset distance exists between the first position point and the second position point, the second preset distance being longer than half the length of the second side in a first predetermined direction, the first predetermined direction being perpendicular to a second predetermined direction, and the second predetermined direction being an extension of the height direction of the first vehicle (1).

5. A method for recognizing a luggage compartment attitude by the luggage compartment attitude recognition system according to any one of claims 1 to 4, comprising: Collecting an image including the sign portion (3) using a binocular camera (4); inputting the image into an image recognition network model to obtain a target image, the target image including a recognition result for each of the plurality of reference points; For each of the reference points, determine the coordinates of the reference point in a pixel coordinate system based on the target image and the recognition result of the reference point; Obtaining the camera parameters of the binocular camera (4); determining the three-dimensional coordinates of the reference point in the target space based on the camera parameters of the binocular camera (4) and the coordinates of the reference point in the pixel coordinate system by a binocular camera stereoscopic geometric vision algorithm; Determine the three-dimensional coordinates of the target space of the luggage compartment (101) based on the three-dimensional coordinates of the target space of each reference point; Determining posture information of the luggage compartment (101) based on three-dimensional coordinates of the target space of the luggage compartment (101); A luggage compartment posture recognition method comprising:

6. 6. The method for recognizing the posture of a luggage compartment according to claim 5, wherein the camera parameters include a distance between a left eye camera and a right eye camera of a binocular camera (4), an installation parameter, and an internal parameter.

7. Determining three-dimensional coordinates of the target space of the reference point based on the camera parameters of the binocular camera (4) and the coordinates of the reference point in a pixel coordinate system, determining three-dimensional coordinates of the target space of the luggage compartment (101) based on the target three-dimensional coordinates of each reference point, and determining attitude information of the luggage compartment (101) based on the target three-dimensional coordinates of the luggage compartment (101) determining three-dimensional coordinates of the reference point in a camera coordinate system based on the distance between the left eye camera and the right eye camera of the binocular camera (4), internal parameters of the binocular camera, and coordinates of the reference point in a pixel coordinate system; determining a first coordinate transformation matrix based on mounting parameters of the binocular camera; transforming the three-dimensional coordinates of the reference point in the camera coordinate system into three-dimensional coordinates in a target vehicle coordinate system according to the first coordinate transformation matrix, the target vehicle coordinate system being a coordinate system in which the second vehicle is located; determining three-dimensional coordinates of the luggage compartment in the target vehicle coordinate system based on three-dimensional coordinates of each reference point in the target vehicle coordinate system; Determining attitude information of the luggage compartment (101) based on three-dimensional coordinates of the luggage compartment in the target vehicle coordinate system; The method for recognizing the attitude of a luggage compartment according to claim 6, further comprising:

8. An electronic device, 6. An electronic device comprising: a processor; and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and wherein the at least one instruction, the at least one program, code set, or instruction set is loaded and executed by the processor to realize the luggage compartment posture recognition method described in claim 5.

9. 1. A computer storage medium, comprising: The computer storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to realize the luggage compartment posture recognition method described in claim 5.

Citation Information

Patent Citations

  • Carriage pose determination method and device

    CN115170648A

  • Automatic positioning device for container truck under quayside bridge

    CN2890845Y

  • Remote operation system and remote operation method

    JP2021151909A

  • Carriage guide device of crane

    JP2021172502A