A method for high-dynamic target recognition and tracking of a micro unmanned aerial vehicle with a three-rotor layout
Through the tri-rotor layout of micro-drone, it is equipped with dual cameras and cloud computing, and uses a random forest training network for target recognition and tracking, solving the problem of limited computing resources of micro-drones and achieving fast and accurate target recognition and tracking.
Patent Information
- Application Number
- CN202310034220.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-01-10
AI Technical Summary
The limited computing resources of existing micro-UAVs have resulted in slow target recognition and tracking speeds, making them unable to be suitable for rapidly changing battlefield environments.
A mini-drone with a tri-rotor layout is equipped with a mini dual camera to obtain color maps and depth maps, and a random forest training network is used to perform distributed parallel computing in the cloud, combining stereo vision and small sample tracking prediction methods to achieve target recognition and tracking.
Implement second-level identification and tracking in complex environments, improve target recognition rate, prevent misleading, and improve the target tracking speed of onboard computers.
Smart Images

Figure CN115953703B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the research field of automatic target recognition (ATR) of a micro-UAV, and relates to a high-dynamic target recognition and tracking method of a micro-UAV with a three-rotor layout. Background Art
[0002] A micro aerial vehicle (MAV) is an unmanned aerial vehicle with size restrictions and can fly autonomously. Compared with traditional aircraft, its main feature is its extremely small size, which allows it to perform tasks in complex environments and narrow spaces, and has the advantage of being difficult to be discovered. MAVs have high research value and application prospects in the military, and can be used to perform a variety of tasks such as search, tracking, detection, and military strikes. The military is increasingly relying on micro-UAV technology to monitor, scout, and strike potential threats to minimize harm to military personnel. The military is exploring and developing the capabilities of micro-UAVs, including enabling micro-UAVs to autonomously identify opponents and their key assets, autonomously decide on action plans, and engage the enemy without direct intervention from central command and control.
[0003] UAV target recognition and tracking technology is an emerging military technology that can provide solutions to many problems on the modern battlefield. Many current target recognition technologies are based on deep learning networks. Although deep learning networks can greatly improve the recognition rate of targets, the computing resources of UAVs are very limited due to the power supply, cabin space and load capacity. Moreover, the computing resources are often provided to key flight equipment such as flight control systems and navigation systems. The computing resources and real-time performance provided to mission equipment cannot be guaranteed. In addition, the frequency of UAV onboard computers is low. According to statistics, it takes at least 10 minutes to identify 10 categories of soldiers, tanks, vehicles and other targets. After the calculation is completed, the enemy target has disappeared, which is not suitable for the rapidly changing battlefield. Summary of the invention
[0004] The purpose of the present invention is to solve the problems in the prior art and to provide a high-dynamic target recognition and tracking method for a micro-UAV with a three-rotor layout.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A high-dynamic target recognition and tracking method for a micro-UAV with a three-rotor layout comprises the following steps:
[0007] Obtain the color image and depth image of the target scene, and calculate the mathematical features and local features of the objects in the target scene;
[0008] Obtain the color images and depth images of multiple types of scenarios, and calculate the reference mathematical features and reference local features;
[0009] Use the obtained reference mathematical features and reference local features as the total features to input into the random forest training network for training, and obtain the reference feature vector;
[0010] Use the mathematical features and local features of the target scenario object as the total features and input them into the random forest training network according to the training method to obtain the target feature vector;
[0011] Compare the target feature vector with the reference feature vector, identify the category of the target scenario, and perform tracking according to the identification result.
[0012] Furthermore, the mathematical features include brightness, color, and texture; the mathematical features are obtained by converting the color image into the form of a pixel matrix and then using a feature analysis tool. The mathematical form of the pixel matrix is:
[0013] X = X(m×n)
[0014] Where X represents the mathematical symbol of the color image, m represents the number of columns of the pixel matrix of the color image, n represents the number of rows of the pixel matrix of the color image, and each element in the pixel matrix represents the rgb value of the pixel at that position.
[0015] Furthermore, the brightness feature of the color image is obtained by inputting the pixel matrix into a brightness analysis tool. The mathematical expression of the brightness analysis tool is:
[0016] LD = ∑rgb(i) / μ
[0017] Where LD represents the brightness feature of the color image, rgb(i) represents the rgb value of the i-th pixel, and μ represents the brightness gradient parameter.
[0018] Furthermore, the color feature of the color image is obtained by inputting the pixel matrix into a color analysis tool. The mathematical expression of the color analysis tool is:
[0019]
[0020] Where YS represents the color feature of the color image, represents the average rgb value of each pixel.
[0021] Furthermore, the texture feature of the color image is obtained by inputting the pixel matrix into a texture analysis tool. The mathematical expression of the texture analysis tool is:
[0022]
[0023] Among them, WL represents the texture feature of the color map, which represents the rgb average value of the i-th pixel.
[0024] Furthermore, the local features include depth gradient, convex normal vector gradient, and concave normal vector gradient.
[0025] Furthermore, the processing layer of the random forest training network is a conditional random field, and the number of layers of the processing layer is greater than or equal to 50 layers.
[0026] Furthermore, the training process of the random forest training network is as follows:
[0027] The first-layer conditional random field in the random forest training network receives the total feature input and outputs the feature vector of the first layer as:
[0028] p(1) = exp{w1∑f1(x i ) + w2∑g1(x i )} + p(0)
[0029] Among them, p(1) represents the feature vector of the first layer, p(0) represents the initial feature vector, w1 and w2 represent model parameters, f1 and g1 respectively represent the feature function of the first layer's feature and label, the relationship function between adjacent features and labels, and x i represents the total feature input;
[0030] The second-layer conditional random field in the random forest training network receives the total feature input and the feature vector of the first layer, and outputs the feature vector of the second layer as:
[0031] p(2) = exp{w1∑f2(x i ) + w2∑g2(x i )} + p(1)
[0032] Among them, p(2) represents the feature vector of the second layer, p(1) represents the feature vector of the first layer, w1 and w2 represent model parameters, f2 and g2 respectively represent the feature function of the second layer's feature and label, the relationship function between adjacent features and labels, and x i represents the total feature input;
[0033] The third-layer conditional random field in the random forest training network receives the total feature input and the feature vector of the second layer, and outputs the feature vector of the third layer as:
[0034] p(3) = exp{w1∑f3(x i ) + w2∑g3(x i )} + p(2)
[0035] Among them, p(3) represents the feature vector of the third layer, p(2) represents the feature vector of the second layer, w1 and w2 represent model parameters, f3 and g3 respectively represent the feature function of the features of the third layer and the label, and the relationship function between adjacent features and the label, x i represents the total feature input;
[0036] The feature vector output by the conditional random field of the nth layer in the random forest training network is:
[0037] p(n) = exp{w1∑f n (x i ) + w2∑g n (x i )} + p(n - 1)
[0038] Among them, p(n) represents the feature vector of the nth layer, p(n - 1) represents the feature vector of the (n - 1)th layer, w1 and w2 represent model parameters, fn and g n respectively represent the feature function of the features of the nth layer and the label, and the relationship function between adjacent features and the label, x i represents the total feature input.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] The present invention provides a method for high-dynamic target recognition and tracking of a micro-unmanned aerial vehicle with a three-rotor layout. The target video captured by the unmanned aerial vehicle is transmitted to the ground station cloud for processing. Using a random forest training network, it supports cloud distributed parallel computing to obtain the target feature vector, and compares and searches the target feature vector with the reference feature vector to identify the target scene category and track it. The present invention can improve the target intelligent perception ability and target recognition ability, and can identify time-sensitive targets in complex environments such as mountains, urban building complexes, forests, sea surfaces, and low altitudes, preventing misidentification and misleading caused by incorrect recognition and missing opportunities due to slow recognition. The recognition efficiency through air-ground cooperation can reach the second level, significantly improving the target recognition rate of the micro-unmanned aerial vehicle with a three-rotor layout. Moreover, through the stereo vision introduced by carrying a micro double camera, the three-dimensional perception of the scene is more accurate. Using the recognition result and stereo vision positioning as prior knowledge, the target tracking introduces a small-sample tracking prediction method, greatly accelerating the tracking speed of the on-board computer and improving the real-time tracking effect of time-sensitive targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 These are multi - perspective views of the micro - unmanned aerial vehicle with a three - rotor layout of the present invention.
[0043] Figure 2 This is a diagram of the micro - optoelectronic sensor carried by the micro - unmanned aerial vehicle with a three - rotor layout of the present invention.
[0044] Figure 3 This is a diagram of the environmental image collected by the optoelectronic sensor carried by the micro - unmanned aerial vehicle with a three - rotor layout of the present invention.
[0045] Figure 4 This is a flowchart of the target scene understanding based on deep structure learning of the present invention.
[0046] Wherein: 10 - thin - walled duct, 12 - fixed rotor, 14 - movable tail rotor, 16 - unmanned aerial vehicle fuselage, 18 - binocular camera, 20 - data transmission interface;
[0047] Figure 1 (a) is the front view of the micro - unmanned aerial vehicle with a three - rotor layout, Figure 1 (b) is the left view of the micro - unmanned aerial vehicle with a three - rotor layout, Figure 1 (c) is the top view of the micro - unmanned aerial vehicle with a three - rotor layout, Figure 1 (d) is the three - dimensional schematic diagram of the micro - unmanned aerial vehicle with a three - rotor layout. Detailed implementation manners
[0048] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well - known functions and structures are omitted below.
[0049] Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.
[0050] It should be noted that the terminals involved in the embodiments of the present application may include, but are not limited to, mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers, personal computers (PCs), MP3 players, MP4 players, wearable devices (such as smart glasses, smart watches, smart bracelets, etc.), smart home devices and other intelligent devices.
[0051] In addition, the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the preceding and following associated objects.
[0052] The following further describes the present invention in detail with reference to the accompanying drawings:
[0053] See Figures 1 to 3 , the present invention provides a method for high-dynamic target recognition and tracking of a micro-unmanned aerial vehicle with a three-rotor layout. The micro-unmanned aerial vehicle with a three-rotor layout is a low-cost military equipment that can receive and download control instructions before being released from a mobile airborne platform and is equipped with a suitable visual sensor, such as a binocular vision sensor, which can identify and track targets. According to the different ways of generating lift, micro-aircraft are mainly divided into three main configurations: fixed-wing, flapping-wing, and rotor. In the present invention, a multi-rotor unmanned aerial vehicle with a three-rotor layout is used as a platform, which has the following advantages:
[0054] 1. The development and application of control algorithms have been quite mature, and the control efficiency is relatively high. It is a relatively practical layout form of micro-aircraft.
[0055] 2. It meets the requirements of high dynamics and aggressiveness, and the setting of the thin-walled duct 10 outside the rotor can improve the efficiency of the rotor system and increase the use safety.
[0056] 3. The three-rotor layout form can improve the maneuverability. The two fixed rotors 12 on the left and right provide most of the lift and roll control forces, and a movable tail rotor 14 provides pitch and yaw control forces and provides a certain amount of lift. Since all three-axis torques are directly generated by the rotor thrust, it is expected to further improve the aircraft's maneuvering ability.
[0057] 4. Compared with conventional multi-rotor unmanned aerial vehicles, the three-rotor layout has a lighter structural weight and is more convenient for carrying and transportation.
[0058] The micro unmanned aerial vehicle with a three-rotor layout of the present invention includes a ground platform and is based on the ground platform. The micro unmanned aerial vehicle with a three-rotor layout is stored and released by the ground platform. A pair of micro optoelectronic sensors are carried on the micro unmanned aerial vehicle with a three-rotor layout, including an image acquisition module and a transmission module. The paired sensors are installed in parallel to collect visible light images. The paired micro optoelectronic sensors are responsible for shooting ground scene images and videos on the one hand, and forming a binocular measurement system to measure the positioning of scene targets in real time on the other hand. The image transmission module provides to the on-board computer on the one hand, and dynamically shoots the scene through two carried cameras on the other hand, and transmits the scene information to the ground station cloud in real time through the air-ground data link. The cloud target database is used to identify the targets appearing in the scene, and the recognition results are uploaded to the on-board computer to perform real-time three-dimensional positioning and high-dynamic tracking of the targets.
[0059] The micro unmanned aerial vehicle with a three-rotor layout is released from the ground platform and flies towards the designated target area, which is calibrated by the ground station. Here, the ground platform is a ground platform capable of storing and deploying small micro unmanned aerial vehicles with a three-rotor layout. The target area can be any enemy target area that needs to be reconnoitered, including radar devices, air defense systems, enemy tanks and vehicles, etc. A micro unmanned aerial vehicle with a three-rotor layout is released within a certain preset distance (such as 1 - 3 km) close to the target engagement area. Before release, the micro unmanned aerial vehicle with a three-rotor layout downloads target information and other data from the ground station, and identifies the specific target positions in the target area. After release, it flies towards the target area and finally reaches a specific position in the target area.
[0060] Refer to Figure 4 , the present invention dynamically shoots the scene through two carried cameras, transmits the scene information to the ground station cloud in real time through the air-ground data link, and uses the cloud target database to identify the targets appearing in the scene. Compared with the limited resources of the on-board computer, the powerful parallel computing function of the ground cloud can be utilized to quickly and automatically identify the targets, and the recognition results are transmitted to the unmanned aerial vehicle. The unmanned aerial vehicle performs target tracking accordingly. When the target disappears briefly in a complex environment, the ground cloud recognition function can be restarted until the target appears again.
[0061] A high-dynamic target recognition and tracking method for a micro unmanned aerial vehicle with a three-rotor layout provided by the present invention includes the following steps:
[0062] Step 1: Obtain a color map of the target scene, and calculate the brightness, color, texture and other mathematical features of the objects in the target scene by using feature analysis tools such as brightness, color and texture.
[0063] (1) Convert the scene color map into the form of a pixel matrix, and the mathematical form is:
[0064] X = X(m×n)(1)
[0065] Among them, X represents the mathematical symbol of the scene color map, m represents the number of columns of the pixel matrix of the scene color map, n represents the number of rows of the pixel matrix of the scene color map, and each element in the pixel matrix represents the rgb value of the pixel at that position.
[0066] (2) Substitute the pixel matrix into the brightness analysis tool to obtain the brightness feature of the scene color map. The mathematical expression of the brightness analysis tool is:
[0067] LD = ∑rgb(i) / μ (2)
[0068] Among them, LD represents the brightness feature of the scene color map, rgb(i) represents the rgb value of the i-th pixel, and μ represents the brightness gradient parameter, usually taken as 500.
[0069] (3) Substitute the pixel matrix into the color analysis tool to obtain the color feature of the scene color map. The mathematical expression of the color analysis tool is:
[0070]
[0071] In the formula, YS represents the color feature of the scene color map, represents the average rgb value of each pixel. Therefore, YS is a matrix of size m×n.
[0072] (4) Substitute the pixel matrix into the texture analysis tool to obtain the texture feature of the scene color map. The mathematical expression of the texture analysis tool is:
[0073]
[0074] Among them, WL represents the texture feature of the scene color map, represents the average rgb value of the i-th pixel.
[0075] Step 2: Use the stereo vision system mounted on the drone to obtain the depth map of the target scene. This depth map contains three local features: depth gradient SD, convex normal vector gradient TF, and concave normal vector gradient AF. These features are all position-independent quantities. Among them, the depth gradient SD represents the discontinuity of the depth value, the convex normal vector gradient TF represents the degree of outward bending of the pixel point, and the concave normal vector gradient AF describes the degree of inward bending of the pixel point from the shooting point, reflecting the surface characteristics of the object. These features can all be directly obtained from the depth map without using mathematical tools.
[0076] Step 3: Collect color images and depth images of multiple types of scenes (no less than 5), ensuring that the scene instances of the same type are as diverse as possible. At the same time, label these color images and depth images, especially the important objects in the images. These labeled color images and depth images will be used as training samples for the subsequent training of the random forest. Perform the content of Step 1 and Step 2 on the labeled color images and depth images to obtain the reference brightness feature, reference color feature, reference texture feature, reference depth gradient SD, reference convex normal vector gradient TF, and reference concave normal vector gradient AF.
[0077] Step 4: Build a random forest training network as shown in Figure 4 and use Conditional Random Fields (CRF) as the processing layer in the network.
[0078] (1) Use the reference brightness feature, reference color feature, reference texture feature, reference depth gradient SD, reference convex normal vector gradient TF, and reference concave normal vector gradient AF obtained in Step 3 as the total features and input them into the random forest training network;
[0079] (2) The first-layer conditional random field CRF in the random forest training network receives the total feature input and outputs the feature vector of the first layer according to the following formula:
[0080] p(1) = exp{w1∑f1(x i ) + w2∑g1(x i )} + p(0) (5)
[0081] where p(1) represents the feature vector of the first layer, p(0) represents the artificially set initial feature vector, w1 and w2 represent model parameters, f1 and g1 are the unary and binary feature functions of the first layer, representing the feature function of the feature and the label, and the relationship function between adjacent features and the label respectively, and x i represents the total feature input.
[0082] (3) The second-layer conditional random field CRF in the random forest training network receives the total feature input and the feature vector of the first layer, and outputs the feature vector of the second layer according to the following formula:
[0083] p(2) = exp{w1∑f2(x i ) + w2∑g2(x i )} + p(1) (6)
[0084] where p(2) represents the feature vector of the second layer, p(1) represents the feature vector of the first layer, w1 and w2 represent model parameters, f2 and g2 are the unary and binary feature functions of the second layer, representing the feature function of the feature and the label, and the relationship function between adjacent features and the label respectively, and xi Represents the total feature input.
[0085] (4) The third-layer conditional random field CRF in the random forest training network receives the total feature input and the feature vectors of the second layer, and outputs the feature vectors of the third layer according to the following formula:
[0086] p(3) = exp{w1∑f3(x i ) + w2∑g3(x i )} + p(2) (7)
[0087] Where p(3) represents the feature vectors of the third layer, p(2) represents the feature vectors of the second layer, w1 and w2 represent model parameters, f3 and g3 are the unary and binary feature functions of the third layer, representing the feature function of the feature and the label, and the relationship function between adjacent features and the label respectively, and x i Represents the total feature input.
[0088] (4) Similarly, the feature vectors output by the nth-layer conditional random field CRF in the random forest training network are obtained as:
[0089] p(n) = exp{w1∑f n (x i ) + w2∑g n (x i )} + p(n - 1) (8)
[0091] Where p(n) represents the feature vectors of the nth layer, p(n - 1) represents the feature vectors of the (n - 1)th layer, w1 and w2 represent model parameters, f n and g n are the unary and binary feature functions of the nth layer, representing the feature function of the feature and the label, and the relationship function between adjacent features and the label respectively, and x i Represents the total feature input.
[0092] The number of layers of the random forest training network should not be less than 50 layers. After n layers of training, reference feature vectors are obtained, and each labeled scene corresponds to a reference feature vector.
[0093] (5) Input the unlabeled target scene into the random forest training network in the same way, obtain the target feature vectors, compare and search the target feature vectors with the reference feature vectors, then the category of the target scene can be identified, and tracking can be performed according to the recognition result.
[0094] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A high-dynamic target recognition and tracking method for a micro unmanned aerial vehicle with a three-rotor layout, characterized in that, It includes the following steps: Obtain the color image and depth image of the target scene, and calculate the mathematical features and local features of the objects in the target scene; Obtain the color images and depth images of multiple types of scenes, and calculate the reference mathematical features and reference local features; Use the obtained reference mathematical features and reference local features as the total features to input into the random forest training network for training to obtain the reference feature vector; Use the mathematical features and local features of the objects in the target scene as the total features and input them into the random forest training network according to the training method to obtain the target feature vector; Compare the target feature vector with the reference feature vector, identify the category of the target scene, and perform tracking according to the recognition result.
2. A method for high-dynamic target recognition and tracking of a micro-unmanned aerial vehicle with a three-rotor layout according to claim 1, characterized in that, The mathematical features include brightness, color, and texture; the mathematical features are obtained by converting the color image into the form of a pixel matrix and then calculating using a feature analysis tool. The mathematical form of the pixel matrix is: X = X(m×n) Where X represents the mathematical symbol of the color image, m represents the number of columns of the pixel matrix of the color image, n represents the number of rows of the pixel matrix of the color image, and each element in the pixel matrix represents the rgb value of the pixel at that position.
3. A high-dynamic target recognition and tracking method for a micro unmanned aerial vehicle with a three-rotor layout according to claim 2, characterized in that, The brightness feature of the color image is obtained by inputting the pixel matrix into the brightness analysis tool. The mathematical expression of the brightness analysis tool is: LD = ∑rgb(i) / μ Where LD represents the brightness feature of the color image, rgb(i) represents the rgb value of the i-th pixel, and μ represents the brightness gradient parameter.
4. A high-dynamic target recognition and tracking method for a micro unmanned aerial vehicle with a three-rotor layout according to claim 2, characterized in that The color feature of the color image is obtained by inputting the pixel matrix into the color analysis tool. The mathematical expression of the color analysis tool is: Among them, YS represents the color feature of the color image, which represents the average value of rgb for each pixel.
5. A method for high-dynamic target recognition and tracking of a micro unmanned aerial vehicle with a three-rotor layout according to claim 2, characterized in that The texture feature of the color image is obtained by inputting the pixel matrix into the texture analysis tool. The mathematical expression of the texture analysis tool is: Among them, WL represents the texture feature of the color image, represents the rgb average value of the i-th pixel.
6. A method for high-dynamic target recognition and tracking of a micro-unmanned aerial vehicle with a three-rotor layout according to claim 1, characterized in that The local features include depth gradient, convex normal vector gradient, and concave normal vector gradient.
7. A method for high-dynamic target recognition and tracking of a micro-unmanned aerial vehicle with a three-rotor layout according to claim 1, characterized in that The processing layer of the random forest training network is a conditional random field, and the number of layers of the processing layer is greater than or equal to 50 layers.
8. A high-dynamic target recognition and tracking method for a micro-unmanned aerial vehicle with a three-rotor layout according to claim 1, characterized in that, The training process of the random forest training network is: The first-layer conditional random field in the random forest training network receives the total feature input and outputs the feature vector of the first layer as: p(1) = exp{w1∑f1(x i ) + w2∑g1(x i )} + p(0) Among them, p(1) represents the feature vector of the first layer, p(0) represents the initial feature vector, w1 and w2 represent model parameters, f1 and g1 respectively represent the feature function of the feature and label of the first layer, the relationship function between adjacent features and labels, and x i represents the total feature input; The second-layer conditional random field in the random forest training network receives the total feature input and the feature vector of the first layer and outputs the feature vector of the second layer as: p(2) = exp{w1∑f2(x i ) + w2∑g2(x i )} + p(1) Among them, p(2) represents the feature vector of the second layer, p(1) represents the feature vector of the first layer, w1 and w2 represent model parameters, f2 and g2 respectively represent the feature function of the features and labels of the second layer, and the relationship function between adjacent features and labels, x i represents the total feature input; The third-layer conditional random field in the random forest training network receives the total feature input and the feature vector of the second layer and outputs the feature vector of the third layer as: p(3) = exp{w1∑f3(x i ) + w2∑g3(x i )} + p(2) Among them, p(3) represents the feature vector of the third layer, p(2) represents the feature vector of the second layer, w1 and w2 represent model parameters, f3 and g3 respectively represent the feature function of the third layer's feature and the label, and the relationship function between adjacent features and the label, x i represents the total feature input; The feature vector output by the n-th layer conditional random field in the random forest training network is: p(n) = exp{w1∑f n (x i ) + w2∑g n (x i )} + p(n - 1) Among them, p(n) represents the feature vector of the nth layer, p(n - 1) represents the feature vector of the (n - 1)th layer, w1 and w2 represent model parameters, f n and g n respectively represent the feature function of the features and labels of the nth layer, and the relationship function between adjacent features and labels. x i represents the total feature input.
Citation Information
Patent Citations
Indoor scene refinement analysis method based on depth map
CN107622244A
Context optimization indoor scene semantic annotation method based on superpixel CRF model
CN110084136A