Perception enhancement method and system applied to unmanned aerial vehicles

Through the 5G base station, the sensing signals are transmitted and the signal source is separated, and combined with the vector reduction mapping network and weighted fusion unit, the problem of low accuracy of drone perception technology in complex environments is solved, and high-precision signal enhancement and perception ability improvement is achieved.

CN119341629BActive Publication Date: 2025-06-06SHENZHEN XIYUE ZHIHUI DATA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411845107.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-06-06
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

Existing drone perception technology is difficult to accurately obtain signals in complex environments, and signal processing technology cannot fully mine the effective information in the signal, resulting in low perception accuracy.

Method used

Through the 5G base station, sense signals are transmitted to the set area, mixed reflected signals are received and separated into multiple independent signal sources, and signal enhancement is performed using vector reduction mapping network and weighted fusion unit to improve the enhanced quality of the signal source set.

Benefits of technology

It improves the perception ability of the drone in complex environments, enhances the quality of the signal source collection, and meets the requirements of high-precision perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119341629B_ABST
    Figure CN119341629B_ABST
Patent Text Reader

Abstract

The present application provides a perception enhancement method and system applied to a drone, by obtaining a target signal representation vector corresponding to a first signal source set; obtaining an nth comparison restoration mapping representation vector output by an nth vector restoration mapping component, obtaining an mth cross-layer identity mapping vector based on the nth comparison restoration mapping representation vector, obtaining a first influencing variable corresponding to the nth comparison restoration mapping representation vector and a second influencing variable corresponding to the mth cross-layer identity mapping vector, wherein the ratio between the first influencing variable and the second influencing variable is not less than a reference critical value; determining an mth comparison restoration mapping representation vector according to the first influencing variable, the second influencing variable, the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector; if m=X, determining a second signal source set corresponding to the first signal source set based on the mth comparison restoration mapping representation vector. The present application can improve the quality of signal perception enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a perception enhancement method and system applied to unmanned aerial vehicles. Background Art

[0002] With the rapid development of drone technology, drones have been widely used in many fields such as aerial photography, environmental monitoring, logistics and transportation, etc. In these application scenarios, the drone's accurate perception of the surrounding environment is crucial.

[0003] Traditional UAV perception technology often faces many challenges. On the one hand, complex environmental factors such as urban areas with dense buildings and mountainous areas with complex terrain will interfere with and reflect the perception signals, resulting in a large amount of noise and confusing information in the acquired signals. On the other hand, existing signal processing technologies are often unable to fully tap the effective information in the signals when processing UAV perception signals. When enhancing the signal source set, some methods lack effective means to accurately perform weighted fusion and feature extraction on information at different levels, resulting in the enhanced signals still not meeting the requirements of high-precision perception. In summary, in order to improve the perception capabilities of UAVs in complex environments and meet the requirements of UAV perception accuracy and reliability in different application scenarios, an innovative UAV perception enhancement technology solution is urgently needed. Summary of the invention

[0004] In view of this, the present application provides a perception enhancement method and system applied to a drone.

[0005] The technical solution of this application is implemented as follows:

[0006] On the one hand, the present application provides a perception enhancement method for a drone, comprising: transmitting a perception signal to a set area through a 5G base station; receiving a mixed reflection signal of the perception signal, and separating the received mixed reflection signal into multiple independent signal sources to obtain a separated signal source set, wherein each signal source in the separated signal source set corresponds to a target or area in the set area; using the separated signal source set as a first signal source set, and obtaining a target signal representation vector obtained by extracting representation information of the first signal source set to be enhanced; iterating the following operations until X vector restoration mapping components in a vector restoration mapping network are walked through, wherein the vector restoration mapping network is used to determine a second signal source set corresponding to the first signal source set based on the target signal representation vector, and the vector restoration mapping component includes a first cross-layer identity connection unit and a first weighted fusion unit, wherein X>1, and X is an integer: obtaining the nth control restoration mapping representation vector output by the nth vector restoration mapping component, and loading the nth control restoration mapping representation vector into the mth vector restoration mapping component. The first cross-layer identity connection unit in the vector restoration mapping component obtains the mth cross-layer identity mapping vector, wherein m=n-1, 1<n≤X, and n is an integer; obtains the first influencing variable and the second influencing variable corresponding to the mth vector restoration mapping component, wherein the first influencing variable is used to characterize the importance of the nth comparison restoration mapping representation vector, the second influencing variable is used to characterize the importance of the mth cross-layer identity mapping vector, and the ratio between the first influencing variable and the second influencing variable is not less than a first reference critical value; in the first weighted fusion unit, based on the first influencing variable and the second influencing variable, a weighted restoration mapping vector of the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector is determined, and the weighted restoration mapping vector is determined as the mth comparison restoration mapping representation vector output by the mth vector restoration mapping component; if m=X, then based on the mth comparison restoration mapping representation vector, the second signal source set corresponding to the first signal source set is determined, and the second signal source set is the enhanced signal source set.

[0007] On the other hand, the present application provides a computer system, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above method when executing the program.

[0008] The beneficial effects of the present application are as follows: in the above perception enhancement method applied to unmanned aerial vehicles, when the target signal representation vector corresponding to the first signal source set is obtained, cascade processing can be performed in multiple vector restoration mapping components in the vector restoration mapping network, for example, the restoration mapping vector of the backbone network and the restoration mapping vector of the cross-layer identity network are fused, and at the same time, based on the constraint of the ratio between the first influencing variable corresponding to the backbone network layer and the second influencing variable corresponding to the cross-layer identity network layer, the influencing variable of the backbone network layer in the above-mentioned vector restoration mapping network is improved, and the quality of enhancing the signal source set based on the above-mentioned vector restoration mapping network is increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A schematic diagram of the implementation flow of a perception enhancement method applied to a drone provided in an embodiment of the present application.

[0010] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The embodiment of the present application provides a perception enhancement method applied to a drone, which can be executed by a processor of a computer system. The computer system can refer to a device with data processing capabilities such as a server, a laptop, a tablet computer, a desktop computer, a mobile device, etc.

[0012] Figure 1 A schematic diagram of the implementation process of a perception enhancement method applied to a drone provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:

[0013] Step S100: transmitting a sensing signal to a set area through a 5G base station;

[0014] Step S200: receiving a mixed reflection signal of a sensing signal, and separating the received mixed reflection signal into a plurality of independent signal sources to obtain a separated signal source set, wherein each signal source in the separated signal source set corresponds to a target or area in a set area.

[0015] In step S100 of the embodiment of the present application, the 5G base station is a wireless communication base station, which has the characteristics of high speed, low latency and large capacity, and can provide a good foundation for the transmission of the perception signal. The perception signal is a specific signal used to obtain relevant information in a set area. For example, this perception signal can be a radio signal in the millimeter wave frequency band. The signal in the millimeter wave frequency band has a higher frequency and can provide a finer resolution, which is helpful for accurate detection of the target. For example, the computer system will send instructions to the 5G base station to transmit a perception signal that meets the requirements through the communication interface with the 5G base station according to pre-set parameters, such as the frequency, power, and transmission direction of the signal. Specifically, if the frequency of the perception signal is set to f, the power is P, and the transmission direction vector is d (the direction vector here can be expressed in spatial coordinates, such as d=(x, y, z)), the computer system will convert these parameters into instruction codes that can be recognized by the base station. After receiving the instruction, the base station transmits the perception signal according to the corresponding parameter configuration. This process can be expressed by an instruction flow formula: SendSignal=f(P,d), where SendSignal represents the operation of transmitting a perception signal, which is a function of power P and direction d, and is related to frequency f.

[0016] In step S200, the computer system receives a mixed reflection signal of the perception signal, and separates the received mixed reflection signal into multiple independent signal sources to obtain a set of separated signal sources, wherein each signal source in the set of separated signal sources corresponds to a target or area in the set area. When the perception signal is transmitted to the set area, it will encounter various objects and surfaces in the area, thereby generating reflection signals, which will be mixed together and received by the computer system. For example, there are different targets such as buildings, trees and moving vehicles in the set area. After the perception signal is transmitted to these targets, it will be reflected back to form a mixed reflection signal. The mixed reflection signal is a superposition of multiple reflection signals, which contains information from different targets or areas, but this information is mixed together and needs to be separated. The computer system uses signal processing technology to separate the mixed reflection signal. A feasible technical means is based on a blind source separation algorithm, such as an independent component analysis (ICA) algorithm. Assuming that the mixed reflection signal is X(t) (t represents time), the ICA algorithm assumes that the mixed reflection signal X(t) is a linear mixture of multiple independent source signals S(t)=[s1(t), s2(t),..., sn(t)] (n is the number of source signals), that is, X(t)=AS(t), where A is the mixing matrix. The goal of the ICA algorithm is to find an unmixing matrix W so that Y(t)=WX(t) is as close to the source signal S(t) as possible when only the mixed signal X(t) is known. In this way, the computer system can separate the mixed reflection signal X(t) into multiple independent signal sources to obtain a set of separated signal sources. Each separated signal source corresponds to a specific target or area in the set area. For example, a signal source may correspond to the reflection signal of a building, and another signal source may correspond to the reflection signal of a vehicle. These signal sources contain relevant feature information of the target or area, such as location, shape, material, etc., which provides basic data for subsequent signal processing and perception enhancement of drones.

[0017] Step S300: taking the separated signal source set as the first signal source set, and obtaining a target signal representation vector obtained by extracting representation information of the first signal source set to be enhanced.

[0018] In step S300 of the embodiment of the present application, the first signal source set is obtained through the previous step S200, and it includes multiple independent signal sources, each signal source corresponds to a target or area in the set area. These signal sources carry a variety of information about the target or area, such as characteristic information such as the intensity, phase, and frequency of the reflected signal. Characterization information extraction is a process of converting these complex signal source information into a more representative and abstract vector. The target signal characterization vector is the result of characterization information extraction for the first signal source set, which can describe the information in the first signal source set in a more compact and representative way. For example, in a scene, there are three targets in the set area, namely a building, a car, and a big tree, and the first signal source set contains the signal sources corresponding to these three targets. Each signal source may have multiple feature dimensions, such as the intensity value of the reflected signal at different frequencies. Assume that the characteristics of each signal source can be represented by a vector. The signal source feature vector corresponding to the building is v1=[a1, a2, a3,...,an], the corresponding one for the car is v2=[b1, b2, b3,..., bn], and the corresponding one for the tree is v3=[c1, c2, c3,..., cn] (here a, b, c are specific values, and n is the number of feature dimensions).

[0019] In order to obtain the target signal representation vector, the computer system can use a variety of technical means. One feasible technical means is principal component analysis (PCA). The goal of PCA is to find the main components in the data, thereby reducing the dimension of the data and extracting the main feature information. For this first signal source set, the computer system combines all signal source feature vectors into a matrix, assuming that M=[v1; v2; v3] (here ";" represents the row connection of the matrix). PCA first calculates the covariance matrix Cov(M) of the matrix M. The calculation formula of the covariance matrix is ​​Cov(M)=E[(ME[M])(ME[M]) T ], where E[ ] represents the expected value. Then, the eigenvalues ​​and eigenvectors of the covariance matrix are calculated. By sorting the eigenvalues ​​and selecting the eigenvectors corresponding to the first k largest eigenvalues ​​(k<n), the matrix composed of these eigenvectors can project the original signal source eigenvector into a low-dimensional space to obtain the target signal representation vector. This target signal representation vector retains the main information in the first signal source set to a certain extent, while reducing the redundancy of the data.

[0020] Another technical means may be an auto-encoder in a neural network. An auto-encoder consists of two parts: an encoder and a decoder. The computer system takes the first signal source set as the input of the auto-encoder, and the encoder compresses the input high-dimensional signal source information into a low-dimensional encoding vector, which is similar to the target signal representation vector. For example, in an auto-encoder structure, the number of neurons in the input layer is equal to the number of feature dimensions of the first signal source set, the number of neurons in the middle hidden layer is small, and the number of neurons in the output layer is restored to the number of the input layer. During the training process, the computer system adjusts the weights of the auto-encoder by minimizing the reconstruction error between the input and the decoder output. When the training is completed, the encoding vector output by the encoder is the target signal representation vector obtained by extracting the representation information of the first signal source set. This method can automatically learn the inherent structure in the data, thereby effectively extracting representative feature vectors, providing an important data basis for the subsequent signal source set enhancement operation, and helping to more accurately process the signal source set in the subsequent steps to achieve the perception enhancement of the drone.

[0021] Iterate the following operation steps S410 to S440 until X vector restoration mapping components in the vector restoration mapping network are walked through, wherein the vector restoration mapping network is used to determine a second signal source set corresponding to the first signal source set based on the target signal representation vector, and the vector restoration mapping component includes a first cross-layer identity connection unit and a first weighted fusion unit, wherein X>1, and X is an integer:

[0022] Step S410: Obtain the nth control restoration mapping representation vector output by the nth vector restoration mapping component, and load the nth control restoration mapping representation vector into the first cross-layer identity connection unit in the mth vector restoration mapping component to obtain the mth cross-layer identity mapping vector, where m=n-1, 1<n≤X, and n is an integer.

[0023] In step S410 of the embodiment of the present application, the computer system obtains the nth control restoration mapping representation vector output by the nth vector restoration mapping component, and loads the nth control restoration mapping representation vector into the first cross-layer identity connection unit in the mth vector restoration mapping component to obtain the mth cross-layer identity mapping vector, where m=n-1, 1<n≤X, and n is an integer.

[0024] The vector restoration mapping component is an important component of the vector restoration mapping network. The purpose of this network is to determine the second signal source set corresponding to the first signal source set based on the target signal representation vector. Each vector restoration mapping component has a specific function in the restoration mapping process. For example, the vector restoration mapping component is a decoder. The reference restoration mapping representation vector is a vector output by the nth vector restoration mapping component, which contains information related to the restoration mapping after being processed by the component.

[0025] For example, assuming that there are 5 vector restoration mapping components in the vector restoration mapping network (i.e., X=5), when n=3, the third vector restoration mapping component will output a third comparison restoration mapping representation vector. This vector is the result obtained after the third vector restoration mapping component processes the input data according to its internal operation rules. This result may contain information about partial restoration mapping of the target signal representation vector, such as a vector representation of the degree and direction of restoration of certain features in the signal source set.

[0026] The first cross-layer identity connection unit is a part of the vector reduction mapping component, which plays a key connection and operation role in this step. For example, the first cross-layer identity connection unit can be a structure similar to the residual computing layer (Resnet). When the computer system loads the nth reference reduction mapping representation vector into the first cross-layer identity connection unit in the mth vector reduction mapping component, a series of operations need to be performed.

[0027] For the computer system to perform this loading operation, matrix operation techniques can be used. Assume that the nth reference restoration mapping representation vector is Vn=[v1, v2, v3, …, vn] (where vi is an element in the vector), and the first cross-layer identity connection unit in the mth vector restoration mapping component has a specific operation matrix A. The process of the computer system loading the vector Vn into this unit can be regarded as performing a matrix multiplication operation, and the resulting vector Vm'=A*Vn (here Vm' is the mth cross-layer identity mapping vector). This operation process is similar to the traditional linear transformation, in which the input vector Vn is linearly transformed through the matrix A to obtain a new vector Vm'.

[0028] This cross-layer operation helps to transfer information in the vector restoration mapping network, so that vector restoration mapping components at different levels can share and integrate information. For example, when processing a complex set of signal sources, the nth vector restoration mapping component may have performed preliminary restoration mapping processing on some high-frequency features in the target signal representation vector. When this processing result (i.e., the nth control restoration mapping representation vector) is loaded into the first cross-layer identity connection unit of the mth vector restoration mapping component, the information about the high-frequency features can be passed to the next layer, and through the operation of the first cross-layer identity connection unit, it is integrated with other information in the mth vector restoration mapping component to obtain the mth cross-layer identity mapping vector. This vector will contain partial restoration mapping information from the previous layer and relevant information of this layer, which provides a basis for further processing in the vector restoration mapping network, and helps to gradually construct a second signal source set corresponding to the first signal source set, thereby achieving the purpose of signal source set enhancement in the drone perception enhancement method.

[0029] Step S420: Obtain the first influencing variable and the second influencing variable corresponding to the mth vector restoration mapping component, wherein the first influencing variable is used to characterize the importance of the nth control restoration mapping representation vector, the second influencing variable is used to characterize the importance of the mth cross-layer identity mapping vector, and the ratio between the first influencing variable and the second influencing variable is not less than the first reference critical value.

[0030] In step S420 of the embodiment of the present application, different vectors and mapping relationships have different importance weights, which is the meaning represented by the first influencing variable and the second influencing variable.

[0031] The first influencing variable is associated with the nth control-reduction mapping representation vector. The nth control-reduction mapping representation vector is a vector output by the nth vector reduction mapping component in the processing flow of the vector reduction mapping network, which contains specific information related to the reduction mapping. This first influencing variable is like an indicator of the influence of this vector in the entire reduction mapping process. For example, suppose that in a simulated drone perception enhancement scenario, a set of signal sources in a set area containing multiple targets is being processed. If the nth control-reduction mapping representation vector mainly involves information reduction mapping of key targets (such as high-priority obstacles or specific navigation markers), then its first influencing variable may be relatively large. This is because the information of these key targets is more important in the drone's perception and decision-making process, and has a greater impact on the entire reduction mapping result.

[0032] The second influencing variable is for the mth cross-layer identity mapping vector. The mth cross-layer identity mapping vector is obtained by loading the nth control restoration mapping representation vector into the first cross-layer identity connection unit in the mth vector restoration mapping component. This vector integrates part of the information from the previous layer and is associated with the structure of this layer. The second influencing variable measures the importance of this vector in the entire restoration mapping process. For example, in some cases, if the mth cross-layer identity mapping vector contains elements that can correct errors in the previous processing process or supplement important missing information, then its second influencing variable may be assigned a higher value.

[0033] The condition that the ratio between the first influencing variable and the second influencing variable is not less than the first reference critical value is to ensure that a reasonable weight relationship is maintained between different levels and types of information during the processing of the entire vector reduction mapping network. The computer system can use a variety of methods to determine the two influencing variables.

[0034] One way is based on predefined rules. For example, in the system initialization phase, based on the prior knowledge of the drone perception task, the importance of different types of vectors and mapping relationships in different components is pre-set. If it is known that the nth control-reduction mapping representation vector has a high importance when processing a certain type of signal source (such as the vector corresponding to the reflected signal of a specific frequency band), then the first influencing variable can be set to a relatively large value, such as 0.8, based on experience. At the same time, based on the functional analysis of the mth cross-layer identity mapping vector in the entire reduction mapping architecture, the second influencing variable is set to 0.5. In this way, the ratio between them is 0.8 / 0.5=1.6, which meets the requirement that the ratio is not less than the first reference critical value (assuming the first reference critical value is 1).

[0035] Another way can be a method based on statistical analysis and machine learning. When processing a large amount of test data or historical data, the computer system can analyze the influence of the nth control restoration mapping representation vector and the mth cross-layer identity mapping vector on the accuracy of the final restoration mapping result under different signal source sets and restoration mapping scenarios. For example, by calculating the correlation coefficient between the vector and the final accurate restoration result, it is used as a reference for the influencing variable. Assuming that after analyzing multiple groups of data, it is found that the average correlation coefficient between the nth control restoration mapping representation vector and the final accurate restoration result is 0.7, and the average correlation coefficient between the mth cross-layer identity mapping vector and the final accurate restoration result is 0.3, then the first influencing variable can be set to 0.7, and the second influencing variable can be set to 0.3. The ratio between them is 0.7 / 0.3≈2.33, which meets the ratio requirement. This data-driven method can adapt to different actual situations more flexibly and improve the accuracy and effectiveness of the entire vector restoration mapping network in the UAV perception enhancement task.

[0036] Step S430: In the first weighted fusion unit, a weighted restoration mapping vector of the nth control restoration mapping representation vector and the mth cross-layer identity mapping vector is determined based on the first influencing variable and the second influencing variable, and the weighted restoration mapping vector is determined as the mth control restoration mapping representation vector output by the mth vector restoration mapping component.

[0037] In step S430 of the embodiment of the present application, the function of the first weighted fusion unit is to perform a weighted fusion operation on different vectors. The first influencing variable and the second influencing variable respectively represent the importance of the nth control restoration mapping representation vector and the mth cross-layer identity mapping vector, and they play a key role in weight allocation in the weighted fusion process.

[0038] The nth reference restoration mapping representation vector is the vector output by the nth vector restoration mapping component, which contains specific information related to the restoration mapping. For example, when the drone perceives multiple targets in a set area, this vector may contain information about the restoration mapping of some target features, such as a representation of the target's reflected signal features after being processed by the nth vector restoration mapping component. The mth cross-layer identity mapping vector is the vector obtained by loading the nth reference restoration mapping representation vector in the previous step into the first cross-layer identity connection unit in the mth vector restoration mapping component. It integrates the information from the nth vector and is associated with the characteristics of the mth vector restoration mapping component.

[0039] When a computer system performs a weighted fusion operation, it can use a variety of technical means. From a mathematical point of view, assuming that the nth comparison restoration mapping representation vector is (where k is the dimension of the vector, is an element in the vector), the first influencing variable is ; The mth cross-layer identity mapping vector is , the second influencing variable is .

[0040] First, the computer system determines a core control representation vector based on the multiplication result of the first influencing variable and the nth control reduction mapping representation vector. That is, the core control representation vector This step actually scales the nth control-reduction mapping representation vector according to its importance. The higher the importance ( The larger the value is, the larger the element value in the core comparison representation vector is.

[0041] Next, the computer system determines the reference cross-layer identity mapping vector based on the multiplication result of the second influencing variable and the mth cross-layer identity mapping vector. Again, this scales the mth cross-layer identity mapping vector according to its importance.

[0042] Finally, the weighted restoration mapping vector is obtained based on the addition of the core contrast representation vector and the contrast cross-layer identity mapping vector. .

[0043] For example, in a specific drone perception scenario, there are different types of targets in the set area, such as buildings, vehicles, etc. The nth control restoration mapping representation vector may reflect more of the restoration mapping information of certain features of the building, while the mth cross-layer identity mapping vector may contain supplementary information on the vehicle features or adjustment information on the overall features. =0.6, indicating that the nth control restoration mapping representation vector is relatively important; the second influencing variable =0.4. Assume .

[0044] First, calculate the core contrast representation vector .

[0045] Then calculate the cross-layer identity mapping vector .

[0046] Finally, we get the weighted restored mapping vector .

[0047] The computer system determines this weighted restoration mapping vector as the mth control restoration mapping representation vector output by the mth vector restoration mapping component. This result will continue to be processed in the subsequent vector restoration mapping network, which will help to gradually build a more accurate second signal source set corresponding to the first signal source set, thereby achieving enhanced perception of the drone. This weighted fusion method can reasonably combine the information in different vectors, integrate them according to their importance, and improve the accuracy and effectiveness of the entire restoration mapping process.

[0048] Step S440: if m=X, then determine a second signal source set corresponding to the first signal source set based on the mth comparison restoration mapping characterization vector, the second signal source set being an enhanced signal source set.

[0049] In step S440 of the embodiment of the present application, m and X are variables defined in the previous step, m represents the serial number of the current vector reduction mapping component, and X represents the total number of vector reduction mapping components in the vector reduction mapping network. When m reaches X, it means that all necessary processing steps in the vector reduction mapping network have been completed.

[0050] The mth reference restoration mapping representation vector is the result of a series of complex operations. It contains the comprehensive information from the original first signal source set after multiple layers of processing. This vector is output by the last component (i.e., the Xth component) of the vector restoration mapping network. It condenses the feature information after the entire network processes the original signal source set.

[0051] Take the perception task of a drone in a complex urban environment as an example. The first signal source set includes various signal sources reflected from the surrounding environment, which correspond to various targets and areas in the city, such as buildings, vehicles, pedestrians, etc. After being processed in sequence by multiple components of the vector restoration mapping network, it reaches the Xth component and outputs the mth control restoration mapping representation vector.

[0052] The process of the computer system determining the second signal source set based on the mth comparison restoration mapping characterization vector can adopt a variety of technical means. One possible way is based on a predefined mapping rule. Assume that the mth comparison restoration mapping characterization vector is (where k is the dimension of the vector, are the elements in the vector).

[0053] A mapping matrix M is pre-stored in the computer system. This matrix is ​​obtained based on a large amount of prior knowledge or previous training. (in is the obtained second signal source set, n is the dimension of the second signal source set), converting the vector into a form corresponding to the first signal source set, thereby obtaining the enhanced second signal source set.

[0054] For example, when a drone perceives building height information in an urban environment, the original signal sources in the first signal source set may have incomplete or inaccurate information due to environmental interference and other factors. After being processed by the vector restoration mapping network, this information has been integrated and optimized in the mth reference restoration mapping representation vector. Assuming that the mapping matrix M has specific coefficients in the dimensions related to the building height information, when matrix multiplication is performed, the comprehensive information in the mth reference restoration mapping representation vector can be converted into an accurate representation of the building height information in the second signal source set.

[0055] Another way can be a decoding operation based on a neural network. If the vector restoration mapping network is regarded as the encoding part of an encoding-decoding structure, then the computer system can use the corresponding decoding network structure to generate a second signal source set based on the mth reference restoration mapping representation vector. This decoding network is pre-trained, and it has learned the mapping relationship from the encoded vector (i.e., the mth reference restoration mapping representation vector) to the enhanced signal source set (i.e., the second signal source set). For example, in a structure based on a convolutional neural network (CNN), the mth reference restoration mapping representation vector is sent as input to a decoding network containing multiple deconvolution layers and fully connected layers, and after being processed layer by layer by the network, the second signal source set is finally output.

[0056] This enhanced second signal source set has higher quality than the first signal source set. In the perception application of the drone, it can provide more accurate and comprehensive information about the surrounding environment. For example, for the obstacle avoidance function of the drone, a more accurate signal source set can enable the drone to more accurately identify the location, shape, size and other information of the obstacle; for the navigation function, it can enable the drone to more reliably determine its position in the environment and plan a safe and efficient flight path. In short, step S440 determines the second signal source set based on the mth control restoration mapping representation vector, which plays a key role in the final conversion and improvement of information quality in the entire drone perception enhancement method.

[0057] As an implementation manner, in step S410, the nth comparison restoration mapping representation vector is loaded into the first cross-layer identity connection unit in the mth vector restoration mapping component to obtain the mth cross-layer identity mapping vector, including:

[0058] Step S411: normalizing the nth reference restoration mapping representation vector to obtain the nth normalized restoration mapping vector;

[0059] Step S412: Determine the mth cross-layer identity mapping vector based on the nth normalized restoration mapping vector.

[0060] In step S411 of the embodiment of the present application, the computer system performs a normalization operation on the nth control-reduction mapping representation vector to obtain the nth normalized restoration mapping vector. Normalization operation is of great significance in data processing, as it helps to map data to a specific range, thereby facilitating subsequent calculations and processing. In this scenario, the nth control-reduction mapping representation vector may contain a variety of values, and the range and distribution of these values ​​may be relatively complex. For example, assuming that the nth control-reduction mapping representation vector (where k is the dimension of the vector, are the elements in the vector), the values ​​of these elements may be of different magnitudes, some values ​​may be very large, while some values ​​may be very small.

[0061] Computer systems can use a variety of technical means to perform normalization operations. One feasible method is Min-Max Normalization. Its formula is , where x is an element in the original vector, are the minimum and maximum values ​​of the elements in the original vector, respectively. are the normalized elements.

[0062] For the nth control restore mapping representation vector Each element in , the computer system can calculate according to this formula. For example, if , then for The normalized calculation of . Similarly, the computer system can process vectors Calculate all elements of to get the nth normalized restored mapping vector .

[0063] Normalized vectors have some advantages. They can make data of different dimensions comparable, avoiding the situation where some dimensions have too large or too small an impact in subsequent calculations due to differences in data magnitude. For example, when performing calculations or comparisons with other vectors, the normalized vectors can participate in the calculations more fairly, so that the results can better reflect the combined impact of each dimension.

[0064] Next is step S412, which can be regarded as a linear or nonlinear transformation operation. Assume that the nth normalized restoration mapping vector , the computer system can calculate through a specific transformation matrix T. This transformation matrix T is predetermined according to the structure and task requirements of the vector reduction mapping network. Calculate the mth cross-layer identity mapping vector The formula can be expressed as , where the multiplication operation follows the rules of matrix multiplication.

[0065] For example, suppose the transformation matrix T is a A matrix (k is the dimension of the vector), whose elements are , is a k-dimensional vector. Then The i-th element of The calculation is This means that the computer system Multiply each element in with the corresponding element in the transformation matrix T and add the results to get Each element in , thereby determining the mth cross-layer identity mapping vector.

[0066] This calculation method based on normalized vectors helps to maintain the consistency and stability of data in the vector restoration mapping network. For example, when processing a complex set of signal sources perceived by a drone, different signal sources may have different characteristics and magnitudes. Through the previous normalization operation and the transformation based on normalized vectors here, it can be ensured that in the process of cross-layer information transmission, the data can be converted and integrated in a predetermined manner to avoid error accumulation or information loss due to data irregularity. At the same time, this process also matches the structure and function of the entire vector restoration mapping network, so that the various components can work together to gradually construct an accurate second signal source set corresponding to the first signal source set, and finally achieve enhanced perception of the drone.

[0067] In addition, when executing this process, the computer system can also use some optimization algorithms according to the actual situation. For example, if the computing resources are limited, an approximate calculation method can be used to simplify the transformation matrix T and the nth normalized restoration mapping vector V n-norm The multiplication operation.

[0068] Alternatively, if it is found that the data has some special distribution pattern, the normalization method or the structure of the transformation matrix T can be adjusted according to this pattern to improve the calculation efficiency and the accuracy of the results.

[0069] As an implementation manner, step S412, determining the mth cross-layer identity mapping vector based on the nth normalized restoration mapping vector, includes:

[0070] Step S4121: if the vector restoration mapping subunit included in the first cross-layer identity connection unit is a masked multi-channel attention unit, then the nth normalized restoration mapping vector is loaded into the masked multi-channel attention unit to obtain the mth cross-layer identity mapping vector;

[0071] Step S4122: if the vector restoration mapping subunit included in the first cross-layer identity connection unit is a multi-channel attention unit, the nth normalized restoration mapping vector is loaded into the multi-channel attention unit to obtain the mth cross-layer identity mapping vector;

[0072] Step S4123: If the vector restoration mapping subunit included in the first cross-layer identity connection unit is a multi-layer perceptron subunit, the nth normalized restoration mapping vector is loaded into the multi-layer perceptron subunit to obtain the mth cross-layer identity mapping vector.

[0073] In step S4121, if the vector restoration mapping subunit included in the first cross-layer identity connection unit is a masked multi-channel attention unit, the computer system loads the nth normalized restoration mapping vector into the masked multi-channel attention unit to obtain the mth cross-layer identity mapping vector. The masked multi-channel attention unit is a special processing unit designed to process information in parallel through multiple attention mechanisms, and a mask is added to prevent information at certain locations from being improperly accessed.

[0074] For example, suppose the nth normalized reduction mapping vector is , the masked multi-channel attention unit has c channels, each of which has its corresponding attention mechanism. For each channel i (i=1, 2,…, c), its attention mechanism can be calculated by calculating an attention weight vector To represent the input vector The degree of attention of each element. Calculate the attention weight vector One way to do this can be based on the dot product attention mechanism, for example, for the jth element , in is the query vector for channel i. The dot product here is Represents the inner product operation of the query vector and the input vector elements. The exp function is used to convert the result to a non-negative real number. The denominator is a normalization factor to ensure that the sum of the elements of the weight vector is 1.

[0075] After calculating the attention weight vector for each channel, the computer system applies these weights to the nth normalized reduction mapping vector To obtain the output of the masked multi-channel attention unit, that is, the mth cross-layer identity mapping vector The specific calculation method is, for each element This process actually performs a weighted summation of the elements of the input vector according to the attention weights of different channels, thereby obtaining a vector processed by a masked multi-channel attention unit. This method allows the computer system to selectively focus on the information in the input vector according to the importance of different elements (represented by the attention weights) when determining the mth cross-layer identity mapping vector, and due to the existence of the masking mechanism, it can prevent some unnecessary information from interfering with the calculation results, thereby improving the accuracy and pertinence of vector processing.

[0076] Next is step S4122, if the vector restoration mapping subunit included in the first cross-layer identity connection unit is a multi-channel attention unit, the computer system loads the nth normalized restoration mapping vector into the multi-channel attention unit to obtain the mth cross-layer identity mapping vector. The multi-channel attention unit is similar to the masked multi-channel attention unit in that it also processes information in parallel through multiple attention mechanisms, but without a masking mechanism.

[0077] Also assume that the nth normalized restoration mapping vector is , the multi-channel attention unit has c channels, and the attention mechanism of each channel calculates the attention weight vector The method can be similar to the dot product attention mechanism in the above-mentioned masked multi-channel attention unit, that is, The difference is that since there is no masking mechanism, the attention weight calculation of all channels will not be subject to additional restrictions.

[0078] After calculating the attention weight vector for each channel, the computer system applies these weights to the nth normalized reduction mapping vector To get the mth cross-layer identity mapping vector For each element , calculated as . This method enables the computer system to re-weight and combine the nth normalized restoration mapping vector according to the degree of attention of different channels to the input vector elements through the multi-channel attention mechanism, thereby obtaining the mth cross-layer identity mapping vector. The multi-channel attention unit can pay attention to the information of the input vector from multiple angles (multiple channels) during the processing process, which helps to explore the relationship between different elements in the vector, improve the effect of vector processing, and provide a more accurate information basis for the entire vector restoration mapping network to process the signal source set.

[0079] Finally, in step S4123, if the vector restoration mapping subunit included in the first cross-layer identity connection unit is a multi-layer perceptron subunit, the computer system loads the nth normalized restoration mapping vector into the multi-layer perceptron subunit to obtain the mth cross-layer identity mapping vector. The multi-layer perceptron is a basic neural network structure, a multi-layer network composed of multiple neurons.

[0080] Assume that the multilayer perceptron subunit has L layers, and the input layer receives the nth normalized restoration mapping vector For each layer , the output of its neuron can be calculated by the following formula: , where x j is the output of the previous layer (for the input layer , is the weight connecting the jth input and the lth layer of neurons, is the bias of the lth layer of neurons, and f is the activation function, such as the feasible ReLU function f(x)=max(0, x).

[0081] In the multi-layer perceptron subunit, starting from the input layer, the output of each layer is used as the input of the next layer. After processing by L-1 layers, the output of the last layer is the mth cross-layer identity mapping vector V m This method uses the multi-layer structure and nonlinear activation function of the multi-layer perceptron to perform complex nonlinear transformations on the nth normalized restoration mapping vector, thereby obtaining the mth cross-layer identity mapping vector. The multi-layer structure of the multi-layer perceptron can learn the complex feature relationships in the input vector. By adjusting the weights and biases, the output vector can better adapt to the needs of the vector restoration mapping network, providing a more representative and effective vector representation for subsequent processing.

[0082] In actual drone perception scenarios, these three different implementations play an important role in different situations. For example, when the information contained in the processed signal source set needs to be finely distinguished and selected, the masked multi-channel attention unit may be more applicable. Assuming that the drone is sensing an area containing multiple types of buildings and obstacles, the features of some buildings or obstacles may need to be focused on, while other information may be interference items. The masked multi-channel attention unit can accurately focus on key information through the masking mechanism and attention weights, thereby obtaining the mth cross-layer identity mapping vector that better meets the needs.

[0083] When the processed information focuses more on the mining of overall relationships and there is no obvious interference information, the multi-channel attention unit can effectively analyze the information in the nth normalized restoration mapping vector from multiple angles. For example, when perceiving multiple targets in a relatively open area, the multi-channel attention unit can better integrate the relationship information between the targets and provide a comprehensive information basis for determining the mth cross-layer identity mapping vector.

[0084] When the information in the signal source set has a complex nonlinear relationship, the multilayer perceptron subunit shows its advantages. For example, when the signal source in the perception environment is affected by multiple factors (such as weather, electromagnetic interference, etc.), the multilayer perceptron subunit can learn the relationship between these complex factors and the signal source through its multi-layer structure and nonlinear activation function, thereby effectively transforming the nth normalized reduction mapping vector and obtaining the accurate mth cross-layer identity mapping vector, providing strong support for the perception enhancement of drones.

[0085] As an implementation manner, step S430, in the first weighted fusion unit, determining a weighted restoration mapping vector of the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector based on the first influencing variable and the second influencing variable, includes:

[0086] Step S431: determining a core comparison representation vector based on a multiplication result of the first influencing variable and the nth comparison restoration mapping representation vector;

[0087] Step S432: determining a reference cross-layer identity mapping vector based on a multiplication result of the second influencing variable and the mth cross-layer identity mapping vector;

[0088] Step S433: Obtain a weighted restoration mapping vector based on the addition result of the core control representation vector and the control cross-layer identity mapping vector.

[0089] In step S431 of the embodiment of the present application, the computer system determines the core comparison representation vector based on the multiplication result of the first influencing variable and the nth comparison restoration mapping representation vector. The first influencing variable is a numerical value used to represent the importance of the nth comparison restoration mapping representation vector, which plays a role in scaling the nth comparison restoration mapping representation vector in this multiplication operation.

[0090] For example, suppose the nth control-reduction mapping representation vector is , the first influencing variable is . Computer system computing core comparison characterization vector When , for the vector Each element in , according to the formula Calculate to get the core control representation vector .

[0091] Next is step S432, where the computer system determines a reference cross-layer identity mapping vector based on the multiplication result of the second influencing variable and the m-th cross-layer identity mapping vector. The second influencing variable represents the importance of the m-th cross-layer identity mapping vector, and similar to step S431, this multiplication operation also scales the m-th cross-layer identity mapping vector based on its importance.

[0092] Assume that the mth cross-layer identity mapping vector is , the second influencing variable is The computer system calculates the cross-layer identity mapping vector When , for the vector Each element in , according to the formula Calculate to get the cross-layer identity mapping vector .

[0093] Finally, in step S433, the computer system obtains a weighted restoration mapping vector based on the addition result of the core contrast representation vector and the contrast cross-layer identity mapping vector. This step adds the two scaled vectors obtained in the previous two steps to obtain a weighted restoration mapping vector.

[0094] The specific calculation method is as follows: for the weighted restoration mapping vector Each element in , according to the formula Calculate, that is .

[0095] For example, continuing with the previous example, ,So This weighted restoration mapping vector combines the information of the nth control restoration mapping representation vector and the mth cross-layer identity mapping vector, and is weighted according to their respective importance.

[0096] This weighted fusion method helps improve the accuracy and effectiveness of the vector restoration mapping network when processing drone perception data. It can reasonably integrate the information contained in different vectors according to their importance.

[0097] As an implementation manner, step S300, obtaining a target signal representation vector obtained by extracting representation information of a first signal source set to be enhanced, includes:

[0098] Step S310: determining a two-dimensional array of signal sources corresponding to the first signal source set based on signal source feature vectors respectively corresponding to a plurality of signal sources included in the first signal source set;

[0099] The following operations are iterated until Y representation information extraction components in the representation information extraction network are walked through, wherein the representation information extraction network is used to determine a target signal representation vector corresponding to the first signal source set based on a two-dimensional array of signal sources, and the representation information extraction component includes a second cross-layer identity connection unit and a second weighted fusion unit, wherein Y>1, and Y is an integer:

[0100] Step S321: obtaining the rth comparison embedding mapping representation vector output by the rth representation information extraction component, and loading the rth comparison embedding mapping representation vector into the second cross-layer identity connection unit in the sth representation information extraction component to obtain the sth cross-layer identity embedding representation vector, where r=s-1, 1<s≤Y, and r is an integer;

[0101] Step S322: obtaining a third influencing variable and a fourth influencing variable corresponding to the s-th representation information extraction component, wherein the third influencing variable is used to characterize the importance of the r-th comparison embedding mapping representation vector, and the fourth influencing variable is used to characterize the importance of the s-th cross-layer identity embedding representation vector, and the ratio between the third influencing variable and the fourth influencing variable is not less than the second reference critical value;

[0102] Step S323: In the second weighted fusion unit, a weighted embedding mapping vector of the rth comparison embedding mapping representation vector and the sth cross-layer identity embedding representation vector is determined based on the third influencing variable and the fourth influencing variable, and the weighted embedding mapping vector is determined as the sth comparison embedding mapping representation vector output by the sth representation information extraction component;

[0103] Step S324: if s<Y, then load the sth control embedding mapping representation vector into the tth representation information extraction component, t=s+1;

[0104] Step S325: If s=Y, the s-th comparison embedding mapping representation vector is determined as the target signal representation vector.

[0105] In step S310 of the embodiment of the present application, the computer system determines a two-dimensional array of signal sources corresponding to the first signal source set based on the signal source feature vectors corresponding to the multiple signal sources included in the first signal source set. The first signal source set is obtained by processing the previous steps, which includes multiple independent signal sources, each of which corresponds to a target or area in a set area, and each signal source has its own signal source feature vector. These signal source feature vectors contain rich information, such as signal strength, frequency, phase and other characteristics.

[0106] For example, suppose that in the scene perceived by the drone, there are three targets in the set area: a building, a car, and a big tree. The signal source feature vector corresponding to each target may be as follows: The signal source feature vector corresponding to the building ,in It may indicate the intensity of a signal reflected from a building at a certain frequency. It may represent the intensity at another frequency, and so on; the signal source feature vector corresponding to the car ; The signal source feature vector corresponding to the big tree The computer system combines these signal source feature vectors to construct a signal source two-dimensional array A. In this example, if each signal source feature vector is regarded as a row, then this two-dimensional array .

[0107] The purpose of constructing a two-dimensional array of signal sources is to represent the characteristic information of multiple signal sources in a structure that is easier to process and analyze. This two-dimensional array format facilitates subsequent operations in the representation information extraction network. It can be regarded as an integrated representation of the first signal source set information, providing a unified data structure foundation for subsequent representation information extraction.

[0108] Next are steps S321-S325, which is an iterative process in the characterization information extraction network for determining the target signal characterization vector corresponding to the first signal source set based on the two-dimensional array of signal sources.

[0109] In step S321, the computer system obtains the rth comparison embedding mapping representation vector output by the rth representation information extraction component, and loads the rth comparison embedding mapping representation vector into the second cross-layer identity connection unit in the sth representation information extraction component to obtain the sth cross-layer identity embedding representation vector, where r=s-1, 1<s≤Y, and r is an integer. The representation information extraction component is an important component in the representation information extraction network. Each component performs specific processing on the input information to gradually extract more representative representation information. The comparison embedding mapping representation vector is a vector output by the rth representation information extraction component, which contains information related to the embedding map after being processed by the component.

[0110] For example, assuming Y = 5 (i.e., there are 5 representation information extraction components in the representation information extraction network), when s = 3, r = 2. The second comparison embedding mapping representation vector output by the second representation information extraction component is (k is the dimension of the vector).

[0111] The computer system loads this vector into the second cross-layer identity connection unit in the third representation information extraction component. This second cross-layer identity connection unit can be regarded as a special connection structure that transmits and integrates information between different representation information extraction components. The computer system uses a specific calculation method to Converted to the third cross-layer identity embedding representation vector .

[0112] One possible way to calculate it is to perform matrix multiplication through a transformation matrix T (this matrix is ​​predetermined according to the structure and function of the representation information extraction network), that is, . Assume that T is a The matrix whose elements are ,So The i-th element of The calculation is .

[0113] In step S322, the computer system obtains the third influencing variable and the fourth influencing variable corresponding to the sth representation information extraction component. The third influencing variable is used to characterize the importance of the rth control embedding mapping representation vector, and the fourth influencing variable is used to characterize the importance of the sth cross-layer identity embedding representation vector, and the ratio between the third influencing variable and the fourth influencing variable is not less than the second reference critical value. These influencing variables play a role in adjusting the importance weights of different vectors in the representation information extraction process.

[0114] For example, in the above drone perception scenario, if the rth control embedding mapping representation vector mainly contains information related to the shape features of the target, and in the current perception task, the target shape features are of high importance for the final representation information extraction, then the third influencing variable may be set to a relatively large value, such as 0.7. At the same time, if the sth cross-layer identity embedding representation vector contains supplementary or adjustment information to the previous information, its importance is relatively low, and the fourth influencing variable may be set to 0.3. Here, the ratio of the third influencing variable to the fourth influencing variable is , assuming that the second reference critical value is 1, the ratio requirement is met. The computer system can determine these influencing variables based on pre-set rules, such as based on the importance analysis of different types of target features in the perception task, or by statistical analysis of a large amount of test data.

[0115] In step S323, the computer system determines the weighted embedding mapping vector of the rth comparison embedding mapping representation vector and the sth cross-layer identity embedding representation vector in the second weighted fusion unit based on the third influencing variable and the fourth influencing variable, and determines the weighted embedding mapping vector as the sth comparison embedding mapping representation vector output by the sth representation information extraction component. The function of the second weighted fusion unit is to perform a weighted fusion operation on the two input vectors.

[0116] Assume that the rth control embedding mapping representation vector is , the third influencing variable is ; The s-th cross-layer identity embedding representation vector is , the fourth influencing variable is The computer system first determines the core contrast embedding vector based on the multiplication result of the third influencing variable and the rth contrast embedding mapping representation vector. That is, for the vector Each element in , according to the formula Calculate and get the core control embedding vector Next, the computer system determines the reference cross-layer identity embedding vector based on the multiplication result of the fourth influencing variable and the sth cross-layer identity embedding representation vector. That is, for the vector Each element in , according to the formula Calculate and get the cross-layer identity embedding vector .

[0117] Finally, the computer system obtains a weighted embedding mapping vector based on the sum of the core contrast embedding vector and the contrast cross-layer identity embedding vector. That is, for the weighted embedding mapping vector V weighted Each element v in weightedi (i=1, 2,…, k), according to the formula Calculate and get .

[0118] The computer system determines this weighted embedding mapping vector as the sth comparison embedding mapping representation vector output by the sth representation information extraction component.

[0119] In step S324, if s<Y, the computer system loads the sth contrast embedding mapping representation vector to the tth representation information extraction component, where t=s+1. This step enables information to be passed sequentially between different components in the representation information extraction network to continue the next round of processing. For example, when s=3, t=4, the computer system passes the third contrast embedding mapping representation vector output by the third representation information extraction component to the fourth representation information extraction component for the next round of processing to gradually extract more representative representation information.

[0120] In step S325, if s=Y, the computer system determines the sth control embedding mapping representation vector as the target signal representation vector. When reaching the last representation information extraction component in the representation information extraction network, the output control embedding mapping representation vector has been processed by the entire network and contains the most representative information gradually extracted from the original signal source two-dimensional array, so it is determined as the target signal representation vector. This target signal representation vector is the result of extracting the representation information of the first signal source set. It can describe the information in the first signal source set in a more compact and representative way, and provides an important data basis for subsequent operations in the vector restoration mapping network.

[0121] As an implementation mode, step S321, loading the rth control embedding mapping representation vector into the second cross-layer identical connection unit in the sth representation information extraction component to obtain the sth cross-layer identical embedding representation vector, includes:

[0122] Step S3211: normalizing the rth control embedding mapping representation vector to obtain the rth normalized embedding mapping vector;

[0123] Step S3212: Determine the sth cross-layer identity embedding representation vector based on the rth normalized embedding mapping vector.

[0124] In step S3211 of the embodiment of the present application, the computer system performs a normalization operation on the rth control embedding mapping representation vector to obtain the rth normalized embedding mapping vector. The purpose of the normalization operation is to map the data to a specific range so that the data has better comparability and stability, which is convenient for subsequent calculations and processing. When it is subsequently associated with other vectors or calculation operations, it can avoid the influence of certain elements being too large or too small due to differences in data magnitude. For example, when the sth cross-layer identity embedding representation vector is subsequently fused or otherwise calculated, if there is no normalization operation, elements with larger values ​​may dominate the calculation and mask the influence of other elements. The normalization operation can solve this problem so that each element can play a more fair role in the calculation.

[0125] Next is step S3212, where the computer system determines the sth cross-layer identical embedding representation vector based on the rth normalized embedding mapping vector. This step is to process the rth normalized embedding mapping vector through a specific calculation method on the basis of completing the normalization operation to obtain the sth cross-layer identical embedding representation vector.

[0126] One possible way to calculate is through linear transformation. Assume that there is a transformation matrix T (this transformation matrix is ​​predetermined according to the structure and function of the representation information extraction network), for the rth normalized embedding mapping vector , the computer system calculates the s-th cross-layer identity embedding representation vector When , we can perform matrix multiplication, that is, .

[0127] Assume that the transformation matrix T is a A matrix (k is the dimension of the vector), whose elements are ,So The i-th element of The calculation is This means that the computer system Multiply each element in with the corresponding element in the transformation matrix T and add the results to get Each element in , thereby determining the s-th cross-layer identity embedding representation vector.

[0128] For example, assuming k=3, the transformation matrix , the rth normalized embedding mapping vector Then calculate The first element for: Similarly, calculate , and finally get (The calculation results here are just examples).

[0129] This calculation method based on normalized vectors helps to maintain data consistency and stability in the representation information extraction network. In the scene perceived by the drone, the signal source feature vectors corresponding to different targets or areas may have different numerical ranges and distributions after processing. Through normalization operations and calculations based on normalized vectors, it can be ensured that when information is transmitted between different representation information extraction components, the data can be converted and integrated in a predetermined manner to avoid error accumulation or information loss caused by data irregularity.

[0130] In addition, when executing this process, the computer system can also optimize the transformation matrix T according to the actual situation. For example, if it is found that after multiple calculations, the relationship between certain elements in the final result does not meet expectations, the element values ​​in the transformation matrix T can be adjusted according to the actual situation. Or, if computing resources are limited, some approximate calculation methods can be used to simplify the matrix multiplication operation while ensuring that the result is within an acceptable error range.

[0131] As an implementation manner, step S3212, determining the sth cross-layer identity embedding representation vector based on the rth normalized embedding mapping vector, includes:

[0132] Step S32121: if the vector embedding mapping subunit included in the second cross-layer identity connection unit is a multi-channel attention unit, the rth normalized embedding mapping vector is loaded into the multi-channel attention unit to obtain the sth cross-layer identity embedding representation vector;

[0133] Step S32122: If the vector embedding mapping subunit included in the second cross-layer identity connection unit is a multi-layer perceptron subunit, then the rth normalized embedding mapping vector is loaded into the multi-layer perceptron subunit to obtain the sth cross-layer identity embedding representation vector.

[0134] In step S32121 of the embodiment of the present application, if the vector embedding mapping subunit included in the second cross-layer identity connection unit is a multi-channel attention unit, the computer system loads the rth normalized embedding mapping vector into the multi-channel attention unit to obtain the sth cross-layer identity embedding representation vector. The multi-channel attention unit is a unit that can perform attention weighted processing on the input vector from multiple channel perspectives, aiming to explore the relationship between different dimensions of the vector in order to more effectively integrate information.

[0135] For example, suppose the rth normalized embedding mapping vector is , the multi-channel attention unit has c channels, each of which has its corresponding attention mechanism to calculate the attention weight vector. For each channel i (i=1, 2,…, c), its attention weight vector The calculation method of can be based on the dot product attention mechanism.

[0136] Specifically, for the jth element (j=1, 2,…, k), ,in is the query vector for channel i, where the dot product Represents the inner product operation of the query vector and the input vector elements. The exp function is used to convert the result to a non-negative real number. The denominator is a normalization factor to ensure that the sum of the elements of the weight vector is 1.

[0137] After calculating the attention weight vector for each channel, the computer system applies these weights to the rth normalized embedding map vector To obtain the s-th cross-layer identity embedding representation vector The specific calculation method is, for each element , This process actually performs weighted summation on the elements of the input vector according to the attention weights of different channels, thereby obtaining a vector processed by the multi-channel attention unit.

[0138] For example, suppose , for channel 1, the query vector . First, calculate the attention weight vector of channel 1 : ; Similarly, .

[0139] Calculate according to the above formula The first element .

[0140] Similarly, other elements of V_s can be calculated to obtain the s-th cross-layer identity embedding representation vector. This method allows the computer system to selectively focus on the information in the input vector according to the importance of different elements (represented by attention weights) when determining the s-th cross-layer identity embedding representation vector, thereby improving the accuracy and pertinence of vector processing.

[0141] Next is step S32122, if the vector embedding mapping subunit included in the second cross-layer identity connection unit is a multi-layer perceptron subunit, the computer system loads the rth normalized embedding mapping vector into the multi-layer perceptron subunit to obtain the sth cross-layer identity embedding representation vector. The multi-layer perceptron is a basic neural network structure, a multi-layer network composed of multiple neurons, which can perform complex nonlinear transformations on input vectors.

[0142] Assume that the multilayer perceptron subunit has L layers, and the input layer receives the rth normalized embedding mapping vector For each layer l (l=1, 2,…, L-1), the output of its neurons can be calculated by the following formula: ,in is the output of the previous layer (for the input layer ) is the weight connecting the jth input and the lth layer of neurons, is the bias of the lth layer of neurons, and f is the activation function, such as the feasible ReLU function f(x)=max(0, x).

[0143] In the multi-layer perceptron subunit, starting from the input layer, the output of each layer is used as the input of the next layer. After processing by L-1 layers, the output of the last layer is the s-th cross-layer identity embedding representation vector V_s.

[0144] For example, suppose , for the first layer, assuming the weights Bias , using ReLU as the activation function.

[0145] Calculate the output y of the first layer 1 : ;

[0146] For the second layer, assuming the weights , bias , calculate the output of the second layer .

[0147] The output of the last layer is the s-th cross-layer identity embedding representation vector (To simplify the example, only the final result is shown here.) This method uses the multi-layer structure and nonlinear activation function of the multi-layer perceptron to perform complex nonlinear transformations on the r-th normalized embedding mapping vector, thereby obtaining the s-th cross-layer identity embedding representation vector. The multi-layer structure of the multi-layer perceptron can learn the complex feature relationships in the input vector, and by adjusting the weights and biases, the output vector can better meet the needs of the representation information extraction network, providing a more representative and effective vector representation for subsequent processing.

[0148] In actual drone perception scenarios, these two different implementations have different advantages.

[0149] When the processed signal source information needs to be comprehensively considered from multiple aspects, and the relationship between different dimensions is complex and needs to be focused on, the multi-channel attention unit may be more applicable. For example, when a drone perceives an area containing multiple types of targets (such as different types of buildings, vehicles, and pedestrians, etc.) and complex environments (such as different areas in a city, including commercial areas, residential areas, etc.), the multi-channel attention unit can focus on different types of targets or regional features through different channels, such as one channel focusing on the height features of buildings, one channel focusing on the speed features of vehicles, and one channel focusing on the density features of pedestrians, etc., and then the normalized embedding mapping vector is weighted according to the importance of these features to obtain the sth cross-layer identity embedding representation vector that is more in line with actual needs.

[0150] When the signal source information has a complex nonlinear relationship, such as being affected by multiple environmental factors (such as weather, electromagnetic interference, etc.), the multi-layer perceptron subunit shows its advantages. For example, in the perception of a target by a drone under severe weather conditions, the multi-layer perceptron subunit can learn the complex relationship between weather factors and target features through its multi-layer structure and nonlinear activation function, and effectively transform the normalized embedding mapping vector to obtain an accurate s-th cross-layer identity embedding representation vector, providing strong support for drone perception enhancement.

[0151] As an implementation manner, step S420, obtaining the first influencing variable and the second influencing variable corresponding to the mth vector restoration mapping component, includes one of the following three implementations:

[0152] Method 1: obtaining a first influencing variable and a second influencing variable pre-set for the mth vector reduction mapping component;

[0153] Method 2: If X is an integer not less than 10, the square root of the logarithm of X is used as the first influencing variable corresponding to the m-th vector reduction mapping component, and the value 1 is used as the second influencing variable corresponding to the m-th vector reduction mapping component;

[0154] Method three: using the square root of the logarithm of m as the first influencing variable corresponding to the mth vector restoration mapping component, and using the value 1 as the second influencing variable corresponding to the mth vector restoration mapping component.

[0155] Method 1 determines the influencing variables based on pre-set rules. When constructing the entire vector restoration mapping network, specific first influencing variables and second influencing variables are set for each vector restoration mapping component based on prior knowledge of the drone perception task and network structure.

[0156] For example, in a specific UAV urban environment perception task, assume that there are 5 vector restoration mapping components in the vector restoration mapping network (i.e., X=5). For the third vector restoration mapping component (m=3), according to the prior analysis, if the nth control restoration mapping representation vector mainly involves the restoration mapping of large fixed target information such as buildings, and this information is considered to be very important in the entire perception task, the computer system may pre-set the first influencing variable α to 0.7. At the same time, if the mth cross-layer identity mapping vector plays an auxiliary but relatively minor role in integrating the previous information and subsequent processing, the second influencing variable β may be pre-set to 0.3. These preset values ​​are determined based on the understanding of the importance of different information in the entire perception task and the functions of the network components.

[0157] The second method is to dynamically determine the first influencing variable according to the position of the vector restoration mapping component in the entire network (the total number of components is represented by X), while the second influencing variable is fixed to 1.

[0158] Assume X = 16, and restore the mapping component for the m = 5th vector. First, calculate the first influencing variable α. When X = 16, log 216=4 (here, calculating the logarithm with base 2 is just an example. In practice, you can also choose a suitable base according to the specific situation), α=2, and the second influencing variable β=1. The principle of this determination method is that when the vector reduction mapping network is large (X is large), the depth or position of different components in the network may have a certain regularity in the importance of information processing. Using the square root of the logarithm to determine the first influencing variable reflects that as the depth of the network increases, the influence of the component on information processing may show a certain growth trend. The second influencing variable is fixed to 1 in order to provide a relatively stable benchmark when compared with the first influencing variable.

[0159] This approach is suitable for larger-scale vector reduction mapping networks, in which dynamically adjusting the influencing variables based on the characteristics of the network scale itself can better adapt to the different requirements of different depth components for information processing. For example, in large-scale urban area drone perception tasks, a large number of complex signal source sets need to be processed. This approach can automatically adjust the influencing variables of each vector reduction mapping component according to the scale of the network structure, thereby optimizing the information processing capabilities of the entire network.

[0160] Method three determines the first influencing variable according to the sequence number m of the current vector restoration mapping component, and also fixes the second influencing variable to 1.

[0161] For example, in a vector reduction mapping network, assuming X=8, for the m=3th vector reduction mapping component. Calculate the first influencing variable , (The calculation of the logarithm with base 2 is just an example.) , the second influencing variable .

[0162] This approach takes into account that as the component number increases, the component's position in the network gradually deepens, and its importance to information processing may be related to the order of the components. The first influencing variable is determined by taking the square root of the logarithm of m, which reflects the influence of the component's order position in the network on its importance.

[0163] The advantage of this approach is that it can flexibly determine the influencing variables based on the order of each component in the network, without relying on the overall size of the network (such as the second approach relying on X). In some cases where the network structure may be dynamically adjusted or some components require special processing, this approach can better adapt to the needs. For example, in drone perception tasks, if special weight adjustments are required for vector reduction mapping components at certain specific locations in the network, the third approach provides a flexible adjustment mechanism based on component order.

[0164] These three methods have their own advantages and disadvantages in different application scenarios and network structure requirements. The computer system can choose an appropriate method to obtain the first influencing variable and the second influencing variable according to factors such as the specific UAV perception task, the scale and structural characteristics of the vector restoration mapping network, thereby ensuring that different vector information can be accurately integrated in the vector restoration mapping network to achieve effective enhancement of the signal source set.

[0165] As an implementation manner, step S322, obtaining the third influencing variable and the fourth influencing variable corresponding to the s-th characterization information extraction component, includes one of the following three implementation methods:

[0166] Mode A: obtaining the third influencing variable and the fourth influencing variable pre-set for the s-th representation information extraction component;

[0167] Mode B: If Y is an integer not less than 10, the square root of the logarithm of Y is used as the third influencing variable corresponding to the s-th representation information extraction component, and the value 1 is used as the fourth influencing variable corresponding to the s-th representation information extraction component;

[0168] Method C: The square root of the logarithm of s is used as the third influencing variable corresponding to the s-th characterization information extraction component, and the value 1 is used as the fourth influencing variable corresponding to the s-th characterization information extraction component.

[0169] Method A determines the influencing variables based on pre-set rules. When constructing the entire representation information extraction network, specific third and fourth influencing variables are set for each representation information extraction component based on prior knowledge of the drone perception task and network structure.

[0170] For example, in a specific UAV mountain environment perception task, assume that there are 7 representation information extraction components in the representation information extraction network (i.e., Y=7). For the 4th representation information extraction component (s=4), according to the prior analysis, if the rth control embedding mapping representation vector mainly involves the embedding mapping of key terrain feature information such as mountain contours, and this information is considered to be very important in the entire perception task, the computer system may pre-set the third influencing variable α to 0.8. At the same time, if the sth cross-layer identity embedding representation vector plays an auxiliary but relatively minor role in integrating the previous information and subsequent processing, the fourth influencing variable β may be pre-set to 0.2. These preset values ​​are determined based on the understanding of the importance of different information in the entire perception task and the functions of the network components.

[0171] The advantage of this method is that it is simple and direct, and does not require complex calculations at runtime to determine the influencing variables. It is suitable for situations where the perception tasks are relatively fixed and the data patterns are relatively stable. For example, in some specific agricultural scenarios, drones are used to monitor farmland and irrigation facilities with fixed layouts. The types of targets and task requirements they perceive do not change much, and setting the influencing variables in advance can effectively meet the needs.

[0172] Method B is to dynamically determine the third influencing variable according to the position of the characterization information extraction component in the entire network (the total number of components is represented by Y), while the fourth influencing variable is fixed to 1.

[0173] Assume that Y=16, for the s=5th characterization information extraction component. First, calculate the third influencing variable α. When Y=16, calculate the logarithm with base 2 (in practice, the appropriate base can also be selected according to the specific situation). , the fourth influencing variable The principle of this determination method is that when the representation information extraction network is large (Y is large), the depth or position of different components in the network may have a certain regularity in the importance of information processing. The square root of the logarithm is used to determine the third influencing variable, which reflects that as the network depth increases, the influence of the component on information processing may show a certain growth trend. The fourth influencing variable is fixed at 1 to provide a relatively stable benchmark when compared with the third influencing variable.

[0174] Method C is to determine the third influencing variable according to the serial number s of the current characterization information extraction component, and also fix the fourth influencing variable to 1.

[0175] For example, in a representation information extraction network, assuming Y = 9, for the s = 3 representation information extraction component. Calculate the third influencing variable , calculate the logarithm with base 2 (this is just an example, you can choose the base as needed), , , the fourth influencing variable This approach takes into account that as the component number increases, the component's position in the network gradually deepens, and its importance to information processing may be related to the order of the components. The third influencing variable is determined by taking the square root of the logarithm of s, which reflects the influence of the component's order position in the network on its importance.

[0176] As an implementation mode, the method further includes a training process of a vector restoration mapping network, specifically including:

[0177] Step S10: obtaining a target training signal representation vector obtained by extracting representation information of the first training signal source set to be enhanced;

[0178] Iterate the following operations until X vector restoration mapping components in the vector restoration mapping network in the calibration process are walked through, wherein the vector restoration mapping network in the calibration process is used to determine a second training signal source set corresponding to the first training signal source set based on the target training signal representation vector, and the vector restoration mapping component includes a first cross-layer identity connection unit and a first weighted fusion unit, X>1, and X is an integer:

[0179] Step S201: obtaining an nth comparison restoration mapping representation vector output by an nth vector restoration mapping component, and loading the nth comparison restoration mapping representation vector into a first cross-layer identity connection unit in an mth vector restoration mapping component to obtain an mth cross-layer identity mapping vector, wherein m=n-1, 1<n≤X, and n is an integer;

[0180] Step S202: obtaining a first influencing variable and a second influencing variable corresponding to the mth vector restoration mapping component, wherein the first influencing variable is used to characterize the importance of the nth control restoration mapping characterization vector, the second influencing variable is used to characterize the importance of the mth cross-layer identity mapping vector, and the ratio between the first influencing variable and the second influencing variable is not less than a first reference critical value;

[0181] Step S203: In the first weighted fusion unit, a weighted restoration mapping vector of the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector is determined based on the first influencing variable and the second influencing variable, and the weighted restoration mapping vector is used as the mth comparison restoration mapping representation vector output by the mth vector restoration mapping component;

[0182] Step S204: if m=X, determining a second training signal source set corresponding to the first training signal source set based on the mth comparison restoration mapping representation vector;

[0183] Step S30: when it is determined based on the second training signal source set that the vector restoration mapping network in the calibration process meets the network stable state, the vector restoration mapping network in the calibration process is determined as the target vector restoration mapping network.

[0184] In step S10 of the embodiment of the present application, the computer system obtains a target training signal representation vector obtained by extracting representation information of the first training signal source set to be enhanced. The first training signal source set is similar to the first signal source set mentioned above, and is a data set for training a vector restoration mapping network, which includes multiple signal sources corresponding to targets or regions in a set area.

[0185] For example, when a drone performs perception training on a simulated urban area, the first training signal source set may include signal sources of different types of buildings, vehicles, and other landmarks. Each signal source has its own characteristic information, and the computer system extracts characterization information from these signal sources. This process can adopt a method similar to that in step S300, such as constructing a two-dimensional array of signal sources, and then processing them through a characterization information extraction network. Assume that there are three signal sources in the first training signal source set, corresponding to a high-rise building, a car, and a street lamp respectively. These signal sources each have different signal characteristics, such as the intensity and frequency of the reflected signal. The computer system converts the characteristic information of these signal sources into a target training signal characterization vector through a specific algorithm. This vector can describe the information in the first training signal source set in a more abstract and representative way.

[0186] Next are steps S201-S204, which is an iterative process in the vector restoration mapping network during the tuning process.

[0187] In step S201, the computer system obtains the nth reference restoration mapping representation vector output by the nth vector restoration mapping component, and loads the nth reference restoration mapping representation vector into the first cross-layer identity connection unit in the mth vector restoration mapping component to obtain the mth cross-layer identity mapping vector, where m=n-1, 1<n≤X, and n is an integer. This process is similar to the operation in step S410, but during the training process, its purpose is to optimize the parameters of the vector restoration mapping network so that the network can better restore the accurate second training signal source set from the target training signal representation vector.

[0188] For example, suppose there are X=5 vector restoration mapping components in the vector restoration mapping network. When n=3, the third vector restoration mapping component outputs the third comparison restoration mapping representation vector V n The computer system will V n The first cross-layer identity connection unit loaded into the m=2th vector reduction mapping component. Assuming that the first cross-layer identity connection unit is a structure similar to the residual calculation layer, the computer system may use some linear transformation or specific calculation rules (such as matrix multiplication operation V through a predefined transformation matrix T) m =T×V n ) to obtain the mth cross-layer identity mapping vector V m This cross-layer operation helps to transfer information during the training process, allowing the network to learn the relationship between different layers and thus adjust the network parameters to better restore the signal source set.

[0189] In step S202, the computer system obtains the first influencing variable and the second influencing variable corresponding to the mth vector restoration mapping component, wherein the first influencing variable is used to characterize the importance of the nth control restoration mapping characterization vector, the second influencing variable is used to characterize the importance of the mth cross-layer identity mapping vector, and the ratio between the first influencing variable and the second influencing variable is not less than the first reference critical value. This step is similar to step S420, but is an operation during the training process.

[0190] For example, during the training process, if the nth control restoration mapping representation vector contains key information about the target shape feature restoration mapping, then the computer system may determine that the value of the first influencing variable α is relatively large, such as α=0.7, based on prior knowledge or by analyzing training data. If the mth cross-layer identity mapping vector plays a role in supplementing other auxiliary information, the second influencing variable β may be relatively small, such as β=0.3, satisfying The first reference critical value requirement. The computer system can determine these influencing variables by pre-setting (similar to method one), according to the network size (similar to method two), or according to the component sequence number (similar to method three), depending on the characteristics of the training task and the network structure.

[0191] In step S203, the computer system determines the weighted restoration mapping vector of the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector in the first weighted fusion unit based on the first influencing variable and the second influencing variable, and uses the weighted restoration mapping vector as the mth comparison restoration mapping representation vector output by the mth vector restoration mapping component. This step is similar to step S430, which is to perform a weighted fusion operation on the vector during the training process.

[0192] Assume that the nth control-reduction mapping representation vector is V n =[v n1 , v n2 ,…, v nk ], the first influencing variable is α; the mth cross-layer identity mapping vector is V m =[v m1 , v m2 ,…, v mk ], and the second influencing variable is β. The computer system first calculates the core control characterization vector V core =α×V n =[α×v n1 , α×v n2 ,…, α×v nk ], and then calculate the cross-layer identity mapping vector V m-map =β×V m =[β×v m1 , β×vm2 ,…, β×v mk ], and finally obtain the weighted restoration mapping vector V weighted =V core +V m-map =[α×v n1 +β×v m1 , α×v n2 +β×v m2 ,…, α×v nk +β×v mk ], and V weighted The mth control restoration mapping representation vector is outputted by the mth vector restoration mapping component. This weighted fusion operation helps to adjust the output of the network according to the importance of the vector during training, so that the network can learn a more accurate mapping relationship.

[0193] In step S204, if m=X, the computer system determines the second training signal source set corresponding to the first training signal source set based on the mth comparison restoration mapping representation vector. When reaching the last component of the vector restoration mapping network, the comparison restoration mapping representation vector output at this time has undergone a series of previous processing and contains sufficient information to restore the second training signal source set corresponding to the first training signal source set.

[0194] For example, suppose the mth control-reduction mapping representation vector is V m =[v m1 , v m2 ,…, v mk ], the computer system converts V through a pre-trained mapping relationship (this mapping relationship can be a matrix M, which is learned through a large amount of training data) m Converted to the second training signal source set S = M × V m , where S=[s 1 , s 2 ,…, s n' ] is the second training signal source set (n' is the dimension of the set). This process is similar to the operation in step S440, but during the training process, it is to enable the network to learn how to accurately restore the second training signal source set from the target training signal representation vector.

[0195] Finally, in step S30, when it is determined based on the second training signal source set that the vector restoration mapping network in the calibration process meets the network stability state, the computer system determines the vector restoration mapping network in the calibration process as the target vector restoration mapping network. The network stability state means that the performance indicators of the network (such as mean square error, accuracy, etc.) no longer change significantly in continuous training iterations or reach a preset convergence standard.

[0196] For example, during the training process, the computer system can calculate the mean square error (MSE) between the second training signal source set and the real target signal source set (known in the training data). When the value of MSE is less than a preset threshold in several consecutive iterations, the network can be considered to have reached a stable state. At this time, the computer system determines that the calibrated vector restoration mapping network is the target vector restoration mapping network, and this target vector restoration mapping network can be used to subsequently process the actual first signal source set to obtain an enhanced second signal source set, thereby achieving enhanced perception of the drone.

[0197] In the training process of the entire vector restoration mapping network, these steps cooperate with each other, and by continuously adjusting the network parameters, the network can learn the accurate mapping relationship from the first training signal source set to the second training signal source set, thereby providing a reliable basis for the perception enhancement of the UAV. This training process requires a large amount of training data and a suitable training algorithm, and needs to be adjusted according to the specific application scenario and performance requirements to ensure that the final target vector restoration mapping network can meet the perception enhancement needs of the UAV in different environments.

[0198] As an implementation manner, before step S10, obtaining the target training signal representation vector obtained by extracting the representation information of the first training signal source set to be enhanced, it also includes a training process of the representation information extraction network:

[0199] Step S1: determining a two-dimensional array of sample signal sources corresponding to the first training signal source set based on signal source feature vectors respectively corresponding to a plurality of signal sources included in the first training signal source set;

[0200] The following operations are iterated until Y representation information extraction components in the representation information extraction network in the calibration process are walked through, wherein the representation information extraction network is used to determine a target training signal representation vector corresponding to the first training signal source set based on a two-dimensional array of sample signal sources, and the representation information extraction component includes a second cross-layer identity connection unit and a second weighted fusion unit, wherein Y>1, and Y is an integer:

[0201] Step S21: obtaining the rth comparison embedding mapping representation vector output by the rth representation information extraction component, and loading the rth comparison embedding mapping representation vector into the second cross-layer identity connection unit in the sth representation information extraction component to obtain the sth cross-layer identity embedding representation vector, wherein r=s-1, 1<s≤Y, and r is an integer;

[0202] Step S22: obtaining a third influencing variable and a fourth influencing variable corresponding to the s-th representation information extraction component, wherein the third influencing variable is used to characterize the importance of the r-th comparison embedding mapping representation vector, and the fourth influencing variable is used to characterize the importance of the s-th cross-layer identity embedding representation vector, and the ratio between the third influencing variable and the fourth influencing variable is not less than the second reference critical value;

[0203] Step S23: in the second weighted fusion unit, a weighted embedding mapping vector of the rth comparison embedding mapping representation vector and the sth cross-layer identity embedding representation vector is determined based on the third influencing variable and the fourth influencing variable, and the weighted embedding mapping vector is determined as the sth comparison embedding mapping representation vector output by the sth representation information extraction component;

[0204] Step S24: if s<Y, load the sth control embedding mapping representation vector into the tth representation information extraction component;

[0205] Step S25: if s=Y, the sth control embedding mapping representation vector is used as the target training signal representation vector;

[0206] Step S3: when it is determined based on the target training signal representation vector that the representation information extraction network in the calibration process meets the network stable state, the representation information extraction network in the calibration process is determined as the target representation information extraction network.

[0207] In step S1 of the embodiment of the present application, the computer system determines a two-dimensional array of sample signal sources corresponding to the first training signal source set based on the signal source feature vectors corresponding to the multiple signal sources included in the first training signal source set. The first training signal source set includes multiple signal sources, each of which corresponds to a target or area in a set area, and has its own specific signal source feature vector. These feature vectors cover information such as signal strength, frequency, phase, etc.

[0208] For example, in a scenario where a drone is performing perception training on an area containing multiple targets (such as buildings, vehicles, trees, etc.), for the building target, its signal source feature vector may contain reflected signal strength values ​​in different frequency bands, such as V building =[a 1 , a 2 ,…, a n ], where a i Represents the reflected signal strength in the i-th frequency band. Similarly, vehicles and trees also have their own corresponding feature vectors V vehicle and V treeThe computer system combines the feature vectors corresponding to these different targets to construct a two-dimensional array of sample signal sources. Assuming that there are three targets in this example (buildings, vehicles, and trees), and each feature vector has a dimension of n, then this two-dimensional array can be expressed as , where the second row is the feature vector element of the vehicle, and the third row is the feature vector element of the tree. This two-dimensional array of sample signal sources provides the initial data structure basis for the subsequent training of the representation information extraction network.

[0209] Next are steps S21-S25, which is an iterative process in the representation information extraction network during the tuning process.

[0210] In step S21, the computer system obtains the rth contrast embedding mapping representation vector output by the rth representation information extraction component, and loads the rth contrast embedding mapping representation vector into the second cross-layer identity connection unit in the sth representation information extraction component to obtain the sth cross-layer identity embedding representation vector, where r=s-1, 1<s≤Y, and r is an integer. The representation information extraction component is the basic building block of the representation information extraction network. Each component gradually processes the input data during the training process to extract more representative representation information. The contrast embedding mapping representation vector is a vector output by the rth representation information extraction component, which contains information related to the embedding map after being processed by the component.

[0211] For example, assuming Y = 5 (i.e., there are 5 representation information extraction components in the representation information extraction network), when s = 3, r = 2. The second representation information extraction component outputs the second contrast embedding mapping representation vector (k is the dimension of the vector).

[0212] The computer system loads this vector into the second cross-layer identity connection unit in the third representation information extraction component. This second cross-layer identity connection unit is similar to a structure for information transmission and integration. The computer system uses a specific calculation method to Converted to the third cross-layer identity embedding representation vector One possible way to calculate is to perform matrix multiplication through a transformation matrix T (this matrix is ​​predetermined according to the structure and function of the representation information extraction network), that is, . Assume that T is a k×k matrix whose elements are ,So The i-th element of The calculation is This cross-layer operation helps to share and integrate information between different components during the training process, allowing the network to learn more complex representation relationships.

[0213] In step S22, the computer system obtains a third influencing variable and a fourth influencing variable corresponding to the s-th representation information extraction component. The third influencing variable is used to characterize the importance of the r-th comparison embedding mapping representation vector, and the fourth influencing variable is used to characterize the importance of the s-th cross-layer identity embedding representation vector, and the ratio between the third influencing variable and the fourth influencing variable is not less than the second reference critical value.

[0214] For example, in a drone perception training scenario, if the rth control embedding mapping representation vector mainly contains the embedding mapping information of the target key features (such as the height information of the building or the speed information of the vehicle), this information may be very important for the entire representation information extraction. The computer system may determine that the value of the third influencing variable α is relatively large, such as α=0.8 based on prior knowledge or analysis of the training data. If the sth cross-layer identity embedding representation vector plays a certain role in supplementing or adjusting the previous information, the fourth influencing variable β may be relatively small, such as β=0.2, satisfying The second reference critical value requirement. The computer system can determine these influencing variables by pre-setting (similar to method A), according to the network size (similar to method B), or according to the component sequence number (similar to method C), depending on the characteristics of the training task and network structure.

[0215] In step S23, the computer system determines the weighted embedding mapping vector of the rth comparison embedding mapping representation vector and the sth cross-layer identity embedding representation vector in the second weighted fusion unit based on the third influencing variable and the fourth influencing variable, and determines the weighted embedding mapping vector as the sth comparison embedding mapping representation vector output by the sth representation information extraction component. The function of the second weighted fusion unit is to perform a weighted fusion operation on the two input vectors to adjust the representation of the vector in the network so that it better meets the requirements of representation information extraction.

[0216] Assume that the rth control embedding mapping representation vector is , the third influencing variable is ; The s-th cross-layer identity embedding representation vector is , the fourth influencing variable is The computer system first determines the core contrast embedding vector based on the multiplication result of the third influencing variable and the rth contrast embedding mapping representation vector. That is, for the vector Each element in (i=1, 2,…, k), according to the formula Calculate and get the core control embedding vector .

[0217] Next, the computer system determines a reference cross-layer identity embedding vector based on the multiplication result of the fourth influencing variable and the sth cross-layer identity embedding representation vector. That is, for the vector V s Each element in (i=1, 2,…, k), according to the formula Calculate and get the cross-layer identity embedding vector .

[0218] Finally, the computer system obtains a weighted embedding mapping vector based on the sum of the core contrast embedding vector and the contrast cross-layer identity embedding vector. That is, for the weighted embedding mapping vector V weighted Each element v in weightedi (i=1, 2,…,k), according to the formula Calculate and get The computer system determines this weighted embedding mapping vector as the sth control embedding mapping representation vector output by the sth representation information extraction component. This weighted fusion operation helps to adjust the output of the network according to the importance of the vector during the training process, so that the network can learn more accurate representation information extraction relationships.

[0219] In step S24, if s<Y, the computer system loads the sth contrast embedding mapping representation vector to the tth representation information extraction component, where t=s+1. This step enables information to be passed sequentially between different components in the representation information extraction network to continue the next round of processing. For example, when s=3, t=4, the computer system passes the third contrast embedding mapping representation vector output by the third representation information extraction component to the fourth representation information extraction component for the next round of processing to gradually extract more representative representation information.

[0220] In step S25, if s=Y, the computer system determines the sth control embedding mapping representation vector as the target training signal representation vector. When reaching the last representation information extraction component in the representation information extraction network, the control embedding mapping representation vector output at this time has been processed by the entire network and contains the most representative information gradually extracted from the original two-dimensional array of sample signal sources, so it is determined as the target training signal representation vector. This target training signal representation vector is the result of extracting the representation information of the first training signal source set. It can describe the information in the first training signal source set in a more compact and representative way, and provides an important data basis for subsequent training in the vector restoration mapping network.

[0221] Finally, in step S3, when it is determined based on the target training signal representation vector that the representation information extraction network in the calibration process meets the network stability state, the computer system determines the representation information extraction network in the calibration process as the target representation information extraction network. The network stability state means that the performance indicators of the network (such as mean square error, accuracy, etc.) no longer change significantly in continuous training iterations or reach a preset convergence standard.

[0222] For example, during the training process, the computer system can calculate the mean square error (MSE) between the target training signal representation vector and the real target representation vector (known in the training data or determined by other means). When the value of MSE is less than a preset threshold in several consecutive iterations, the network can be considered to have reached a stable state. At this time, the computer system determines that the calibrated representation information extraction network is the target representation information extraction network, and this target representation information extraction network can be used to subsequently process the actual first signal source set to obtain an accurate target signal representation vector, thereby providing a reliable representation information extraction basis for the entire UAV perception enhancement process.

[0223] In the training process of the entire representation information extraction network, these steps cooperate with each other, and by continuously adjusting the network parameters, the network can learn the accurate mapping relationship from the first training signal source set to the target training signal representation vector, thereby providing a reliable basis for the perception enhancement of the drone. This training process requires a large amount of training data and a suitable training algorithm, and needs to be adjusted according to the specific application scenario and performance requirements to ensure that the final target representation information extraction network can meet the perception enhancement needs of drones in different environments.

[0224] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.

[0225] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A perception enhancement method applied to an unmanned aerial vehicle, characterized in that: include: Transmit sensing signals to the set area through 5G base stations; receiving a mixed reflection signal of the sensing signal, and separating the received mixed reflection signal into a plurality of independent signal sources to obtain a separated signal source set, wherein each signal source in the separated signal source set corresponds to a target or area in the set area; Taking the separated signal source set as the first signal source set, obtaining a target signal representation vector obtained by extracting representation information of the first signal source set to be enhanced; Iterate the following operations until X vector restoration mapping components in the vector restoration mapping network are walked through, wherein the vector restoration mapping network is used to determine a second signal source set corresponding to the first signal source set based on the target signal representation vector, and the vector restoration mapping component includes a first cross-layer identity connection unit and a first weighted fusion unit, wherein X>1, and X is an integer: the vector restoration mapping component is a decoder; Obtaining an nth comparison restoration mapping representation vector output by an nth vector restoration mapping component, and performing a normalization operation on the nth comparison restoration mapping representation vector to obtain an nth normalized restoration mapping vector; If the vector restoration mapping subunit included in the first cross-layer identity connection unit is a masked multi-channel attention unit, the nth normalized restoration mapping vector is loaded into the masked multi-channel attention unit to obtain the mth cross-layer identity mapping vector; wherein the masked multi-channel attention unit is a unit that processes information in parallel through multiple attention mechanisms and adds a mask to prevent information at certain positions from being improperly accessed; If the vector restoration mapping subunit included in the first cross-layer identity connection unit is a multi-channel attention unit, the nth normalized restoration mapping vector is loaded into the multi-channel attention unit to obtain the mth cross-layer identity mapping vector; If the vector restoration mapping subunit included in the first cross-layer identity connection unit is a multi-layer perceptron subunit, the nth normalized restoration mapping vector is loaded into the multi-layer perceptron subunit to obtain the mth cross-layer identity mapping vector; wherein m=n-1, 1<n≤X, and n is an integer; the reference restoration mapping representation vector is a vector output by the nth vector restoration mapping component, and contains information related to the restoration mapping after being processed by the vector restoration mapping component; Obtaining a first influencing variable and a second influencing variable corresponding to the mth vector restoration mapping component, wherein the first influencing variable is used to characterize the importance of the nth control restoration mapping characterization vector, the second influencing variable is used to characterize the importance of the mth cross-layer identity mapping vector, and the ratio between the first influencing variable and the second influencing variable is not less than a first reference critical value; In the first weighted fusion unit, a weighted restoration mapping vector of the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector is determined based on the first influencing variable and the second influencing variable, and the weighted restoration mapping vector is determined as the mth comparison restoration mapping representation vector output by the mth vector restoration mapping component; If n=X, then the second signal source set corresponding to the first signal source set is determined based on the nth comparison restoration mapping characterization vector, and the second signal source set is an enhanced signal source set.

2. The method according to claim 1, characterized in that The determining the mth cross-layer identity mapping vector based on the nth normalized restoration mapping vector comprises: The step of determining, in the first weighted fusion unit, a weighted restoration mapping vector of the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector based on the first influencing variable and the second influencing variable comprises: Determining a core control representation vector based on a multiplication result of the first influencing variable and the nth control reduction mapping representation vector; Determine a reference cross-layer identity mapping vector based on a multiplication result of the second influencing variable and the mth cross-layer identity mapping vector; The weighted restoration mapping vector is obtained based on the addition result of the core control representation vector and the control cross-layer identity mapping vector.

3. The method according to claim 1, characterized in that The step of obtaining a target signal representation vector obtained by extracting representation information of the first signal source set to be enhanced includes: Determine a two-dimensional array of signal sources corresponding to the first signal source set based on signal source feature vectors respectively corresponding to a plurality of signal sources included in the first signal source set; Iterate the following operations until Y representation information extraction components in the representation information extraction network are walked through, wherein the representation information extraction network is used to determine the target signal representation vector corresponding to the first signal source set based on the signal source two-dimensional array, and the representation information extraction component includes a second cross-layer identity connection unit and a second weighted fusion unit, wherein Y>1, and Y is an integer: Obtaining an rth comparison embedding mapping representation vector output by the rth representation information extraction component, and loading the rth comparison embedding mapping representation vector into the second cross-layer identity connection unit in the sth representation information extraction component to obtain an sth cross-layer identity embedding representation vector, wherein r=s-1, 1<s≤Y, and r is an integer; Obtaining a third influencing variable and a fourth influencing variable corresponding to the s-th representation information extraction component, wherein the third influencing variable is used to characterize the importance of the r-th control embedding mapping representation vector, the fourth influencing variable is used to characterize the importance of the s-th cross-layer identity embedding representation vector, and the ratio between the third influencing variable and the fourth influencing variable is not less than a second reference critical value; In the second weighted fusion unit, a weighted embedding mapping vector of the rth comparison embedding mapping representation vector and the sth cross-layer identity embedding representation vector is determined based on the third influencing variable and the fourth influencing variable, and the weighted embedding mapping vector is determined as the sth comparison embedding mapping representation vector output by the sth representation information extraction component; If s<Y, then the s-th control embedding mapping representation vector is loaded into the t-th representation information extraction component, t=s+1; If s=Y, then determining the s-th control embedding mapping representation vector as the target signal representation vector; The step of loading the rth control embedding mapping representation vector into the second cross-layer identical connection unit in the sth representation information extraction component to obtain the sth cross-layer identical embedding representation vector includes: Normalizing the r-th control embedding mapping representation vector to obtain an r-th normalized embedding mapping vector; The s-th cross-layer identity embedding representation vector is determined based on the r-th normalized embedding mapping vector.

4. The method according to claim 3, characterized in that The determining the sth cross-layer identity embedding representation vector based on the rth normalized embedding mapping vector includes: If the vector embedding mapping subunit included in the second cross-layer identity connection unit is a multi-channel attention unit, the rth normalized embedding mapping vector is loaded into the multi-channel attention unit to obtain the sth cross-layer identity embedding representation vector; If the vector embedding mapping subunit included in the second cross-layer identity connection unit is a multi-layer perceptron subunit, the rth normalized embedding mapping vector is loaded into the multi-layer perceptron subunit to obtain the sth cross-layer identity embedding representation vector.

5. The method according to claim 1, characterized in that The obtaining of the first influencing variable and the second influencing variable corresponding to the mth vector restoration mapping component comprises: Obtaining the first influencing variable and the second influencing variable pre-set for the mth vector restoration mapping component; or; If X is an integer not less than 10, the square root of the logarithm of X is used as the first influencing variable corresponding to the m-th vector restoration mapping component, and the value 1 is used as the second influencing variable corresponding to the m-th vector restoration mapping component; or; The square root of the logarithm of m is used as the first influencing variable corresponding to the mth vector restoration mapping component, and the value 1 is used as the second influencing variable corresponding to the mth vector restoration mapping component.

6. The method according to claim 3, characterized in that The obtaining of the third influencing variable and the fourth influencing variable corresponding to the s-th characterization information extraction component includes: Acquire the third influencing variable and the fourth influencing variable pre-set for the s-th characterization information extraction component; or; If Y is an integer not less than 10, the square root of the logarithm of Y is used as the third influencing variable corresponding to the s-th characterization information extraction component, and the value 1 is used as the fourth influencing variable corresponding to the s-th characterization information extraction component; or; The square root of the logarithm of s is used as the third influencing variable corresponding to the s-th characterization information extraction component, and the value 1 is used as the fourth influencing variable corresponding to the s-th characterization information extraction component.

7. The method according to claim 1, characterized in that The method further comprises: Acquire a target training signal representation vector obtained by extracting representation information of the first training signal source set to be enhanced; Iterate the following operations until X vector restoration mapping components in the vector restoration mapping network in the calibration process are walked through, wherein the vector restoration mapping network in the calibration process is used to determine a second training signal source set corresponding to the first training signal source set based on the target training signal representation vector, and the vector restoration mapping component includes a first cross-layer identity connection unit and a first weighted fusion unit, and X>1, and X is an integer: Obtaining an nth comparison restoration mapping representation vector output by the nth vector restoration mapping component, and loading the nth comparison restoration mapping representation vector into the first cross-layer identity connection unit in the mth vector restoration mapping component to obtain an mth cross-layer identity mapping vector, wherein m=n-1, 1<n≤X, and n is an integer; Obtaining a first influencing variable and a second influencing variable corresponding to the mth vector restoration mapping component, wherein the first influencing variable is used to characterize the importance of the nth control restoration mapping characterization vector, the second influencing variable is used to characterize the importance of the mth cross-layer identity mapping vector, and the ratio between the first influencing variable and the second influencing variable is not less than a first reference critical value; In the first weighted fusion unit, a weighted restoration mapping vector of the nth comparison restoration mapping representation vector and the mth cross-layer identity mapping vector is determined based on the first influencing variable and the second influencing variable, and the weighted restoration mapping vector is used as the mth comparison restoration mapping representation vector output by the mth vector restoration mapping component; If n=X, determining the second training signal source set corresponding to the first training signal source set based on the nth comparison restoration mapping representation vector; When it is determined based on the second training signal source set that the vector restoration mapping network in the calibration process meets the network stable state, the vector restoration mapping network in the calibration process is determined as the target vector restoration mapping network.

8. The method according to claim 7, characterized in that Before obtaining the target training signal representation vector obtained by extracting the representation information of the first training signal source set to be enhanced, the method further includes: Determine a two-dimensional array of sample signal sources corresponding to the first training signal source set based on signal source feature vectors respectively corresponding to a plurality of signal sources included in the first training signal source set; Iterate the following operations until Y representation information extraction components in the representation information extraction network in the calibration process are walked through, wherein the representation information extraction network is used to determine the target training signal representation vector corresponding to the first training signal source set based on the two-dimensional array of sample signal sources, and the representation information extraction component includes a second cross-layer identity connection unit and a second weighted fusion unit, wherein Y>1, and Y is an integer: Obtaining an rth comparison embedding mapping representation vector output by the rth representation information extraction component, and loading the rth comparison embedding mapping representation vector into the second cross-layer identity connection unit in the sth representation information extraction component to obtain an sth cross-layer identity embedding representation vector, wherein r=s-1, 1<s≤Y, and r is an integer; Obtaining a third influencing variable and a fourth influencing variable corresponding to the s-th representation information extraction component, wherein the third influencing variable is used to characterize the importance of the r-th control embedding mapping representation vector, the fourth influencing variable is used to characterize the importance of the s-th cross-layer identity embedding representation vector, and the ratio between the third influencing variable and the fourth influencing variable is not less than a second reference critical value; In the second weighted fusion unit, a weighted embedding mapping vector of the rth comparison embedding mapping representation vector and the sth cross-layer identity embedding representation vector is determined based on the third influencing variable and the fourth influencing variable, and the weighted embedding mapping vector is determined as the sth comparison embedding mapping representation vector output by the sth representation information extraction component; If s<Y, loading the s-th control embedding mapping representation vector into the t-th representation information extraction component; If s=Y, taking the s-th control embedding mapping representation vector as the target training signal representation vector; When it is determined based on the target training signal representation vector that the representation information extraction network in the calibration process meets the network stable state, the representation information extraction network in the calibration process is determined as the target representation information extraction network.

9. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Target detection method and target detection model training method and device

    CN118230084A

  • Compressed video quality enhancement method based on hybrid difference equation heuristic

    CN118646881A