Method for Matching User Visual Signal and Communication Signal Based on Object Detection

By adopting a target detection method in the communication system, the distribution characteristics and beam information of scattered objects in the communication scenario are obtained, and the location of the target user is identified using a neural network model, which solves the problem that the target user cannot be accurately identified in the prior art, and achieves a more accurate estimation of communication parameters.

CN116074864BActive Publication Date: 2025-06-10TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211690122.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-06-10
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

The prior art cannot accurately identify the position of the target user in the image from multiple communication scenarios through communication signals, making it difficult to achieve accurate communication parameter estimation.

Method used

Using a method based on object detection, a sequence of scattered object distribution characteristics of the target communication scene at multiple consecutive moments and an optimal transmitting and receiving beam pair sequence between the base station and the user is obtained, and the position probability distribution of the target user is obtained through the user matching neural network model, and the bounding box of the target user is determined.

Benefits of technology

It realizes the precise identification of the target user's position in the image from multiple communication scenarios, and improves the accuracy of communication parameter estimation of the visual perception auxiliary communication system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116074864B_ABST
    Figure CN116074864B_ABST
Patent Text Reader

Abstract

The present invention provides a matching method for user visual signals and communication signals based on target detection. The method includes: obtaining a sequence of scattering object distribution features corresponding to a target communication scenario at a plurality of consecutive moments; obtaining an optimal transceiver beam pair between a base station and a user at a plurality of consecutive moments, and forming a sequence of optimal transceiver beam pairs; obtaining a position probability distribution of a target user based on the sequence of scattering object distribution features and the sequence of optimal transceiver beam pairs; and determining a bounding box corresponding to the target user based on the position probability distribution of the target user. The method uses the beam direction information included in the optimal beam pair sequence to represent the azimuth information of the user relative to the base station, and jointly uses the sequence of scattering object distribution features corresponding to the communication scenario for the position estimation of the target user, and then matches the visual signal and the communication signal of the user, and can accurately identify the position of the target user in the image from multiple communication scenarios, so as to realize more accurate communication parameter design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and in particular, to a method for matching user visual signals and communication signals based on object detection. Background Art

[0002] Scene images obtained by visual perception can effectively reflect information such as the shape, position, and movement speed of various scattering objects in the environment, and thus can characterize the spatio-temporal characteristics of a wireless signal path, such as the fading coefficient, propagation direction, and correlation time, and can further be used to estimate or predict parameters such as the beam, occlusion, or received power of a wireless communication system. Therefore, a communication system assisted by visual perception has become one of the important research directions in the integrated sensing and communication technology.

[0003] However, when visual perception is used to assist a communication system, the images captured by a camera usually contain a variable number of scene objects, and it is difficult for the base station side to accurately identify the user corresponding to the current communication signal from these multiple scene objects, resulting in difficulty in achieving accurate communication parameter estimation. Summary of the Invention

[0004] The present invention provides a method for matching user visual signals and communication signals based on object detection to overcome the defect that the prior art cannot accurately identify the position of a target user in an image from multiple communication scenes through communication signals, thereby facilitating more accurate communication parameter estimation assisted by visual perception.

[0005] On the one hand, the present invention provides a method for matching user visual signals and communication signals based on object detection, including: obtaining a sequence of scattering object distribution features corresponding to a target communication scene at a continuous plurality of moments; obtaining an optimal transceiver beam pair between a base station and a user at the continuous plurality of moments, and forming a sequence of optimal transceiver beam pairs; obtaining a position probability distribution of a target user based on the sequence of scattering object distribution features and the sequence of optimal transceiver beam pairs; and determining a bounding box corresponding to the target user based on the position probability distribution of the target user.

[0006] Further, the obtaining a sequence of scattering object distribution features corresponding to a target communication scene at a continuous plurality of moments includes: obtaining multi-view images of the target communication scene at the continuous plurality of moments; using an object detection algorithm to extract bounding boxes of all scattering objects including the user from the multi-view images at each moment; performing feature design based on the bounding boxes of the scattering objects at each moment to obtain a sequence of scattering object distribution features; wherein the bounding box of the scattering object includes size information, position information, and attitude information of the corresponding scattering object.

[0007] Further, the feature design based on the bounding boxes of the scattering objects at each moment to obtain a scattering object distribution feature sequence includes: dividing the base station coverage plane area into multiple grids of equal size; determining the normalized average length, width, height, and average azimuth angle corresponding to the bounding boxes contained in each grid at each moment to obtain a four-dimensional feature vector corresponding to the grid at the corresponding moment; splicing the four-dimensional feature vectors corresponding to the multiple grids of equal size to obtain the scattering object distribution feature corresponding to each moment; obtaining the scattering object distribution feature sequence according to the scattering object distribution features corresponding to multiple consecutive moments; wherein, in the case where it is determined that a grid does not contain any bounding boxes, the four-dimensional feature vector corresponding to the grid is a zero vector.

[0008] Further, the obtaining of the position probability distribution of the target user based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence includes: inputting the scattering object distribution feature sequence and the optimal transceiver beam sequence into a pre-trained user matching neural network model to obtain the position probability distribution of the target user; wherein, the user matching neural network includes a first sub-neural network and a second sub-neural network, and the first sub-neural network is used to process the scattering object distribution feature sequence, and the second sub-neural network is used to process the optimal transceiver beam sequence.

[0009] Further, the inputting of the scattering object distribution feature sequence and the optimal transceiver beam sequence into a pre-trained user matching neural network model to obtain the position probability distribution of the target user includes: inputting the scattering object distribution feature sequence into the first sub-neural network to obtain a first output tensor; inputting the optimal transceiver beam sequence into the second sub-neural network to obtain a second output tensor; performing fusion processing on the first output tensor and the second output tensor to obtain a corresponding fusion tensor; inputting the fusion tensor into a number of pooling layers and two-dimensional convolutional layers to obtain a target heat map; wherein, the first output tensor and the second output tensor have the same dimension, and the target heat map represents the position probability distribution of the target user.

[0010] Further, the first sub-neural network includes a number of two-dimensional convolutional layers, and the second sub-neural network includes an embedding layer, a long short-term memory layer, a fully connected layer, a reshaping layer, and a two-dimensional convolutional layer.

[0011] Further, the determining of the bounding box corresponding to the target user based on the position probability distribution of the target user includes: estimating the user position of the target user at the current moment according to the position probability distribution of the target user; matching the bounding box corresponding to the target user from the scattering object bounding boxes according to the estimated user position.

[0012] Further, estimating the user location of the target user at the current moment according to the location probability distribution of the target user includes: dividing the base station coverage plane area into a plurality of heatmap grids, and the plurality of heatmap grids form the target heatmap; determining the coordinate serial number of the largest element in the target heatmap; determining the central plane position coordinate of the heatmap grid corresponding to the coordinate serial number, and the central plane position coordinate is the estimated user location of the target user at the current moment.

[0013] Further, the expression of the central plane position coordinate is as follows:

[0014]

[0015] where k is the heatmap grid label, is the heatmap grid The coordinate of the lower left corner vertex in the global coordinate system, and are the length and width of the heatmap grid respectively.

[0016] Further, matching the bounding box corresponding to the target user from the scattering object bounding box according to the estimated user location includes: determining the target bounding box with the minimum distance from the central plane position coordinate in the scattering object bounding box through a preset formula; where the target bounding box is the bounding box corresponding to the target user, and the preset formula is as follows:

[0017]

[0018] where Box U is the target bounding box, is the heatmap grid The coordinate of the lower left corner vertex in the global coordinate system, (x G , y G ) is the plane coordinate of the center position of the bounding box in the global coordinate system, and are the length and width of the heatmap grid respectively.

[0019] Further, training the user matching neural network model specifically includes: based on multi-perspective pictures of the communication scenario at multiple consecutive moments collected, obtaining a corresponding training set of scattering object distribution feature sequences; obtaining the optimal transceiver beam pair between the base station and the user at the multiple consecutive moments, and forming a training set of optimal transceiver beam pair sequences; generating a corresponding two-dimensional Gaussian probability distribution on the training heat map according to the actual position coordinates of the user at the multiple consecutive moments. Using the optimal transceiver beam pair sequence training set and the scattering object distribution feature sequence training set as sample inputs, and using the two-dimensional Gaussian probability distribution as the true sample label, training the user matching neural network model until convergence.

[0020] In a second aspect, the present invention further provides a matching device for user visual signals and communication signals based on target detection, including: a feature sequence acquisition module, configured to acquire a scattering object distribution feature sequence corresponding to a target communication scenario at multiple consecutive moments; an optimal transceiver beam pair sequence acquisition module, configured to acquire the optimal transceiver beam pair between the base station and the user at the multiple consecutive moments, and form an optimal transceiver beam pair sequence; a position probability distribution acquisition module, configured to acquire the position probability distribution of the target user based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence; and a user bounding box determination module, configured to determine the bounding box corresponding to the target user based on the position probability distribution of the target user.

[0021] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the matching method for user visual signals and communication signals based on target detection as described in any one of the above.

[0022] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the matching method for user visual signals and communication signals based on target detection as described in any one of the above.

[0023] A method for matching user visual signals and communication signals based on object detection provided by the present invention obtains a sequence of scattering object distribution features corresponding to a target communication scenario at a continuous plurality of moments, and a sequence of optimal transceiver beam pairs between a base station and a user at a continuous plurality of moments, and based on the sequence of scattering object distribution features and the sequence of optimal transceiver beam pairs, obtains a position probability distribution of a target user, so as to determine a bounding box corresponding to the target user according to the position probability distribution of the target user. This method uses the beam direction information included in the optimal beam pair sequence to characterize the azimuth information of the user relative to the base station, and jointly uses the sequence of scattering object distribution features corresponding to the target communication scenario for the position estimation of the target user, and then matches the visual signal and the communication signal of the user, and can accurately identify the position of the target user in the image from multiple communication scenarios, so as to achieve more accurate communication parameter design. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0025] Figure 1 It is a schematic flowchart of the method for matching user visual signals and communication signals based on object detection provided by the present invention;

[0026] Figure 2 It is a schematic diagram of a communication scenario of the method for matching user visual signals and communication signals based on object detection provided by the present invention;

[0027] Figure 3 It is a schematic diagram of a 3D bounding box of the method for matching user visual signals and communication signals based on object detection provided by the present invention;

[0028] Figure 4 It is a schematic diagram of the design of scattering object distribution features of the method for matching user visual signals and communication signals based on object detection provided by the present invention;

[0029] Figure 5 It is a schematic diagram of the inference of the user matching neural network model provided by the present invention;

[0030] Figure 6 It is a schematic diagram of the structure of the device for matching user visual signals and communication signals based on object detection provided by the present invention;

[0031] Figure 7 It is a schematic diagram of the structure of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] Figure 1 The flowchart of the method for matching user visual signals and communication signals based on object detection provided by the present invention is shown. As Figure 1 shown, the method includes:

[0034] S110, obtaining a sequence of scattering object distribution features corresponding to a target communication scenario at a plurality of consecutive moments.

[0035] It should be noted that in the target communication scenario, a plurality of RGB cameras with different viewing angle ranges are equipped, which can ensure that the viewing angles of all cameras can completely cover the entire area range of the target communication scenario.

[0036] It should also be noted that based on the target communication scenario, a global coordinate system GCS is defined, and the global coordinate system has coordinate axes X G -Y G -Z G . Among them, the X G -Y G plane is the ground plane of the communication scenario, and the Z G axis is perpendicular to the ground plane axially.

[0037] It can be understood that to obtain the scattering object distribution features corresponding to the target communication scenario at a plurality of consecutive moments, specifically, first, use the plurality of RGB cameras equipped in the target communication scenario to collect multi-view pictures of the target communication scenario at a plurality of consecutive moments, and obtain a corresponding multi-view picture sequence.

[0038] For example, in a specific embodiment, obtain a multi-view picture set at M consecutive moments t k-M+1 , t k-M+2 , …, t k .

[0039] Based on the collected multi-view picture sequence, use a 3D object detection algorithm to extract the scattering object feature information corresponding to the scattering objects in the multi-view pictures at each moment. The scattering object feature information includes, but is not limited to, the bounding box where the corresponding scattering object is located, position information, length, width, and height information, and pose information.

[0040] Further, based on the scattering object feature information corresponding to each moment obtained by extraction, feature design is performed, and thus a scattering object distribution feature sequence corresponding to the target communication scenario at consecutive moments can be obtained.

[0041] S120. Obtain the optimal transceiver beam pair between the base station and the user at consecutive moments, and form an optimal transceiver beam pair sequence.

[0042] It can be understood that when using multiple cameras to collect multi-view images of the target communication scenario at consecutive moments, the optimal transceiver beam pair between the base station and the user at these consecutive moments is also obtained simultaneously, thereby forming an optimal transceiver beam pair sequence.

[0043] Among them, the optimal transceiver beam pair can be measured by a beam measurement device.

[0044] It should be noted that the consecutive moments corresponding to the optimal transceiver beam pair sequence are the same consecutive moments as those corresponding to the scattering object distribution feature sequence.

[0045] In a specific embodiment, the serial numbers corresponding to the formed optimal transceiver beam pair sequence are I k-M+1 , I k-M+2 , …, I k .

[0046] S130. Based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence, obtain the position probability distribution of the target user;

[0047] S140. Based on the position probability distribution of the target user, determine the bounding box corresponding to the target user.

[0048] It can be understood that based on obtaining the scattering object distribution feature sequence in step S110 and obtaining the optimal transceiver beam pair sequence in step S120, according to the scattering object distribution feature sequence and the optimal transceiver beam pair sequence, the position probability distribution of the target user can be obtained, and then the bounding box corresponding to the target user can be determined, that is, the position of the target user corresponding to the current communication signal.

[0049] Preferably, to obtain the position probability distribution of the target user, it can be predicted by a pre-trained user matching neural network model. This user matching neural network model takes the scattering object distribution feature sequence and the optimal transceiver beam pair sequence as inputs and the position probability distribution of the target user as the output.

[0050] Based on the position probability distribution of the target user, the user position of the target user at the current moment can be estimated, and then the bounding box corresponding to the target user can be matched from the scattering object feature information.

[0051] Among them, the position probability distribution of the target user can be represented by a heat map.

[0052] In this embodiment, by obtaining the scattering object distribution feature sequence corresponding to the target communication scenario at a continuous plurality of moments, and the optimal transceiver beam pair sequence between the base station and the user at a continuous plurality of moments, and based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence, the position probability distribution of the target user is obtained, so as to determine the bounding box corresponding to the target user according to the position probability distribution of the target user. This method uses the beam direction information included in the optimal beam pair sequence to characterize the azimuth information of the user relative to the base station, and jointly uses the scattering object distribution feature sequence corresponding to the target communication scenario for the position estimation of the target user, and then matches the visual signal and the communication signal of the user, and can accurately identify the position of the target user in the image from multiple communication scenarios, so as to achieve more accurate communication parameter design.

[0053] On the basis of the above embodiment, further, obtaining the scattering object distribution feature sequence corresponding to the target communication scenario at a continuous plurality of moments includes: obtaining multi-view pictures of the target communication scenario at a continuous plurality of moments; using a target detection algorithm to extract the 3D bounding boxes of all scattering objects including the user from the multi-view pictures at each moment; performing feature design based on the scattering object bounding boxes at each moment to obtain the scattering object distribution feature sequence; wherein, the scattering object bounding box includes the size information, position information and attitude information of the corresponding scattering object.

[0054] It can be understood that to obtain the scattering object distribution feature sequence corresponding to the target communication scenario at a continuous plurality of moments, specifically, first use multiple RGB cameras equipped in the target communication scenario to collect multi-view pictures of the target communication scenario at a continuous plurality of moments to obtain a corresponding multi-view picture sequence.

[0055] Then, use a 3D target detection algorithm to extract the 3D bounding boxes of all scattering objects including the user from the multi-view pictures at each moment, and perform feature design for the scattering object bounding boxes at each moment, then the feature sequence that can describe the scattering object distribution of the scene can be obtained, that is, the scattering object distribution feature sequence corresponding to the target communication scenario at a continuous plurality of moments.

[0056] It should be noted that the 3D bounding box of each extracted scattering object includes but is not limited to the size information, position information and attitude information of the corresponding scattering object, wherein the size information refers to the length, width and height of the corresponding scattering object, the position information is the coordinate information relative to the global coordinate system, and the attitude information is the azimuth angle information.

[0057] Specifically, as described in the previous embodiment, in the target communication scenario, there are multiple RGB cameras with different viewing ranges, and based on the target communication scenario, a global coordinate system GCS is also defined, and the global coordinate system has coordinate axes X G -Y G -Z G . Among them, the X G -Y G plane is the ground plane of the communication scenario, and the Z G axis is perpendicular to the ground plane.

[0058] Specifically, Figure 2 shows a schematic diagram of the communication scenario of the method for matching user visual signals and communication signals based on object detection provided by the present invention. As Figure 2 shown, there are multiple base stations in the target communication scenario, and multiple cameras with different viewing ranges are equipped. These cameras will transmit the multi-view pictures collected to a common central processing unit, and the central processing unit will process the multi-view pictures collected.

[0059] That is to say, at each moment, the multi-view images captured by all the cameras equipped in the target communication scenario will be transmitted to the central processing unit, and the central processing unit will perform multi-camera 3D object detection, that is, extract the distributed feature information within the scattered objects from the multi-view pictures captured at each moment through the multi-camera 3D object detection algorithm, specifically the 3D bounding boxes of the scene scattered objects.

[0060] Figure 3 shows a schematic diagram of the 3D bounding box of the method for matching user visual signals and communication signals based on object detection provided by the present invention. As Figure 3 shown, the 3D bounding box information of each scattered object includes the coordinates of its center position in the global coordinate system GCS, its length, width, height, and the azimuth angle in the global coordinate system GCS.

[0061] In a specific embodiment, obtain the multi-view picture sets at consecutive M moments t k-M+1 , t k-M+2 , …, t k The set of 3D bounding boxes of the scattered objects detected from the multi-view picture sets will be respectively denoted as x , x k-M+1 , x k-M+2 , … x k .

[0062] Further, based on the scattering object bounding boxes at each moment, feature design is performed to obtain a scattering object distribution feature sequence. Specifically, the base station coverage plane area is divided into multiple grids of equal size; the normalized average length, width, height, and average azimuth angle corresponding to the bounding boxes contained in each grid at each moment are determined, obtaining a four-dimensional feature vector corresponding to the grid at the corresponding moment; the four-dimensional feature vectors corresponding to multiple grids of equal size are concatenated to obtain the scattering object distribution feature corresponding to each moment; according to the scattering object distribution features corresponding to the consecutive multiple moments, a scattering object distribution feature sequence is obtained; wherein, in the case where it is determined that a grid does not contain any bounding boxes, the four-dimensional feature vector corresponding to the grid is a zero vector.

[0063] It can be understood that the feature design is performed separately for the set of scattering object 3D bounding boxes at each moment. For the convenience of description hereinafter, it is all represented by x, indicating the set composed of the scattering object 3D bounding boxes detected in the multi-view image set taken at a certain moment t.

[0064] First, the base station coverage plane area is divided into multiple grids of equal size. Specifically, Figure 4 FIG. shows a schematic diagram of the scattering object distribution feature design of the user visual signal and communication signal matching method based on object detection provided by the present invention.

[0065] As Figure 4 shown, the base station coverage area on the X G -Y G plane in the global coordinate system GCS is divided into multiple grids of equal size. The length and width of each grid are L D and W D . The number of columns and rows of the grids in the entire base station coverage area are N X and N Y .

[0066] The grid in the n X th column and the n Y th row is denoted as the (n X , n Y )th grid, where n X = 1, 2,... N X , and n Y = 1, 2,... N Y . All 3D bounding boxes whose center plane positions are included in the (n X , n Y )th grid are statistically counted from the scattering object 3D bounding box χ, and the set composed of these 3D bounding boxes is denoted as:

[0067]

[0068] Wherein, is the coordinate of the lower left corner vertex of the (n X , n Y )-th grid in the global coordinate system GCS, and (x G , y G ) is the X G -Y G plane coordinate of the center position of the 3D bounding box Box in the global coordinate system GCS.

[0069] On the above basis, the average length , width , height and azimuth angle of all 3D bounding boxes in can be further obtained. At the same time, the maximum length l max , width w max and height h max of all possible scattering objects in the target communication scenario are obtained, so as to obtain the normalized average length, width and height of all 3D bounding boxes in the 3D bounding box set of the scattering objects.

[0070] Among them, the normalized average length is The normalized average width is The normalized average height is

[0071] According to the above, the four-dimensional feature vector corresponding to the (n X , n Y )-th grid can be obtained. This four-dimensional feature vector includes the above-mentioned normalized average length, width, height and average azimuth angle, and can be specifically expressed as

[0072] Furthermore, by splicing the four-dimensional feature vectors corresponding to all equal-sized networks in the base station coverage plane area, an N X ×N Y ×4-dimensional matrix D can be obtained. This matrix D is the scattering object distribution feature corresponding to time t.

[0073] It should be noted that when it is determined that there is a grid in the base station coverage plane area that does not contain any 3D bounding boxes, the four-dimensional feature vector corresponding to this grid is recorded as a zero vector.

[0074] It should also be noted that the above matrix D is only the scattering object distribution feature sequence corresponding to time t. For times t k-M+1 , t k-M+2 , …, t k , through x k-M+1 , x k-M+2 , …, x k, the corresponding scattering object distribution feature sequence D can be obtained in the same manner as above k-M+1 , D k-M+2 , …, D k , which will not be elaborated one by one here.

[0075] In this embodiment, by acquiring multi-view images of the target communication scenario at consecutive moments and using a target detection algorithm to extract the bounding boxes of all scattering objects including users from the multi-view images at each moment, feature design is performed based on the scattering object bounding boxes at each moment to obtain the corresponding scattering object distribution feature sequence, and jointly with the optimal transceiver beam pair sequence that can be used to characterize the azimuth information of the user relative to the base station, it is used for the position estimation of the target user. Furthermore, by matching the visual signal and communication signal of the user, the position of the target user in the image can be accurately identified from multiple communication scenarios, thus enabling more accurate communication parameter design.

[0076] On the basis of the above embodiment, further, based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence, the position probability distribution of the target user is obtained, including: inputting the scattering object distribution feature sequence and the optimal transceiver beam sequence into a pre-trained user matching neural network model to obtain the position probability distribution of the target user.

[0077] It can be understood that a user matching neural network model that can fuse multi-modal information is pre-designed and trained until convergence for predicting the position probability distribution of the target user.

[0078] Specifically, the scattering object distribution feature sequence D k-M+1 , D k-M+2 , …, D k and the optimal transceiver beam pair sequence I k-M+1 , I k-M+2 , …, I k are simultaneously input into the trained user matching neural network model, and the corresponding output, that is, the two-dimensional position probability distribution of the target user at the current moment, can be obtained.

[0079] It should be noted that the pre-trained user matching neural network model includes a first sub-neural network and a second sub-neural network. The first sub-neural network is used to process the scattering object distribution feature sequence, and the second sub-neural network is used to process the optimal transceiver beam sequence.

[0080] The scattering object distribution feature sequence and the optimal transceiver beam sequence are input into a pre-trained user matching neural network model to obtain the position probability distribution of the target user. Specifically, the scattering object distribution feature sequence is input into the first sub-neural network to obtain a first output tensor; the optimal transceiver beam sequence is input into the second sub-neural network to obtain a second output tensor; the first output tensor and the second output tensor are fused to obtain a corresponding fusion tensor; the fusion tensor is input into a number of pooling layers and two-dimensional convolutional layers to obtain a target heat map; wherein, the first output tensor and the second output tensor have the same dimension, and the target heat map represents the position probability distribution of the target user.

[0081] Specifically, Figure 5 Fig. shows the inference schematic diagram of the user matching neural network model provided by the present invention. As Figure 5 shown, the network model respectively uses two sub-neural networks to extract features from the scattering object distribution feature sequence and the optimal transceiver beam pair sequence, and then adds and fuses the output tensors of the two sub-neural networks, and then uses the fused features as the input of other network layers to obtain the position probability distribution of the target user at time t k below.

[0082] Specifically, for the first sub-neural network used to process the scattering object distribution feature sequence, the scattering object distribution feature sequence D k-M+1 , D k-M+2 , …, D k is concatenated into an N X ×N Y ×4M-dimensional tensor Then it is input into a number of two-dimensional convolutional layers to obtain a first output tensor

[0083] For the second sub-neural network used to process the optimal transceiver beam pair sequence, the optimal transceiver beam pair sequence I k-M+1 , I k-M+2 , …, I k will first be input into an embedding layer to transform it into a vector sequence, and then the vector sequence will be input into a number of long short-term memory network layers, and then the output tensor of the last long short-term memory network layer will be input into multiple fully connected layers to obtain a vector of length N X N Y The vector is input into a reshaping layer to reshape it into an N ×N X ×1-dimensional tensor, and then, this N Y ×N X ×N YThe 1D tensor is input into several 2D convolutional layers, and finally the second output tensor is obtained.

[0084] It should be noted that the first output tensor and the second output tensor have the same dimension, and the specific dimension can be set according to the actual situation, and no specific limitation is made here.

[0085] Based on the first output tensor and the second output tensor, the first output tensor and the second output tensor are added and fused to obtain a fused tensor, and the fused tensor is input into several pooling layers and 2D convolutional layers to obtain the final output, that is, a -dimensional target heat map F k , and the target heat map can represent the position probability distribution of the user.

[0086] In addition, it should be noted that before using the user matching neural network model for inference, the user matching neural network model also needs to be trained. Specifically, based on the multi-view images of the communication scenarios at multiple consecutive times collected, the corresponding scattering object distribution feature sequence training set is obtained; the optimal transceiver beam pairs between the base station and the user at multiple consecutive times are obtained, and the optimal transceiver beam pair sequence training set is formed; according to the actual position coordinates of the user at multiple consecutive times, a corresponding two-dimensional Gaussian probability distribution is generated on the training heat map. Using the optimal transceiver beam pair sequence training set and the scattering object distribution feature sequence training set as sample inputs, and using the two-dimensional Gaussian probability distribution as the true sample label, the user matching neural network model is trained until convergence.

[0087] It is easy to understand that at multiple consecutive times, multiple configured cameras are used to collect multi-view images, and the optimal transceiver beam pairs between the base station and the user at each time are measured, and the actual position coordinates of the user at this time are recorded.

[0088] Based on multiple consecutive times t k-M+1 , t k-M+2 , …, t k The scattering object distribution feature sequence is obtained from the collected multi-view images, and the scattering object distribution feature sequence and the corresponding optimal transceiver beam pair sequence are used as sample inputs of the neural network. At the same time, according to the actual center plane position coordinates of the user at the current time t k a two-dimensional Gaussian probability distribution is generated on the heat map , and is used as the sample label, so as to train the user matching neural network model.

[0089] Among them, is -dimensional heat map.

[0090] Denote the serial number of the heatmap grid containing as Then a two-dimensional Gaussian distribution can be obtained:

[0091]

[0092] where σ is the standard deviation of this Gaussian distribution, which is related to the length and width of the target user. Preferably, σ can be set to where then The elements of are:

[0093]

[0094] where

[0095] In this embodiment, the first sub-neural network of the user matching neural network model is used to extract features from the scattering object distribution feature sequence to obtain a first output tensor, and the second sub-neural network of the user matching neural network model is used to extract features from the corresponding optimal transceiver beam pair sequence to obtain a second output tensor. The first output tensor and the second output tensor are fused to obtain a corresponding fusion tensor, and this fusion tensor is input into several pooling layers and two-dimensional convolutional layers, then a target heatmap for characterizing the position probability distribution of the target user can be obtained. By predicting the position probability distribution of the target user through the pre-trained user matching neural network model, the position of the target user in the image can be accurately identified from multiple communication scenarios, thereby realizing more accurate communication parameter design.

[0096] Based on the above embodiment, further, based on the position probability distribution of the target user, determining the bounding box corresponding to the target user includes: estimating the user position of the target user at the current moment according to the position probability distribution of the target user; matching the bounding box corresponding to the target user from the scattering object bounding boxes according to the estimated user position.

[0097] It can be understood that determining the bounding box corresponding to the target user based on the position probability distribution of the target user, specifically, estimating the user position of the target user at the current moment according to the position probability distribution of the target user output by the user matching neural network model, that is, the target heatmap, so as to match the bounding box corresponding to the target user from the extracted scattering object bounding boxes.

[0098] Estimate the user location of the target user at the current moment according to the location probability distribution of the target user. Specifically, divide the base station coverage plane area into multiple heatmap grids, and the multiple heatmap grids form a target heatmap; determine the coordinate number of the largest element in the target heatmap; determine the central plane position coordinates of the heatmap grid corresponding to the coordinate number, and the central plane position coordinates are the estimate of the user location of the target user at the current moment.

[0099] Specifically, cut the base station coverage plane area on the X G -Y G plane of the global coordinate system GCS into equally sized heatmap grids, and are the number of columns and rows of each heatmap grid respectively, and the length and width of each heatmap grid are and

[0100] The (n k , n X , n Y )-th element of the target heatmap F X , n Y ) represents the probability that the central plane position of the 3D bounding box corresponding to the target user will fall within the (n k )-th heatmap grid. Thus, the most likely heatmap grid where the target user exists can be estimated by determining the serial number of the largest element in the target heatmap D

[0101] Among them, the central plane position coordinates of the heatmap grid corresponding to the coordinate number of the largest element in the target heatmap F k can be expressed as follows:

[0102]

[0103] where k is the heatmap grid label, is the coordinate of the lower left vertex of the heatmap grid in the global coordinate system GCS, and are the length and width of the heatmap grid respectively.

[0104] That is, the central plane position coordinates of the heatmap grid corresponding to the coordinate number of the above largest element can be used as the location estimate of the target user at time t k .

[0105] Based on the estimated user location of the target user, according to the estimated user location, match the bounding box corresponding to the target user from the scattering object bounding boxes. Specifically, determine the target bounding box with the smallest distance from the center plane position coordinates in the scattering object bounding boxes through a preset formula; where the target bounding box is the bounding box corresponding to the target user, and the preset formula is as follows:

[0106]

[0107] where Box U is the target bounding box, is the heatmap grid coordinates of the lower left vertex in the global coordinate system, (x G , y G ) is the plane coordinates of the center position of the bounding box in the global coordinate system, and are the length and width of the heatmap grid respectively.

[0108] In this embodiment, by estimating the user location of the target user according to the position probability distribution of the target user, and according to the estimated user location, match the bounding box corresponding to the target user from the scattering object bounding boxes. This method uses the beam direction information included in the optimal beam pair sequence to represent the azimuth information of the user relative to the base station, and jointly uses the scattering object distribution feature sequence corresponding to the target communication scenario for the position estimation of the target user, and then matches the visual signal and communication signal of the user, and can accurately identify the position of the target user in the image from multiple communication scenarios, so as to realize more accurate communication parameter design.

[0109] Figure 6 shows the structural schematic diagram of the matching device for the user visual signal and communication signal based on target detection provided by the present invention. As Figure 6 shown, the device includes: a feature sequence acquisition module 610, configured to acquire a scattering object distribution feature sequence corresponding to a target communication scenario at a continuous plurality of moments; an optimal transceiver beam pair sequence acquisition module 620, configured to acquire an optimal transceiver beam pair between the base station and the user at the continuous plurality of moments, and form an optimal transceiver beam pair sequence; a position probability distribution acquisition module 630, configured to acquire a position probability distribution of a target user based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence; and a user bounding box determination module 640, configured to determine a bounding box corresponding to the target user based on the position probability distribution of the target user.

[0110] In this embodiment, the scattering object distribution feature sequence corresponding to the target communication scenario is obtained by the feature sequence acquisition module 610, and the optimal transceiver beam pair sequence between the base station and the user is obtained by the optimal transceiver beam pair sequence acquisition module 620 at consecutive moments. The position probability distribution acquisition module 630 obtains the position probability distribution of the target user based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence. Then, the user bounding box determination module 640 determines the bounding box corresponding to the target user according to the position probability distribution of the target user. The device uses the beam direction information included in the optimal beam pair sequence to represent the azimuth information of the user relative to the base station, and jointly uses the scattering object distribution feature sequence corresponding to the target communication scenario for the position estimation of the target user, and then matches the visual signal and the communication signal of the user, so as to accurately identify the position of the target user in the image from multiple communication scenarios, thereby realizing more accurate communication parameter design.

[0111] The matching device for user visual signals and communication signals based on target detection provided by the embodiments of the present invention can be correspondingly referred to the matching method for user visual signals and communication signals based on target detection described above, and will not be elaborated here.

[0112] Figure 7 An entity structure diagram of an electronic device is exemplified, as Figure 7 shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the matching method for user visual signals and communication signals based on target detection, and the method includes: obtaining the scattering object distribution feature sequence corresponding to the target communication scenario at consecutive moments; obtaining the optimal transceiver beam pair between the base station and the user at the consecutive moments and forming an optimal transceiver beam pair sequence; obtaining the position probability distribution of the target user based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence; and determining the bounding box corresponding to the target user based on the position probability distribution of the target user.

[0113] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0114] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the matching method of the user visual signal and the communication signal based on target detection provided by the above-mentioned various methods. The method includes: obtaining a sequence of scattering object distribution characteristics corresponding to a target communication scenario at a continuous plurality of moments; obtaining an optimal transceiver beam pair between a base station and a user at the continuous plurality of moments and forming a sequence of optimal transceiver beam pairs; obtaining a position probability distribution of a target user based on the sequence of scattering object distribution characteristics and the sequence of optimal transceiver beam pairs; and determining a bounding box corresponding to the target user based on the position probability distribution of the target user.

[0115] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A matching method for user visual signals and communication signals based on object detection, characterized in that, it includes: Obtaining a sequence of scattering object distribution features corresponding to a target communication scenario at consecutive multiple moments, including: Obtaining multi-view images of the target communication scenario at the consecutive multiple moments; Using an object detection algorithm to extract the bounding boxes of all scattering objects including the user from the multi-view images at each moment; Performing feature design based on the scattering object bounding boxes at each moment to obtain a sequence of scattering object distribution features; wherein, the scattering object bounding box contains size information, position information, and attitude information of the corresponding scattering object; Obtaining the optimal transceiver beam pair between the base station and the user at the consecutive multiple moments, and forming a sequence of optimal transceiver beam pairs; Based on the sequence of scattering object distribution features and the sequence of optimal transceiver beam pairs, obtaining the position probability distribution of the target user; Based on the position probability distribution of the target user, determining the bounding box corresponding to the target user, including: Estimating the user position of the target user at the current moment according to the position probability distribution of the target user; According to the estimated user position, matching the bounding box corresponding to the target user from the scattering object bounding boxes.

2. The matching method for user visual signals and communication signals based on object detection according to claim 1, characterized in that, The performing feature design based on the scattering object bounding boxes at each moment to obtain a sequence of scattering object distribution features includes: Dividing the base station coverage plane area into multiple grids of equal size; Determining the normalized average length, width, and height, and the average azimuth angle corresponding to the bounding boxes contained in each grid at each moment, to obtain a four-dimensional feature vector corresponding to the grid at the corresponding moment; Concatenating the four-dimensional feature vectors corresponding to the multiple grids of equal size to obtain the scattering object distribution feature corresponding to each moment; According to the scattering object distribution features corresponding to the consecutive multiple moments, obtaining the sequence of scattering object distribution features; wherein, in the case where it is determined that the grid does not contain any bounding boxes, the four-dimensional feature vector corresponding to the grid is a zero vector.

3. The matching method for user visual signals and communication signals based on object detection according to claim 1, characterized in that, The obtaining the position probability distribution of the target user based on the sequence of scattering object distribution features and the sequence of optimal transceiver beam pairs includes: Inputting the sequence of scattering object distribution features and the sequence of optimal transceiver beam pairs into a pre-trained user matching neural network model to obtain the position probability distribution of the target user; wherein, the user matching neural network includes a first sub-neural network and a second sub-neural network, the first sub-neural network is used to process the sequence of scattering object distribution features, and the second sub-neural network is used to process the sequence of optimal transceiver beam pairs.

4. The matching method for user visual signals and communication signals based on object detection according to claim 3, characterized in that, Inputting the scattering object distribution feature sequence and the optimal transceiver beam pair sequence into a pre-trained user matching neural network model to obtain the position probability distribution of the target user includes: Inputting the scattering object distribution feature sequence into the first sub-neural network to obtain a first output tensor; Inputting the optimal transceiver beam pair sequence into the second sub-neural network to obtain a second output tensor; Performing fusion processing on the first output tensor and the second output tensor to obtain a corresponding fusion tensor; Inputting the fusion tensor into a number of pooling layers and two-dimensional convolutional layers to obtain a target heat map; Wherein, the first output tensor has the same dimension as the second output tensor, and the target heat map represents the position probability distribution of the target user.

5. The method for matching user visual signals and communication signals based on target detection according to claim 4, wherein, The first sub-neural network includes a number of two-dimensional convolutional layers, and the second sub-neural network includes an embedding layer, a long short-term memory layer, a fully connected layer, a reshaping layer, and a two-dimensional convolutional layer.

6. The method for matching user visual signals and communication signals based on target detection according to claim 5, wherein, Estimating the user position of the target user at the current moment according to the position probability distribution of the target user includes: Dividing the base station coverage plane area into multiple heat map grids, and the multiple heat map grids form the target heat map; Determining the coordinate serial number of the largest element in the target heat map; Determining the central plane position coordinate of the heat map grid corresponding to the coordinate serial number, and the central plane position coordinate is the estimated user position of the target user at the current moment.

7. The method for matching user visual signals and communication signals based on target detection according to any one of claims 3-6, wherein, Training the user matching neural network model specifically includes: Based on multiple multi-view images of the communication scenario at multiple consecutive moments, obtaining a corresponding training set of scattering object distribution feature sequences; Obtaining the optimal transceiver beam pairs between the base station and the user at the multiple consecutive moments, and forming a training set of optimal transceiver beam pair sequences; Generating a corresponding two-dimensional Gaussian probability distribution on the training heat map according to the actual position coordinates of the user at the multiple consecutive moments; Using the training set of optimal transceiver beam pair sequences and the training set of scattering object distribution feature sequences as sample inputs, and using the two-dimensional Gaussian probability distribution as the true sample label, training the user matching neural network model until convergence.

8. A device for matching user visual signals and communication signals based on target detection, wherein, includes: A feature sequence acquisition module for acquiring a scattering object distribution feature sequence corresponding to a target communication scenario at multiple consecutive moments, including: Obtaining the multi-view images of the target communication scenario at the multiple consecutive moments; Using a target detection algorithm to extract the bounding boxes of all scattering objects including the user from the multi-view images at each moment; Performing feature design based on the scattering object bounding boxes at each moment to obtain a scattering object distribution feature sequence; Among them, the scattering object bounding box contains size information, position information, and attitude information of the corresponding scattering object; An optimal transceiver beam pair sequence acquisition module, configured to acquire the optimal transceiver beam pairs between the base station and the user at the consecutive multiple moments, and form an optimal transceiver beam pair sequence; A position probability distribution acquisition module, configured to acquire the position probability distribution of the target user based on the scattering object distribution feature sequence and the optimal transceiver beam pair sequence; A user bounding box determination module, configured to determine the bounding box corresponding to the target user based on the position probability distribution of the target user, including: Estimate the user position of the target user at the current moment according to the position probability distribution of the target user; Match the bounding box corresponding to the target user from the scattering object bounding boxes according to the estimated user position.

9. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the program, the steps of the method for matching user visual signals and communication signals based on target detection according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by the processor, the steps of the method for matching user visual signals and communication signals based on target detection according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Systems and methods for visual target tracking

    CN108351654A

  • Method for visual target tracking

    CN113589833A