An article detection method, device and equipment
By compressing and extracting features from millimeter-wave holographic data, a second feature sequence is generated for detection, solving the problems of noise interference and storage resource consumption caused by dense data, and achieving accurate detection and resource saving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2025-11-27
- Publication Date
- 2026-07-21
AI Technical Summary
When performing detection based on millimeter-wave holographic data, the dense data leads to significant noise interference, making it impossible to obtain accurate detection results. Furthermore, storing a large amount of holographic data consumes a significant amount of storage resources.
By generating a first feature sequence and performing information compression and feature extraction, C first feature matrices are converted into K second feature matrices, where K is less than C. The target feature values are determined using the parameter metric matrix and the state matrix, and the second feature sequence is generated for detection.
It improves the accuracy of detection results, reduces storage resource consumption and training costs, enhances target resolution, and alleviates the pressure on holographic data storage and training resources.
Smart Images

Figure CN121634026B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent recognition technology, and in particular to an article detection method, apparatus and equipment. Background Technology
[0002] Millimeter waves are electromagnetic waves with wavelengths of 1-10 mm. They can penetrate most fabrics. When they encounter metals or non-metals that cannot be penetrated by clothing, the signal is reflected, ultimately yielding millimeter-wave holographic data. This data includes amplitude, phase, and distance information. Based on these characteristics, security inspection equipment (such as security scanners) can actively emit millimeter-wave signals and acquire the echo signals. By extracting key wavebands from the echo signals, they can reconstruct millimeter-wave holographic data. Based on this, it is possible to detect whether an object is carrying a specified type of item, and to detect the category and location of that item.
[0003] However, when performing detection based on millimeter-wave holographic data, the high density of this data leads to significant noise interference, making accurate detection results impossible. Furthermore, the high density of millimeter-wave holographic data necessitates the storage of large amounts of data, consuming substantial storage resources. Summary of the Invention
[0004] This application provides a method for detecting articles, the method comprising: A first feature sequence is generated based on the holographic data of the object to be detected. The first feature sequence includes C first feature matrices, and the first feature matrix includes initial feature values of multiple spatial locations. For any spatial location, based on the initial eigenvalues corresponding to that spatial location in C first feature matrices, K target eigenvalues for that spatial location are determined, where K is less than C; wherein, for any initial eigenvalue corresponding to that spatial location, based on the acquired parameter metric matrix, the initial eigenvalue, and the state matrix corresponding to the previous initial eigenvalue, a state matrix corresponding to that initial eigenvalue is determined, and the state matrix includes K eigenvalues; wherein, the K eigenvalues in the state matrix corresponding to the last initial eigenvalue are taken as the K target eigenvalues for that spatial location; A second feature sequence is generated based on K target feature values at each spatial location. The second feature sequence includes K second feature matrices, each of which includes the target feature values at each spatial location. The detection result of the specified type of item of the object to be detected is determined based on the second feature sequence.
[0005] This application provides an object detection device, the device comprising: a generation module, configured to generate a first feature sequence based on the holographic data of the object to be detected, the first feature sequence comprising C first feature matrices, the first feature matrices comprising initial feature values of multiple spatial locations; The determination module is used to determine K target feature values for any spatial location based on the initial feature values corresponding to that spatial location in C first feature matrices, where K is less than C; wherein, for any initial feature value corresponding to that spatial location, the state matrix corresponding to that initial feature value is determined based on the acquired parameter metric matrix, the initial feature value, and the state matrix corresponding to the previous initial feature value, and the state matrix includes K feature values; wherein, the K feature values in the state matrix corresponding to the last initial feature value are used as the K target feature values for that spatial location; The detection module is used to generate a second feature sequence based on K target feature values at each spatial location. The second feature sequence includes K second feature matrices, and the second feature matrices include the target feature values at each spatial location. Based on the second feature sequence, the module determines the detection result of a specified type of item of the object to be detected.
[0006] This application provides an electronic device, including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the article detection method of the above example of this application.
[0007] This application provides a computer program product, which includes a computer program that, when executed by a processor, implements the item detection method described above in this application.
[0008] This application provides a machine-readable storage medium storing machine-executable instructions that can be executed by a processor; wherein the processor is configured to execute the machine-executable instructions to implement the article detection method of the example above in this application.
[0009] As can be seen from the above technical solutions, in this embodiment, after generating the first feature sequence based on the holographic data to be detected, the detection result of a specified type of item is not directly determined based on the first feature sequence. Instead, the first feature sequence is compressed and its features are refined, converting the C first feature matrices of the first feature sequence into K second feature matrices of the second feature sequence, where K is less than C. This maintains the effectiveness of the holographic data, increases data utilization, and reduces the dimensionality of the holographic data, avoiding dense holographic data and noise interference. When determining the detection result of a specified type of item based on the second feature sequence, accurate detection results can be obtained. By compressing and refining the first feature sequence, storing large amounts of holographic data is avoided, saving storage resources. While enhancing target resolution, this effectively alleviates the pressure on holographic data storage and training resource overhead, significantly reducing storage resource consumption and training costs. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating an article detection method according to one embodiment of this application; Figure 2 This is a 2D projection image corresponding to millimeter-wave holographic data in one embodiment of this application; Figure 3 This is a schematic diagram of the preprocessing process for holographic data in one embodiment of this application; Figure 4 This is a schematic diagram illustrating the construction of input data in one embodiment of this application; Figure 5 This is a schematic diagram of the target detection network in one embodiment of this application; Figure 6 This is a schematic diagram of contrastive learning based on positive sample regions in one embodiment of this application; Figure 7 This is a schematic diagram of a multi-stage training method in one embodiment of this application; Figure 8 This is a schematic diagram of the structure of an article detection device according to one embodiment of this application; Figure 9 This is a hardware structure diagram of an electronic device according to one embodiment of this application. Detailed Implementation
[0011] This application proposes an article detection method, which can be applied to electronic devices, such as security inspection devices (e.g., security scanners), personal computers, laptops, and smart terminals. See also... Figure 1 The diagram shown is a flowchart of the method, which may include: Step 101: Generate a first feature sequence based on the holographic data of the object to be detected. The first feature sequence includes C first feature matrices, and each first feature matrix includes initial feature values for multiple spatial locations.
[0012] Step 102: For any spatial location, based on the initial eigenvalues corresponding to that spatial location in C first feature matrices, determine K target eigenvalues for that spatial location, where K is less than C. Specifically, for any initial eigenvalue corresponding to that spatial location, based on the acquired parameter metric matrix, the initial eigenvalue, and the state matrix corresponding to the previous initial eigenvalue, determine the state matrix corresponding to that initial eigenvalue, and this state matrix includes K eigenvalues; wherein the K eigenvalues in the state matrix corresponding to the last initial eigenvalue are taken as the K target eigenvalues for that spatial location.
[0013] Step 103: Generate a second feature sequence based on K target feature values at each spatial location. The second feature sequence includes K second feature matrices, which contain the target feature values at each spatial location.
[0014] Step 104: Determine the detection result of the specified type of item of the object to be detected based on the second feature sequence.
[0015] For example, generating a first feature sequence based on the holographic data of the object to be detected may include, but is not limited to, generating a third feature sequence based on the holographic data to be detected. The third feature sequence includes C first feature matrices, and the holographic data to be detected may include, but is not limited to, at least one of amplitude, phase, and distance. The C first feature matrices in the third feature sequence are sorted in ascending order of the number of candidates for each first feature matrix to obtain a sorted first feature sequence. Specifically, for any first feature matrix, based on the initial feature values of each spatial location within the first feature matrix, the total number of spatial locations whose initial feature values are greater than a preset threshold is determined, and this number of candidates is determined based on this total number.
[0016] For example, the parameter metric matrix includes a first metric matrix and a second metric matrix. Based on the acquired parameter metric matrix, the initial eigenvalue, and the state matrix corresponding to the previous initial eigenvalue, the state matrix corresponding to the initial eigenvalue is determined, which may include, but is not limited to, using the following formula to determine the state matrix corresponding to the initial eigenvalue: ;in, This can represent the state matrix corresponding to the initial eigenvalue. It can represent the first metric matrix. It can represent the second metric matrix. It can represent the state matrix corresponding to the previous initial eigenvalue. The first metric matrix can represent the initial eigenvalue; it can be a K×K dimension matrix. For the i-th row of the first metric matrix, if the sort index of the initial eigenvalue is less than i, then all elements in the i-th row are 0. If the sort index is equal to i, then the elements in the i-th row are determined based on the sort index. If the sort index is greater than i, then the elements in the i-th row are determined based on the sort index and the row number of the i-th row. The second metric matrix is a K×1 dimension matrix, and it is determined based on the sort index.
[0017] For example, determining the detection result of a specified type of item of the target object based on the second feature sequence may include, but is not limited to: inputting the second feature sequence into a trained target detection network, and having the target detection network output the detection result of the specified type of item of the target object based on the second feature sequence, wherein the detection result of the specified type of item includes the predicted category and predicted location of the specified type of item. Alternatively, inputting a fourth feature sequence into a target detection network, and having the target detection network output the detection result of the specified type of item of the target object based on the fourth feature sequence; wherein the fourth feature sequence includes the second feature sequence, a two-dimensional projection feature matrix, and a spatial index feature matrix; the two-dimensional projection feature matrix includes two-dimensional projection values of multiple spatial locations, and the spatial index feature matrix includes spatial index values of multiple spatial locations; wherein, for any spatial location, the two-dimensional projection value of that spatial location is the maximum initial feature value corresponding to that spatial location in C first feature matrices, and the spatial index value of that spatial location is the sorting number corresponding to the maximum initial feature value.
[0018] For example, an object detection network may include a backbone network and a detection head network. The training process of the object detection network may include, but is not limited to: generating a sample feature sequence based on the sample holographic data of the sample object, wherein the sample feature sequence is a second feature sequence or a fourth feature sequence corresponding to the sample holographic data; inputting the sample feature sequence into the backbone network of the initial detection network to obtain an intermediate feature matrix; inputting the intermediate feature matrix into the detection head network of the initial detection network to obtain a predicted category and a predicted location; determining a first loss value based on the predicted category and a second loss value based on the predicted location; inputting the intermediate feature matrix and the calibrated feature matrix into a configured contrastive learning classification network to obtain the target spatial locations corresponding to at least two categories of specified type items in the intermediate feature matrix; determining a third loss value based on the feature values of the target spatial locations; wherein the calibrated feature matrix includes the calibrated spatial locations corresponding to at least two categories of specified type items; determining a target loss value based on the first loss value, the second loss value, and the third loss value; and adjusting the network parameters of the initial detection network based on the target loss value to obtain the object detection network.
[0019] For example, determining the third loss value based on the feature values of the target spatial location may include, but is not limited to: determining the contrastive learning loss for each category based on the feature values of the target spatial location, and determining the third loss value based on the contrastive learning loss for each category. Specifically, for any category, the contrastive learning loss for that category is determined using the following formula: ;in, This represents the contrastive learning loss for category i. The feature value representing the target spatial location of a specified type of item of category i in the intermediate feature matrix. The feature value represents the target spatial location of a specified type of item of category j in the intermediate feature matrix. Category j is different from category i, and the value of j ranges from 1 to N, where N represents the total number of categories other than category i; where, Represents cosine similarity. This represents the learnable temperature coefficient.
[0020] For example, the object detection network may include a backbone network and a detection head network. The detection head network may include a location prediction network and a category prediction network. The training process of the object detection network may include, but is not limited to: generating a sample feature sequence based on the sample holographic data of the sample object, wherein the sample feature sequence is a second feature sequence or a fourth feature sequence corresponding to the sample holographic data; inputting the sample feature sequence into the backbone network of the first initial detection network to obtain a first intermediate feature matrix; inputting the first intermediate feature matrix into the location prediction network of the first initial detection network to obtain a first predicted position; determining a fourth loss value based on the first predicted position; and adjusting the network parameters of the backbone network and the location prediction network based on the fourth loss value to obtain a second initial detection network. The sample feature sequence is input into the backbone network of the second initial detection network to obtain the second intermediate feature matrix. The second intermediate feature matrix is then input into the category prediction network of the second initial detection network to obtain the first predicted category. Based on the first predicted category, a fifth loss value is determined. The second intermediate feature matrix is then input into the contrastive learning classification network to obtain the first target spatial location corresponding to the specified type of item in the second intermediate feature matrix. Based on the feature value of the first target spatial location, a sixth loss value is determined. Based on the fifth and sixth loss values, the network parameters of the category prediction network are adjusted to obtain the third initial detection network. The sample feature sequence is then input into the backbone network of the third initial detection network to obtain the third intermediate feature matrix. The third intermediate feature matrix is then input into the detection head network of the third initial detection network to obtain the second predicted location and the second predicted category. Based on the second predicted location, a seventh loss value is determined, and based on the second predicted category, an eighth loss value is determined. The third intermediate feature matrix is then input into the contrastive learning classification network to obtain the second target spatial location corresponding to the specified type of item in the third intermediate feature matrix. Based on the feature value of the second target spatial location, a ninth loss value is determined. Based on the seventh, eighth, and ninth loss values, the network parameters of the third initial detection network are adjusted to obtain the target detection network.
[0021] As can be seen from the above technical solutions, in this embodiment, after generating the first feature sequence based on the holographic data to be detected, the detection result of a specified type of item is not directly determined based on the first feature sequence. Instead, the first feature sequence is compressed and its features are refined, converting the C first feature matrices of the first feature sequence into K second feature matrices of the second feature sequence, where K is less than C. This maintains the effectiveness of the holographic data, increases data utilization, and reduces the dimensionality of the holographic data, avoiding dense holographic data and noise interference. When determining the detection result of a specified type of item based on the second feature sequence, accurate detection results can be obtained. By compressing and refining the first feature sequence, storing large amounts of holographic data is avoided, saving storage resources. While enhancing target resolution, this effectively alleviates the pressure on holographic data storage and training resource overhead, significantly reducing storage resource consumption and training costs.
[0022] The technical solutions described above in the embodiments of this application will be explained below in conjunction with specific application scenarios.
[0023] Security inspection equipment (such as security scanners) can actively emit millimeter-wave signals and acquire echo signals. By extracting key wavebands from the echo signals, millimeter-wave holographic data can be reconstructed. This data can then be used to detect whether an object is carrying a specified type of item and to identify the category and location of that item. However, when performing detection based on millimeter-wave holographic data, the high density of the data leads to significant noise interference, resulting in inaccurate detection results. Furthermore, the high density of millimeter-wave holographic data necessitates the storage of large amounts of data, consuming substantial storage resources.
[0024] It can also generate 2D projection maps based on millimeter-wave holographic data, and use these 2D projection maps to detect whether an object is carrying a specified type of item, as well as the category and location of the specified item. See also Figure 2 As shown, this is a 2D projection image corresponding to millimeter-wave holographic data. However, projection operations introduce information loss, compounded by the low signal-to-noise ratio and blurred features inherent in millimeter-wave holographic data. Furthermore, 2D projection images lack color and texture information, exhibiting inherent defects such as low feature discrimination, making it difficult to effectively capture the fine structure and material properties of a specified type of object. Therefore, detection based on 2D projection images is poor and cannot yield accurate detection results.
[0025] To address the aforementioned findings, this application proposes an item detection method capable of detecting and classifying specified item types based on holographic data (such as millimeter-wave holographic data). The specified item type can be any pre-planned item, such as metal or non-metal, or even contraband, without restriction. By compressing and extracting features from the holographic data, the method enhances target resolution while effectively alleviating the pressure on holographic data storage and training resource consumption. Adaptive contrastive learning supervision of the holographic data effectively improves the model's classification ability. A multi-stage progressive training method gradually enhances the model's fine-grained classification capability, effectively accelerating model convergence and improving its resolution and generalization performance.
[0026] This embodiment involves a preprocessing process for holographic data, an inference process based on the preprocessed data, an adaptive contrastive learning classification process, and a multi-stage training process. The following describes this process.
[0027] First, the preprocessing of holographic data. In the preprocessing of holographic data, holographic data can be reconstructed based on echo signals. Pixel-level information compression and feature extraction are performed on the holographic data in the range direction to obtain a second feature sequence (as input to the first channel). The second feature sequence, the two-dimensional projection feature matrix (as input to the second channel), and the spatial index feature matrix (as input to the third channel) are concatenated to obtain a fourth feature sequence, enhancing the collaborative representation capability of multi-dimensional features. Then, the fourth feature sequence is input to the object detection network. The object detection network is a large model based on Transformer. With its powerful global feature modeling capability, the object detection network effectively extracts multi-level features of the target.
[0028] See Figure 3 The diagram shown illustrates the preprocessing steps for holographic data, which may include: Step 301: Obtain the holographic data of the object to be detected.
[0029] For example, security inspection equipment (such as security scanners) can emit millimeter-wave signals towards the object to be inspected (such as the human body to be inspected) and acquire the echo signals. Key wavebands in the echo signals can be extracted to reconstruct the holographic data to be inspected. This holographic data can be millimeter-wave holographic data or other types of holographic data. The holographic data to be inspected can include, but is not limited to, at least one of amplitude (observed amplitude), phase, and distance, such as holographic data including amplitude, phase, and distance information.
[0030] Step 302: Generate a third feature sequence based on the holographic data to be detected. This third feature sequence includes C first feature matrices, where C is a positive integer greater than 1. The first feature matrices can be feature maps. For any given first feature matrix, it can include initial feature values for multiple spatial locations.
[0031] For example, see Figure 4 The diagram illustrates the construction of input data corresponding to the holographic data to be detected. C×H×W represents the third feature sequence, where C represents the number of channels in the third feature sequence, meaning the third feature sequence includes C first feature matrices. H×W represents the number of spatial positions in the first feature matrix, meaning the first feature matrix includes H×W spatial positions. For each spatial position, the first feature matrix includes the initial feature value for that spatial position, meaning the first feature matrix includes the initial feature values for H×W spatial positions.
[0032] For example, when generating a third feature sequence based on the holographic data to be detected, the number of channels and spatial locations can be determined based on holographic data such as phase and distance, resulting in C first feature matrices, each containing H×W spatial locations. The initial feature value for each spatial location can be determined based on the amplitude (observed amplitude), which can be either an amplitude value or an energy value.
[0033] In summary, a third feature sequence can be generated based on the holographic data to be detected. The third feature sequence includes C first feature matrices, and each first feature matrix includes H×W initial feature values at spatial locations.
[0034] Step 303: Based on the order of the number of candidates of each first feature matrix from smallest to largest, sort the C first feature matrices in the third feature sequence to obtain the sorted first feature sequence. The first feature sequence includes C first feature matrices, and the first feature matrix includes initial feature values of multiple spatial locations.
[0035] For example, the third feature sequence includes C first feature matrices. For any first feature matrix, based on the initial feature values of each spatial position within the first feature matrix, the total number of spatial positions with initial feature values greater than a preset threshold can be determined, and the candidate number of the first feature matrix can be determined based on the total number. For example, the candidate number can be equal to the total number, the candidate number can be equal to the sum of the total number and the first value, or the candidate number can be equal to the difference between the total number and the second value. There are no restrictions on this.
[0036] For example, for any first feature matrix, which includes H×W initial feature values at spatial locations, if 100 initial feature values are greater than a preset threshold, then the candidate number of the first feature matrix is 100. This process continues until the candidate number of each first feature matrix is obtained. Based on this, the C first feature matrices in the third feature sequence are sorted according to their candidate number in ascending order to obtain the first feature sequence, which includes C first feature matrices.
[0037] See Figure 4 As shown, t×H×W represents the first feature sequence, which still includes C first feature matrices. t represents the time step, which is a different channel. Here, t indicates that the C first feature matrices are sorted in order, and the first feature matrix includes the initial feature values of H×W spatial locations.
[0038] In the above process, a value function is constructed for the first feature matrix of each channel. The C first feature matrices are then sorted according to the value function to form a pseudo-time series (i.e., the first feature sequence), with channels having higher values appearing later in the time sequence. For example, the value function can be described as the number of candidates relative to a preset threshold in a certain channel (i.e., the total number of spatial locations where the initial feature value is greater than the preset threshold).
[0039] Step 304: For any spatial location, determine the time feature sequence corresponding to that spatial location. The time feature sequence includes the initial feature value corresponding to that spatial location in C first feature matrices.
[0040] See Figure 4 As shown, taking the upper right corner as an example, This represents the temporal feature sequence corresponding to this spatial location. This represents the t-th initial feature value in the time feature sequence. For example, the time feature sequence sequentially includes the initial feature value corresponding to the spatial location in the first feature matrix. The initial eigenvalue corresponding to this spatial location in the second first characteristic matrix. ..., the initial eigenvalue corresponding to this spatial location in the Cth first characteristic matrix. That is, there are a total of C initial eigenvalues.
[0041] Step 305: For any spatial location, based on the initial feature values corresponding to the spatial location in the C first feature matrices, determine the K target feature values of the spatial location, where K is less than C.
[0042] For example, for any spatial location, based on the corresponding time feature sequence, K target feature values for that spatial location can be determined. For instance, at any given time... Based on historical signals that have been observed and the current input Find a best approximation (under some metric). This allows for a concise summary of the entire historical information. Based on this, for any initial eigenvalue corresponding to that spatial location (i.e., any initial eigenvalue within the time feature sequence), the state matrix corresponding to that initial eigenvalue can be determined using the acquired parameter metric matrix, the initial eigenvalue itself, and the state matrix corresponding to the previous initial eigenvalue. This state matrix includes K eigenvalues. For example, the parameter metric matrix can include a first metric matrix and a second metric matrix. The first metric matrix can serve as a weighting coefficient for the state matrix corresponding to the previous initial eigenvalue, and the second metric matrix can serve as a weighting coefficient for the initial eigenvalue. Thus, based on the state matrix corresponding to the previous initial eigenvalue, the first metric matrix, the initial eigenvalue, and the second metric matrix, the state matrix corresponding to the initial eigenvalue is determined.
[0043] For example, the state matrix corresponding to the initial eigenvalue can be determined using the following formula (1): Formula (1) In formula (1), This can represent the state matrix corresponding to the initial eigenvalue. It can represent the first metric matrix. It can represent the second metric matrix. It can represent the state matrix corresponding to the previous initial eigenvalue. This can represent the initial feature value. For example, the initial feature value within a time feature sequence. , It can be a default state matrix, such as K eigenvalues all being 1. Represents the initial eigenvalues The corresponding state matrix. Initial eigenvalues within the time feature sequence. , can be the initial eigenvalue The corresponding state matrix, Represents the initial eigenvalues The corresponding state matrix. And so on, for the initial eigenvalues within the time feature sequence. , can be the initial eigenvalue The corresponding state matrix, Represents the initial eigenvalues The corresponding state matrix. It is important to note that the state matrix... It is the last initial eigenvalue The corresponding state matrix, and the state matrix Includes K eigenvalues, state matrix The K feature values are selected as the K target feature values for this spatial location. Thus, the K target feature values for this spatial location are obtained.
[0044] In summary, after completing the time sorting, each spatial location is regarded as an independent "time series", that is, a time feature series. In this way, the time feature series can be compressed. The core parameter measures are the first metric matrix A and the second metric matrix B, which yields K target feature values for that spatial location.
[0045] For example, the first metric matrix A can be a K×K dimension matrix, where K can be configured according to actual needs, such as K being 3, 4, etc. For the i-th row of the first metric matrix A, where i ranges from 1 to K (e.g., from 1 to 4), if the sort index (from 0 to C-1) of the initial eigenvalue is less than i, then all elements in the i-th row (e.g., the K elements in the i-th row) are 0; if the sort index is equal to i, then the elements in the i-th row are determined based on that sort index, such as all elements in the i-th row (e.g., the K elements in the i-th row) are 0. , where n represents the sorting index; if the sorting index is greater than i, then the elements in the i-th row (such as the K elements in the i-th row) are determined based on the sorting index and the row number of the i-th row, such as all elements in the i-th row being... .
[0046] For example, for the initial feature values within a time feature sequence The sorting index is 0, and the elements in rows 1, 2, 3, and 4 of the first metric matrix A are 0. This is for the initial eigenvalues within the time feature sequence. The sorting index is 1, and the elements in the first row of the first metric matrix A are determined based on this sorting index, such as... The elements in rows 2, 3, and 4 of the first metric matrix A are 0. This is for the initial eigenvalues within the time feature sequence. The sorting index is 2. The elements of the first row of the first metric matrix A are determined based on this sorting index and the row number of the i-th row. The elements in the second row of the first metric matrix A are determined based on this sorting index, such as... The elements in the 3rd and 4th rows of the first metric matrix A are 0, and so on.
[0047] For example, the second metric matrix B can be a K×1 dimension matrix, and the second metric matrix is determined based on the sorting index of the initial eigenvalues. For instance, if the K elements of the second metric matrix are all... .
[0048] As can be seen from formula (1), the first metric matrix A is a K×K dimension matrix, the second metric matrix B is a K×1 dimension matrix, and the state matrix... It is a K×1 dimension matrix with initial eigenvalues The eigenvalues are 1×1 dimensional; therefore, the state matrix... It is a K×1 dimension matrix. Obviously, the state matrix corresponding to the last initial eigenvalue is also a K×1 dimension matrix, that is, the state matrix corresponding to the last initial eigenvalue includes K eigenvalues, which serve as the K target eigenvalues for spatial location.
[0049] Step 306: Generate a second feature sequence based on the K target feature values of each spatial location. The second feature sequence includes K second feature matrices, which include the target feature values of each spatial location.
[0050] For example, the first target feature value based on all spatial locations can form a second feature matrix 1, which includes target feature values of H×W spatial locations; the second target feature value based on all spatial locations can form a second feature matrix 2, which includes target feature values of H×W spatial locations; and so on, the Kth target feature value based on all spatial locations can form a second feature matrix K, which includes target feature values of H×W spatial locations.
[0051] In summary, a second feature sequence can be obtained, which may include K second feature matrices. This transforms the C first feature matrices of the first feature sequence into the K second feature matrices of the second feature sequence, such as transforming 128 first feature matrices into 4 second feature matrices, ultimately resulting in a second feature sequence of H×W×K. The second feature sequence can be a feature map, such as a state feature map.
[0052] Step 307: Obtain the two-dimensional projection feature matrix and the spatial index feature matrix based on the third feature sequence or the first feature sequence. For example, the two-dimensional projection feature matrix may include two-dimensional projection values of multiple spatial locations, and the spatial index feature matrix may include spatial index values of multiple spatial locations.
[0053] For example, the third feature sequence (or the first feature sequence) may include C first feature matrices. For any spatial location, based on the initial feature values corresponding to that spatial location in the C first feature matrices, the largest initial feature value among the C initial feature values can be determined, and the two-dimensional projection value of that spatial location is the largest initial feature value corresponding to that spatial location. Based on this, the largest initial feature values corresponding to all spatial locations are combined to form a two-dimensional projection feature matrix. That is, the two-dimensional projection feature matrix may include H×W largest initial feature values corresponding to spatial locations, and the two-dimensional projection feature matrix can be a 2D projection map.
[0054] For example, the third feature sequence (or the first feature sequence) may include C first feature matrices. For any spatial location, the spatial index value of that location is the sorting number corresponding to the largest initial feature value. For instance, if the largest initial feature value is in the 6th first feature matrix, i.e., the sorting number is 5, then the spatial index value of that spatial location is 5, and so on. Based on this, the spatial index values corresponding to all spatial locations are combined to form a spatial index feature matrix. That is, the spatial index feature matrix may include H×W spatial index values corresponding to spatial locations, and the spatial index feature matrix can be a spatial index feature map.
[0055] Step 308: Obtain the fourth feature sequence. The fourth feature sequence may include the second feature sequence, the two-dimensional projection feature matrix, and the spatial index feature matrix. That is, the fourth feature sequence has a dimension of (K+2)×H×W.
[0056] exist Figure 4 middle, 2D The projection diagram represents the two-dimensional projection feature matrix, and Depth represents the spatial index feature matrix. The two-dimensional projection feature matrix, the spatial index feature matrix, and the second feature sequence are combined to obtain the fourth feature sequence. (For...) 2D The projection image can be a BEV (Birds Eye View).
[0057] In summary, the second feature sequence can be concatenated with the two-dimensional projection feature matrix and the spatial index feature matrix to form the model input data. This requires storing only (K+2) images (feature maps), significantly reducing storage requirements and training costs compared to the original holographic data, such as C images (feature maps). Simultaneously, it allows for information compression and extraction from the complete holographic data, increasing data utilization.
[0058] Second, regarding the inference process based on preprocessed data. In the inference process based on preprocessed data, the detection result of a specific type of item of the object to be detected can be determined based on the second feature sequence. For example, the second feature sequence is input into the target detection network, and the target detection network outputs the detection result of the specific type of item of the object to be detected based on the second feature sequence. The detection result of a specific type of item of the object to be detected can also be determined based on the fourth feature sequence. For example, the fourth feature sequence is input into the target detection network, and the target detection network outputs the detection result of the specific type of item of the object to be detected based on the fourth feature sequence.
[0059] For example, the detection result for a specified type of item includes information on whether the object to be detected is carrying a specified type of item. If the object to be detected is carrying a specified type of item, the detection result for a specified type of item may also include the predicted category of the specified type of item (indicating what category the specified type of item is) and the predicted location.
[0060] For example, the object detection network can be a large model based on the Transformer architecture, or a millimeter-wave holographic large model network using the Transformer architecture. In this way, after the fourth feature sequence is input into the object detection network, the powerful global feature modeling capability of the Transformer can effectively extract the multi-level features of the object to be detected, and thus obtain accurate detection results of the specified type of item.
[0061] For example, before inputting the fourth feature sequence into the object detection network, the fourth feature sequence can be augmented, such as by using random cropping and scaling, or CutMix. Then, the augmented fourth feature sequence can be input into the object detection network.
[0062] For example, an object detection network may include a backbone network and a detection head network. The backbone network may include a deep self-attention encoder, a downsampling network, and a feature fusion network. See [link to relevant documentation]. Figure 5 The diagram shown illustrates the structure of an object detection network. This embodiment does not impose restrictions on the structure of this network. In the object detection network, a Feature Pyramid Network (FPN) structure is introduced into the feature fusion network. This significantly enhances the semantic representation capability of small-scale targets through a multi-scale feature fusion strategy. Finally, the network connects to the detection head network (two-dimensional detection head) to achieve accurate localization and classification of items of a specified type.
[0063] For example, assuming the object detection network outputs features at N scales, where N is a positive integer, then the feature fusion network includes N fusion subnetworks. Figure 5 Taking three fusion sub-networks as an example), the detection head network includes N two-dimensional detection heads (2DHead, ...). Figure 5 (Taking three 2D detection heads as an example). Furthermore, the target detection network includes N self-attention sub-networks (taking three self-attention sub-networks as an example), each of which includes M deep self-attention encoders and a downsampling network. For the deep self-attention encoder, it can sequentially include a deep self-attention network layer, a residual connection (Add) and a layer normalization (Norm) network layer, a feedforward network layer, a residual connection (Add) and a layer normalization (Norm) network layer. For the downsampling network, it can be a downsampling network implemented using patchmerging, or it can be implemented using other methods.
[0064] Based on the above object detection network, the processing steps of the object detection network include: The fourth feature sequence is input into self-attention sub-network 1 (M deep self-attention encoders and a downsampling network) to obtain intermediate feature a1. For example, the scale of intermediate feature a1 is 256×256. Intermediate feature a1 is input into self-attention sub-network 2 (M deep self-attention encoders and a downsampling network) to obtain intermediate feature a2. The scale of intermediate feature a2 is smaller than that of intermediate feature a1. For example, the scale of intermediate feature a2 is 128×128. Intermediate feature a2 is input into self-attention sub-network 3 (M deep self-attention encoders and a downsampling network) to obtain intermediate feature a3. The scale of intermediate feature a3 is smaller than that of intermediate feature a2. For example, the scale of intermediate feature a3 is 64×64.
[0065] The intermediate feature a3 is input into the fusion subnetwork 1 to obtain the intermediate feature a3-1 with a scale of 64×64.
[0066] The intermediate feature a2 is input into the fusion sub-network 2 to obtain the intermediate feature a2-1 with a scale of 128×128. The intermediate feature a3-1 is upsampled to obtain the feature with a scale of 128×128, and then it is operated with the intermediate feature a2-1 (such as addition) to obtain the intermediate feature a2-2 with a scale of 128×128.
[0067] The intermediate feature a1 is input into the fusion sub-network 3 to obtain the intermediate feature a1-1 with a scale of 256×256. The intermediate feature a2-2 is upsampled to obtain a feature with a scale of 256×256, and then it is operated with the intermediate feature a1-1 (such as addition) to obtain the intermediate feature a1-2 with a scale of 256×256.
[0068] In summary, intermediate features at multiple scales can be obtained, and the detection result for a specified type of item can then be determined based on these intermediate features. For example, an intermediate feature a3-1 with a scale of 64×64 can be input into 2D detector head 1, an intermediate feature a2-2 with a scale of 128×128 can be input into 2D detector head 2, and an intermediate feature a1-2 with a scale of 256×256 can be input into 2D detector head 3. By combining the output features from 2D detector head 1, 2D detector head 2, and 3, the detection result for the specified type of item can be obtained.
[0069] Third, for the adaptive contrastive learning classification process, an adaptive contrastive learning strategy can be applied to the representation of positive sample regions. An adaptive contrastive learning mechanism based on positive samples is introduced. By increasing the distance between samples of different categories in the representation space, the inter-class discrimination ability of the model is improved, and the classification performance is further enhanced.
[0070] For example, due to the lack of color and texture information dimensions in millimeter waves, different categories of items exhibit highly overlapping distribution patterns in the feature space, significantly reducing feature separability. Therefore, in this embodiment, a contrastive learning method based on positive sample sampling regions improves classification accuracy by increasing the similarity between representations of different categories. Furthermore, to address the issue of discretization of intra-class feature distribution caused by pose changes in millimeter-wave imaging, a processing mechanism more adapted to the characteristics of millimeter-wave data is designed. This mechanism introduces a dynamically weighted region representation contrast loss to improve the model's robustness to pose disturbances.
[0071] Based on this, in this embodiment, the target detection network can be trained using the following steps: Step S11: Obtain the holographic data of the sample object.
[0072] Step S12: Generate a sample feature sequence based on the sample holographic data of the sample object. The sample feature sequence can be the second feature sequence or the fourth feature sequence corresponding to the sample holographic data.
[0073] For example, the holographic data during the detection process can be referred to as the holographic data to be detected, and the holographic data during the training process can be referred to as the sample holographic data. During the training process, sample feature sequences can be generated based on the sample holographic data. The generation method can be found in steps 302-308, and will not be repeated here.
[0074] Step S13: Obtain the initial detection network to be trained. The initial detection network may include a backbone network and a detection head network. The backbone network may include a deep self-attention encoder, a downsampling network, and a feature fusion network, etc. For example, see... Figure 5 The diagram shown is a schematic of the initial detection network.
[0075] Step S14: Input the sample feature sequence into the backbone network of the initial detection network to obtain the intermediate feature matrix, and input the intermediate feature matrix into the detection head network of the initial detection network to obtain the predicted category and predicted position; determine the first loss value based on the predicted category, and determine the second loss value based on the predicted position.
[0076] For example, after inputting the sample feature sequence into the initial detection network, the initial detection network can output detection results for a specified type of item. These results can include the predicted category and predicted location. Based on the predicted category and the labeled category (i.e., the category of the specified type of item is pre-labeled in the sample feature sequence), a first loss value can be calculated, such as using a loss function like cross-entropy. Based on the predicted location and the labeled location (i.e., the location of the specified type of item is pre-labeled in the sample feature sequence), a second loss value can be calculated, such as using a loss function like cross-entropy.
[0077] Step S15: Input the intermediate feature matrix and the labeled feature matrix into the contrastive learning classification network to obtain the target spatial positions of at least two categories of specified type items in the intermediate feature matrix. The labeled feature matrix includes the labeled spatial positions of at least two categories of specified type items.
[0078] For example, after inputting the sample feature sequence into the backbone network (deep self-attention encoder, downsampling network, and feature fusion network), the intermediate feature matrix can include intermediate features at multiple scales, such as intermediate feature a3-1 with a scale of 64×64, intermediate feature a2-2 with a scale of 128×128, and intermediate feature a1-2 with a scale of 256×256. Furthermore, the labeled feature matrix can be a sample feature sequence (i.e., multiple feature matrices within the sample feature sequence) or a feature matrix constructed based on the sample feature sequence. There are no restrictions on the labeled feature matrix, as long as it corresponds to multiple categories of specified type items. Based on this, for multiple categories of specified type items, the labeled spatial positions corresponding to at least two categories of specified type items are pre-labeled in the labeled feature matrix, such as labeling the labeled spatial position corresponding to the specified type item in category 1, the labeled spatial position corresponding to the specified type item in category 2, and so on.
[0079] During pre-calibration, it is necessary to calibrate the spatial locations of different categories. That is, for the same category, only one spatial location is calibrated, rather than multiple spatial locations. In this way, if there are P spatial locations, then the P spatial locations correspond to P categories. In subsequent processing, it is not necessary to know which category each spatial location corresponds to; it is sufficient that different spatial locations correspond to different categories.
[0080] For example, after inputting the intermediate feature matrix and the labeled feature matrix into the contrastive learning classification network, based on the labeled spatial locations in the labeled feature matrix, the contrastive learning classification network can find the target spatial locations corresponding to the labeled spatial locations from the intermediate feature matrix. For instance, for P labeled spatial locations, the contrastive learning classification network can find P target spatial locations from the intermediate feature matrix. For example, since P labeled spatial locations correspond to P categories, the P target spatial locations correspond to P categories, where P is a positive integer greater than 1, meaning different target spatial locations correspond to different categories.
[0081] Considering that the intermediate feature matrix can include intermediate features at multiple scales, the contrastive learning classification network can find P target spatial locations from the intermediate features for each scale.
[0082] Step S16: Determine the third loss value based on the feature value of the target spatial location.
[0083] For example, for any category, the contrastive learning loss for that category is determined based on the feature values of the target spatial location, and a third loss value is determined based on the contrastive learning losses of all categories. For instance, the sum of the contrastive learning losses of all categories can be used as the third loss value, or a weighted average of the contrastive learning losses of all categories can be performed to obtain the third loss value. There are no restrictions on how the third loss value is determined.
[0084] For example, considering that the intermediate feature matrix may include intermediate features at multiple scales, a third loss value for each intermediate feature can be determined based on the feature value of the target spatial location. Based on this, the sum of the third loss values of all intermediate features can be used as the final output third loss value; alternatively, a weighted operation can be performed on the third loss values of all intermediate features, and the weighted loss value can be used as the final output third loss value.
[0085] For example, suppose a contrastive learning classification network finds P target spatial locations from an intermediate feature matrix. These P target spatial locations correspond to P categories. It's not necessary to know the actual categories of these P categories; simply label them, such as category 1, category 2, ..., category P. For any category i, the contrastive learning loss for category i can be determined based on the cosine similarity between the feature values of the target spatial locations corresponding to category i and those corresponding to category j. This contrastive learning loss is negatively correlated with the cosine similarity. For instance, category j could be any category other than category i.
[0086] For example, for any class i, the contrastive learning loss for class i can be determined using the following formula: ;in, This can represent the contrastive learning loss for category i. The feature value representing the target spatial location of a specified type of item of category i in the intermediate feature matrix. This represents the feature value of the target spatial location of a specified type of item in category j within the intermediate feature matrix. Category j differs from category i, and its value ranges from 1 to N, where N can represent the total number of categories other than category i, such as N = P - 1. Represents cosine similarity. This represents the learnable temperature coefficient. Eigenvalues and eigenvalues These represent the feature embeddings from positive sample sampling regions of different categories.
[0087] For example, when i is 1 and j is 2, Let cosine similarity represent the feature values of the target spatial locations corresponding to category 1 and category 2, where i = 1 and j = 3. Let represent the cosine similarity between the feature values of the target spatial location corresponding to category 1 and the feature values of the target spatial location corresponding to category 3. Similarly, the cosine similarity between the feature values of the target spatial location corresponding to category 1 and the feature values of the target spatial locations corresponding to each of the remaining categories can be obtained. Then, the contrastive learning loss for category 1 can be determined using the above formula. Likewise, the contrastive learning loss for category 2 can be obtained, and so on, to obtain the contrastive learning loss for each category.
[0088] The learnable temperature coefficient is used to adaptively adjust the contribution weights of different class samples in the loss function, thereby optimizing the gradient update process and enhancing classification discriminativeness. The learnable temperature coefficients for all classes can be the same, while the learnable temperature coefficients for different classes can be different. Based on this, the above formula can be updated to: Thus, when j is 2, This represents the learnable temperature coefficient for category 2, when j is 3. This represents the learnable temperature coefficient for category 3, and so on.
[0089] For example, a contrastive learning classification network can include learnable temperature coefficients for each category. This network might include a temperature coefficient matrix containing P temperature coefficients, which correspond to the learnable temperature coefficients for each of the P categories. During the initial adjustment of the detection network's parameters, the contrastive learning classification network's parameters also need to be adjusted, specifically the P temperature coefficients. Thus, these P temperature coefficients become the learnable temperature coefficients.
[0090] For example, see Figure 6 The diagram shown illustrates contrastive learning based on positive sample regions, with feature values... and eigenvalues These represent the feature embeddings from the sampling regions of positive samples from different categories, i.e. The feature value representing the target spatial location corresponding to category j. The feature value represents the target spatial location corresponding to category i. Based on this, we can use the feature value... and eigenvalues Cosine similarity and learnable temperature coefficient The contrastive learning loss is determined; it is negatively correlated with cosine similarity; and it is also correlated with the learnable temperature coefficient. Positive correlation, therefore no restrictions are placed on how the learning loss is determined.
[0091] Step S17: Determine the target loss value based on the first loss value, the second loss value, and the third loss value. For example, the target loss value is obtained by performing a weighted operation on the first loss value, the second loss value, and the third loss value, and the sum of the first loss value, the second loss value, and the third loss value is used as the target loss value.
[0092] Step S18: Adjust the network parameters of the initial detection network based on the target loss value to obtain the target detection network. For example, based on the target loss value, the network parameters of the initial detection network can be adjusted using methods such as gradient descent to obtain the adjusted network model. If this adjusted network model does not converge, update the adjusted network model as the initial detection network and repeat the above steps. If this adjusted network model has converged, use this adjusted network model as the target detection network.
[0093] Fourth, regarding the multi-stage training process, the strategy involves first learning a general feature representation on single-class data, then freezing a portion of the network and focusing on optimizing the classification head, and finally fine-tuning based on hard examples selected according to confidence levels. This strategy effectively accelerates model convergence and significantly improves detection resolution and generalization performance. For example, the first stage involves training the entire model for a single class, the second stage involves freezing the backbone and training multi-class heads, and the third stage involves reducing the proportion of loss from low-confidence hard examples and increasing the proportion of loss from important classes.
[0094] For example, millimeter-wave holographic data suffers from low signal-to-noise ratio due to background clutter, resulting in a large amount of redundant noise unrelated to the target information in the original feature space. If global end-to-end learning is used in the early stages of training, the model is prone to overfitting to local imaging artifacts or noise patterns. Furthermore, due to the non-convexity of the loss function and the presence of numerous local minima, the optimization process easily gets trapped in suboptimal solutions, making effective convergence difficult. This phenomenon not only weakens the model's ability to identify key target features but also significantly reduces its generalization performance on unseen data. Based on these characteristics of millimeter-wave data, this embodiment proposes a multi-stage training method, see [link to relevant documentation]. Figure 7 The diagram shown is a schematic of a multi-stage training method.
[0095] In the first stage, training does not further subdivide the categories; instead, all targets are uniformly mapped to a single category, i.e., global feature learning is performed based on single-category training data. The core objective of this stage is to guide the model to effectively distinguish foreground targets from background interference, extract common features that are universally applicable to millimeter-wave imaging, thereby laying a stable detection foundation for subsequent fine-grained classification tasks and improving the model's target detection robustness.
[0096] In the second stage, multi-class classification heads are trained using multi-class training data, while the backbone network parameters are frozen to maintain the general feature extraction capabilities learned in the first stage. The core objective of this stage is to endow the model with fine-grained multi-class recognition capabilities while maintaining robustness in foreground-background discrimination.
[0097] For the third stage, a global fine-tuning strategy is adopted, using a relatively small learning rate to achieve stable convergence. To alleviate the overfitting of the model to low-confidence difficult samples, a dynamic weighted loss mechanism is introduced, appropriately reducing the weight of samples with confidence below a preset threshold in the loss function; and increasing the loss contribution ratio for key category samples to enhance the model's ability to identify and detect key categories, achieving a balance between accuracy and robustness.
[0098] Based on this, in this embodiment, the target detection network can be trained using the following steps: Step S21: Obtain the sample holographic data of the sample object.
[0099] Step S22: Generate a sample feature sequence based on the sample holographic data of the sample object. The sample feature sequence can be the second feature sequence or the fourth feature sequence corresponding to the sample holographic data.
[0100] Step S23: Obtain the first initial detection network to be trained. The first initial detection network may include a backbone network and a detection head network. The backbone network may include a deep self-attention encoder, a downsampling network, and a feature fusion network, etc. For example, see... Figure 5 The diagram shown is a schematic representation of the structure of the first initial detection network. In this embodiment, the detection head network may include a location prediction network and a category prediction network. The location prediction network is used to predict the location of a specified type of item, and the category prediction network is used to predict the category of the specified type of item. In other words, the location prediction network and the category prediction network are separated to predict location and category information respectively.
[0101] Step S24: Input the sample feature sequence into the backbone network of the first initial detection network to obtain the first intermediate feature matrix, input the first intermediate feature matrix into the position prediction network of the first initial detection network to obtain the first predicted position, determine the fourth loss value based on the first predicted position, and adjust the network parameters of the backbone network and the position prediction network based on the fourth loss value to obtain the second initial detection network.
[0102] For example, after inputting the sample feature sequence into the first initial detection network, the location prediction network of the first initial detection network can output the first predicted position. Based on the first predicted position and the calibration position (i.e., the position of a specified type of item is pre-calibrated in the sample feature sequence), a fourth loss value can be calculated. For example, a loss function such as cross-entropy can be used to calculate the fourth loss value. There are no restrictions on how the fourth loss value is determined.
[0103] Based on the fourth loss value, the network parameters of the category prediction network are frozen. Gradient descent and other methods can be used to adjust the network parameters of the backbone network and the location prediction network to obtain the adjusted network model. If this adjusted network model does not converge, it is updated as the first initial detection network, and step S24 is repeated. If this adjusted network model has converged, it is used as the second initial detection network. This completes the first stage of training. During the first stage of training, no category distinction is made; that is, the first initial detection network does not output a predicted category. Therefore, the sample feature sequences can be used as single-class training data without considering the labeled category of the sample feature sequences.
[0104] Step S25: Input the sample feature sequence into the backbone network of the second initial detection network to obtain the second intermediate feature matrix. Input the second intermediate feature matrix into the category prediction network of the second initial detection network to obtain the first predicted category. Determine the fifth loss value based on the first predicted category. Adjust the network parameters of the category prediction network based on the fifth loss value to obtain the third initial detection network.
[0105] Alternatively, after obtaining the second intermediate feature matrix, the second intermediate feature matrix is input into the contrastive learning classification network to obtain the first target spatial location corresponding to the specified type of item in the second intermediate feature matrix. The sixth loss value is determined based on the feature value of the first target spatial location. The network parameters of the category prediction network are adjusted based on the fifth and sixth loss values to obtain the third initial detection network.
[0106] For example, after inputting the sample feature sequence into the second initial detection network, the category prediction network of the second initial detection network can output the first predicted category. Based on the first predicted category and the labeled category (i.e., the category of a specified type of item is pre-labeled in the sample feature sequence), a fifth loss value can be calculated. For example, a loss function such as cross-entropy can be used to calculate the fifth loss value. There are no restrictions on how the fifth loss value is determined.
[0107] For example, after obtaining the second intermediate feature matrix, the second intermediate feature matrix and the labeled feature matrix can be input into the contrastive learning classification network to obtain the first target spatial location corresponding to the specified type of item in the second intermediate feature matrix. Based on the feature values of the first target spatial locations of multiple categories, the sixth loss value can be determined. This process can be referred to in steps S15-S16, and will not be repeated here.
[0108] For example, the target loss value is determined based on the fifth and sixth loss values, such as by weighting the fifth and sixth loss values. Based on the target loss value, the network parameters of the backbone network and the location prediction network are frozen. The network parameters of the category prediction network can be adjusted using methods such as gradient descent to obtain the adjusted network model. If this adjusted network model does not converge, it is updated as the second initial detection network, and step S25 is repeated. If this adjusted network model has converged, it is used as the third initial detection network. This completes the second stage of training. In the second stage of training, it is necessary to distinguish categories, i.e., the second initial detection network outputs the predicted category. This allows for the differentiation of the category of the sample feature sequence. The labeled category of the sample feature sequence needs to be considered; therefore, the sample feature sequence is used as multi-class training data.
[0109] Step S26: Input the sample feature sequence into the backbone network of the third initial detection network to obtain the third intermediate feature matrix. Input the third intermediate feature matrix into the detection head network (position prediction network and category prediction network) of the third initial detection network to obtain the second predicted position and the second predicted category. Determine the seventh loss value based on the second predicted position and the eighth loss value based on the second predicted category. Adjust the network parameters of the third initial detection network based on the seventh and eighth loss values to obtain the target detection network.
[0110] Alternatively, after obtaining the third intermediate feature matrix, the third intermediate feature matrix is input into the contrastive learning classification network to obtain the second target spatial location corresponding to the specified type of item in the third intermediate feature matrix. The ninth loss value is determined based on the feature value of the second target spatial location. The network parameters of the third initial detection network are adjusted based on the seventh, eighth and ninth loss values to obtain the target detection network.
[0111] For example, after inputting the sample feature sequence into the third initial detection network, the third initial detection network can output the second predicted position and the second predicted class. Based on the second predicted position and the labeled position, a seventh loss value can be calculated, such as using a loss function like cross-entropy. There are no restrictions on how this seventh loss value is determined. Based on the second predicted class and the labeled class, an eighth loss value can be calculated, such as using a loss function like cross-entropy. There are no restrictions on how this eighth loss value is determined.
[0112] For example, after obtaining the third intermediate feature matrix, the third intermediate feature matrix and the labeled feature matrix can be input into the contrastive learning classification network to obtain the second target spatial location corresponding to the specified type of item in the third intermediate feature matrix. Based on the feature values of the second target spatial locations of multiple categories, the ninth loss value can be determined. This process can be referred to in steps S15-S16, and will not be repeated here.
[0113] For example, the target loss value is determined based on the seventh, eighth, and ninth loss values, such as by weighting these values. Based on the target loss value, the network parameters of the third initial detection network can be adjusted using methods such as gradient descent to obtain the adjusted network model. If the adjusted network model does not converge, it is updated as the third initial detection network, and step S26 is repeated. If the adjusted network model converges, it is used as the object detection network. This completes the third stage of training.
[0114] For example, the training data may include multiple sample feature sequences. For each sample feature sequence, step S26 can be used to obtain the seventh, eighth, and ninth loss values corresponding to that sample feature sequence. Then, the seventh loss values corresponding to all sample feature sequences can be weighted to obtain the seventh loss value for all sample feature sequences, the eighth loss values can be weighted to obtain the eighth loss value for all sample feature sequences, and the ninth loss values can be weighted to obtain the ninth loss value for all sample feature sequences. Based on this, the target loss value can be determined based on the seventh, eighth, and ninth loss values corresponding to all sample feature sequences. For any sample feature sequence, after inputting the third intermediate feature matrix into the detection head network of the third initial detection network, in addition to outputting the second predicted position and the second predicted class, a confidence score can also be output, representing the confidence level of the second predicted position and the second predicted class. Based on this, if the second predicted class is a configured priority class, the sample feature sequence is used as a boosted sample. If the second predicted class is not a configured priority class, then if the confidence score is not greater than a preset threshold, the sample feature sequence is used as a hard sample; if the confidence score is greater than the preset threshold, the sample feature sequence is used as a regular sample.
[0115] When weighting the seventh loss value corresponding to all sample feature sequences to obtain the seventh loss value, the weighting coefficient of the enhanced sample is greater than that of the regular sample, and the weighting coefficient of the regular sample is greater than that of the difficult sample. Similarly, when weighting the eighth loss value corresponding to all sample feature sequences to obtain the eighth loss value, the weighting coefficient of the enhanced sample is greater than that of the regular sample, and the weighting coefficient of the regular sample is greater than that of the difficult sample. Likewise, when weighting the ninth loss value corresponding to all sample feature sequences to obtain the ninth loss value, the weighting coefficient of the enhanced sample is greater than that of the regular sample, and the weighting coefficient of the regular sample is greater than that of the difficult sample. Based on these processing steps, the overfitting of the model to low-confidence difficult samples can be alleviated, the weight of samples with confidence below a preset threshold can be reduced, and the loss contribution ratio of key category samples can be increased, thus strengthening the model's ability to identify and detect key categories.
[0116] As can be seen from the above technical solutions, in this embodiment, by effectively compressing information and extracting features from complete holographic data, higher information utilization efficiency can be achieved. In the data acquisition stage, only K+2 key images need to be stored, significantly reducing storage resource consumption and overall training costs. By adopting an adaptive contrastive learning strategy, addressing the lack of information dimensions in millimeter-wave imaging data, the similarity measurement scale of negative sample pairs can be adaptively adjusted based on the distribution density and inter-class separability of samples in the feature space. This enhances the model's ability to model inter-class separability, effectively improving the model's classification robustness in complex scenes, while accelerating the convergence stability of the training process. By using a multi-stage training strategy, through a three-stage progressive training of "detection-classification-fine-tuning," the model's robustness in detecting foreground targets is ensured while its fine-grained classification ability is gradually enhanced. Furthermore, a dynamic weighted loss mechanism is used to balance overfitting of difficult examples and reinforcement of key categories, achieving a synergistic improvement in accuracy, robustness, and training efficiency in millimeter-wave target recognition tasks.
[0117] Based on the same concept as the method described above, this application proposes an article detection device, see [link to relevant documentation]. Figure 8 The diagram shown is a structural schematic of the item detection device, which may include: The generation module 81 is used to generate a first feature sequence based on the holographic data of the object to be detected. The first feature sequence includes C first feature matrices, and the first feature matrices include initial feature values for multiple spatial locations. The determination module 82 is used to determine K target feature values for any spatial location based on the initial feature values corresponding to that spatial location in the C first feature matrices, where K is less than C. Specifically, for any initial feature value corresponding to that spatial location, the state matrix corresponding to that initial feature value is determined based on the acquired parameter metric matrix, the initial feature value, and the state matrix corresponding to the previous initial feature value. The state matrix includes K feature values. The K feature values in the state matrix corresponding to the last initial feature value are used as the K target feature values for that spatial location. The detection module 83 is used to generate a second feature sequence based on the K target feature values for each spatial location. The second feature sequence includes K second feature matrices, and the second feature matrices include the target feature values for each spatial location. The detection result of a specified type of item of the object to be detected is determined based on the second feature sequence.
[0118] For example, when the generation module 81 generates a first feature sequence based on the holographic data of the object to be detected, it is specifically used to: generate a third feature sequence based on the holographic data to be detected, the third feature sequence including C first feature matrices, the holographic data to be detected including at least one of amplitude, phase, and distance; sort the C first feature matrices in the third feature sequence according to the order of the number of candidates of each first feature matrix from smallest to largest, to obtain a sorted first feature sequence; wherein, for any first feature matrix, based on the initial feature value of each spatial position in the first feature matrix, the total number of spatial positions with initial feature values greater than a preset threshold is determined, and the number of candidates is determined based on the total number.
[0119] For example, the parameter metric matrix includes a first metric matrix and a second metric matrix; when determining the state matrix corresponding to the initial eigenvalue based on the acquired parameter metric matrix, the initial eigenvalue, and the state matrix corresponding to the previous initial eigenvalue, the determining module 82 specifically uses the following formula to determine the state matrix corresponding to the initial eigenvalue: ; in, This represents the state matrix corresponding to the initial eigenvalue. Represents the first metric matrix. Represents the second metric matrix. This represents the state matrix corresponding to the previous initial eigenvalue. The first metric matrix is a K×K dimension matrix. For the i-th row of the first metric matrix, if the sorting index of the initial eigenvalue is less than i, then all elements in the i-th row are 0. If the sorting index is equal to i, then the elements in the i-th row are determined based on the sorting index. If the sorting index is greater than i, then the elements in the i-th row are determined based on the sorting index and the row number of the i-th row. The second metric matrix is a K×1 dimension matrix, and the second metric matrix is determined based on the sorting index.
[0120] For example, when the detection module 83 determines the detection result of a specified type of item of the object to be detected based on the second feature sequence, it is specifically used to: input the second feature sequence to a trained target detection network, and have the target detection network output the detection result of the specified type of item of the object to be detected based on the second feature sequence, wherein the detection result of the specified type of item includes the predicted category and predicted location of the specified type of item; or, input the fourth feature sequence to the target detection network, and have the target detection network output the detection result of the specified type of item of the object to be detected based on the fourth feature sequence; wherein the fourth feature sequence includes the second feature sequence, a two-dimensional projection feature matrix, and a spatial index feature matrix; the two-dimensional projection feature matrix includes two-dimensional projection values of multiple spatial locations, and the spatial index feature matrix includes spatial index values of multiple spatial locations; wherein, for any spatial location, the two-dimensional projection value of the spatial location is the maximum initial feature value corresponding to the spatial location in C first feature matrices, and the spatial index value of the spatial location is the sorting number corresponding to the maximum initial feature value.
[0121] For example, the target detection network includes a backbone network and a detection head network. The device further includes a training module for training the target detection network. Specifically, the training module trains the target detection network by: generating a sample feature sequence based on the sample holographic data of a sample object, wherein the sample feature sequence is a second feature sequence or a fourth feature sequence corresponding to the sample holographic data; inputting the sample feature sequence to the backbone network of the initial detection network to obtain an intermediate feature matrix; inputting the intermediate feature matrix to the detection head network of the initial detection network to obtain a predicted category and a predicted position; determining a first loss value based on the predicted category; determining a second loss value based on the predicted position; inputting the intermediate feature matrix and the calibrated feature matrix to a configured contrastive learning classification network to obtain the target spatial positions corresponding to at least two categories of specified type items in the intermediate feature matrix; determining a third loss value based on the feature values of the target spatial positions; wherein the calibrated feature matrix includes the calibrated spatial positions corresponding to the at least two categories of specified type items; determining a target loss value based on the first loss value, the second loss value, and the third loss value; and adjusting the network parameters of the initial detection network based on the target loss value to obtain the target detection network.
[0122] For example, when the training module determines the third loss value based on the feature values of the target spatial location, it specifically performs the following steps: determining the contrastive learning loss for each category based on the feature values of the target spatial location, and determining the third loss value based on the contrastive learning loss for each category; wherein, for any category, the training module uses the following formula to determine the contrastive learning loss for that category: ; in, Indicate category i Contrastive learning loss, Indicate category i The feature value of the target spatial location corresponding to the specified type of item in the intermediate feature matrix. Indicate category j The feature value of the target spatial location corresponding to the specified type of item in the intermediate feature matrix, the category j With the category i different, j The value ranges from 1 to N, where N represents the value excluding the stated category. i Total number of categories other than those mentioned above; in, Represents cosine similarity. This represents the learnable temperature coefficient.
[0123] For example, the detection head network includes a location prediction network and a category prediction network; the training module trains the target detection network by: generating a sample feature sequence based on the sample holographic data of the sample object, wherein the sample feature sequence is a second feature sequence or a fourth feature sequence corresponding to the sample holographic data; inputting the sample feature sequence into the backbone network of the first initial detection network to obtain a first intermediate feature matrix, inputting the first intermediate feature matrix into the location prediction network of the first initial detection network to obtain a first predicted position, and determining a fourth loss value based on the first predicted position; adjusting the network parameters of the backbone network and the location prediction network based on the fourth loss value to obtain a second initial detection network; inputting the sample feature sequence into the backbone network of the second initial detection network to obtain a second intermediate feature matrix, inputting the second intermediate feature matrix into the category prediction network of the second initial detection network to obtain a first predicted category, and determining a fifth loss value based on the first predicted category; inputting the second intermediate feature matrix into the target detection network to obtain a second intermediate feature matrix; inputting the second intermediate feature matrix into the target detection network to obtain a second predicted category; and inputting the second intermediate feature matrix into the target detection network to obtain a third intermediate feature matrix. A contrastive learning classification network is used to obtain the first target spatial position corresponding to a specified type of item in the second intermediate feature matrix. A sixth loss value is determined based on the feature values of the first target spatial position. The network parameters of the category prediction network are adjusted based on the fifth and sixth loss values to obtain a third initial detection network. The sample feature sequence is input into the backbone network of the third initial detection network to obtain a third intermediate feature matrix. The third intermediate feature matrix is input into the detection head network of the third initial detection network to obtain a second predicted position and a second predicted category. A seventh loss value is determined based on the second predicted position, and an eighth loss value is determined based on the second predicted category. The third intermediate feature matrix is input into the contrastive learning classification network to obtain the second target spatial position corresponding to the specified type of item in the third intermediate feature matrix. A ninth loss value is determined based on the feature values of the second target spatial position. The network parameters of the third initial detection network are adjusted based on the seventh, eighth, and ninth loss values to obtain the target detection network.
[0124] Based on the same concept as the above method, this application proposes an electronic device, see [link to previous application]. Figure 9 As shown, the electronic device includes a processor 91 and a machine-readable storage medium 92, the machine-readable storage medium 92 storing machine-executable instructions that can be executed by the processor 91; the processor 91 is used to execute the machine-executable instructions to implement the article detection method disclosed in the above example of this application.
[0125] Based on the same concept as the above method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the article detection method disclosed in the above examples of this application.
[0126] The aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0127] Based on the same concept as the method described above, this application also provides a computer program product, which includes a computer program. When executed by a processor, the computer program implements the article detection method disclosed in the examples above.
[0128] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0129] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for detecting an item, characterized in that, The method includes: A first feature sequence is generated based on the holographic data of the object to be detected. The first feature sequence includes C first feature matrices, and the first feature matrix includes initial feature values of multiple spatial locations. For any spatial location, based on the initial eigenvalues corresponding to that spatial location in C first feature matrices, K target eigenvalues for that spatial location are determined, where K is less than C; wherein, for any initial eigenvalue corresponding to that spatial location, based on the acquired parameter metric matrix, the initial eigenvalue, and the state matrix corresponding to the previous initial eigenvalue, a state matrix corresponding to that initial eigenvalue is determined, and the state matrix includes K eigenvalues; wherein, the K eigenvalues in the state matrix corresponding to the last initial eigenvalue are taken as the K target eigenvalues for that spatial location; A second feature sequence is generated based on K target feature values at each spatial location. The second feature sequence includes K second feature matrices, each of which includes the target feature values at each spatial location. The detection result of the specified type of item of the object to be detected is determined based on the second feature sequence.
2. The method according to claim 1, characterized in that, The generation of the first feature sequence based on the holographic data of the object to be detected includes: A third feature sequence is generated based on the holographic data to be detected. The third feature sequence includes C first feature matrices. The holographic data to be detected includes at least one of amplitude, phase, and distance. Based on the candidate number of each first feature matrix in ascending order, the C first feature matrices in the third feature sequence are sorted to obtain the sorted first feature sequence; wherein, for any first feature matrix, based on the initial feature value of each spatial position in the first feature matrix, the total number of spatial positions with an initial feature value greater than a preset threshold is determined, and the candidate number is determined based on the total number.
3. The method according to claim 1, characterized in that, The parameter metric matrix includes a first metric matrix and a second metric matrix; The step of determining the state matrix corresponding to the initial eigenvalue based on the acquired parameter metric matrix, the initial eigenvalue, and the state matrix corresponding to the previous initial eigenvalue includes: The state matrix corresponding to the initial eigenvalue is determined using the following formula: ; in, This represents the state matrix corresponding to the initial eigenvalue. Represents the first metric matrix. Represents the second metric matrix. This represents the state matrix corresponding to the previous initial eigenvalue. The first metric matrix is a K×K dimension matrix. For the i-th row of the first metric matrix, if the sorting index of the initial eigenvalue is less than i, then all elements in the i-th row are 0. If the sorting index is equal to i, then the elements in the i-th row are determined based on the sorting index. If the sorting index is greater than i, then the elements in the i-th row are determined based on the sorting index and the row number of the i-th row. The second metric matrix is a K×1 dimension matrix, and the second metric matrix is determined based on the sorting index.
4. The method according to claim 1, characterized in that, The step of determining the detection result of the specified type of item of the object to be detected based on the second feature sequence includes: The second feature sequence is input into a trained object detection network, which then outputs a detection result for a specified type of item of the target object based on the second feature sequence. The detection result includes the predicted category and predicted location of the specified type of item. Alternatively... The fourth feature sequence is input to the target detection network, which outputs the detection result of the specified type of item of the object to be detected based on the fourth feature sequence. The fourth feature sequence includes the second feature sequence, a two-dimensional projection feature matrix, and a spatial index feature matrix. The two-dimensional projection feature matrix includes two-dimensional projection values for multiple spatial locations, and the spatial index feature matrix includes spatial index values for multiple spatial locations. For any given spatial location, the two-dimensional projection value is the maximum initial feature value corresponding to that spatial location in C first feature matrices, and the spatial index value is the sorting number corresponding to the maximum initial feature value.
5. The method according to claim 4, characterized in that, The target detection network includes a backbone network and a detection head network. The training process of the target detection network includes: A sample feature sequence is generated based on the sample holographic data of the sample object, wherein the sample feature sequence is the second feature sequence or the fourth feature sequence corresponding to the sample holographic data; The sample feature sequence is input into the backbone network of the initial detection network to obtain an intermediate feature matrix. The intermediate feature matrix is then input into the detection head network of the initial detection network to obtain the predicted category and predicted position. A first loss value is determined based on the predicted category, and a second loss value is determined based on the predicted position. The intermediate feature matrix and the calibrated feature matrix are input into the configured contrastive learning classification network to obtain the target spatial locations of at least two categories of specified type items in the intermediate feature matrix. A third loss value is determined based on the feature values of the target spatial locations. The calibrated feature matrix includes the calibrated spatial locations corresponding to the at least two categories of specified type items. The target loss value is determined based on the first loss value, the second loss value, and the third loss value. The network parameters of the initial detection network are adjusted based on the target loss value to obtain the target detection network.
6. The method according to claim 5, characterized in that, The step of determining the third loss value based on the feature value of the target spatial location includes: determining the contrastive learning loss of each category based on the feature value of the target spatial location, and determining the third loss value based on the contrastive learning loss of each category. For any given category, the contrastive learning loss is determined using the following formula: ; in, Indicate category i Contrast learning loss, Indicate category i The feature values of the target spatial location corresponding to the specified type of item in the intermediate feature matrix. Indicate category j The feature value of the target spatial location corresponding to the specified type of item in the intermediate feature matrix, the category j With the category i different, j The value ranges from 1 to N, where N represents the value excluding the stated category. i Total number of categories other than those mentioned above; Represents cosine similarity. This represents the learnable temperature coefficient.
7. The method according to claim 4, characterized in that, The target detection network includes a backbone network and a detection head network, and the detection head network includes a location prediction network and a category prediction network. The training process of the target detection network specifically includes: A sample feature sequence is generated based on the sample holographic data of the sample object, wherein the sample feature sequence is the second feature sequence or the fourth feature sequence corresponding to the sample holographic data; The sample feature sequence is input into the backbone network of the first initial detection network to obtain the first intermediate feature matrix. The first intermediate feature matrix is input into the position prediction network of the first initial detection network to obtain the first predicted position. The fourth loss value is determined based on the first predicted position. The network parameters of the backbone network and the position prediction network are adjusted based on the fourth loss value to obtain the second initial detection network. The sample feature sequence is input into the backbone network of the second initial detection network to obtain the second intermediate feature matrix. The second intermediate feature matrix is then input into the category prediction network of the second initial detection network to obtain the first predicted category. A fifth loss value is determined based on the first predicted category. The second intermediate feature matrix is then input into the contrastive learning classification network to obtain the first target spatial location corresponding to the specified type of item in the second intermediate feature matrix. A sixth loss value is determined based on the feature value of the first target spatial location. The network parameters of the category prediction network are adjusted based on the fifth and sixth loss values to obtain the third initial detection network. The sample feature sequence is input into the backbone network of the third initial detection network to obtain the third intermediate feature matrix. The third intermediate feature matrix is then input into the detection head network of the third initial detection network to obtain the second predicted position and the second predicted category. A seventh loss value is determined based on the second predicted position, and an eighth loss value is determined based on the second predicted category. The third intermediate feature matrix is then input into the contrastive learning classification network to obtain the second target spatial position corresponding to the specified type of item in the third intermediate feature matrix. A ninth loss value is determined based on the feature value of the second target spatial position. The network parameters of the third initial detection network are adjusted based on the seventh, eighth, and ninth loss values to obtain the target detection network.
8. An item detection device, characterized in that, The device includes: The generation module is used to generate a first feature sequence based on the holographic data of the object to be detected. The first feature sequence includes C first feature matrices, and the first feature matrix includes initial feature values of multiple spatial locations. The determination module is used to determine K target feature values for any spatial location based on the initial feature values corresponding to that spatial location in C first feature matrices, where K is less than C; wherein, for any initial feature value corresponding to that spatial location, the state matrix corresponding to that initial feature value is determined based on the acquired parameter metric matrix, the initial feature value, and the state matrix corresponding to the previous initial feature value, and the state matrix includes K feature values; wherein, the K feature values in the state matrix corresponding to the last initial feature value are used as the K target feature values for that spatial location; The detection module is used to generate a second feature sequence based on K target feature values at each spatial location. The second feature sequence includes K second feature matrices, and the second feature matrices include the target feature values at each spatial location. Based on the second feature sequence, the module determines the detection result of a specified type of item of the object to be detected.
9. The apparatus according to claim 8, characterized in that, When the generation module generates a first feature sequence based on the holographic data of the object to be detected, it is specifically used to: generate a third feature sequence based on the holographic data to be detected, the third feature sequence including C first feature matrices, the holographic data to be detected including at least one of amplitude, phase, and distance; sort the C first feature matrices in the third feature sequence according to the order of the number of candidates of each first feature matrix from smallest to largest, to obtain a sorted first feature sequence; wherein, for any first feature matrix, based on the initial feature value of each spatial position in the first feature matrix, the total number of spatial positions with initial feature values greater than a preset threshold is determined, and the number of candidates is determined based on the total number; Alternatively, the parameter metric matrix includes a first metric matrix and a second metric matrix; when the determining module determines the state matrix corresponding to the initial eigenvalue based on the acquired parameter metric matrix, the initial eigenvalue, and the state matrix corresponding to the previous initial eigenvalue, it specifically uses the following formula to determine the state matrix corresponding to the initial eigenvalue: ; in, This represents the state matrix corresponding to the initial eigenvalue. Represents the first metric matrix. Represents the second metric matrix. This represents the state matrix corresponding to the previous initial eigenvalue. The first metric matrix is a K×K dimension matrix. For the i-th row of the first metric matrix, if the sort index of the initial eigenvalue is less than i, then all elements in the i-th row are 0. If the sort index is equal to i, then the elements in the i-th row are determined based on the sort index. If the sort index is greater than i, then the elements in the i-th row are determined based on the sort index and the row number of the i-th row. The second metric matrix is a K×1 dimension matrix, and the second metric matrix is determined based on the sort index. Alternatively, when the detection module determines the detection result of a specified type of item of the object to be detected based on the second feature sequence, it is specifically used to: input the second feature sequence to a trained target detection network, and have the target detection network output the detection result of the specified type of item of the object to be detected based on the second feature sequence, wherein the detection result of the specified type of item includes the predicted category and predicted location of the specified type of item; or, input the fourth feature sequence to the target detection network, and have the target detection network output the detection result of the specified type of item of the object to be detected based on the fourth feature sequence; wherein the fourth feature sequence includes the second feature sequence, a two-dimensional projection feature matrix, and a spatial index feature matrix; the two-dimensional projection feature matrix includes two-dimensional projection values of multiple spatial locations, and the spatial index feature matrix includes spatial index values of multiple spatial locations; wherein, for any spatial location, the two-dimensional projection value of the spatial location is the maximum initial feature value corresponding to the spatial location in C first feature matrices, and the spatial index value of the spatial location is the sorting number corresponding to the maximum initial feature value; Alternatively, the target detection network includes a backbone network and a detection head network. The device further includes a training module for training the target detection network. Specifically, the training module trains the target detection network by: generating a sample feature sequence based on the sample holographic data of the sample object, wherein the sample feature sequence is a second feature sequence or a fourth feature sequence corresponding to the sample holographic data; inputting the sample feature sequence to the backbone network of the initial detection network to obtain an intermediate feature matrix; inputting the intermediate feature matrix to the detection head network of the initial detection network to obtain a predicted category and a predicted position; determining a first loss value based on the predicted category; determining a second loss value based on the predicted position; inputting the intermediate feature matrix and the calibrated feature matrix to a configured contrastive learning classification network to obtain the target spatial positions of at least two categories of specified type items in the intermediate feature matrix; determining a third loss value based on the feature values of the target spatial positions; wherein the calibrated feature matrix includes the calibrated spatial positions corresponding to the at least two categories of specified type items; determining a target loss value based on the first loss value, the second loss value, and the third loss value; and adjusting the network parameters of the initial detection network based on the target loss value to obtain the target detection network. Alternatively, when the training module determines the third loss value based on the feature values of the target spatial location, it specifically performs the following steps: determining the contrastive learning loss for each category based on the feature values of the target spatial location, and determining the third loss value based on the contrastive learning loss for each category; wherein, for any category, the training module uses the following formula to determine the contrastive learning loss for that category: ; in, Indicate category i Contrast learning loss, Indicate category i The feature values of the target spatial location corresponding to the specified type of item in the intermediate feature matrix. Indicate category j The feature value of the target spatial location corresponding to the specified type of item in the intermediate feature matrix, the category j With the category i different, j The value ranges from 1 to N, where N represents the value excluding the stated category. i Total number of categories other than those mentioned above; in, Represents cosine similarity. This represents the learnable temperature coefficient; Alternatively, the detection head network includes a location prediction network and a category prediction network; the training module trains the target detection network by: generating a sample feature sequence based on the sample holographic data of the sample object, wherein the sample feature sequence is a second feature sequence or a fourth feature sequence corresponding to the sample holographic data; inputting the sample feature sequence into the backbone network of the first initial detection network to obtain a first intermediate feature matrix, inputting the first intermediate feature matrix into the location prediction network of the first initial detection network to obtain a first predicted position, and determining a fourth loss value based on the first predicted position; adjusting the network parameters of the backbone network and the location prediction network based on the fourth loss value to obtain a second initial detection network; inputting the sample feature sequence into the backbone network of the second initial detection network to obtain a second intermediate feature matrix, inputting the second intermediate feature matrix into the category prediction network of the second initial detection network to obtain a first predicted category, and determining a fifth loss value based on the first predicted category; inputting the second intermediate feature matrix into the comparison... A learning classification network is used to obtain the first target spatial position corresponding to a specified type of item in the second intermediate feature matrix. A sixth loss value is determined based on the feature values of the first target spatial position. The network parameters of the category prediction network are adjusted based on the fifth and sixth loss values to obtain a third initial detection network. The sample feature sequence is input into the backbone network of the third initial detection network to obtain a third intermediate feature matrix. The third intermediate feature matrix is input into the detection head network of the third initial detection network to obtain a second predicted position and a second predicted category. A seventh loss value is determined based on the second predicted position, and an eighth loss value is determined based on the second predicted category. The third intermediate feature matrix is input into a contrastive learning classification network to obtain the second target spatial position corresponding to the specified type of item in the third intermediate feature matrix. A ninth loss value is determined based on the feature values of the second target spatial position. The network parameters of the third initial detection network are adjusted based on the seventh, eighth, and ninth loss values to obtain the target detection network.
10. An electronic device, characterized in that, include: A processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method of any one of claims 1-7.