Basic group detection method and device, electronic equipment and storage medium
By using a preset detection model to intercept, convert, extract features and reduce dimensionality of image data, the problems of high computational cost and low accuracy of existing base detection algorithms are solved, and more efficient base detection is achieved.
Patent Information
- Application Number
- CN202410256356.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-09
Smart Images

Figure CN120611230A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a method and device for base detection, an electronic device, and a storage medium. Background Art
[0002] Base calling is an algorithm (software) that uses computer vision to identify base types (DNA sequences) from row images (raw images), writes the results to a cal file, and ultimately generates a sequencing report and FastQ data. Each shot of the sequencer's camera captures a square area called the field of view (FOV). This FOV displays the fluorescent image of tens of thousands of DNA nanoballs (DNBs) on the sequencing chip, called base fluorescence. The base calling algorithm processes this fluorescent image and converts the fluorescent signal of each DNB on the sequencing chip into a base sequence.
[0003] Currently, there are multiple base calling algorithms for base detection based on different sequencing platforms and technologies, such as base calling algorithms based on statistical models and machine learning methods (Bustard and Alta-Cyclic), algorithms in TorrentSuite software, and algorithms in Zebracall software. In summary, the above methods often encounter problems such as signal crosstalk and high error rates. Therefore, machine learning models or electrical signal processing are currently mainly used to solve noise factors.
[0004] However, the existing base calling algorithms based on machine learning models require the use of support vector machines (SVM) or the maximum expectation algorithm and clustering algorithm to detect bases. When using SVM, supervised learning is required to optimize the SVM parameters on the reference sequence, and the SVM is trained for each grid point on the reference sequence. When using the maximum expectation algorithm and clustering algorithm, accuracy is sacrificed in order to improve speed. As a result, the existing base calling algorithms based on machine learning models have high computational costs and low accuracy. Summary of the Invention
[0005] The present disclosure provides a method and apparatus for base calling, an electronic device, and a storage medium. The main purpose is to address the problems of high computational cost and low accuracy in existing base calling algorithms based on machine learning models.
[0006] According to a first aspect of the present disclosure, a method for base detection is provided, comprising:
[0007] Acquire a first preset amount of image data in each field of view, and intercept and process the image data according to position information of the data to be detected in the image data to obtain a second preset amount of image block data, wherein each of the image block data contains a third preset amount of data to be detected;
[0008] Performing data conversion processing on the second preset number of image block data respectively according to a pre-trained preset detection model to obtain a to-be-detected conversion vector of a first preset dimension corresponding to each of the to-be-detected data;
[0009] Based on the pre-trained preset detection model, feature extraction processing is performed on the conversion vectors to be detected to obtain a feature vector to be detected corresponding to each conversion vector to be detected;
[0010] Based on the pre-trained preset detection model, the feature vector to be detected is subjected to dimensionality reduction processing to obtain a corresponding vector to be classified of a second preset dimension, and base classification is performed according to the vector to be classified to obtain a base detection result.
[0011] Optionally, the intercepting the image data according to the position information of the data to be detected in the image data to obtain the second preset number of image block data includes:
[0012] Obtaining an index of each of the to-be-detected data in the image data, wherein the index is used to determine position information of the to-be-detected data in the image data;
[0013] Determine position information of the data to be detected in the image data according to each index;
[0014] Based on the position information, with each of the data to be detected as the center, in the image data, unit image data containing the third preset number of data to be detected is intercepted until all the data to be detected are processed to obtain the second preset number of image block data.
[0015] Optionally, performing data conversion processing on the second preset number of image block data respectively according to the pre-trained preset detection model to obtain a to-be-detected conversion vector of the first preset dimension corresponding to each of the to-be-detected data includes:
[0016] After the image block data is convolved through a convolution layer, activation processing is performed through a preset activation function to obtain a data vector to be detected of a first preset dimension corresponding to each data to be detected;
[0017] A position coding vector is randomly assigned to each of the data to be detected, and the data vector to be detected corresponding to the data to be detected is added to the assigned position coding vector to obtain the conversion vector to be detected, and the position coding vector is used to determine the spatial feature information in the conversion vector to be detected.
[0018] Optionally, performing feature extraction processing on the to-be-detected conversion vectors based on the pre-trained preset detection model to obtain a to-be-detected feature vector corresponding to each to-be-detected conversion vector includes:
[0019] Performing self-attention calculation on the transformation vector to be detected through a preset attention algorithm to obtain an attention vector;
[0020] Inputting the attention vector into a first preset fully connected structure, performing data operation with a first fully connected parameter, and activating the structure through a second preset function to obtain a first feature vector, wherein the first fully connected parameter is used to extract features from the attention vector;
[0021] Inputting the first feature vector into a second preset fully connected structure, performing data operation with a second fully connected parameter to obtain a second feature vector, wherein the second fully connected parameter is used to further extract features from the first feature vector;
[0022] After the second eigenvector is calculated using a preset residual function, the second eigenvector is added to the first eigenvector to obtain the eigenvector to be detected.
[0023] Optionally, performing dimensionality reduction processing on the feature vector to be detected based on the pre-trained preset detection model to obtain a corresponding vector to be classified of a second preset dimension includes:
[0024] Performing a regularization operation on the feature vector to be detected by a preset regularization algorithm to obtain a regularized feature vector;
[0025] The regularized feature vector is subjected to dimensionality reduction processing through a third preset fully connected structure to obtain a vector to be classified of the second preset dimension.
[0026] According to a second aspect of the present disclosure, a method for training a preset detection model is provided, wherein the preset detection model can be applied to the method described in the first aspect, including:
[0027] Acquiring a fourth preset quantity of training image data in each field of view, and intercepting the training image data based on position information of the training data to be detected in the training image data to obtain a fifth preset quantity of training image block data, wherein each of the image block data contains a sixth preset quantity of training data to be detected;
[0028] Obtaining a true label corresponding to each of the training data to be detected, where the true label is a known base detection result of the training data to be detected;
[0029] Performing data conversion processing on the fifth preset number of training image block data according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each training to-be-detected data;
[0030] Based on the preset detection model, feature extraction processing is performed on the training conversion vectors to be detected to obtain a training feature vector to be detected corresponding to each training conversion vector to be detected;
[0031] Based on the preset detection model, dimensionality reduction processing is performed on the training feature vector to be detected to obtain a corresponding training vector to be classified of a fourth preset dimension, and base classification is performed according to the training vector to be classified to obtain a training base detection result;
[0032] Based on the training base detection results and the true labels, the preset detection model is optimized by a preset loss function and a preset optimization function until a preset convergence condition is met, thereby obtaining a trained preset detection model.
[0033] Optionally, the intercepting the training image data according to the position information of the training data to be detected in the training image data to obtain a fifth preset number of training image block data includes:
[0034] Obtaining an index of each of the training data to be detected in the training image data, wherein the index is used to determine position information of the training data to be detected in the training image data;
[0035] Determine position information of the training data to be detected in the training image data according to each index;
[0036] Based on the position information, with each of the training data to be detected as the center, the training unit image data containing the sixth preset number of training data to be detected is intercepted from the training image data until all the training data to be detected are processed to obtain the fifth preset number of training image block data.
[0037] Optionally, performing data conversion processing on the fifth preset number of training image block data according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each training to-be-detected data includes:
[0038] After the training image block data is convolved through a convolution layer, activation processing is performed through a preset activation function to obtain a training data vector to be detected of a third preset dimension corresponding to each training data to be detected;
[0039] A training position coding vector is randomly assigned to each of the training data to be detected, and the training data vector to be detected corresponding to the training data to be detected is added to the assigned training position coding vector to obtain the training transformation vector to be detected, and the training position coding vector is used to determine the spatial feature information in the training transformation vector to be detected.
[0040] Optionally, performing feature extraction processing on the training to-be-detected conversion vectors based on the preset detection model to obtain a training to-be-detected feature vector corresponding to each training to-be-detected conversion vector includes:
[0041] Inputting the training attention vector into a first preset fully connected structure, performing data operation with the first training fully connected parameters, and activating the training attention vector through a second preset function to obtain a first training feature vector, wherein the first training fully connected parameters are used to extract features from the training attention vector;
[0042] Inputting the first training feature vector into a second preset fully connected structure, performing data operation with a second training fully connected parameter to obtain a second training feature vector, wherein the second training fully connected parameter is used to further extract features from the first training feature vector;
[0043] After the second training feature vector is calculated using a preset residual function, the feature vector is added to the second training feature vector to obtain the feature vector to be detected for training.
[0044] Optionally, performing dimensionality reduction processing on the training feature vector to be detected based on the preset detection model to obtain a corresponding training vector to be classified of a fourth preset dimension includes:
[0045] Performing a regularization operation on the training feature vector to be detected by a preset regularization algorithm to obtain a training regularized feature vector;
[0046] The training regularized feature vector is subjected to dimensionality reduction processing through a third preset fully connected structure to obtain a training vector to be classified of the fourth preset dimension.
[0047] According to a third aspect of the present disclosure, there is provided a base detection device, comprising:
[0048] a capture unit, configured to acquire a first preset amount of image data in each field of view, and perform capture processing on the image data based on position information of the data to be detected in the image data to obtain a second preset amount of image block data, wherein each of the image block data contains a third preset amount of data to be detected;
[0049] a conversion unit, configured to perform data conversion processing on each of the second preset number of image block data according to a pre-trained preset detection model, to obtain a to-be-detected conversion vector of a first preset dimension corresponding to each of the to-be-detected data;
[0050] An extraction unit, configured to perform feature extraction processing on the conversion vectors to be detected based on the pre-trained preset detection model, to obtain a feature vector to be detected corresponding to each conversion vector to be detected;
[0051] A dimensionality reduction unit, configured to perform dimensionality reduction processing on the feature vector to be detected based on the pre-trained preset detection model to obtain a corresponding vector to be classified of a second preset dimension;
[0052] The classification unit is used to perform base classification according to the vector to be classified to obtain a base detection result.
[0053] Optionally, the interception unit includes:
[0054] an acquisition module, configured to acquire an index of each of the data to be detected in the image data, wherein the index is used to determine position information of the data to be detected in the image data;
[0055] a determination module, configured to determine position information of the data to be detected in the image data according to each of the indexes;
[0056] The interception module is used to intercept the unit image data containing the third preset number of data to be detected in the image data based on the position information and with each of the data to be detected as the center, until all the data to be detected are processed to obtain the second preset number of image block data.
[0057] Optionally, the conversion unit includes:
[0058] a processing module, configured to perform a convolution operation on the image block data through a convolution layer, and then perform activation processing through a preset activation function to obtain a data vector to be detected of a first preset dimension corresponding to each data to be detected;
[0059] A configuration module is used to randomly assign a position coding vector to each of the data to be detected, and add the data vector to be detected corresponding to the data to be detected and the assigned position coding vector to obtain the conversion vector to be detected, and the position coding vector is used to determine the spatial feature information in the conversion vector to be detected.
[0060] Optionally, the extraction unit includes:
[0061] A first calculation module is used to perform self-attention calculation on the conversion vector to be detected through a preset attention algorithm to obtain an attention vector;
[0062] a processing module, configured to input the attention vector into a first preset fully connected structure, perform data operation on the attention vector and a first fully connected parameter, and then activate the attention vector through a second preset function to obtain a first feature vector, wherein the first fully connected parameter is used to perform feature extraction on the attention vector;
[0063] The processing module is further configured to input the first feature vector into a second preset fully connected structure, perform data operation on the first feature vector and a second fully connected parameter to obtain a second feature vector, wherein the second fully connected parameter is used to further extract features from the first feature vector;
[0064] The second calculation module is configured to calculate the second eigenvector using a preset residual function, and then add the calculated second eigenvector to the second eigenvector to obtain the eigenvector to be detected.
[0065] Optionally, the dimensionality reduction unit includes:
[0066] An operation module, configured to perform a regularization operation on the feature vector to be detected by using a preset regularization algorithm to obtain a regularized feature vector;
[0067] A dimensionality reduction module is used to perform dimensionality reduction processing on the regularized feature vector through a third preset fully connected structure to obtain a vector to be classified of the second preset dimension.
[0068] According to a fourth aspect of the present disclosure, a training device for a preset detection model is provided, comprising:
[0069] a clipping unit, configured to obtain a fourth preset number of training image data in each field of view, and clip the training image data based on position information of the training data to be detected in the training image data to obtain a fifth preset number of training image block data, wherein each of the image block data contains a sixth preset number of training data to be detected;
[0070] An acquiring unit, configured to acquire a true label corresponding to each of the training data to be detected, wherein the true label is a known base detection result of the training data to be detected;
[0071] a conversion unit, configured to perform data conversion processing on the fifth preset number of training image block data respectively according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each of the training to-be-detected data;
[0072] an extraction unit, configured to perform feature extraction processing on the training transformation vectors to be detected based on the preset detection model, to obtain a training feature vector to be detected corresponding to each training transformation vector to be detected;
[0073] a dimensionality reduction unit, configured to perform dimensionality reduction processing on the training feature vector to be detected based on the preset detection model to obtain a corresponding training vector to be classified of a fourth preset dimension;
[0074] a classification unit, configured to perform base classification according to the training vector to be classified to obtain a training base detection result;
[0075] An optimization unit is used to optimize the preset detection model based on the training base detection results and the true labels through a preset loss function and a preset optimization function until a preset convergence condition is met, thereby obtaining a trained preset detection model.
[0076] Optionally, the interception unit includes:
[0077] an acquisition module, configured to acquire an index of each of the training data to be detected in the training image data, wherein the index is used to determine position information of the training data to be detected in the training image data;
[0078] a determination module, configured to determine position information of the training data to be detected in the training image data according to each index;
[0079] The interception module is used to intercept the training unit image data containing the sixth preset number of training data to be detected in the training image data based on the position information and with each of the training data to be detected as the center, until all the training data to be detected are processed to obtain the fifth preset number of training image block data.
[0080] Optionally, the conversion unit includes:
[0081] a processing module, configured to perform a convolution operation on the training image block data through a convolution layer, and then perform activation processing through a preset activation function to obtain a training data vector to be detected of a third preset dimension corresponding to each training data to be detected;
[0082] A configuration module is provided for randomly assigning a training position coding vector to each of the training data to be detected, and adding the training data vector to be detected corresponding to the training data to be detected to the assigned training position coding vector to obtain the training transformation vector to be detected, wherein the training position coding vector is used to determine the spatial feature information in the training transformation vector to be detected.
[0083] Optionally, the extraction unit includes:
[0084] A first calculation module is used to perform self-attention calculation on the training to-be-detected conversion vector through a preset attention algorithm to obtain a training attention vector;
[0085] a processing module, configured to input the training attention vector into a first preset fully connected structure, perform data calculation with the first training fully connected parameters, and then activate the training attention vector through a second preset function to obtain a first training feature vector, wherein the first training fully connected parameters are used to perform feature extraction on the training attention vector;
[0086] The processing module is further configured to input the first training feature vector into a second preset fully connected structure, perform data operation on the first training feature vector and a second training fully connected parameter to obtain a second training feature vector, wherein the second training fully connected parameter is used to further extract features from the first training feature vector;
[0087] The second calculation module is configured to calculate the second training feature vector using a preset residual function, and then perform addition calculation on the second training feature vector to obtain the feature vector to be detected for training.
[0088] Optionally, the dimensionality reduction unit includes:
[0089] An operation module, configured to perform a regularization operation on the training feature vector to be detected by using a preset regularization algorithm to obtain a training regularized feature vector;
[0090] A dimensionality reduction module is used to perform dimensionality reduction processing on the training regularized feature vector through a third preset fully connected structure to obtain the training vector to be classified of the fourth preset dimension.
[0091] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0092] at least one processor; and
[0093] a memory communicatively connected to the at least one processor; wherein,
[0094] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect or the method described in the second aspect.
[0095] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect or the method described in the second aspect.
[0096] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method described in the first aspect or the method described in the second aspect.
[0097] The method and apparatus, electronic device, and storage medium for base detection provided by the present disclosure obtain a first preset number of image data in each field of view, and perform interception processing on the image data according to the position information of the data to be detected in the image data to obtain a second preset number of image block data, wherein each of the image block data contains a third preset number of data to be detected; perform data conversion processing on the second preset number of image block data according to a pre-trained preset detection model to obtain a first preset dimension of a detection conversion vector corresponding to each of the data to be detected; perform feature extraction processing on the detection conversion vector based on the pre-trained preset detection model to obtain a detection feature vector corresponding to each detection conversion vector; perform dimensionality reduction processing on the detection feature vector based on the pre-trained preset detection model to obtain a corresponding second preset dimension of a classification vector, and perform base classification based on the classification vector to obtain a base detection result. Compared with the related art, the embodiment of the present disclosure realizes data modeling processing of the original image data by adopting a preset detection model, enhances the feature extraction capability of data features, reduces computing costs, and improves the accuracy of base detection.
[0098] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.
[0100] Figure 1 A schematic flow chart of a base detection method provided in an embodiment of the present disclosure;
[0101] Figure 2 A schematic diagram of an example of base detection provided by an embodiment of the present disclosure;
[0102] Figure 3 A schematic diagram of a data preprocessing process provided by an embodiment of the present disclosure;
[0103] Figure 4 A schematic diagram of a data conversion process provided by an embodiment of the present disclosure;
[0104] Figure 5 A schematic diagram of a data feature extraction process provided by an embodiment of the present disclosure;
[0105] Figure 6 A schematic diagram of the structure of a multi-head attention mechanism layer provided in an embodiment of the present disclosure;
[0106] Figure 7 A schematic diagram of a data dimensionality reduction process provided by an embodiment of the present disclosure;
[0107] Figure 8 A flow chart showing the overall structure of a base detection method provided by an embodiment of the present disclosure;
[0108] Figure 9 A flowchart of a method for training a preset detection model provided in an embodiment of the present disclosure;
[0109] Figure 10 A schematic structural diagram of a base detection device provided in an embodiment of the present disclosure;
[0110] Figure 11 A schematic structural diagram of another base detection device provided by an embodiment of the present disclosure;
[0111] Figure 12 A schematic diagram of the structure of a training device for a preset detection model provided in an embodiment of the present disclosure;
[0112] Figure 13 A schematic diagram of the structure of another training device for a preset detection model provided in an embodiment of the present disclosure;
[0113] Figure 14 A schematic block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0114] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0115] The following describes the base detection method and apparatus, electronic device, and storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.
[0116] Figure 1 A schematic flow chart of a base detection method provided in an embodiment of the present disclosure.
[0117] like Figure 1 As shown, the method comprises the following steps:
[0118] Step 101: Acquire a first preset amount of image data in each field of view, and intercept and process the image data according to position information of the data to be detected in the image data to obtain a second preset amount of image block data, wherein each of the image block data contains a third preset amount of data to be detected.
[0119] In the embodiment of the present disclosure, the image data is a fluorescent image of the data to be detected on the sequencing chip displayed in the field of view (FOV). The first preset number is the number of all image data in the FOV, and the second preset number is the number of all image block data obtained after intercepting the image data. Specifically, the size of the first preset number and the second preset number depends on actual conditions and is not limited in the embodiment of the present disclosure.
[0120] Among them, the third preset number is a custom-set number, for example: 9, 16, etc., the expression form of the image block data includes but is not limited to: matrix form, vector form, etc., the data to be detected includes but is not limited to: DNA nanoball (DNB), etc., the DNB is a spherical structure self-assembled by DNA molecules, which is usually used in high-throughput sequencing technology. Specifically, the size of the third preset number, the expression form of the image block data and the data to be detected are not limited in the embodiment of the present disclosure.
[0121] In order to facilitate understanding of the implementation process of the embodiment of the present disclosure, an example is provided for illustration: on the image data in each FOV, a 3x3 patch data (image block data) X∈R centered on the DNB (data to be detected) is intercepted. N ×T×3×3×2 , N is the number of DNBs on each image data, T is the length of the sequence Cycles, the number of light source intensity channels is 2 or 4, wherein the 3x3 patch data (image block data) indicates that the image block data contains 9 DNBs (data to be detected), that is, the third preset number is 9. For example, the third preset number is 16, which means that the 4x4 patch data is intercepted, and the image block data is expressed as X∈R N×T×4×4×2 The length of the Cycles sequence is the number of image data in the FOV (the first preset number); the number of light source intensity channels is the attribute of the image data itself. For example, if the image data has three light source intensity channels and the number of light source intensity channels is 3, the image block data is expressed as X∈R N×T×3×3×3 , specifically, the embodiments of the present disclosure are not limited.
[0122] Step 102 : performing data conversion processing on the second preset number of image block data respectively according to a pre-trained preset detection model to obtain a to-be-detected conversion vector of a first preset dimension corresponding to each of the to-be-detected data.
[0123] In an embodiment of the present disclosure, the first preset dimension is a custom-set vector dimension, for example: 64, 128, etc. The pre-trained preset detection model contains multiple module structures, and the module structure includes but is not limited to: Embedding module, ViT module, classification module, etc. Specifically, the embodiment of the present disclosure does not limit the internal module structure of the pre-trained preset detection model and the size of the first preset dimension.
[0124] The conversion of the image block data is performed in the Embedding module of the pre-trained preset detection model. The Embedding module mainly maps the input X into a DNB vector of embed_size dimension, that is, converting the input image block data into a conversion vector to be detected of the first preset dimension. For example, the image block data X∈R N ×T×3×3×2 , converted into a detection conversion vector X of the first preset dimension e ∈R N×T×d , d is the first preset dimension.
[0125] Step 103 : Based on the pre-trained preset detection model, feature extraction processing is performed on the conversion vectors to be detected to obtain a feature vector to be detected corresponding to each conversion vector to be detected.
[0126] In the disclosed embodiment, feature extraction of the conversion vector to be detected is performed in the ViT module of the pre-trained preset detection model. The ViT module is mainly composed of a Transformer encoding structure. After the conversion vector to be detected passes through the self-attention mechanism of the Transformer, it can learn the correlation between time series. At the same time, compared with the existing recurrent neural network, the self-attention mechanism in the Transformer encoding structure can learn the correlation between different cycles (i.e., the distribution of A, T, G, C, and N in the first base of each sequencing read obtained by a single biochemical reaction) without being restricted by distance, thereby avoiding the defect that recurrent neural networks are prone to forgetting.
[0127] The internal structure of the Transformer encoding structure consists of two important parts: the multi-head attention mechanism and the feedforward neural network. Multi-head attention introduces the concept of multi-head on the basis of self-attention, indicating that the model can observe from multiple angles. After the transformation vector to be detected is processed by the multi-head attention mechanism and the feedforward neural network, and then undergoes data processing of a residual structure, the feature vector X to be detected can be obtained. vit ∈R N×T×d , for example: X e After the self-attention mechanism of Transformer, the correlation between time series is learned, and the output result is recorded as X vit ∈R N×T×d , if the process of feature extraction of the transformation vector to be detected is represented by a function, it can be expressed as X vit =f vit (X e ), the feature vector to be detected is also a vector of the first preset dimension.
[0128] Step 104: Based on the pre-trained preset detection model, the feature vector to be detected is subjected to dimensionality reduction processing to obtain a corresponding vector to be classified of a second preset dimension, and base classification is performed according to the vector to be classified to obtain a base detection result.
[0129] In an embodiment of the present disclosure, the second preset dimension is a custom-set vector dimension, and the setting of the second preset dimension corresponds to the type of base. For example, if there are 4 types of bases in total, the second preset dimension is also set to 4. The base detection result refers to the base type corresponding to each data to be detected.
[0130] Among them, the dimensionality reduction processing of the feature vector to be detected and the classification of the vector to be classified are both carried out in the classification module of the pre-trained preset detection model. Since the feature vector to be detected has learned the relationship between DNB (data to be detected) and its neighbors, as well as the association relationship between DNB in the Cycles dimension, after the feature vector to be detected is reduced to the dimension of base type, that is, the second preset dimension, classification can be performed to obtain the base type result of each DNB (data to be detected), that is, the base detection result.
[0131] In order to facilitate the understanding of the embodiments of the present disclosure, a schematic diagram of an example of base detection is provided, such as Figure 2As shown, base Fluorescence is the image data, Cycle is the number of image data, DNB is the data to be detected, patch input is the image block data, ViT model is the pre-trained preset detection model, patch output is the base detection result, N is the number of data to be detected on each image data, for the convenience of display, Figure 2 There is 1 data to be detected in each image data, that is, N=1.
[0132] The method for base detection provided by the present disclosure obtains a first preset number of image data in each field of view, and intercepts the image data according to the position information of the data to be detected in the image data to obtain a second preset number of image block data, wherein each of the image block data contains a third preset number of data to be detected; the second preset number of image block data are respectively subjected to data conversion processing according to a pre-trained preset detection model to obtain a first preset dimension of a detection conversion vector corresponding to each of the data to be detected; based on the pre-trained preset detection model, feature extraction processing is performed on the detection conversion vector to obtain a detection feature vector corresponding to each detection conversion vector; based on the pre-trained preset detection model, dimensionality reduction processing is performed on the detection feature vector to obtain a corresponding second preset dimension of a classification vector, and base classification is performed according to the classification vector to obtain a base detection result. Compared with the related art, the embodiment of the present disclosure realizes data modeling processing of the original image data by adopting a preset detection model, enhances the feature extraction capability of data features, reduces computing costs, and improves the accuracy of base detection.
[0133] In one possible implementation of the embodiment of the present disclosure, as a refinement of the above step 101, regarding the interception processing of the image data, the embodiment of the present disclosure provides a flow chart of data preprocessing, such as Figure 3 Shown, including:
[0134] Step 301: Obtain an index of each of the to-be-detected data in the image data, where the index is used to determine position information of the to-be-detected data in the image data.
[0135] In the embodiment of the present disclosure, the index is a kind of positioning information carried by each of the data to be detected. The position of the corresponding data to be detected can be located in the image data through the index. The expression form of the index includes but is not limited to: numerical value, coordinates, etc. Specifically, the embodiment of the present disclosure does not limit the expression form of the index.
[0136] Step 302: Determine the position information of the data to be detected in the image data according to each index.
[0137] In the embodiment of the present disclosure, the position information is used to indicate the position of the data to be detected in the image data, including but not limited to: coordinate information of the data to be detected in the image data, etc. The positions of all the data to be detected in the image data can be determined by the index.
[0138] Step 303, based on the position information, with each of the data to be detected as the center, intercept the unit image data containing the third preset number of data to be detected in the image data, until all the data to be detected are processed to obtain the second preset number of image block data.
[0139] In the embodiment of the present disclosure, since it is necessary to perform base detection on each data to be detected, each data to be detected needs to be used as the center of an image block data to improve the accuracy of base detection on each data to be detected. It should be noted that for the data to be detected existing at the edge of the image data, when the data to be detected is used as the center of the image block data, a certain amount of virtual data will be supplemented in the area outside the edge of the image data to ensure that the data to be detected is the center of the image block data. The virtual data includes but is not limited to: blank data, data with true base type labels, etc. Specifically, the embodiment of the present disclosure does not impose any restrictions.
[0140] In one possible implementation of the embodiment of the present disclosure, as a refinement of the above step 102, regarding the data conversion of the image block data, the embodiment of the present disclosure provides a flow chart of data conversion processing, such as Figure 4 Shown, including:
[0141] In step 401 , after the image block data is subjected to a convolution operation through a convolution layer, activation processing is performed through a preset activation function to obtain a data vector to be detected of a first preset dimension corresponding to each data to be detected.
[0142] In the embodiment of the present disclosure, the convolution kernel of the convolution layer is custom-set, for example: 1x3x3, 1x4x4, etc. The convolution layer includes but is not limited to: 1x3x3 3D convolution layer, 1x4x4 3D convolution layer, etc. The preset activation function is a custom-set activation function, for example: RELU activation function, etc. Specifically, the embodiment of the present disclosure does not limit the convolution kernel of the convolution layer, the convolution layer, and the preset activation function.
[0143] It should be noted that, when performing the convolution operation on the image block data, no padding operation is required, and the step size setting in the convolution operation can be customized, for example, the step size is set to 1.
[0144] Therefore, the implementation process of the embodiment of the present disclosure includes but is not limited to: inputting patch data (the image block data) X∈R N×T×3×3×2 , the convolution operation is performed through a 3D convolution layer with a convolution kernel of 1x3x3 (no padding, step size of 1), and then the RELU activation function is used to obtain the data vector e∈R N×T×d , if this process is recorded as f e , then the process is f e Convert the patch data into a DNB vector (data vector to be detected) e∈R with a dimension of d N×T×d , wherein the data vector to be detected is a first preset dimension vector.
[0145] Step 402: randomly assign a position coding vector to each of the data to be detected, and add the data vector to be detected corresponding to the data to be detected and the assigned position coding vector to obtain the transformation vector to be detected, wherein the position coding vector is used to determine the spatial feature information in the transformation vector to be detected.
[0146] In the embodiment of the present disclosure, the position encoding vector is a position encoding about T, that is, the position encoding vector is related to the amount of the image data in the field of view, and can be expressed as: e ∈R 1×T×d The position encoding vector can make up for the lack of perception of data order in the self-attention mechanism when processing the transformation vector to be detected. Since the self-attention mechanism is adopted in the VIT module (Transformer structure), this mechanism focuses on the relationship between each element in the input sequence, that is, the transformation vector to be detected, but does not take into account the position information of the transformation vector to be detected in the sequence. The position encoding vector can enable the VIT module to use the position information to distinguish the transformation vectors to be detected at different positions in the sequence, so that when the pre-trained preset detection model processes the image block data, it can better understand the order and relationship of the data to be detected in the image block data.
[0147] It should be noted here that the position coding vector is a trainable vector, the initial position coding vector is an all-zero matrix, and the position coding vector will be trained synchronously with the training of the preset detection model. Therefore, in the pre-trained preset detection model, the position coding vector is also a trained position coding vector and can be directly called for allocation and use.
[0148] The overall implementation process of data conversion in the embodiment of the present disclosure includes but is not limited to: inputting patch data (the image block data) X∈R N×T×3×3×2 , the convolution operation is performed through a 3D convolution layer with a convolution kernel of 1x3x3 (no padding, step size of 1), and then the RELU activation function is used to obtain the data vector e∈R N×T×d , assign a position encoding vector p about T to each DNB vector e ∈R 1×T×d , the position encoding vector is added to the DNB vector to obtain the conversion vector to be detected. This process can be expressed by formula (1):
[0149] X e =f e (X)+p e Formula (1)
[0150] Among them, X e ∈R N×T×d ,The Embedding module uses 3D convolution to process image block data, which is better than the full connection method in related technologies.
[0151] In one possible implementation of the embodiment of the present disclosure, as a refinement of the above step 103, regarding the feature extraction process of the conversion vector to be detected, the embodiment of the present disclosure provides a flow chart of data feature extraction process, such as Figure 5 Shown, including:
[0152] Step 501: Perform self-attention calculation on the conversion vector to be detected through a preset attention algorithm to obtain an attention vector.
[0153] In the embodiment of the present disclosure, the attention algorithm is all the algorithms used in the process of using Muti-Head Attention. When performing self-attention calculation on the conversion vector to be detected, the conversion vector to be detected is split into n_heads parts on the d dimension (the first preset dimension), where n_heads represents the number of self-attentions (attention is calculated separately). Finally, the results of n_heads self-attentions are re-spliced back along the d dimension to obtain the attention vector.
[0154] In order to facilitate understanding of the implementation process of the embodiment of the present disclosure, a structural diagram of a multi-head attention mechanism layer is provided, as shown in FIG. Figure 6 As shown, Figure 6 The left is the structure diagram of self-attention. Figure 6The right is the structure diagram of Multi-Head Attention. Specifically, the process of calculating the self-attention result of the transformation vector to be detected by Multi-Head Attention includes but is not limited to: e The (transformation vector to be detected) is input into three fully connected layers (FC) to obtain Q, K, and V respectively. This process can be expressed by formula (2), formula (3), and formula (4) respectively:
[0155] Q=FC(X e )Q∈R N×T×d Formula (2)
[0156] K=FC(X e )K∈R N×T×d Formula (3)
[0157] V=FC(X e )V∈R N×T×d Formula (4)
[0158] Then, Q, K, and V are split into n_heads parts in the d dimension. This process can be expressed by formula (5), formula (6), and formula (7) respectively:
[0159]
[0160]
[0161]
[0162] Then, the self-attention results are calculated for n_heads parts Q, K, and V respectively. This process can be expressed by formula (8):
[0163]
[0164] in,
[0165] Finally, the n_heads self-attentions are concatenated along the d dimension to obtain the attention vector. This process can be expressed by formula (9):
[0166] Attention=Concat(Attention n_heads ) Formula (9)
[0167] Among them, Attention is the attention vector, Attention∈R N×T×d, Concat refers to the concatenate operation in the d dimension.
[0168] In step 502, the attention vector is input into a first preset fully connected structure, and after performing data operation with a first fully connected parameter, it is activated by a second preset function to obtain a first feature vector, wherein the first fully connected parameter is used to extract features from the attention vector.
[0169] In the embodiment of the present disclosure, the first preset fully connected structure is a fully connected structure of a custom setting, such as: a feedforward neural network (Feed forward), etc., and the second preset function is a function of a custom setting, such as: a GELU activation function, etc. The role of the second preset function is to learn more abstract features so that the interaction relationship between the data to be detected (DNB) is strengthened. In the Muti-Head Attention calculation process, most of the operations are performed using matrix multiplication, and these operations are linear transformations. The introduction of Feed forward increases the model's ability to perform nonlinear transformations.
[0170] The calculation process of the first eigenvector includes but is not limited to obtaining it through formula (10):
[0171] FFN(x1)=GELU(AttentionW1+b1) Formula (10)
[0172] Among them, x1 is the first eigenvector, FFN represents a feedforward neural network, b1 is a bias value, and the method for determining the bias value includes but is not limited to: custom setting, random generation, etc. W1 is a parameter of the feedforward neural network, namely the first fully connected parameter, which mainly plays the role of connection and weight. The first fully connected parameter exists in the form of weight between each level in the neural network, and determines the mapping relationship between the input and output of the neuron. By learning and adjusting the first fully connected parameter, the neural network can realize various complex nonlinear function approximation and classification tasks, that is, the first fully connected parameter plays the role of connecting input and output, extracting features, realizing nonlinear mapping, fitting optimization model and providing interpretability in the feedforward neural network.
[0173] It should be noted that the first fully connected parameter is a trainable parameter. The initial first fully connected parameter is a randomly generated parameter, and the first fully connected parameter will be trained synchronously with the training of the preset detection model. Therefore, in the pre-trained preset detection model, the first fully connected parameter is also a trained fully connected parameter and can be directly called for use.
[0174] Step 503: Input the first feature vector into a second preset fully connected structure, perform data operation with a second fully connected parameter, and obtain a second feature vector, wherein the second fully connected parameter is used to further extract features from the first feature vector.
[0175] In the embodiment of the present disclosure, the second preset fully connected structure is a fully connected structure with custom settings, such as a feedforward neural network (Feed forward), etc. It can be the same structure as the first fully connected structure or a different structure. Specifically, the embodiment of the present disclosure does not limit it.
[0176] The calculation process of the second eigenvector includes but is not limited to obtaining it through formula (11):
[0177] FFN(x2)=GELU(xW1+b1)W2+b2 Formula (11)
[0178] Among them, x2 is the second eigenvector, GELU(xW1+b1) is the calculation formula of the first eigenvector, that is, the second eigenvector is calculated by the first eigenvector, b2 is the bias value, W2 is the parameter of the feedforward neural network, that is, the second fully connected parameter. Specifically, for the specific description of b2 and W2, please refer to the description of b1 and W1 in the above step 502, so they will not be repeated here.
[0179] Step 504 : After calculating the second eigenvector using a preset residual function, the second eigenvector is added to the second eigenvector to obtain the eigenvector to be detected.
[0180] In the embodiment of the present disclosure, the preset residual function is all functions in the residual structure, such as: identity mapping function (Identity Mapping), skip connection (Skip Connectiohs), batch normalization (BatchNormalization), weighted residual (Weighted Residual), depthwise separable convolution (Depthwise SeparableConvolution), etc. Specifically, regarding the preset residual function, it can be determined according to the actual residual structure, and the embodiment of the present disclosure is not limited.
[0181] The calculation process of the feature vector to be detected includes but is not limited to obtaining it through formula (12):
[0182] X vit =F(x2)+x2 Formula (12)
[0183] Among them, X vitis the feature vector to be detected, x2 is the second feature vector, and F(x2) is the result of calculating the second feature vector using a preset residual function.
[0184] In one possible implementation of the embodiment of the present disclosure, as a refinement of the above step 104, regarding the dimensionality reduction processing of the feature vector to be detected, the embodiment of the present disclosure provides a flow chart of data dimensionality reduction processing, such as Figure 7 Shown, including:
[0185] Step 701 : performing a regularization operation on the feature vector to be detected by using a preset regularization algorithm to obtain a regularized feature vector.
[0186] In the embodiment of the present disclosure, the preset regularization algorithm is a custom-set algorithm, and the feature vector to be detected can be calculated by a formula to complete the regularization operation (LayerNorm). Therefore, the method of performing LayerNorm on the feature vector to be detected includes but is not limited to performing it by formula (13):
[0187]
[0188] Among them, μ, σ are respectively vit The mean and standard deviation of , γ, β are learnable parameters of LayerNorm, and θ is the minimum value to prevent division by 0.
[0189] Step 702: Perform dimensionality reduction processing on the regularized feature vector through a third preset fully connected structure to obtain a vector to be classified of the second preset dimension.
[0190] In the embodiment of the present disclosure, the third preset fully connected structure is a fully connected structure with custom settings, such as: feed forward neural network (Feed forward), recurrent neural network (Recurrent Neural Network, RNN), multilayer perceptron (Multilayer Perceptron, MLP), etc. Specifically, the embodiment of the present disclosure does not limit the third preset fully connected structure.
[0191] Regarding the second preset dimension, please refer to the description in the above step 104, so it will not be described in detail here.
[0192] Regarding the dimensionality reduction processing of the feature vector to be detected, it can be expressed as: Y = f class (X vit ), where Y∈R N ×T×4 .
[0193] Regarding the method of base detection, the embodiment of the present disclosure provides an overall structural flow chart of the method of base detection, such as Figure 8 As shown, the image block data is the input data (N, 150, 3, 3, 2) obtained by intercepting 3x3 patch data centered on each data to be detected, the first preset dimension is 64, the second preset dimension is 4, the number of data to be detected in each image data is n, and the number of image data in each field of view is 150.
[0194] Corresponding to the above-mentioned base detection method, Figure 9 A flow chart of a method for training a preset detection model provided in an embodiment of the present disclosure is shown as follows: Figure 9 Shown, including:
[0195] Step 901: Acquire a fourth preset number of training image data in each field of view, and intercept and process the training image data according to position information of the training data to be detected in the training image data to obtain a fifth preset number of training image block data, wherein each of the image block data contains a sixth preset number of training data to be detected.
[0196] Specifically, regarding the implementation process of the embodiment of the present disclosure, please refer to the description in the above step 101, so it will not be described here one by one.
[0197] Step 902: Obtain a true label corresponding to each of the training data to be detected, where the true label is a known base detection result of the training data to be detected.
[0198] In the embodiment of the present disclosure, when training the preset prediction model, a training set and a validation set are required. The training set is used to train the preset prediction model, and the validation set is used to optimize the preset prediction model. Therefore, it is necessary to obtain the true label corresponding to each of the training data to be tested as a validation set.
[0199] Step 903 : performing data conversion processing on the fifth preset number of training image block data respectively according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each of the training to-be-detected data.
[0200] Step 904 : Based on the preset detection model, feature extraction processing is performed on the training transformation vectors to be detected to obtain a training feature vector to be detected corresponding to each training transformation vector to be detected.
[0201] Step 905: Based on the preset detection model, the training feature vector to be detected is subjected to dimensionality reduction processing to obtain a corresponding training vector to be classified of a fourth preset dimension, and base classification is performed according to the training vector to be classified to obtain a training base detection result.
[0202] Specifically, regarding the implementation process of steps 903 to 905, please refer to the description of steps 102 to 104 above, so they will not be described here one by one.
[0203] Step 906: Based on the training base detection results and the true labels, the preset detection model is optimized by a preset loss function and a preset optimization function until a preset convergence condition is met, thereby obtaining a trained preset detection model.
[0204] In the embodiment of the present disclosure, the preset loss function is a custom-selected loss function, such as the cross entropy function, etc., and the preset optimization function is a custom-selected optimization function, such as the AdamW function, etc. Specifically, the embodiment of the present disclosure does not limit the selection of the preset loss function and the preset optimization function.
[0205] Calculating classification loss with the preset loss function The method includes but is not limited to formula (14):
[0206]
[0207] in, is the output of the model prediction, i.e., the training base detection result, and y is the one-hot form of the true category (true label).
[0208] When training a model, an example is provided to illustrate the selection of data:
[0209] Dataset: Sequencing image data is a matrix with a shape of (300, 6229, 4608, 2). The first dimension represents the sum of the cycle lengths of read1 and read2, both 150. The second and third dimensions represent the width and height of the image, respectively. The last dimension represents the pixel intensity. The input data is (N, 150, 3, 3, 2) based on the index and coordinates of the DNB and a 3x3 patch centered on the DNB. N is the sum of the DNBs of read1 and read2.
[0210] Parameter Settings: Both the spatial ViT module and the temporal ViT module in this paper use a 4-layer, 8-head TransformerBlock, with an embedding layer dimension of 64 and a class size of 4. The batch size during training is 384, the learning rate is set to 5e-4, and a linear learning rate decay strategy with warmup is introduced during training. The warmup step size is set to 0.1 of the total step size. The AdamW optimizer is used for parameter updates, with an epoch of 1000 and an early stopping round number of 20.
[0211] In one implementation of the embodiment of the present disclosure, the intercepting and processing the training image data according to the position information of the training data to be detected in the training image data to obtain the fifth preset number of training image block data includes:
[0212] Obtaining an index of each of the training data to be detected in the training image data, wherein the index is used to determine position information of the training data to be detected in the training image data;
[0213] Determine position information of the training data to be detected in the training image data according to each index;
[0214] Based on the position information, with each of the training data to be detected as the center, the training unit image data containing the sixth preset number of training data to be detected is intercepted from the training image data until all the training data to be detected are processed to obtain the fifth preset number of training image block data.
[0215] Specifically, regarding the implementation process of the embodiment of the present disclosure, reference may be made to the description of steps 301-303 above, and therefore will not be repeated here.
[0216] In one implementation of the embodiment of the present disclosure, performing data conversion processing on the fifth preset number of training image block data according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each training to-be-detected data includes:
[0217] After the training image block data is convolved through a convolution layer, activation processing is performed through a preset activation function to obtain a training data vector to be detected of a third preset dimension corresponding to each training data to be detected;
[0218] A training position coding vector is randomly assigned to each of the training data to be detected, and the training data vector to be detected corresponding to the training data to be detected is added to the assigned training position coding vector to obtain the training transformation vector to be detected, and the training position coding vector is used to determine the spatial feature information in the training transformation vector to be detected.
[0219] Specifically, regarding the implementation process of the embodiment of the present disclosure, reference may be made to the description in steps 401-402 above, and therefore will not be repeated here.
[0220] In one implementation of the embodiment of the present disclosure, performing feature extraction processing on the training to-be-detected conversion vectors based on the preset detection model to obtain a training to-be-detected feature vector corresponding to each training to-be-detected conversion vector includes:
[0221] Performing self-attention calculation on the training transformation vector to be detected through a preset attention algorithm to obtain a training attention vector;
[0222] Inputting the training attention vector into a first preset fully connected structure, performing data operation with the first training fully connected parameters, and activating the training attention vector through a second preset function to obtain a first training feature vector, wherein the first training fully connected parameters are used to extract features from the training attention vector;
[0223] Inputting the first training feature vector into a second preset fully connected structure, performing data operation with a second training fully connected parameter to obtain a second training feature vector, wherein the second training fully connected parameter is used to further extract features from the first training feature vector;
[0224] After the second training feature vector is calculated using a preset residual function, the feature vector is added to the second training feature vector to obtain the feature vector to be detected for training.
[0225] Specifically, regarding the implementation process of the embodiment of the present disclosure, reference may be made to the description in steps 501 to 504 above, and therefore details will not be repeated here.
[0226] In one implementation of the embodiment of the present disclosure, performing dimensionality reduction processing on the training feature vector to be detected based on the preset detection model to obtain a corresponding training vector to be classified of the fourth preset dimension includes:
[0227] Performing a regularization operation on the training feature vector to be detected by a preset regularization algorithm to obtain a training regularized feature vector;
[0228] The training regularized feature vector is subjected to dimensionality reduction processing through a third preset fully connected structure to obtain a training vector to be classified of the fourth preset dimension.
[0229] Specifically, regarding the implementation process of the embodiment of the present disclosure, please refer to the description in the above steps 701-702, so they will not be described here one by one.
[0230] In summary, the embodiments of the present disclosure can achieve the following effects:
[0231] 1. The disclosed embodiment implements data modeling processing of raw image data by adopting a preset detection model, thereby enhancing the feature extraction capability of data features, reducing computational costs, and improving the accuracy of base detection.
[0232] 2. The disclosed embodiment achieves global perception capability through the self-attention mechanism of a preset detection model. Compared with traditional convolutional neural networks that mainly process images through local perception and local feature extraction, the preset detection model can simultaneously consider the global information in the image, thereby better capturing the long-distance dependencies in the image.
[0233] 3. The preset detection model provided in the embodiment of the present disclosure has a relatively simple structure and is composed of multiple attention modules. The scale of the model can be expanded by increasing the number of modules and adjusting the size of the modules. The high scalability enables the preset detection model to adapt to image data of different sizes and complexities.
[0234] 4. The embodiments of the present disclosure can balance the computing and storage requirements of the model by adjusting the number and size of attention modules through the preset detection model, thereby optimizing under different resource constraints. At the same time, the preset detection model can also perform transfer learning through pre-training and fine-tuning to adapt to different image tasks.
[0235] 5. The preset detection model provided in the embodiment of the present disclosure can better model the position information in the image by introducing a position encoding vector into the input data (image block data). Compared with traditional convolutional neural networks for image processing, the preset detection model is more effective when processing tasks that require consideration of positional relationships, such as image segmentation and target detection.
[0236] Corresponding to the aforementioned base detection method and method for training a preset detection model, the present invention also provides a base detection device and a device for training a preset detection model. Since the device embodiments of the present invention correspond to the aforementioned method embodiments, details not disclosed in the device embodiments can be referred to the aforementioned method embodiments and will not be further described in this invention.
[0237] Figure 10 A schematic diagram of the structure of a base detection device provided in an embodiment of the present disclosure is shown in FIG. Figure 10 Shown, including:
[0238] The interception unit 1001 is configured to obtain a first preset amount of image data in each field of view, and intercept the image data based on position information of the data to be detected in the image data to obtain a second preset amount of image block data, wherein each of the image block data contains a third preset amount of data to be detected;
[0239] The conversion unit 1002 is configured to perform data conversion processing on the second preset number of image block data respectively according to a pre-trained preset detection model to obtain a to-be-detected conversion vector of a first preset dimension corresponding to each of the to-be-detected data;
[0240] An extraction unit 1003 is configured to perform feature extraction processing on the to-be-detected conversion vectors based on the pre-trained preset detection model to obtain a to-be-detected feature vector corresponding to each to-be-detected conversion vector;
[0241] A dimensionality reduction unit 1004 is configured to perform dimensionality reduction processing on the feature vector to be detected based on the pre-trained preset detection model to obtain a corresponding vector to be classified of a second preset dimension;
[0242] The classification unit 1005 is configured to perform base classification according to the vector to be classified to obtain a base detection result.
[0243] The base detection device provided by the present disclosure obtains a first preset number of image data in each field of view, and intercepts the image data according to the position information of the data to be detected in the image data to obtain a second preset number of image block data, wherein each of the image block data contains a third preset number of data to be detected; the second preset number of image block data are respectively subjected to data conversion processing according to a pre-trained preset detection model to obtain a first preset dimension of the to-be-detected conversion vector corresponding to each of the data to be detected; based on the pre-trained preset detection model, the to-be-detected conversion vector is subjected to feature extraction processing to obtain a feature vector to be detected corresponding to each of the to-be-detected conversion vectors; based on the pre-trained preset detection model, the to-be-detected feature vector is subjected to dimensionality reduction processing to obtain a corresponding second preset dimension of the to-be-detected vector to be classified, and base classification is performed according to the to-be-classified vector to obtain a base detection result. Compared with the related art, the embodiment of the present disclosure realizes data modeling processing of the original image data by adopting a preset detection model, enhances the feature extraction capability of the data features, reduces the computational cost, and improves the accuracy of base detection.
[0244] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 11As shown, the interception unit 1001 includes:
[0245] An acquisition module 10011 is configured to acquire an index of each of the to-be-detected data in the image data, wherein the index is used to determine position information of the to-be-detected data in the image data;
[0246] A determination module 10012 is configured to determine position information of the data to be detected in the image data according to each index;
[0247] The interception module 10013 is used to intercept the unit image data containing the third preset number of data to be detected in the image data based on the position information and with each of the data to be detected as the center, until all the data to be detected are processed to obtain the second preset number of image block data.
[0248] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 11 As shown, the conversion unit 1002 includes:
[0249] The processing module 10021 is configured to perform a convolution operation on the image block data through a convolution layer, and then perform activation processing using a preset activation function to obtain a data vector to be detected of a first preset dimension corresponding to each data to be detected;
[0250] Configuration module 10022 is used to randomly assign a position coding vector to each of the data to be detected, and add the data vector to be detected corresponding to the data to be detected and the assigned position coding vector to obtain the transformation vector to be detected, and the position coding vector is used to determine the spatial feature information in the transformation vector to be detected.
[0251] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 11 As shown, the extraction unit 1003 includes:
[0252] A first calculation module 10031 is configured to perform self-attention calculation on the to-be-detected conversion vector through a preset attention algorithm to obtain an attention vector;
[0253] Processing module 10032, configured to input the attention vector into a first preset fully connected structure, perform data operation on the attention vector and first fully connected parameters, and then activate the attention vector using a second preset function to obtain a first feature vector, wherein the first fully connected parameters are used to extract features from the attention vector;
[0254] The processing module 10032 is further configured to input the first feature vector into a second preset fully connected structure, perform data operation on the first feature vector and a second fully connected parameter to obtain a second feature vector, wherein the second fully connected parameter is used to further extract features from the first feature vector;
[0255] The second calculation module 10033 is configured to calculate the second eigenvector using a preset residual function, and then add the calculated second eigenvector to the second eigenvector to obtain the eigenvector to be detected.
[0256] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 11 As shown, the dimension reduction unit 1004 includes:
[0257] An operating module 10041 is configured to perform a regularization operation on the feature vector to be detected using a preset regularization algorithm to obtain a regularized feature vector;
[0258] The dimensionality reduction module 10042 is configured to perform dimensionality reduction processing on the regularized feature vector through a third preset fully connected structure to obtain a vector to be classified of the second preset dimension.
[0259] Figure 12 A schematic diagram of a structure of a training device for a preset detection model provided in an embodiment of the present disclosure, such as Figure 12 Shown, including:
[0260] a clipping unit 1201 configured to obtain a fourth preset number of training image data in each field of view, and clip the training image data based on position information of the training data to be detected in the training image data to obtain a fifth preset number of training image block data, wherein each of the image block data contains a sixth preset number of training data to be detected;
[0261] An acquiring unit 1202 is configured to acquire a true label corresponding to each of the training data to be detected, where the true label is a known base detection result of the training data to be detected;
[0262] The conversion unit 1203 is configured to perform data conversion processing on the fifth preset number of training image block data according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each training to-be-detected data;
[0263] The extraction unit 1204 is configured to perform feature extraction processing on the training transformation vectors to be detected based on the preset detection model to obtain a training feature vector to be detected corresponding to each training transformation vector to be detected;
[0264] A dimensionality reduction unit 1205 is configured to perform dimensionality reduction processing on the training feature vector to be detected based on the preset detection model to obtain a corresponding training vector to be classified of a fourth preset dimension;
[0265] A classification unit 1206 is configured to perform base classification based on the training vector to be classified to obtain a training base detection result;
[0266] The optimization unit 1207 is used to optimize the preset detection model based on the training base detection results and the true labels through a preset loss function and a preset optimization function until a preset convergence condition is met, thereby obtaining a trained preset detection model.
[0267] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 13 As shown, the interception unit 1201 includes:
[0268] An acquisition module 12011 is configured to acquire an index of each of the training data to be detected in the training image data, wherein the index is used to determine position information of the training data to be detected in the training image data;
[0269] a determination module 12012, configured to determine position information of the training data to be detected in the training image data according to each index;
[0270] The interception module 12013 is used to intercept the training unit image data containing the sixth preset number of training data to be detected in the training image data based on the position information and with each of the training data to be detected as the center, until all the training data to be detected are processed to obtain the fifth preset number of training image block data.
[0271] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 13 As shown, the conversion unit 1203 includes:
[0272] The processing module 12031 is configured to perform a convolution operation on the training image block data through a convolution layer, and then perform activation processing using a preset activation function to obtain a training data vector to be detected of a third preset dimension corresponding to each training data to be detected;
[0273] Configuration module 12032 is used to randomly assign a training position coding vector to each of the training data to be detected, and add the training data vector to be detected corresponding to the training data to be detected and the assigned training position coding vector to obtain the training transformation vector to be detected, and the training position coding vector is used to determine the spatial feature information in the training transformation vector to be detected.
[0274] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 13 As shown, the extraction unit 1204 includes:
[0275] A first calculation module 12041 is configured to perform self-attention calculation on the training to-be-detected conversion vector through a preset attention algorithm to obtain a training attention vector;
[0276] Processing module 12042 is configured to input the training attention vector into a first preset fully connected structure, perform data operation with the first training fully connected parameters, and then activate the structure using a second preset function to obtain a first training feature vector, wherein the first training fully connected parameters are used to perform feature extraction on the training attention vector;
[0277] The processing module 12042 is further configured to input the first training feature vector into a second preset fully connected structure, perform data operation on the first training feature vector and a second training fully connected parameter to obtain a second training feature vector, wherein the second training fully connected parameter is used to further extract features from the first training feature vector.
[0278] The second calculation module 12043 is configured to calculate the second training feature vector using a preset residual function, and then perform addition calculation on the second training feature vector to obtain the training feature vector to be detected.
[0279] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 13 As shown, the dimension reduction unit 1205 includes:
[0280] An operation module 12051 is configured to perform a regularization operation on the training feature vector to be detected using a preset regularization algorithm to obtain a training regularized feature vector;
[0281] The dimensionality reduction module 12052 is configured to perform dimensionality reduction processing on the training regularized feature vector through a third preset fully connected structure to obtain a training vector to be classified of the fourth preset dimension.
[0282] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present disclosure, and the principles are the same, which is no longer limited in the embodiment of the present disclosure.
[0283] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0284] Figure 14 A schematic block diagram of an example electronic device 1400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0285] like Figure 14 As shown, the device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1402 or a computer program loaded from a storage unit 1408 into a RAM (Random Access Memory) 1403. Various programs and data required for the operation of the device 1400 can also be stored in the RAM 1403. The computing unit 1401, the ROM 1402, and the RAM 1403 are connected to each other via a bus 1404. An I / O (Input / Output) interface 1405 is also connected to the bus 1404.
[0286] Various components in device 1400 are connected to I / O interface 1405, including: an input unit 1406, such as a keyboard, mouse, etc.; an output unit 1407, such as various types of displays, speakers, etc.; a storage unit 1408, such as a magnetic disk, optical disk, etc.; and a communication unit 1409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1409 allows device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0287] The computing unit 1401 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units for running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 1401 performs the various methods and processes described above, such as a method for base detection or a training method for a preset detection model. For example, in some embodiments, the method for base detection or a training method for a preset detection model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 1408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed by computing unit 1401, one or more steps of the method described above may be performed. Alternatively, in other embodiments, computing unit 1401 may be configured to perform the aforementioned base detection method or the method for training a preset detection model by any other appropriate means (e.g., via firmware).
[0288] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0289] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0290] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0291] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0292] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0293] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0294] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0295] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0296] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for base detection, characterized in that: include: Acquire a first preset amount of image data in each field of view, and intercept and process the image data according to position information of the data to be detected in the image data to obtain a second preset amount of image block data, wherein each of the image block data contains a third preset amount of data to be detected; Performing data conversion processing on the second preset number of image block data respectively according to a pre-trained preset detection model to obtain a to-be-detected conversion vector of a first preset dimension corresponding to each of the to-be-detected data; Based on the pre-trained preset detection model, feature extraction processing is performed on the conversion vectors to be detected to obtain a feature vector to be detected corresponding to each conversion vector to be detected; Based on the pre-trained preset detection model, the feature vector to be detected is subjected to dimensionality reduction processing to obtain a corresponding vector to be classified of a second preset dimension, and base classification is performed according to the vector to be classified to obtain a base detection result.
2. The method according to claim 1, characterized in that The intercepting and processing the image data according to the position information of the data to be detected in the image data to obtain the second preset number of image block data includes: Obtaining an index of each of the to-be-detected data in the image data, wherein the index is used to determine position information of the to-be-detected data in the image data; Determine position information of the data to be detected in the image data according to each index; Based on the position information, with each of the data to be detected as the center, in the image data, unit image data containing the third preset number of data to be detected is intercepted until all the data to be detected are processed to obtain the second preset number of image block data.
3. The method according to claim 1, characterized in that The step of performing data conversion processing on the second preset number of image block data respectively according to the pre-trained preset detection model to obtain a detection conversion vector of the first preset dimension corresponding to each of the detection data includes: After the image block data is convolved through a convolution layer, activation processing is performed through a preset activation function to obtain a data vector to be detected of a first preset dimension corresponding to each data to be detected; A position coding vector is randomly assigned to each of the data to be detected, and the data vector to be detected corresponding to the data to be detected is added to the assigned position coding vector to obtain the conversion vector to be detected, and the position coding vector is used to determine the spatial feature information in the conversion vector to be detected.
4. The method according to claim 1, wherein The performing feature extraction processing on the to-be-detected conversion vector based on the pre-trained preset detection model to obtain a to-be-detected feature vector corresponding to each to-be-detected conversion vector includes: Performing self-attention calculation on the transformation vector to be detected through a preset attention algorithm to obtain an attention vector; Inputting the attention vector into a first preset fully connected structure, performing data operation with a first fully connected parameter, and activating the structure through a second preset function to obtain a first feature vector, wherein the first fully connected parameter is used to extract features from the attention vector; Inputting the first feature vector into a second preset fully connected structure, performing data operation with a second fully connected parameter to obtain a second feature vector, wherein the second fully connected parameter is used to further extract features from the first feature vector; After the second eigenvector is calculated using a preset residual function, the second eigenvector is added to the first eigenvector to obtain the eigenvector to be detected.
5. The method according to claim 1, wherein The step of performing dimensionality reduction processing on the feature vector to be detected based on the pre-trained preset detection model to obtain a corresponding vector to be classified of a second preset dimension includes: Performing a regularization operation on the feature vector to be detected by a preset regularization algorithm to obtain a regularized feature vector; The regularized feature vector is subjected to dimensionality reduction processing through a third preset fully connected structure to obtain a vector to be classified of the second preset dimension.
6. A method for training a preset detection model, characterized in that: The preset detection model can be applied to the method according to any one of claims 1 to 5, including: Acquiring a fourth preset quantity of training image data in each field of view, and intercepting the training image data based on position information of the training data to be detected in the training image data to obtain a fifth preset quantity of training image block data, wherein each of the image block data contains a sixth preset quantity of training data to be detected; Obtaining a true label corresponding to each of the training data to be detected, where the true label is a known base detection result of the training data to be detected; Performing data conversion processing on the fifth preset number of training image block data according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each training to-be-detected data; Based on the preset detection model, feature extraction processing is performed on the training conversion vectors to be detected to obtain a training feature vector to be detected corresponding to each training conversion vector to be detected; Based on the preset detection model, dimensionality reduction processing is performed on the training feature vector to be detected to obtain a corresponding training vector to be classified of a fourth preset dimension, and base classification is performed according to the training vector to be classified to obtain a training base detection result; Based on the training base detection results and the true labels, the preset detection model is optimized by a preset loss function and a preset optimization function until a preset convergence condition is met, thereby obtaining a trained preset detection model.
7. The method according to claim 6, characterized in that The intercepting and processing the training image data according to the position information of the training data to be detected in the training image data to obtain a fifth preset number of training image block data includes: Obtaining an index of each of the training data to be detected in the training image data, wherein the index is used to determine position information of the training data to be detected in the training image data; Determine position information of the training data to be detected in the training image data according to each index; Based on the position information, with each of the training data to be detected as the center, the training unit image data containing the sixth preset number of training data to be detected is intercepted from the training image data until all the training data to be detected are processed to obtain the fifth preset number of training image block data.
8. The method according to claim 6, characterized in that The performing data conversion processing on the fifth preset number of training image block data according to the preset detection model to obtain a training to-be-detected conversion vector of a third preset dimension corresponding to each training to-be-detected data comprises: After the training image block data is convolved through a convolution layer, activation processing is performed through a preset activation function to obtain a training data vector to be detected of a third preset dimension corresponding to each training data to be detected; A training position coding vector is randomly assigned to each of the training data to be detected, and the training data vector to be detected corresponding to the training data to be detected is added to the assigned training position coding vector to obtain the training transformation vector to be detected, and the training position coding vector is used to determine the spatial feature information in the training transformation vector to be detected.
9. The method according to claim 6, characterized in that The performing feature extraction processing on the training to-be-detected conversion vectors based on the preset detection model to obtain a training to-be-detected feature vector corresponding to each training to-be-detected conversion vector comprises: Performing self-attention calculation on the training transformation vector to be detected through a preset attention algorithm to obtain a training attention vector; Inputting the training attention vector into a first preset fully connected structure, performing data operation with the first training fully connected parameters, and activating the training attention vector through a second preset function to obtain a first training feature vector, wherein the first training fully connected parameters are used to extract features from the training attention vector; Inputting the first training feature vector into a second preset fully connected structure, performing data operation with a second training fully connected parameter to obtain a second training feature vector, wherein the second training fully connected parameter is used to further extract features from the first training feature vector; After the second training feature vector is calculated using a preset residual function, the feature vector is added to the second training feature vector to obtain the feature vector to be detected for training.
10. The method according to claim 6, characterized in that The step of performing dimensionality reduction processing on the training feature vector to be detected based on the preset detection model to obtain a corresponding training vector to be classified of a fourth preset dimension includes: Performing a regularization operation on the training feature vector to be detected by a preset regularization algorithm to obtain a training regularized feature vector; The training regularized feature vector is subjected to dimensionality reduction processing through a third preset fully connected structure to obtain a training vector to be classified of the fourth preset dimension.
11. A base detection device, characterized in that: include: a capture unit, configured to acquire a first preset amount of image data in each field of view, and perform capture processing on the image data based on position information of the data to be detected in the image data to obtain a second preset amount of image block data, wherein each of the image block data contains a third preset amount of data to be detected; a conversion unit, configured to perform data conversion processing on each of the second preset number of image block data according to a pre-trained preset detection model, to obtain a to-be-detected conversion vector of a first preset dimension corresponding to each of the to-be-detected data; An extraction unit, configured to perform feature extraction processing on the conversion vectors to be detected based on the pre-trained preset detection model, to obtain a feature vector to be detected corresponding to each conversion vector to be detected; A dimensionality reduction unit, configured to perform dimensionality reduction processing on the feature vector to be detected based on the pre-trained preset detection model to obtain a corresponding vector to be classified of a second preset dimension; The classification unit is used to perform base classification according to the vector to be classified to obtain a base detection result.
12. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 5 or the method of any one of claims 6 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5 or the method according to any one of claims 6 to 10.
14. A computer program product, characterized in that The method comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 5 or the method according to any one of claims 6 to 10.