Hash code generation method and device, equipment, storage medium and product

The hash encoding is generated through regular autoencoder and multi-head gated attention network model, which solves the problems of high hash index conflict rate and low data calculation efficiency in the database all-in-one environment, and realizes efficient data retrieval and resource utilization.

CN120179651APending Publication Date: 2025-06-20CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510323302.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing technology faces the problems of high hash index conflict rate and low data calculation efficiency in the database all-in-one environment, and cannot fully utilize the multi-dimensional characteristics of the original data, resulting in a decrease in data retrieval efficiency and insufficient system resource utilization.

Method used

Data feature pre-extraction and discretization are performed by pre-constructing regular autoencoder, combining the pre-trained multi-head gated attention network model to learn multi-dimensional subspace data features, form common and personal feature encoding, and perform linear stitching to generate hash encoding and store it in a hash bucket.

Benefits of technology

It improves the discriminant nature of data sample hash encoding, reduces hash index conflicts, optimizes data computing efficiency, and improves the resource utilization rate of database all-in-one system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179651A_ABST
    Figure CN120179651A_ABST
Patent Text Reader

Abstract

The invention discloses a Hash code generation method and device, equipment, a storage medium and a product, and relates to the technical field of data processing, and the method comprises the steps: carrying out the data feature pre-extraction and discretization processing of an original data sample obtained in advance through a regular auto-encoder, and obtaining a common feature coded value; learning multi-dimensional subspace data features of the sampled data samples through a multi-head gating attention network model to obtain subspace personality feature codes of the sampled data samples; and performing linear splicing on the common feature coding value and the subspace personalized feature coding value corresponding to the sampled data sample to form a hash code and storing the hash code in the preset hash bucket, thereby obtaining the feature discrete value with high difference to form the hash code through the method, improving the discrimination of the data sample hash code, reducing the hash index conflict, and improving the accuracy of the data sample hash code. Therefore, the effects of improving data calculation efficiency and reducing resource consumption are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method, apparatus, device, storage medium, and product for generating hash codes. Background Art

[0002] In the fields of cloud computing, big data, and artificial intelligence, data retrieval and index optimization are key technical challenges. With the explosive growth of data volume, traditional data retrieval methods face problems of low efficiency and high hash conflict rate. Existing data retrieval methods are mainly divided into three categories: heuristic rules, machine learning, and deep learning methods. The heuristic rules method achieves good retrieval results by iteratively optimizing individual bits of the hash code, but this method often relies on specific optimization algorithms and lacks flexibility and generalization ability. The machine learning method constructs a hash index model by learning the cumulative distribution function of keywords, but this method may not be able to fully capture the deep features of the data. The deep learning method extracts the deep semantics of multi-modal data through a neural network and encodes them into hash codes. Although there is progress in feature extraction, there are still problems of hash conflict and low data calculation efficiency.

[0003] Existing solutions face challenges of high hash index conflict rate and low data calculation efficiency when dealing with large-scale data, especially in the database all-in-one environment. For example, although the prediction method based on structured data improves the accuracy and efficiency of target prediction, it does not solve the problems of hash conflict and data retrieval. The learning-based index method based on time series data features solves the problems of memory occupation and indexing time, but fails to extract high-quality features of the original data and is limited to time series data. The database query processing and optimization method based on artificial intelligence technology improves the hash coding accuracy and query efficiency, but is mainly oriented to image-text multi-modal data and is not applicable to large-scale data environments.

[0004] In summary, existing solutions either cannot make full use of the multi-dimensional features of the original data, are inefficient when dealing with large-scale data, or have limited effect in reducing hash conflicts. These problems lead to a decline in data retrieval efficiency and insufficient utilization of system resources.

[0005] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a method, apparatus, device, storage medium, and computer program product for generating hash codes, aiming to reduce hash conflicts in the database all-in-one environment and improve data calculation efficiency.

[0007] To achieve the above object, this application proposes a method for generating hash codes, and the method includes:

[0008] Pre - extract and discretize the data features of the pre - obtained original data samples through a pre - constructed regular auto - encoder to obtain the common feature coding values of the original data samples;

[0009] Through a pre - trained multi - head gated attention network model, learn the multi - dimensional subspace data features of the sampled data samples in the original data samples to obtain the subspace personality feature coding of the sampled data samples;

[0010] Linearly splice the common feature coding values and the subspace personality feature coding values corresponding to the sampled data samples to form a hash code and store it in a preset hash bucket.

[0011] In one embodiment, the multi - head gated attention network model includes a multi - head gated attention layer and a single - head attention layer. The step of learning the multi - dimensional subspace data features of the sampled data samples in the original data samples through the pre - trained multi - head gated attention network model to obtain the subspace personality feature coding of the sampled data samples includes:

[0012] Sample the original data samples according to a pre - set data sample sampling method to obtain the sampled data samples;

[0013] Based on the hidden - layer space representation corresponding to the sampled data samples, calculate the output vector of the multi - head attention layer of the sampled data samples through the multi - head gated attention layer. The hidden - layer space representation is obtained by compressing and mapping the sampled data samples through a pre - constructed regular auto - encoder;

[0014] Based on the output vector of the multi - head attention layer, calculate the aggregated information representation of the sampled data samples through the single - head attention layer. The aggregated information representation includes several feature dimension values;

[0015] Discretize the several feature dimension values of the aggregated information representation to obtain the subspace personality feature coding of the sampled data samples.

[0016] In one embodiment, the step of calculating the output vector of the multi - head attention layer of the sampled data samples through the multi - head gated attention layer based on the hidden - layer space representation corresponding to the sampled data samples includes:

[0017] Based on the hidden - layer space representation corresponding to the sampled data samples, calculate the multi - head gated vector corresponding to the sampled data samples;

[0018] Based on the hidden - layer space representation, calculate the weighted summation coefficient in the convolution process;

[0019] Calculate the output vector of the multi-head attention layer corresponding to the sampled data sample according to the multi-head gating vector and the weighted summation coefficient.

[0020] In one embodiment, the step of sampling the original data sample according to a preset data sample sampling method to obtain the sampled data sample includes:

[0021] Traverse the original data sample to obtain the sample index number of the current sampled sample;

[0022] If the sample index number is not greater than a preset sampling threshold, save the current sampled sample in a preset sample array;

[0023] If the sample index number is greater than the preset sampling threshold, generate a current sample random number not greater than the sample index number;

[0024] If the current sample random number is not greater than the preset sampling threshold, replace the array element corresponding to the current sample random number in the sample array with the current sampled sample;

[0025] If the sample index number is the maximum sample index number, end the traversal to obtain the sampled data sample.

[0026] In one embodiment, the step of pre-extracting data features and discretizing the pre-obtained original data sample through a pre-constructed regular autoencoder to obtain the common feature coding value of the original data sample includes:

[0027] Represent the original data sample as a high-dimensional input matrix;

[0028] Map the high-dimensional input matrix to a real number vector of a preset dimension through the input layer of the regular autoencoder to obtain the hidden layer space representation corresponding to the original data sample;

[0029] Discretize the hidden layer space representation corresponding to the original data sample to obtain the common feature coding value of the original data sample.

[0030] In one embodiment, the method further includes:

[0031] Receive and parse a system operation instruction to perform a syntax accuracy analysis on the input text corresponding to the operation instruction;

[0032] After the syntax accuracy analysis passes, convert the system operation instruction into a symbol stream;

[0033] Build an abstract syntax tree based on the symbol stream to establish an operation task by traversing the abstract syntax tree;

[0034] Transfer the operation task to the operation function for execution and return the operation result to the user through the user interface layer.

[0035] In addition, to achieve the above object, the present application also proposes a hash code generation device, which includes:

[0036] A processing module, configured to perform data feature pre-extraction and discretization processing on the pre-obtained original data samples through a pre-constructed regular autoencoder to obtain the common feature coding values of the original data samples;

[0037] A learning module, configured to learn the multi-dimensional subspace data features of the sampled data samples in the original data samples through a pre-trained multi-head gated attention network model to obtain the subspace personality feature coding of the sampled data samples;

[0038] A splicing module, configured to linearly splice the common feature coding values and the subspace personality feature coding values corresponding to the sampled data samples to form a hash code and store it in a preset hash bucket.

[0039] In addition, to achieve the above object, the present application also proposes a hash code generation device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the hash code generation method as described above.

[0040] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the hash code generation method as described above.

[0041] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the hash code generation method as described above.

[0042] One or more technical solutions proposed in this application pre - extract and discretize data features from pre - acquired original data samples through a pre - constructed regular auto - encoder to obtain the common feature encoding values of the original data samples; through a pre - trained multi - head gated attention network model, learn the multi - dimensional subspace data features of the sampled data samples in the original data samples to obtain the subspace individual feature encoding of the sampled data samples; linearly splice the common feature encoding values and the subspace individual feature encoding values corresponding to the sampled data samples to form a hash code and store it in a preset hash bucket. By extracting and learning the features of data samples based on the regular auto - encoder and the multi - head gated attention mechanism network model, making full use of the multi - dimensional subspace features of the original data and the parallelism of matrix operations, obtaining highly - differentiated feature discrete values to form a hash code, enhancing the discriminability of the hash code of data samples, reducing hash index conflicts, and thus achieving the effect of improving data calculation efficiency and reducing resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and, together with the specification, used to explain the principles of the present application.

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the hash code generation method of the present application;

[0046] Figure 2 It is a schematic diagram of the system implementation architecture of a hash code generation method provided in an embodiment of the present application;

[0047] Figure 3 It is a schematic flowchart provided for Embodiment 2 of the hash code generation method of the present application;

[0048] Figure 4 It is a schematic diagram of the module structure of the hash code generation device in an embodiment of the present application;

[0049] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the hash code generation method in an embodiment of the present application.

[0050] The realization of the purpose, functional features, and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not used to limit the present application.

[0052] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0053] The main solution of the embodiments of the present application is: to propose a learning-based hash indexing method and device based on a multi-head gated attention network, aiming to solve the problems of large-scale data index construction and optimization in the database all-in-one machine environment. The solution of the present application first uses a regular autoencoder to pre-extract features from the original input data, retains the core information of the data, and obtains the common feature coding values of the data samples through discretization processing. This step helps to reduce data redundancy and noise, improve the training effect of the hash indexing model, reduce storage occupancy, and reduce system pressure.

[0054] Among them, this solution uses the reservoir sampling method to sample the data and constructs a multi-head gated attention network model to train the learning-based hash function. This model can learn the data features of multi-dimensional subspaces from different spatial perspectives during the training process, obtain the unique and rich topological structure information of the data, and through the discretization processing of the multi-dimensional subspace data features, obtain the subspace individual feature coding values of the data samples, which helps to improve the discriminability of the hash coding of the data samples and reduce the hash index conflict.

[0055] In addition, the solution provided by the present application linearly concatenates the common feature coding values and the subspace individual feature coding values to form the complete hash coding of the data samples and stores them in the hash bucket. During the query and insertion operations, by calculating the hash values and handling hash conflicts, the fast retrieval and update of the data are realized. This method not only optimizes the data calculation efficiency but also improves the utilization rate of the system resources of the database all-in-one machine, effectively solving the problems of high hash conflict rate and low data calculation efficiency in the prior art. Through this deep learning-based technical solution, index optimization is achieved, the data query performance is improved, and the high-efficiency operation of the system is maintained at the same time.

[0056] Since the existing technologies face challenges of high hash index conflict rate and low data calculation efficiency when dealing with large-scale data, especially in the database all-in-one machine environment. For example, although the prediction method based on structured data improves the accuracy and efficiency of target prediction, it does not solve the problems of hash conflict and data retrieval. Although the learning-based index method based on time series data features solves the problems of memory occupation and indexing time, it fails to extract high-quality features of the original data and is limited to time series data. Although the database query processing and optimization method based on artificial intelligence technology improves the hash coding accuracy and query efficiency, it is mainly oriented to image-text multimodal data and is not applicable to large-scale data environments.

[0057] In summary, the existing solutions either cannot fully utilize the multi-dimensional features of the original data, or are inefficient in processing large-scale data, or have limited effects in reducing hash conflicts. These problems lead to a decline in data retrieval efficiency and insufficient utilization of system resources.

[0058] This application provides a solution. Through a pre-constructed regular autoencoder, pre-extraction and discretization processing of data features are performed on pre-acquired original data samples to obtain the common feature coding values of the original data samples; through a pre-trained multi-head gated attention network model, multi-dimensional subspace data features of the sampled data samples in the original data samples are learned to obtain the subspace individual feature coding of the sampled data samples; the common feature coding values and the subspace individual feature coding values corresponding to the sampled data samples are linearly spliced to form hash codes and stored in a preset hash bucket. By extracting and learning the features of data samples based on the regular autoencoder and the multi-head gated attention mechanism network model, the multi-dimensional subspace features of the original data and the parallelism of matrix operations are fully utilized, and highly discriminative feature discrete values are obtained, thereby improving the discriminability of the hash codes of data samples. While reducing hash conflicts and optimizing data calculation efficiency, it also improves the utilization rate of system resources of the database all-in-one machine system.

[0059] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a hash code generation system, etc. that can implement the above functions. Hereinafter, taking the hash code generation system as an example, this embodiment and the following embodiments will be described.

[0060] Based on this, an embodiment of this application provides a hash code generation method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the hash code generation method of this application.

[0061] In this embodiment, the hash code generation method includes steps S1000 to S3000:

[0062] Step S1000: Perform data feature pre-extraction and discretization processing on the pre-acquired original data samples through a pre-constructed regular autoencoder to obtain the common feature coding values of the original data samples;

[0063] It should be noted that in order to solve some existing technical solutions for index optimization based on machine learning and deep learning algorithms, due to the large data scale in the database all-in-one machine environment, problems such as high hash index conflict rate and low data calculation efficiency are brought. In the embodiments of this application, a learning-based hash index method based on a multi-head gated attention network is proposed, which makes full use of the multi-dimensional subspace features of the original data and the parallelism of matrix operations to obtain highly different feature discrete values, thereby improving the discriminability of the hash codes of data samples, reducing hash index conflicts, and improving data calculation efficiency. At the same time, the method in the embodiments of this application also improves the utilization rate of system resources of the database all-in-one machine system to a certain extent.

[0064] Specifically, referring to Figure 2 , Figure 2 which is a schematic diagram of the system implementation architecture of a hash code generation method provided by the embodiments of this application. Briefly speaking, the technical solution of the embodiments of this application can be realized in the following way. First, use a regular autoencoder to perform data feature pre-extraction on the original data samples and retain the core information of the most important part of the original data. After the original input data (i.e., the original data samples) undergo feature extraction, discretization processing is performed to obtain the common feature coding values of the data samples. This compressed and compact data feature can not only distinguish the main difference information of the original data, but also help reduce storage occupancy and reduce system pressure.

[0065] Secondly, use the reservoir sampling method to sample the original data samples and use a multi-head gated attention network to train a learning-based hash index model, so that the model can obtain data features in multi-dimensional subspaces through learning from different spatial perspectives for the same data sample during the training process, and maximize the acquisition of unique and rich topological structure information of the data. After discretization processing of the multi-dimensional subspace data features, the subspace individual feature coding values of the data samples are obtained, and at the same time, the data in the hash table is dynamically sampled to re-train the model before the hash table needs to be expanded to obtain a more accurate model.

[0066] Then, perform a linear splicing operation on the common feature coding values of the data samples learned through the regular autoencoder and the subspace individual feature coding values of the data samples learned through the multi-head gated attention network model, and finally form the complete hash code of the data sample and store it in the hash bucket.

[0067] Then, receive and process the query instructions passed by the system query, and perform read or write operations on the data in the hash bucket. The query parser is responsible for parsing the operation instructions passed by the system query. First, analyze the query text to verify the accuracy of the syntax, and then extract the operation type and data. The parser passes the parsed operation type, after data encapsulation, to the operation function to perform the actual query processing.

[0068] Finally, perform operations such as data loading, data updating, data deletion, and data query for interaction with system users. After the query execution is completed, the user interface layer returns the query result to the user, providing necessary feedback information.

[0069] It should be noted that in the embodiments of the present application, the regular autoencoder is actually a neural network structure. It compresses high-dimensional input data into a low-dimensional representation through an encoder, and then restores the low-dimensional representation into high-dimensional data through a decoder. The purpose of the regular autoencoder is to extract a low-dimensional representation from the original data samples that can represent their core features, that is, the common feature coding value. The common feature coding value refers to the common features of the data samples extracted by the regular autoencoder. These features can represent the main information of the data samples and are used for subsequent hash code generation. The discretization process refers to converting continuous numerical features into discrete binary codes for the generation and comparison of hash codes.

[0070] In addition, it should be noted that in a possible implementation manner, the regular autoencoder can be pre-trained through an unsupervised learning method to ensure that it can capture the main features of the data samples. The pre-extracted data features will then be converted into binary code form through specific discretization algorithms, such as k-means clustering or threshold segmentation. For example, in a specific implementation manner, the system represents the received original data samples as a high-dimensional input matrix, and then maps the high-dimensional input matrix to a real number vector of a preset dimension through the input layer of the regular autoencoder to obtain the hidden layer space representation. Then, the system performs discretization processing on the hidden layer space representation to finally obtain the common feature coding value of the original data samples.

[0071] Step S2000: Learn the multi-dimensional subspace data features of the sampled data samples in the original data samples through the pre-trained multi-head gated attention network model to obtain the subspace individual feature coding of the sampled data samples;

[0072] It should be noted that in the embodiments of the present application, the multi-head gated attention network model refers to a deep learning model containing multiple attention mechanisms, where the multi-head gated attention layer is responsible for extracting features from multiple subspaces, and the single-head attention layer further aggregates these features to obtain a more discriminative feature representation. The sampled data sample refers to a part of the data selected from the original data sample according to certain rules for training the multi-head gated attention network model. The subspace personality feature encoding refers to the encoding value learned from the sampled data sample that can represent its personality features in the multi-dimensional subspace.

[0073] In addition, it should be noted that in a possible implementation manner, the multi-head gated attention network model can be trained by the backpropagation algorithm to minimize the difference between the predicted value and the actual value. After the model training is completed, it can extract features from the sampled data sample to obtain the subspace personality feature encoding.

[0074] For example, in a specific implementation manner, the system first samples the original data sample according to the preset data sample sampling method to obtain the sampled data sample. Then, based on the hidden layer space representation corresponding to the sampled data sample, the multi-head attention layer output vector corresponding to the sampled data sample is calculated through the multi-head gated attention layer. Next, based on the multi-head attention layer output vector, the aggregated information representation corresponding to the sampled data sample is calculated through the single-head attention layer. Finally, the discrete processing is performed on several feature dimension values of the aggregated information representation to obtain the subspace personality feature encoding of the sampled data sample.

[0075] Step S3000: Linearly splice the common feature encoding value and the subspace personality feature encoding value corresponding to the sampled data sample to form a hash code and store it in a preset hash bucket.

[0076] It should be noted that linear splicing refers to connecting two or more sequences in a certain order to form a longer sequence. In this step, the purpose of linear splicing is to combine the common feature encoding value and the subspace personality feature encoding value to form a hash code that can uniquely identify the data sample. The hash bucket refers to a data structure used to store and manage hash codes for quick retrieval and access.

[0077] In addition, it should be noted that in a possible implementation manner, linear splicing can be achieved through a simple concatenation operation, or more complex methods such as weighted summation can be used to enhance the discriminability of the encoding. After splicing is completed, the system will store the generated hash code in a preset hash bucket for subsequent data retrieval and management.

[0078] For example, in a specific embodiment, the system linearly concatenates the common feature encoding values and subspace individual feature encoding values corresponding to the sampled data samples to form hash codes. Then, the system stores these hash codes in a preset hash bucket to enable quick location and retrieval of data when a query instruction is received.

[0079] In a feasible embodiment, step S1000 may include steps S1100 to S1300:

[0080] Step S1100: Represent the original data sample as a high-dimensional input matrix;

[0081] Step S1200: Map the high-dimensional input matrix into a real number vector of a preset dimension through the input layer of the regular autoencoder to obtain the hidden layer space representation corresponding to the original data sample;

[0082] Step S1300: Discretize the hidden layer space representation corresponding to the original data sample to obtain the common feature encoding value of the original data sample.

[0083] It should be noted that in this embodiment, the original data sample refers to an unprocessed data set collected from actual applications, and these data sets may contain various types of data, such as images, texts, or structured data, etc. The high-dimensional input matrix refers to representing these original data samples in matrix form, where each row or column represents a feature of the data. The regular autoencoder is a deep learning model that compresses high-dimensional data into a low-dimensional representation through an encoder and then reconstructs the original data as much as possible through a decoder, and its purpose is to learn an effective and compact representation of the data. The hidden layer space representation refers to the representation of the data in the hidden layer after being encoded by the regular autoencoder, which is a compressed form of the original data and retains the core features of the data. The discretization process refers to converting a continuous real number vector into a discrete binary encoding value, and this step is to generate hash codes for easy retrieval and comparison.

[0084] In addition, it should be noted that in this embodiment, the process of pre-extracting and discretizing the features of the original data sample through the regular autoencoder aims to extract the common feature encoding values of the data, and these encoding values can represent the main information of the data sample and are used for subsequent hash code generation. In a possible embodiment, the regular autoencoder can be pre-trained in an unsupervised learning manner to ensure that it can capture the main features of the data sample. The pre-extracted data features will then be converted into binary encoding form through specific discretization algorithms, such as k-means clustering or threshold segmentation, etc.

[0085] In a feasible embodiment, the system first represents the received original data samples as a high-dimensional input matrix, and then maps the high-dimensional input matrix into a real number vector of a preset dimension through the input layer of the regular autoencoder to obtain a hidden layer space representation. Then, the system discretizes the hidden layer space representation to finally obtain the common feature encoding values of the original data samples. These common feature encoding values can not only distinguish the main difference information of the original data, but also help reduce storage occupancy, reduce system pressure, and reserve a subspace personality feature encoding with sufficient bits for the overall hash encoding of a single data sample, which plays an important role in reducing hash conflicts in the overall hash index model. In this way, the system can effectively extract high-quality and highly differentiated data features from large-scale data, reduce hash conflicts, improve data query performance, and improve the resource utilization rate of the database all-in-one system.

[0086] In the embodiments of the present application, the original input data is preprocessed by a feature extraction algorithm, which can reduce data redundancy and data noise to a certain extent to improve the training effect of the overall hash index model. In addition, after feature extraction, "compressed" data features can be obtained. Such compact data features are more conducive to reducing storage occupancy and reducing system pressure. In the embodiments of the present application, an initial feature representation of sample data is obtained by constructing a regular autoencoder. The regular autoencoder can encourage the model to learn other features, in addition to copying the input to the output, without having to limit the use of shallow encoders and decoders and small coding dimensions to limit the capacity of the model, and the latent space representation generated by the regular autoencoder can ensure that the core information of the data is not lost.

[0087] In another feasible embodiment, an encoder is constructed to compress each original data sample into a hidden space representation Z, thereby retaining the core information of the data sample. The input data is represented as an input matrix X. After mapping the high-dimensional input matrix X into a D-dimensional real number vector through the input layer of the encoder θ, the hidden layer space representation Z is obtained. The encoding process is as follows:

[0088] Z = θ(W * X + b)

[0089] Where, W represents the learnable weight matrix between the hidden layer and the output of the encoder, b represents the corresponding bias term, and θ represents a preset encoder function.

[0090] In addition, in the embodiments of the present application, the information of the output layer of the encoder can also be received through the hidden layer of the decoder, and the hidden layer space representation Z is mapped back to the original representation layer, so as to learn and train the reconstructed original data representation M of the data (such as the original data sample). The decoding process is as follows:

[0091] M = σ(W' * Z + b')

[0092] Among them, W′ represents the learnable weight matrix from the decoder hidden layer to the decoder's original representation layer, b′ represents the corresponding bias term, and σ represents the activation function.

[0093] Moreover, the embodiment of the present application can also optimize the regular autoencoder by reconstructing the original data representation M of the original data sample and minimizing the mean square error. In the process of constructing the autoencoder in this article, the mean square error is used as the loss function, and the training process of the autoencoder is optimized by minimizing the mean square error. The specific loss function is as follows:

[0094]

[0095] where sign is the sign function, ||·||2 is the 2-norm, X is the input matrix of the original data sample, and M is the reconstructed original data representation of the original data sample.

[0096] More specifically, due to the fact that the hidden layer space representation Z generated by the regular encoder retains the core features of the original data sample input, and this core feature covers the main information of the sample and can be used as the main discriminant basis for the differences between samples, the embodiment of the present application proposes to discretize the latent space representation Z to obtain the common feature encoding h, and define the common feature encoding as the header encoding value of the final hash encoding of the data sample.

[0097] Discretize each dimension value of the hidden layer representation Z i to obtain the common feature encoding h(Z i ), which is specifically represented as follows.

[0098]

[0099] where, Z i represents the feature value of the i-th dimension of the hidden layer feature Z, and h j represents the encoding value of the j-th dimension of the common feature encoding, that is, the common feature encoding h is finally mapped to a binary encoding string of 0 / 1.

[0100] In summary, in the embodiment of the present application, the regular autoencoder is used to extract the core feature representation of the original data sample, and the hidden layer feature representation is used to be transformed into the common feature encoding to reflect the main differences between data samples. The advantage of using the encoder hidden layer features here is that it not only uses the most compact data to feedback the most differential representation, saving memory resources, but also reserves enough bits of subspace personality feature encoding for the overall hash encoding of a single data sample, which plays an important role in reducing hash conflicts in the overall hash index model.

[0101] In a feasible implementation manner, the multi-head gated attention network model includes a multi-head gated attention layer and a single-head attention layer, and step S2000 may include steps S2100 to S2400:

[0102] Step S2100: Sample the original data sample according to a preset data sample sampling method to obtain the sampled data sample;

[0103] Step S2200: Based on the hidden layer spatial representation corresponding to the sampled data sample, calculate the output vector of the multi-head attention layer corresponding to the sampled data sample through the multi-head gated attention layer, where the hidden layer spatial representation is obtained by compressing and mapping the sampled data sample by a pre-constructed regular autoencoder;

[0104] Step S2300: Based on the output vector of the multi-head attention layer, calculate the aggregated information representation corresponding to the sampled data sample through the single-head attention layer, where the aggregated information representation includes a plurality of feature dimension values;

[0105] Step S2400: Discretize the plurality of feature dimension values of the aggregated information representation to obtain the subspace personality feature encoding of the sampled data sample.

[0106] It should be noted that in this embodiment, the multi-head gated attention network model is a deep learning model that combines a multi-head gated attention layer and a single-head attention layer to extract multi-dimensional features of data. The design of this model aims to capture the features of data in different subspaces through the attention mechanism to enhance the discriminative ability of the model for data features. The sampled data sample refers to a part of the data selected from the entire dataset according to certain rules for training the multi-head gated attention network model. The hidden layer spatial representation is the result of feature extraction of the sampled data sample by the regular autoencoder, which is a low-dimensional data representation that can capture the core features of the data. The output vector of the multi-head attention layer is the result of processing by the multi-head gated attention layer, which contains the features extracted from multiple subspaces. The aggregated information representation is the result of further processing of the output vector of the multi-head attention layer by the single-head attention layer, which is a data representation that synthesizes multiple feature dimensions. The discretization process is to convert the continuous feature dimension values into discrete binary encoding values to facilitate the generation of hash codes.

[0107] In addition, it should be noted that in this embodiment, the process of learning the multi-dimensional subspace data features of the sampled data samples through the multi-head gated attention network model aims to obtain data feature encodings with high discriminability, so as to reduce hash conflicts and improve data retrieval efficiency. In a possible implementation manner, the multi-head gated attention layer can aggregate the data features in different subspaces by calculating the importance of data relationships, and the single-head attention layer further refines these features to obtain a more discriminative feature representation.

[0108] In a feasible implementation manner, the system first samples the original data samples according to a pre-set data sample sampling method to obtain sampled data samples. Then, the system uses a regular autoencoder to extract features from the sampled data samples and compressively maps the sampled data samples to the hidden layer space representation. Next, the system calculates the output vector of the multi-head attention layer corresponding to the sampled data samples through the multi-head gated attention layer, and then calculates the aggregated information representation through the single-head attention layer. Finally, the system discretizes the feature dimension values of the aggregated information representation to obtain the subspace personality feature encoding of the sampled data samples. This process enables the system to effectively extract high-quality and highly differentiated data features from large-scale data, reduce hash conflicts, improve data query performance, and enhance the utilization rate of the system resources of the database all-in-one machine. Through this method, the system can generate hash encodings with high discriminability, thereby accelerating the data retrieval process and optimizing the data index construction and index optimization in the database all-in-one machine environment.

[0109] For example, in another feasible implementation manner, according to the training data set (i.e., the sampled data samples) obtained in advance based on the original data samples, the embodiment of the present application constructs a multi-head gated attention network model to train a learning-based hash function. In this stage, a multi-head gated attention network is designed to aggregate other important data features in different subspaces by calculating the importance of data relationships, making the target data vector more distinguishable. The multi-head gated attention network mainly includes two layers of attention networks, namely the multi-head gated attention layer and the single-head attention layer.

[0110] The classic multi-head attention mechanism captures the local topological structure information of the complex network by aggregating the features of the data in multiple subspaces and performing an averaging operation on them. Since this method uses an averaging operation in the aggregation process and ignores the fact that the importance of the target data and other data in different subspaces is different, it may affect the quality of the data features after the unimportant subspaces are removed.

[0111] To address the above deficiencies, the embodiments of the present application calculate the weight values of the target data and other data in different subspaces by introducing a gating mechanism, enabling more important data information to be aggregated to the target data with higher weights in the subspaces. Specifically, after representing the original data samples as the input matrix X, the single target data Di in X and other data undergo max pooling and average pooling, thereby constructing a convolutional sub-network. The gating value g for each head of the target data Di is obtained through a fully connected layer. i It controls the interaction degree between the target data Di and other data in different subspaces, and learns the importance degree of the relationship between the target data and other data except the target data in different head gating mechanisms, helping to mine important data relationships. The calculation formula of the gating value is as follows:

[0112]

[0113] Where, maps the data features to a d m -dimensional vector, d m is a preset dimensional parameter, σ represents a non-linear activation function, θv is a mapping function that maps the connected features to the gate. When setting a smaller d m , the computational overhead of the gating mechanism can be ignored. represents the data set in the training data T except the target data Di, gi represents the gating vector, Z i represents the hidden layer features of the target data, and Z j are the hidden layer features of the data other than the target data Di.

[0114] Then, the multi-head attention mechanism aggregates data information through multiple subspaces. Each subspace calculates more discriminative features with stronger expressive ability for each data through a linear layer, obtaining the coefficients for weighted summation in each convolutional process. The formula is as follows:

[0115]

[0116] Where, W is the weight matrix, a is the attention coefficient, Z j and Z k are the hidden layer features of the data other than the target data Di.

[0117] Finally, by adding a gating mechanism to the traditional multi-head attention mechanism, the importance degree of the relationship between the data and other data under each subspace is controlled, and data information with different importance levels is aggregated from different subspaces to obtain the data embedding of the first layer network. The calculation method is as follows:

[0118]

[0119] Among them, σ is the non-linear activation function RELU, || represents the concatenation operation, K represents the number of attention heads, and qi is the output vector of the multi-head attention layer.

[0120] The first layer of the network learns the local topological structure information of the data from multiple subspaces, aggregates other multi-descriptive information according to the gating mechanism, and then uses a single-head attention layer to further learn the low-dimensional embedding representation of the data and aggregate the information representation p i The calculation formula is as follows:

[0121]

[0122] In the embodiments of the present application, the purpose of designing the single-head attention layer mainly has three points: First, it can make the aggregated data information in complex dimensions have a slow decline process during dimension compression, ensuring that as little non-critical information as possible is lost and guaranteeing the performance of the overall index model; Second, it lies in ensuring that the model can better focus on the discriminative information in the aggregated features through the attention mechanism; Third, it is convenient to control the discretization process of the feature representation, and the output can be made according to the required dimension of the subspace personality feature coding value.

[0123] In addition, it can be understood that the designed learning-based hash function model in this article is a two-layer multi-head gated attention network, that is, the trained learning-based hash function model is the multi-head gated attention network model, which includes a multi-head attention layer and a single-head attention layer. During the training process, the cross-entropy loss function is used to train the multi-head gated attention network model, and the parameters of the model are iteratively updated:

[0124]

[0125] In the formula, y is the predicted data value, class is the current data category label, and k represents the number of attention heads. The probability that the input data belongs to each type label is obtained through the formula, and then the parameters in the model are iteratively updated by minimizing the loss function.

[0126] Specifically, through the above steps, the model extracts the aggregated information representation p of the data sample i , because this low-dimensional representation retains the most personalized representation learned by the data sample from different subspace perspectives and plays a core role in distinguishing the features between data samples. Therefore, in the embodiments of the present application, p i can be discretized to obtain the subspace personality feature coding value s, and the subspace personality feature coding is defined as the tail part coding value of the final hash coding of the data sample.

[0127] Discretize each dimension value of the aggregated information representation p i to obtain the subspace personality feature coding s(p i ) which is specifically represented as follows.

[0128]

[0129] where p i represents the aggregated information representing p i the eigenvalue of the i-th dimension, s j represents the j-th dimensional coding value of the subspace personality feature coding, that is, the subspace personality feature coding S is finally mapped to a binary coding string of 0 / 1.

[0130] In summary, a multi-head gated attention mechanism network hashing index model proposed in the embodiments of the present application is used to extract the subspace personality features of data samples, comprehensively considers the local topological structure information of samples in different subspaces in multiple aspects and aggregates and refines them according to the importance degree, effectively improving the difference between data features. The subspace personality feature coding obtained by discretization plays a key role in constructing the hashing coding value of data samples subsequently, reducing hash conflicts, effectively improving the discriminability of the hashing coding value of data samples, and thus accelerating the data query efficiency.

[0131] This embodiment provides a hashing coding generation method. Through a pre-constructed regular autoencoder, data feature pre-extraction and discretization processing are performed on pre-acquired original data samples to obtain the common feature coding values of the original data samples; through a pre-trained multi-head gated attention network model, multi-dimensional subspace data features of sampled data samples in the original data samples are learned to obtain the subspace personality feature coding of the sampled data samples; the common feature coding values and the subspace personality feature coding values corresponding to the sampled data samples are linearly spliced to form a hashing code and stored in a preset hash bucket. By extracting and learning the features of data samples based on the regular autoencoder and the multi-head gated attention mechanism network model, the multi-dimensional subspace features of the original data and the parallelism of matrix operations are fully utilized to obtain highly different feature discrete values to form a hashing code, improving the discriminability of the hashing coding of data samples, reducing hash index conflicts, and thus achieving the effects of improving data calculation efficiency and reducing resource consumption.

[0132] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , the step S2100 of the hashing coding generation method includes steps S2110 to S2150:

[0133] Step S2110: Traverse the original data samples to obtain the sample index number of the current sampled sample;

[0134] Step S2120: If the sample index number is not greater than the preset sampling threshold, save the current sampled sample in a preset sample array.

[0135] Step S2130: If the sample index number is greater than the preset sampling threshold, generate a current sample random number not greater than the sample index number.

[0136] Step S2140: If the current sample random number is not greater than the preset sampling threshold, replace the array element corresponding to the current sample random number in the sample array with the current sampled sample.

[0137] Step S2150: If the sample index number is the maximum sample index number, end the traversal and obtain the sampled data sample.

[0138] It should be noted that in this embodiment, in order to deeply process the data characteristics of the original data sample, a multi-head gated attention network is used to train a learning-based hash function. First, a uniform and reasonable sample set is extracted by reservoir sampling of the data in the system; then, a multi-head gated attention model with a double-layer attention network is constructed to train the learning-based hash function; finally, the trained model is used to generate the feature representation in the multi-dimensional subspace of the data sample, and it is discretized to obtain the subspace personality space coding value.

[0139] To ensure that the sampled data sample can represent the characteristics of the overall data set, the embodiment of this application samples the original data sample by the reservoir sampling method to obtain the sampled data sample. Among them, the sample index number refers to the unique identifier of each sample when traversing the original data sample, and the preset sampling threshold is a pre-determined value used to control the selection of samples during the sampling process. The sample array is a data structure used to store the samples selected during the sampling process. The current sample random number is a random value generated in each iteration and is used to determine whether to replace the existing sample in the sample array. The maximum sample index number refers to the total number of samples in the original data sample set, and when the traversal reaches this value, the sampling process ends.

[0140] In addition, it should be noted that the sampling method in this embodiment is a random sampling technique, which ensures that the probability of each sample being selected is equal through random selection, so that the sampled data sample is representative. This method is particularly suitable for large-scale data sets because it does not require storing the entire data set and only needs to maintain a sample array of a fixed size. In a possible implementation manner, the preset sampling threshold can be dynamically adjusted according to the size and characteristics of the data set to adapt to different sampling requirements.

[0141] In a feasible implementation, the system first traverses the original data sample set and assigns an index number to each sample. If the index number of a sample is less than or equal to a preset sampling threshold, the system saves the sample into the sample array. For samples with index numbers greater than the preset sampling threshold, the system generates a random number not greater than the index number of the sample and decides whether to replace an existing sample in the sample array with the current sample based on this random number. This process continues until all the original data samples are traversed. At this time, the sample index number reaches the maximum value, and the system ends the traversal and obtains the sampled data samples. Through this method, the system can efficiently select representative samples from a large-scale data set, providing support for subsequent hash code generation and data retrieval. This method not only improves the efficiency of data processing but also ensures the randomness and fairness of the samples, thereby improving the quality of hash codes and the accuracy of data retrieval.

[0142] For example, first set the sampling size to k. The first k elements that appear in the reservoir are directly saved in the array A. The probability of each of the first k numbers being selected is the same, which is 1.

[0143] Then, when processing the (k + 1)-th element, there are two cases: One is that the (k + 1)-th element is not selected, and no element in the array is replaced; at this time, the probability of each element in the array appearing is the same, and the probability that the (k + 1)-th element is not selected is:

[0144]

[0145] The other is that the (k + 1)-th element is selected, and an element in the array is replaced by the (k + 1)-th element. The probability that the (k + 1)-th element is selected is k / (k + 1), so the probability that this new element appears must be k / (k + 1), and the probability that the original data is replaced is equal for all, which is 1 / k. The probability that any one element is replaced is:

[0146]

[0147] When the (k + 1)-th element is selected and it is selected among the k elements, the probability that it is not replaced is:

[0148]

[0149] More specifically, for the above process, first, an array A of size k is set up. This array will be used to store the final sample set. The first k elements in the data stream are directly added to array A. Since these are the initial k elements, the probability of their being selected is 100%. For the (k + 1)-th element and subsequent data, the algorithm operates as follows: for each new element, a random number is generated within the range [1, k + 1], and then this random number is compared with the index of the currently processed element (i.e., its position in the data stream). If the random number is less than or equal to k, then the current new element replaces the element corresponding to the random number in array A. If the random number is greater than k, then no element is replaced and the new element is ignored. For each subsequent element in the data stream, the above process is repeated until all data is processed. After all data is processed, the array A stores the final sample set T, which has a size of k. It should be understood that in the embodiments of the present application, through the above sampling process, the probability of each element being selected is equal. For the first k elements, the probability of their being selected is 100%. For the (k + 1)-th element and subsequent elements, the probability of their being selected is k / (k + 1), because there are k positions where they can be replaced, and there are a total of k + 1 possible positions (including the new element itself). As the data stream continues, each new element has an equal chance of replacing any one of the old elements in the pond, ensuring the randomness and fairness of the samples.

[0150] In summary, in the reservoir sampling algorithm, the selection probability of each element is equal, ensuring equal opportunities for old and new elements to be selected. Starting from the (k + 1)-th element, each new element has a chance to be selected into the reservoir: under the current total amount, random numbers are generated so that the occurrence probability of all numbers is equal. Based on these random numbers, it is determined whether a certain element in the reservoir should be replaced. If the random number points to a position within the reservoir, the element at that position is replaced; otherwise, the traversal continues to the next element until all elements are processed, obtaining the training data T. This process ensures the randomness and fairness of sample selection.

[0151] In a feasible implementation manner, step S2200 may include steps S2210 to S2230:

[0152] Step S2210: Calculate the multi-head gating vector corresponding to the sampled data sample based on the hidden layer spatial representation corresponding to the sampled data sample;

[0153] Step S2220: Calculate the weighted summation coefficient for the convolution process based on the hidden layer spatial representation;

[0154] Step S2230: Calculate the output vector of the multi-head attention layer corresponding to the sampled data sample according to the multi-head gating vector and the weighted summation coefficient.

[0155] It should be noted that in this embodiment, the multi-head gated attention network model is used to deeply explore the feature representation of the sampled data samples in the multi-dimensional subspace. The hidden layer spatial representation is obtained by the regular autoencoder extracting features from the sampled data samples. It is a low-dimensional data compression representation that can capture the core features of the data. The multi-head gated vector is calculated based on the hidden layer spatial representation, which reflects the attention distribution of different heads in the multi-head attention mechanism. The weighted summation coefficient of the convolution process is calculated based on the hidden layer spatial representation and is used to determine the weights of each feature in the multi-head attention mechanism. The output vector of the multi-head attention layer is calculated by combining the multi-head gated vector and the weighted summation coefficient of the convolution process. It synthesizes the feature information of the multi-dimensional subspace and provides a rich feature representation for the subsequent single-head attention layer processing.

[0156] In addition, it should be noted that in this embodiment, through the combined action of the multi-head gated attention layer and the single-head attention layer, the model can learn the personalized features of the sampled data samples in the multi-dimensional subspace, which is crucial for generating highly discriminative hash codes. In one possible implementation, the multi-head gated attention layer can learn the features of the data in different subspaces through different heads, and the single-head attention layer further refines these features to obtain a more discriminative feature representation. This method not only improves the efficiency of data processing but also enhances the discriminative ability of the model for data features.

[0157] For example, in a feasible implementation, the system first calculates the multi-head gated vector based on the hidden layer spatial representation of the sampled data samples. This process involves in-depth analysis of the hidden layer spatial representation to determine the attention distribution of each head. Then, the system calculates the weighted summation coefficient of the convolution process, which is to determine the feature weights based on the hidden layer spatial representation. Finally, the system combines the multi-head gated vector and the weighted summation coefficient to calculate the output vector of the multi-head attention layer. This output vector synthesizes the feature information of the multi-dimensional subspace and provides a basis for generating the subspace personalized feature code. Through this series of operations, the system can effectively extract the features of the multi-dimensional subspace from the sampled data samples, thereby generating highly discriminative hash codes, optimizing the data retrieval process, and improving the efficiency of data index construction and index optimization in the database all-in-one machine environment.

[0158] In a feasible implementation, the hash code generation method in the embodiments of the present application may further include steps A1000 to A4000:

[0159] Step A1000: Receive and parse the system operation instruction to perform a syntax accuracy analysis on the input text corresponding to the operation instruction;

[0160] Step A2000: After the syntax accuracy analysis passes, convert the system operation instructions into a symbol stream;

[0161] Step A3000: Based on the symbol stream, construct an abstract syntax tree, and establish operation tasks by traversing the abstract syntax tree;

[0162] Step A4000: Pass the operation tasks to the operation function for execution and return the operation result to the user through the user interface layer.

[0163] It should be noted that in this embodiment, the hash code generation method not only involves the extraction and encoding of data features, but also includes the parsing and execution of system operation instructions, realizing a complete process from data processing to user interaction. System operation instructions refer to the specific requirements sent by users or external systems to the data processing system, such as querying, updating, or deleting data, etc. The syntax accuracy analysis of the input text refers to checking whether the instructions input by the user conform to the predetermined syntax rules to ensure that the instructions can be correctly understood by the system. The symbol stream is the process of converting text instructions into a series of symbols or tokens that can be further processed by the system. The abstract syntax tree is a tree-like structure used to represent the syntax structure of instructions, which enables the system to more clearly understand and execute instructions. The operation task refers to the specific operation steps generated according to the user instructions, and the operation function execution refers to the process in which the system executes these steps and returns the results to the user.

[0164] In addition, it should be noted that in this embodiment, by receiving and parsing system operation instructions, the system can understand the user's needs and perform corresponding processing. The purpose of this process is to improve the interactivity and flexibility of the system, enabling the system to respond to various complex user requests. In a possible implementation manner, the system can adopt natural language processing technology to parse and understand the user's natural language instructions, or use a predefined command-line interface to receive and process instructions.

[0165] In a feasible implementation manner, the system first receives and parses the operation instructions sent by the user, performs a syntax accuracy analysis on the input text to ensure the correctness of the instructions. Once the instructions are confirmed to be valid, the system converts the instructions into a symbol stream, and then constructs an abstract syntax tree based on these symbols. The system traverses this tree to establish specific operation tasks and passes these tasks to the corresponding operation functions for execution. The execution result is returned to the user through the user interface layer, enabling the user to see the final result of the operation. This complete process not only improves the efficiency and accuracy of data processing, but also enhances the user experience of the system, enabling the system to flexibly process various user requests and provide timely feedback. In this way, the system can achieve a seamless connection from data feature extraction to user interaction, providing a powerful and flexible data processing platform for users.

[0166] In yet another feasible embodiment, in the hash coding generation system on which the embodiments of the present application rely for implementation, a storage query layer is further provided, which is responsible for receiving and processing data query and deletion instructions. The query parser is responsible for parsing the query statement, extracting the operation type and data, and then passing them to the operation function to execute the actual query processing. The operation function will call the model layer interface to obtain the binary hash code of the data sample, and then operate on the hash table to complete the query task.

[0167] In the embodiments of the present application, general query scenarios are abstracted based on common query types of the index system, and a query language dedicated to the learning-based hash index system, namely the learning-based hash index query language, is adopted. First, the input data query command is converted into a symbol stream; then the symbol stream is extracted and an abstract syntax tree is constructed; finally, a query task is established by traversing the abstract syntax tree. By default, the traverser traverses the syntax parse tree in a depth-first search manner and executes the corresponding business logic when entering and leaving nodes.

[0168] It can be understood that in the hash index model, different keywords may map to the same binary hash code, resulting in hash collisions. A common solution is to use the chaining method, that is, storing the conflicting elements in the same linked list. When querying, if different keywords map to the same value, a hash collision will be triggered, and it is necessary to traverse the linked list and compare the nodes to find the record corresponding to the keyword. However, too many hash collisions will affect the retrieval speed of the hash index data. The embodiments of the present application optimize the traditional hash function by generating hash codes through learning the commonalities of data and the individual characteristics of subspaces, reducing the collision rate and improving the efficiency of the hash index, thereby accelerating the data retrieval process.

[0169] Specifically, in the embodiments of the present application, when querying data using the hash index, first calculate the hash value of the item to be queried, and then take the remainder of the length of the hash table to obtain the index of the data in the hash table. This index can be used to find the corresponding data. If the linked list corresponding to the index contains only the data to be queried and has only one node, the node can be directly returned. However, if there are multiple nodes in the linked list, it means that different original data have generated the same hash code, resulting in a hash collision. In this case, it is necessary to traverse all the nodes and compare them one by one until the element that matches the item to be queried is found.

[0170] When performing a data insertion operation, simply add a pointer to the current node at the corresponding position in the hash table. If there are already other elements at that position, a hash collision occurs, and a pointer to the current node needs to be added at the end of the linked list. At the same time, sampling is required when inserting a key. If the number of elements in the hash table exceeds the product of the load factor and the number of hash buckets, the original hash table needs to be expanded to twice its original size. In addition, the sampling container is used to retrain the model to obtain a new model that better conforms to the query key distribution. The model is learned through an improved reservoir sampling method, which improves the model accuracy while reducing the time cost of full-scale data training.

[0171] This embodiment provides a method for generating hash codes. By traversing the original data samples, the sample index number of the current sampled sample is obtained; if the sample index number is not greater than the preset sampling threshold, the current sampled sample is saved in the preset sample array; if the sample index number is greater than the preset sampling threshold, a current sample random number not greater than the sample index number is generated; if the current sample random number is not greater than the preset sampling threshold, the array element corresponding to the current sample random number in the sample array is replaced with the current sampled sample; if the sample index number is the maximum sample index number, the traversal ends, and the sampled data samples are obtained. Then, the multi-dimensional subspace data features of the sampled data samples are learned to obtain the subspace individual feature codes of the sampled data samples; the common feature code values and subspace individual feature code values corresponding to the sampled data samples are linearly concatenated to form hash codes and stored in the preset hash buckets. By using the above-mentioned network model based on a regular autoencoder and a multi-head gated attention mechanism to extract and learn the features of data samples, the multi-dimensional subspace features of the original data and the parallelism of matrix operations are fully utilized to obtain highly different feature discrete values to form hash codes, improving the discriminability of the hash codes of data samples and reducing hash index conflicts, thereby achieving the effects of improving data calculation efficiency and reducing resource consumption.

[0172] Please refer to Figure 4 , this application also provides a hash code generation device, and the hash code generation device includes:

[0173] A processing module 10, configured to perform data feature pre-extraction and discretization processing on the pre-obtained original data samples through a pre-constructed regular autoencoder to obtain the common feature code values of the original data samples;

[0174] A learning module 20, configured to learn the multi-dimensional subspace data features of the sampled data samples in the original data samples through a pre-trained multi-head gated attention network model to obtain the subspace individual feature codes of the sampled data samples;

[0175] The splicing module 30 is used to linearly splice the common feature encoding value and the subspace individual feature encoding value corresponding to the sampled data sample to form a hash code and store it in a preset hash bucket.

[0176] The hash code generation device provided by this application adopts the hash code generation method in the above embodiment and can solve the technical problems of hash code generation. Compared with the prior art, the beneficial effects of the hash code generation device provided by this application are the same as those of the hash code generation method provided by the above embodiment, and other technical features in the hash code generation device are the same as the features disclosed in the method of the above embodiment, which will not be elaborated here.

[0177] This application provides a hash code generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the hash code generation method in the first embodiment above.

[0178] Next, refer to Figure 5 , which shows a schematic structural diagram of a hash code generation device suitable for implementing the embodiments of this application. The hash code generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The hash code generation device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of this application.

[0179] As Figure 5As shown, the hash code generation device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. In the RAM 1004, various programs and data required for the operation of the hash code generation device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the hash code generation device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a hash code generation device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.

[0180] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0181] The hash code generation device provided by the present application adopts the hash code generation method in the above embodiments. Compared with the prior art, the beneficial effects of the hash code generation device provided by the present application are the same as those of the hash code generation method provided by the above embodiments, and other technical features in the hash code generation device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0182] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0183] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0184] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the hash coding generation method in the above embodiments.

[0185] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0186] The above computer-readable storage medium can be included in the hash coding generation device; it can also exist separately without being assembled into the hash coding generation device.

[0187] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0188] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0189] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0190] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned hash coding generation method, and can solve the technical problems of hash coding generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the hash coding generation method provided in the above embodiments, and will not be elaborated here.

[0191] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the hash code generation method as described above.

[0192] The computer program product provided by the present application can solve the technical problem of hash code generation. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the hash code generation method provided in the above embodiments, and will not be elaborated herein.

[0193] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A hash code generation method, characterized in that: The method comprises: The pre-constructed regular autoencoder is used to pre-extract and discretize the data features of the pre-acquired original data samples to obtain the common feature encoding values ​​of the original data samples; Learning the multi-dimensional subspace data features of the sampled data samples in the original data samples through a pre-trained multi-head gated attention network model to obtain the subspace individual feature encoding of the sampled data samples; The common feature code value and the subspace individual feature code value corresponding to the sampled data sample are linearly concatenated to form a hash code and stored in a preset hash bucket.

2. The method according to claim 1, characterized in that The multi-head gated attention network model includes a multi-head gated attention layer and a single-head attention layer. The multi-dimensional subspace data features of the sampled data samples in the original data samples are learned by the pre-trained multi-head gated attention network model to obtain the subspace individual feature encoding of the sampled data samples. The steps include: Sampling the original data sample according to a preset data sample sampling method to obtain the sampled data sample; Based on the hidden layer spatial representation corresponding to the sampled data sample, calculating the multi-head attention layer output vector corresponding to the sampled data sample through the multi-head gated attention layer, wherein the hidden layer spatial representation is obtained by compressing and mapping the sampled data sample by a pre-constructed regularized autoencoder; Based on the output vector of the multi-head attention layer, through the single-head attention layer, calculate the aggregate information representation corresponding to the sampled data sample, the aggregate information representation including several feature dimension values; Discretization is performed on the plurality of characteristic dimension values ​​represented by the aggregation information to obtain the subspace individual characteristic coding of the sampled data sample.

3. The method according to claim 2, characterized in that The step of calculating the multi-head attention layer output vector corresponding to the sampled data sample through the multi-head gated attention layer based on the hidden layer spatial representation corresponding to the sampled data sample comprises: Calculate a multi-head gating vector corresponding to the sampled data sample based on the hidden layer spatial representation corresponding to the sampled data sample; Based on the hidden layer spatial representation, calculating the weighted sum coefficient of the convolution process; According to the multi-head gating vector and the weighted sum coefficient, the multi-head attention layer output vector corresponding to the sampled data sample is calculated.

4. The method according to claim 2, characterized in that The step of sampling the original data sample according to a preset data sample sampling method to obtain the sampled data sample comprises: Traversing the original data samples to obtain the sample index number of the current sample; If the sample index number is not greater than the preset sampling threshold, the current sample is stored in a preset sample array; If the sample index number is greater than the preset sampling threshold, generating a current sample random number that is not greater than the sample index number; If the current sample random number is not greater than the preset sampling threshold, replacing the array element corresponding to the current sample random number in the sample array with the current sample; If the sample index number is the maximum sample index number, the traversal is terminated to obtain the sampled data sample.

5. The method according to claim 1, characterized in that The step of performing data feature pre-extraction and discretization processing on the pre-acquired original data samples by using the pre-constructed regular autoencoder to obtain the common feature coding values ​​of the original data samples comprises: Representing the raw data samples as a high-dimensional input matrix; Mapping the high-dimensional input matrix into a real vector of a preset dimension through the input layer of the regularized autoencoder to obtain a hidden layer spatial representation corresponding to the original data sample; The hidden layer space representation corresponding to the original data sample is discretized to obtain the common feature encoding value of the original data sample.

6. The method according to claim 1, characterized in that The method further comprises: Receiving and parsing system operation instructions to perform grammatical accuracy analysis on input text corresponding to the operation instructions; After the grammatical accuracy analysis is passed, converting the system operation instruction into a symbol stream; Building an abstract syntax tree based on the symbol stream, and traversing the abstract syntax tree to establish an operation task; The operation task is transferred to the operation function for execution and the operation result is returned to the user through the user interface layer.

7. A hash code generating device, characterized in that: The device comprises: A processing module is used to perform data feature pre-extraction and discretization processing on the pre-acquired original data samples through a pre-built regular autoencoder to obtain common feature coding values ​​of the original data samples; A learning module, used for learning the multi-dimensional subspace data features of the sampled data samples in the original data samples through a pre-trained multi-head gated attention network model, so as to obtain the subspace individual feature encoding of the sampled data samples; The splicing module is used to linearly splice the common feature code value and the subspace individual feature code value corresponding to the sampled data sample to form a hash code and store it in a preset hash bucket.

8. A hash code generating device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the hash code generation method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the hash code generation method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the hash code generation method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Multi-modal data approximate query method and system based on hybrid block chain

    CN121560962A