An artificial intelligence-based data search method and system
By preprocessing and feature extraction of multimodal data, deep neural networks are constructed to calculate feature similarity, the problem of insufficient feature fusion in the existing technology is solved, efficient and accurate data search and query are achieved, and the effectiveness of the search algorithm is improved.
Patent Information
- Application Number
- CN202411757433.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The existing data search methods fail to effectively utilize the deep relationship between the various modal features during feature fusion, resulting in insufficient feature representation ability, affecting the accuracy and recall of searches. The lack of historical data and lack of optimization strategies of deep learning models, limiting the effectiveness of the search algorithm.
By collecting multimodal data for preprocessing, eigenvectors are extracted, topological features of distance matrix and simple shapes are constructed, feature vectors are calculated using deep neural networks, and a visual interface is constructed to display search results. Data processing and feature extraction are used for BERT model, Mel filter, LSTM and other technologies are used for data processing and feature extraction, and loss function is optimized in combination with an adaptive adjustment mechanism.
It significantly improves the similarity discrimination ability, enhances the efficiency of data processing and query response, improves the accuracy and relevance of search results, provides more personalized and accurate search results, simplifies the feature fusion process, and reduces the computational complexity.
Smart Images

Figure CN119227013B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular to an artificial intelligence-based data search method and system. Background Art
[0002] With the rapid development of information technology, the scale of data generation and storage has increased sharply. Especially the popularization of multi-modal data has brought unprecedented challenges and opportunities to data search and information retrieval technologies. Traditional data search methods mainly focus on structured data and adopt keyword-based retrieval strategies, which cannot effectively process unstructured or semi-structured data forms. With the continuous progress of artificial intelligence technologies, especially the rise of deep learning and natural language processing technologies, researchers have gradually applied these advanced technologies to the processing and search of multi-modal data.
[0003] Existing data search methods still face many challenges. When performing feature fusion, most existing methods only rely on simple weighted averaging or concatenation, and fail to effectively utilize the deep relationships between features of each modality, resulting in insufficient feature representation ability, which in turn affects the search precision and recall rate. For deep learning-based models, the lack of historical data and optimization strategies during the model training process further limits the effectiveness of search algorithms. Summary of the Invention
[0004] In view of the problems existing in the above-mentioned existing artificial intelligence-based data search methods, the present invention is proposed.
[0005] Therefore, the problems to be solved by the present invention are that when performing feature fusion, most existing methods only rely on simple weighted averaging or concatenation, and fail to effectively utilize the deep relationships between features of each modality, resulting in insufficient feature representation ability, which in turn affects the search precision and recall rate. For deep learning-based models, the lack of historical data and optimization strategies during the model training process further limits the effectiveness of search algorithms.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: An artificial intelligence-based data search method and system, which includes collecting multi-modal data for preprocessing, extracting feature vectors of data points after preprocessing, receiving user queries for query parsing and calculating the similarity with the feature vectors of data points; constructing a distance matrix, constructing a simplex according to the distance matrix, and calculating the topological features of the simplex through a boundary operator; converting the topological features into topological feature vectors and fusing them with the feature vectors of data points, constructing a deep neural network DNN to calculate the similarity between the fused feature vectors and the query feature vectors, and generating search results; constructing a visualization interface to display the optimized search results.
[0007] As a preferred embodiment of the artificial intelligence-based data search method of the present invention, the steps of collecting multi-modal data for preprocessing and extracting feature vectors of data points after preprocessing include:
[0008] Collect text data, voice data, and video data through API interfaces, and perform preprocessing on each of them separately;
[0009] The preprocessing includes using regular expressions to remove HTML tags and special symbols from text data, using a band-pass filter to denoise voice data, and using a Gaussian filter to remove noise from video data;
[0010] Use the BERT model to convert text data into text feature vectors;
[0011] Perform frame segmentation on voice data to obtain voice signals, calculate the power spectrum of the voice signals using Fourier transform, extract MFCC features of the voice data using Mel filters, and splice the MFCC features to obtain voice feature vectors;
[0012] Use VGG to extract spatial features of video data, and combine LSTM to process the spatial features to obtain video feature vectors;
[0013] For the feature vectors of data points Perform normalization processing to obtain the k-th normalized feature vector of the data points ; The feature vectors of the data points include text feature vectors, voice feature vectors, and video feature vectors.
[0014] As a preferred embodiment of the artificial intelligence-based data search method of the present invention, the steps of receiving a user query, performing query parsing, and calculating the similarity with the feature vectors of data points include:
[0015] Receive user query information, and load the pre-trained BERT model to generate a query feature vector U for the query information ’ Perform normalization processing on the query feature vector;
[0016] Use cosine similarity to calculate the similarity between the normalized query feature vector U and the k-th normalized feature vector of the data points .
[0017] As a preferred embodiment of the artificial intelligence-based data search method of the present invention, the steps of constructing a distance matrix, constructing a simplex based on the distance matrix, and calculating the topological features of the simplex through boundary operators include:
[0018] Use Euclidean distance to calculate the distances between the feature vectors of data points ;
[0019] Form a distance matrix D for the distances between all feature vectors;
[0020] Extract the non-zero distance values from the distance matrix D, sort them in ascending order to generate a sorted distance sequence, use the percentile method to extract the minimum distance of the distance sequence to set the initial distance threshold, set an incremental step size based on the difference between distance sequences, and use the step-by-step incremental method to gradually increase the distance threshold according to the distance sequence, and gradually construct a simplex through the distance threshold ;
[0021] Calculate the topological feature of the a-th dimension of the simplex through the boundary operator .
[0022] As a preferred solution of the data search method based on artificial intelligence according to the present invention, wherein: the conversion of the topological feature into a topological feature vector and the fusion with the data point feature vector include:
[0023] Use persistent entropy to convert the topological feature into a topological feature vector E i ’ And perform normalization processing, and use weighted pooling to combine the normalized topological feature vector E i Into a comprehensive topological feature vector E;
[0024] Use weighted average to calculate the comprehensive data point feature vector F of the feature vectors of the normalized data;
[0025] Calculate the fused feature vector between the comprehensive topological feature vector E and the k-th normalized data point feature vector through non-linear transformation Between , the formula is:
[0026] , where Is the hyperbolic tangent function, Is the minimum value, Is the dot product of the k-th normalized data point feature vector and the comprehensive topological feature vector, And Are the L2 norms of the k-th normalized data point feature vector and the comprehensive topological feature vector respectively.
[0027] As a preferred solution of the data search method based on artificial intelligence according to the present invention, wherein: the construction of the deep neural network DNN to calculate the similarity between the fused feature vector and the query feature vector and generate the search result includes:
[0028] Collect historical data, perform preprocessing and extract historical feature vectors to generate a training set;
[0029] Construct a deep neural network DNN model, including an input layer, a hidden layer, and an output layer;
[0030] Set the format of the input layer as a fused feature vector and the normalized query feature vector U;
[0031] Use the training set to train the deep neural network DNN model;
[0032] Use an adaptive adjustment mechanism to improve the existing loss function, construct an improved loss function to optimize the model parameters, and the formula is:
[0033] , where is the improved loss function, is the dot product of the normalized query feature vector and the fused feature vector, and are the L2 norms of the normalized query feature vector and the fused feature vector respectively, w is the number of samples, is the output of the deep neural network DNN model, is the true label, is the Euclidean distance between the normalized query feature vector and the fused feature vector;
[0034] Bring the fused feature vector and the query feature vector into the trained deep neural network DNN model to obtain the similarity between the fused feature vector and the normalized query feature vector U;
[0035] Compare the similarity between the fused feature vector and the normalized query feature vector U with the similarity between the normalized query feature vector U and the k-th normalized data point feature vector . Retain the data point feature vectors with the similarity greater than the similarity and extract the corresponding multimodal data to generate the final search result.
[0036] As a preferred solution of the data search method based on artificial intelligence described in the present invention, where: the construction of a visualization interface to display the optimized search result includes:
[0037] Use React.js to construct a visualization interface to display the user's query information and the final search result, and use D3.js to dynamically display multimodal data;
[0038] Allow users who have passed real-name verification to view.
[0039] Another object of the present invention is to provide an artificial intelligence-based data search system, which includes.
[0040] A computer device, comprising: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned artificial intelligence-based data search method are realized.
[0041] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned artificial intelligence-based data search method are realized.
[0042] The beneficial effects of the present invention are as follows: by collecting multi-modal data for preprocessing, extracting the feature vectors of data points after preprocessing, receiving user queries for query parsing and calculating the similarity with the feature vectors of data points; constructing a distance matrix, constructing a simplex according to the distance matrix, and calculating the topological features of the simplex through a boundary operator; converting the topological features into topological feature vectors and fusing them with the feature vectors of data points, constructing a deep neural network DNN to calculate the similarity between the fused feature vectors and the query feature vectors, and generating search results; significantly improving the discrimination ability of similarity differences, enhancing the efficiency of data processing and query response, and improving the accuracy and relevance of search results. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0044] Figure 1 It is a schematic flowchart of an artificial intelligence-based data search method.
[0045] Figure 2 It is a schematic structural diagram of an artificial intelligence-based data search system. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings of the specification.
[0047] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention, but the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.
[0048] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separate or selectively exclusive of other embodiments.
[0049] Embodiment 1, referring to Figure 1 , is the first embodiment of the present invention. This embodiment provides an artificial intelligence-based data search method. The artificial intelligence-based data search method includes
[0050] S1. Collect multimodal data for preprocessing to extract the feature vectors of data points after preprocessing, receive a user query for query parsing, and calculate the similarity with the feature vectors of data points;
[0051] Specifically, collecting multimodal data for preprocessing and extracting the feature vectors of data points after preprocessing includes:
[0052] Collect text data, voice data, and video data through an API interface, and perform preprocessing respectively;
[0053] The preprocessing includes using regular expressions to remove HTML tags and special symbols from text data, using a band-pass filter to denoise voice data, and using a Gaussian filter to remove noise from video data;
[0054] Use the BERT model to convert text data into text feature vectors;
[0055] Perform frame splitting on voice data to obtain voice signals, use Fourier transform to calculate the power spectrum of voice signals, use Mel filters to extract MFCC features of voice data, and splice the MFCC features to obtain voice feature vectors;
[0056] Use VGG to extract the spatial features of video data, and combine LSTM to process the spatial features to obtain video feature vectors;
[0057] Perform normalization processing on the feature vectors of data points to obtain the k-th normalized feature vector of data points ; The feature vectors of data points include text feature vectors, voice feature vectors, and video feature vectors.
[0058] By collecting multimodal data from different sources through the API interface, the system can achieve unified management and processing of data. This integration ability ensures the comprehensiveness and accuracy of data analysis, providing a solid data foundation for subsequent searches and queries. The multimodal data that has been effectively denoised and formatted can more accurately reflect real information, reducing errors in subsequent model processing. By using advanced technologies such as the BERT model, Mel filters, and LSTM, the present invention can extract feature vectors with rich semantic and context information. These high-quality feature vectors will significantly improve the accuracy and relevance of the search engine during queries, thereby enhancing the user experience. The standardized data point feature vectors can effectively shorten the calculation time for subsequent queries and analyses, providing a feasible solution for application scenarios of real-time data analysis and rapid response.
[0059] Further, receiving a user query for query parsing and calculating the similarity with the data point feature vectors includes:
[0060] Receiving user query information, loading the pre-trained BERT model to generate a query feature vector U for the query information ’ Performing standardization processing on the query feature vector to obtain the standardized query feature vector U;
[0061] Using cosine similarity to calculate the similarity between the standardized query feature vector U and the standardized data point feature vectors therebetween .
[0062] By loading the pre-trained BERT model, the system can quickly understand the user's query intention. The efficient understanding ability can provide more accurate and relevant search results for users, enhancing the user experience. The combination of the generation and standardization processing of feature vectors ensures that the query information and data points are compared on the same scale. The standardization process significantly improves the accuracy of cosine similarity calculation, enabling the system to more effectively identify relevant data points. Using cosine similarity to calculate the similarity between the query feature vector and the data point feature vector can quickly locate the data most relevant to the user's query. The precise matching ability is particularly important for handling complex queries, especially in a multimodal data environment.
[0063] S2. Construct a distance matrix, construct a simplex according to the distance matrix, and calculate the topological features of the simplex through boundary operators;
[0064] Specifically, constructing a distance matrix, constructing a simplex according to the distance matrix, using persistent homology to analyze the topological features of the simplex, and plotting the birth and death times of each topological feature as a persistence diagram includes:
[0065] Using Euclidean distance to calculate the distances between data point feature vectors A distance matrix D is formed, and the formula is:
[0066] ,
[0067] where n is the total number of feature vectors;
[0068] Extract the non-zero distance values in the distance matrix D and sort them from smallest to largest to generate a sorted distance sequence. Use the percentile method to extract the minimum distance of the distance sequence to set the initial distance threshold, and set the incremental step size based on the distance difference between the first and the second in the distance sequence. Use the step-by-step incremental method to gradually increase the distance threshold according to the distance sequence, and gradually construct a simplex through the distance threshold. The formula is: , where is the distance threshold gradually increased the a-th dimensional simplex constructed under is the feature vector of a + 1 data points participating in the construction of the a-th dimensional simplex;
[0069] Starting from the initial distance threshold, if the distance between each pair of feature vectors is less than or equal to the gradually increased distance threshold, connect the feature vectors to form a one-dimensional simplex, and the line segments form a constructed topological structure;
[0070] If the distance between each pair of three feature vectors is less than or equal to the gradually increased distance threshold, connect the feature vectors to form a two-dimensional simplex, a triangle;
[0071] If the distance between each pair of four feature vectors is less than or equal to the gradually increased distance threshold, connect the feature vectors to form a three-dimensional simplex, a tetrahedron;
[0072] ……
[0073] If the distance between each pair of n feature vectors is less than or equal to the gradually increased distance threshold, connect the feature vectors to form an n - 1 dimensional simplex;
[0074] Calculate the topological feature of the a-th dimension of the simplex through the boundary operator , and the formula is:
[0075] , , where is the boundary operator of the a-th dimensional simplex, is the sign term used to control the positive and negative signs in the summation, is the kernel of the a-th dimensional boundary operator to capture the topological structure, is the boundary image of the (a + 1)-th dimensional simplex to remove non-independent topological structures.
[0076] By calculating the distances between the feature vectors of data points using the Euclidean distance, the accuracy of the distance matrix is ensured. The precise distance metric is the basis for subsequent analysis, making the similarity evaluation of data more reliable. By constructing a simplex through the step-by-step incremental method, the topological features of the data set can be effectively reflected. The construction of the simplex makes the high-dimensional data become visual during the analysis process, promoting the understanding of the data structure. Setting the initial distance threshold and incremental step size enables this method to flexibly adapt to the characteristics of different data sets. This adaptability ensures the effectiveness of the algorithm in various application scenarios and can perform well in both dense data and sparse data environments. In a multi-modal data environment, the construction of the distance matrix can handle complex data relationships, enabling the system to remain efficient and accurate when processing multi-modal data such as text, speech, and video. By extracting the topological features from the distance matrix, the potential structures and patterns in the data can be revealed. This process is crucial for data mining and knowledge discovery, helping researchers and enterprises identify key data points and their relationships. The weighted merging and sorting processing of the distance matrix significantly reduces the computational complexity. Through optimizing the algorithm design, the present invention improves the computational efficiency while ensuring the accuracy, and is suitable for real-time data processing requirements. Through the profound understanding of the relationships between data points and the extraction of topological features, the present invention provides strong decision-making support for users. Whether in a recommendation system, a classification model, or predictive analysis, accurate topological information can effectively enhance the scientific nature of decision-making.
[0077] S3. Convert the topological features into topological feature vectors and fuse them with the data point feature vectors, construct a deep neural network DNN to calculate the similarity between the fused feature vectors and the query feature vectors, and generate search results.
[0078] Specifically, converting the topological features into topological feature vectors and fusing them with the data point feature vectors includes:
[0079] Use persistent entropy to convert the topological features into topological feature vector E i ’ And perform normalization processing, use weighted pooling to process the normalized topological feature vector E i Combine them into a comprehensive topological feature vector E;
[0080] Based on the existing technology, use linear weighted fusion to calculate the fused feature vector between the comprehensive topological feature vector E and the k-th normalized data point feature vector The formula is:
[0081] where and are the weighted coefficients respectively;
[0082] Calculate the integrated topological feature vector E and the feature vector of the k-th data point after standardization through a non-linear transformation The integrated feature vector , and the formula is: , where is the hyperbolic tangent function, which is used to non-linearly compress the similarity between the topological feature vector and the data point feature vector, is the minimum value, which is used to avoid division by zero errors or very small numerical values being ignored, ensuring numerical stability during the calculation process, is the dot product of the feature vector of the k-th data point after standardization and the integrated topological feature vector, and are the L2 norms of the feature vector of the k-th data point after standardization and the integrated topological feature vector respectively;
[0083] The hyperbolic tangent function tanh non-linear activation function can compress eigenvalues and limit the output range to -1 to 1. It can capture the non-linear relationships between features. Linear methods such as simple weighting or addition may lead to over-simplification of the relationships between features and cannot fully represent the structure of complex data. Through the non-linear activation function, the subtle differences between features can be better retained, especially in scenarios with large data magnitudes. By introducing non-linear transformations, the model's ability to capture the relationships between complex features is enhanced, which is particularly suitable for high-dimensional and complex data sets. The dot product can calculate the similarity degree of two vectors in the same direction, ensuring that the maximum degree of correlation is captured in the mutual comparison of integrated feature vectors. The normalization operation after the dot product further eliminates the influence of vector length, making the fusion result only reflect the direction and similarity of the two feature vectors. The normalization operation ensures that the fusion process will not be unbalanced due to the overly large or small modulus of the vector. The summation operation is often used to calculate the sum of multiple feature vectors, but this may increase the computational complexity, especially in multi-dimensional feature spaces. By first integrating the feature vectors and then performing the fusion, the complex element-by-element summation calculation is avoided, simplifying the calculation process. This method not only improves efficiency but also ensures the balance of feature fusion.
[0084] Persistent entropy can quantify the complexity and persistence of topological features, especially suitable for extracting topological structures in high-dimensional data. The topological feature vectors after standardization can eliminate the scale differences between features, enabling topological features to have a consistent expression in subsequent fusion processes. Standardization can also prevent features in specific dimensions from having an excessive impact on the final result during fusion, improving the reliability and accuracy of feature fusion. Through nonlinear transformation using the hyperbolic tangent function, the similarity results between feature vectors can be effectively compressed, avoiding the influence of extreme values, increasing the nonlinear expression ability of the model, enabling more complex data relationships to be captured more fully, and ensuring that the similarity between features is mapped to the output space more smoothly, enhancing the generalization ability of the model. The construction of fused feature vectors improves the model's generalization ability for unknown data. Through the effective fusion of different types of features, the model can still maintain a high prediction accuracy when facing new data. The fused feature vectors can provide more comprehensive information for the decision-making system, helping analysts obtain more accurate reference bases when making data-driven decisions. The automated feature fusion method significantly reduces the complexity of manual feature engineering, making data analysis work more efficient. This feature is particularly important for startups or teams with limited resources, enabling them to obtain results more quickly in data analysis.
[0085] Furthermore, construct a deep neural network DNN to calculate the similarity between the fused feature vector and the query feature vector, and the generated search results include:
[0086] Collect historical data, preprocess it, and extract historical feature vectors to generate a training set;
[0087] Construct a deep neural network DNN model, including an input layer, a hidden layer, and an output layer;
[0088] Set the format of the input layer as the fused feature vector and the standardized query feature vector U;
[0089] Use the training set to train the deep neural network DNN model with an existing loss function , the formula is: , where w is the number of samples, is the output of the deep neural network DNN model, is the true label;
[0090] Use an adaptive adjustment mechanism to improve the existing loss function and construct an improved loss function to optimize the model parameters. The formula is: , where is the improved loss function, is the dot product of the normalized query feature vector and the fused feature vector, and are the L2 norms of the normalized query feature vector and the fused feature vector respectively, representing the magnitude of the vector, is the Euclidean distance between the query feature vector and the fused feature vector, w is the number of samples, is the output of the deep neural network DNN model, is the true label, is the Euclidean distance between the normalized query feature vector and the fused feature vector;
[0091] In the prior art, the cosine similarity is usually directly used to measure the similarity between two vectors. Only the change of the original cosine similarity is relatively smooth when it is close to 1. When the similarity of two vectors is close, the difference is not obvious, and the model is not sensitive enough to this small change, and may not be able to effectively distinguish vectors with similar but not exactly the same similarities. Using the squared difference enhances the sensitivity of the model to the similarity difference, and the small difference will be amplified, ensuring that the model more actively optimizes the similarity of the feature vectors and significantly improving the discrimination ability of the similarity difference. The cosine similarity only considers the directional similarity of two vectors and ignores the absolute distance between them. Two vectors may be very similar in direction but very different in magnitude. Introducing the Euclidean distance model reduces the similarity score when dealing with vectors that are similar in direction but far apart. The model can dynamically adjust the loss value, ensuring that the model considers the absolute distance of the vectors and does not solely rely on directional similarity. MSE is the optimal option when dealing with continuous value predictions and cannot be replaced by a simple other error metric such as MAE. Normalization makes the denominator increase when there are extremely large errors, thereby reducing the contribution of these points to the total loss. While maintaining the amplification effect of MSE on large errors, MAE also reduces the impact of outliers. If the cosine similarity and the MSE loss are simply added together, the model cannot adaptively adjust the influence of each part according to different scenarios. Through the fractional structure design and the normalization mechanism of each loss term, the model can adaptively adjust the weights of each loss term, avoiding the limitations of traditional linear combinations. The fractional structure achieves adaptive adjustment, and each loss term is dynamically adjusted through its own normalization mechanism, ensuring that the loss optimization of the model is more balanced and reasonable in different scenarios.
[0092] Bring the fused feature vector and the query feature vector into the trained deep neural network DNN model to obtain the similarity between the fused feature vector and the normalized query feature vector U;
[0093] Bring the fused feature vector and the normalized query feature vector U to obtain the similarity Compare with the standardized query feature vector U and the feature vector of the k-th data point after standardization for similarity and retain the data point feature vectors with similarity greater than the similarity to extract the corresponding multimodal data to generate the final search result;
[0094] The collection of historical data and preprocessing and extraction of historical feature vectors include operating using the same methods for preprocessing and feature extraction of multimodal data, constructing a historical distance matrix, constructing a historical simplex based on the historical distance matrix, calculating the topological features of the historical simplex through a boundary operator, and fusing the historical topological features into the historical data point feature vectors to obtain historical fusion feature vectors.
[0095] By calculating the similarity between the fusion feature vector and the query feature vector, the present invention can effectively improve the accuracy and relevance of search results. Compared with traditional methods, the deep learning-based model can better capture complex relationships in the data, thereby providing more personalized and accurate search results. By constructing a deep neural network and adopting an adaptive adjustment mechanism to optimize the loss function, the model can adapt to different data distributions, thus performing more stably on new data. This feature enables the present invention to maintain a high level of accuracy when facing dynamic data. By collecting historical data and extracting feature vectors, the model can make full use of existing data resources, enhancing the practicality and effectiveness of the model in practical applications. The utilization of this historical data not only improves the learning effect of the model but also reduces the cost of data collection and processing, enabling users to find the required information more quickly. When dealing with large-scale data, the deep neural network model can improve efficiency through parallel computing. The processing of fusion feature vectors can significantly reduce the consumption of computing resources, making it more efficient in practical applications. The search results based on real-time relevance scoring can provide a more intelligent user interaction experience, and users can obtain search results that better meet their needs, thereby improving satisfaction and usage stickiness. By exploring the application of deep neural networks in feature fusion and similarity, the present invention provides a new research direction for the field of machine learning, promoting the progress and development of related technologies.
[0096] S4. Construct a visualization interface to display the optimized search results;
[0097] Specifically, constructing a visualization interface to display the optimized search results includes: using React.js to construct a visualization interface to display user query information and the final search results, and using D3.js to dynamically display multimodal data;
[0098] Allow users who have passed real-name verification to view.
[0099] The interface built with React.js provides an efficient user interaction experience. The componentized structure enables the interface to quickly respond to user operations, enhancing the overall usability and fluency. This dynamic response ability increases user engagement and improves user satisfaction. With the powerful graphing capabilities of D3.js, data points can be presented dynamically, allowing users to intuitively understand search results in various graphical forms (such as line charts, bar charts, scatter plots, etc.). This intuitive visualization method helps users quickly identify data trends and patterns, enabling them to make more effective decisions. The design of the visualization interface can transform complex data sets into easily understandable information, assisting users in quickly finding key information among vast amounts of data. By graphically presenting data relationships, users can more clearly identify correlations and enhance the efficiency of data analysis. Through the real-name verification mechanism, only verified users can access and query the optimized search results. This security measure effectively protects user data and prevents unauthorized access, thus enhancing the credibility of the system and user trust. Through the combination of front-end visualization and back-end search optimization technologies, users can view the optimized search results in real time. This instant feedback mechanism helps users understand the search process and increases user trust in the system.
[0100] Example 2, referring to Figure 2 , is the second embodiment of the present invention. This embodiment is different from the previous one and provides an artificial intelligence-based data search system, including,
[0101] A collection query module for collecting multimodal data for preprocessing, extracting the feature vectors of data points after preprocessing, receiving user queries for query parsing, and calculating the similarity with the feature vectors of data points;
[0102] A construction and analysis module for constructing a distance matrix, constructing a simplex based on the distance matrix, and calculating the topological features of the simplex through boundary operators;
[0103] A fusion and optimization module for converting topological features into topological feature vectors and fusing them with the feature vectors of data points, constructing a deep neural network DNN to calculate the similarity between the fused feature vectors and the query feature vectors, and generating search results;
[0104] A visualization module for constructing a visualization interface to display the optimized search results.
[0105] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.
[0106] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0107] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber device, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0108] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
Claims
1. A data search method based on artificial intelligence, characterized in that: including, collecting multimodal data for preprocessing, extracting feature vectors of data points after preprocessing, receiving user queries for query parsing and calculating the similarity with the feature vectors of data points; constructing a distance matrix, constructing a simplex according to the distance matrix, and calculating the topological features of the simplex through a boundary operator; converting topological features into topological feature vectors and fusing them with data point feature vectors, constructing a deep neural network DNN to calculate the similarity between the fused feature vectors and the query feature vectors, and generating search results; constructing a visualization interface to display the optimized search results; The collecting multimodal data for preprocessing, extracting feature vectors of data points after preprocessing and constructing a distance matrix includes collecting text data, speech data, and video data through an API interface and performing preprocessing on them respectively; The preprocessing includes using regular expressions to remove HTML tags and special symbols from text data, using a band-pass filter to denoise speech data, and using a Gaussian filter to remove noise from video data; using a BERT model to convert text data into text feature vectors; performing frame processing on speech data to obtain speech signals, calculating the power spectrum of the speech signals using Fourier transform, extracting MFCC features of speech data using Mel filters, and concatenating the MFCC features to obtain speech feature vectors; using VGG to extract spatial features of video data and combining LSTM to process the spatial features to obtain video feature vectors; Normalize the data point feature vectors wherein the data point feature vectors include text feature vectors, speech feature vectors, and video feature vectors.
2. The data search method based on artificial intelligence according to claim 1, characterized in that: The receiving user queries for query parsing and calculating the similarity with the feature vectors of data points includes receiving user query information, loading a pre-trained BERT model to generate a query feature vector U' and normalizing the query feature vector; Calculate the similarity A between the normalized query feature vector U and the normalized feature vector F of the k-th data point using cosine similarity k between k1 .
3. The data search method based on artificial intelligence according to claim 2, wherein: The construction of the distance matrix, the construction of the simplex based on the distance matrix, and the calculation of the topological features of the simplex through the boundary operator use the Euclidean distance to calculate the distance D between the eigenvectors of the data points ij ; forming a distance matrix D for the distances between all feature vectors; Extract the non-zero distance values in the distance matrix D and sort them in ascending order to generate a sorted distance sequence. Use the percentile method to extract the minimum distance of the distance sequence to set the initial distance threshold. Set the incremental step size based on the difference between distance sequences. Use the step-by-step incremental method to gradually increase the distance threshold according to the distance sequence in turn, and gradually construct the simplex V through the distance threshold a (ε a ); Calculating the topological feature H of the a-th dimension of a simplex through a boundary operator a .
4. The data search method based on artificial intelligence according to claim 3, wherein: The conversion of topological features into topological feature vectors and the fusion with data point feature vectors refer to using persistent entropy to convert topological features into topological feature vector E i ’, and performing normalization processing, and using weighted pooling to combine the normalized topological feature vector E i into a comprehensive topological feature vector E; Calculate the fused feature vector F' between the comprehensive topological feature vector E and the feature vector F of the k-th data point through non-linear transformation k The formula is as follows: k , formula is: where tanh(·) is the hyperbolic tangent function, μ is the minimum value, ||F k || and ||E|| are the norms of the data point feature vector and the comprehensive topological feature vector, respectively.
5. The data search method based on artificial intelligence according to claim 4, wherein: The constructing a deep neural network DNN to calculate the similarity between the fused feature vectors and the query feature vectors and generating search results means collecting historical data and performing preprocessing and extracting historical feature vectors to generate a training set; constructing a deep neural network DNN model, including an output layer, hidden layers, and an output layer; Set the format of the input layer as the fused feature vector F'. k and the normalized query feature vector U; using the training set to train the deep neural network DNN model; using an adaptive adjustment mechanism to improve the existing loss function, constructing an improved loss function to optimize model parameters, and the formula is: Where Q' is the improved loss function, U·F' k is the dot product of the normalized query feature vector and the fused feature vector, ||U|| and ||F' k || are the L2 norms of the normalized query feature vector and the fused feature vector respectively, w is the number of samples, is the output of the deep neural network DNN model, y i is the true label; Bring the fused feature vector and the query feature vector into the trained deep neural network DNN model to obtain the fused feature vector F'. k The similarity A between the fused feature vector F' and the normalized query feature vector U k2 ; The fused feature vector F' k The similarity A between the fused feature vector F' and the normalized query feature vector U k2 The similarity A between the normalized query feature vector U and the k-th normalized data point feature vector F k are compared, and the similarity A k1 is compared. The data point feature vectors with similarity A k2 greater than the similarity A k1 are retained, and the corresponding multimodal data is extracted to generate the final search result.
6. The data search method based on artificial intelligence according to claim 5, wherein: The constructing a visualization interface to display the optimized search results means using React.js to construct a visualization interface to display user query information and the final search results, and using D3.js to dynamically display multimodal data; allowing users who have passed real-name verification to view.
7. An artificial intelligence-based data search system based on the artificial intelligence-based data search method according to any one of claims 1-6, characterized in that: including, a collection query module for collecting multimodal data for preprocessing, extracting feature vectors of data points after preprocessing, receiving user queries for query parsing and calculating the similarity with the feature vectors of data points; a construction analysis module for constructing a distance matrix, constructing a simplex according to the distance matrix, and calculating the topological features of the simplex through a boundary operator; a fusion optimization module for converting topological features into topological feature vectors and fusing them with data point feature vectors, constructing a deep neural network DNN to calculate the similarity between the fused feature vectors and the query feature vectors, and generating search results; A visualization module for constructing a visualization interface to display the optimized search results.
8. A computer device, comprising: A memory and a processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, the steps of the data search method based on artificial intelligence according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the data search method based on artificial intelligence according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and device for searching and editing objects in NeRF
CN118916510A