Search engine-based reference material and standard product retrieval and sorting method and system

By using a search engine-based method in the standard matter search system, using a bidirectional encoder and tensor decomposition algorithm to generate field adaptive language models and distributed representations, combining heterogeneous information networks and metapathic path methods to determine semantic associations, the problem of insufficient understanding of semantic associations in traditional search methods is solved, and higher relevance and accuracy of search results are achieved.

CN119127970BActive Publication Date: 2025-05-16TAN-MO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411258733.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-05-16
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

Traditional standard substance search methods cannot deeply understand the semantic relationship between user's search intention and standard substance knowledge, resulting in low correlation and accuracy of search results.

Method used

A search engine-based method is used to build a common language model through a bidirectional encoder, and fine-tune it in combination with the standard material domain corpus to generate a domain adaptive language model. Combining ontology and knowledge graphs, a standard material standard knowledge base is constructed, and a tensor decomposition algorithm is used to map entities, relationships and attributes in the knowledge base to low-dimensional semantic space to generate a distributed representation. The semantic correlation between search keywords and standard material knowledge is calculated, and the semantic correlation is determined through heterogeneous information network and metapathic path methods, and a comprehensive correlation measure is generated for multi-grained semantic matching to obtain candidate search results.

Benefits of technology

It significantly improves the relevance and accuracy of search results, can better understand the semantic relationship between user's search intention and standard material knowledge, and provides more accurate and comprehensive search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119127970B_ABST
    Figure CN119127970B_ABST
Patent Text Reader

Abstract

The present invention provides a search engine-based reference material and standard product retrieval and sorting, which relates to the technical field of data processing, and includes: obtaining search keywords input by a user, constructing a reference material and standard product knowledge base, mapping to a low-dimensional semantic space, calculating the semantic relevance of the search keywords and reference material knowledge, constructing a heterogeneous information network, generating a comprehensive relevance measure, performing multi-granularity semantic matching, and obtaining candidate search results; assigning adaptive feature weights to structured features corresponding to the candidate search results, performing independent clustering, obtaining subspace clustering results, performing optimization, obtaining a global optimal clustering result, mining high-order semantic association features and performing cross-category semantic association, and generating high-quality search results; performing low-rank decomposition, obtaining an initial comprehensive score, modeling the sorting problem as a Markov decision process, determining an action space and a state space, constructing a reward function and determining a sorting position, and obtaining an optimal sorting result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for searching and sorting reference materials and reference products based on a search engine. Background Art

[0002] Traditional reference material retrieval mainly relies on keyword matching and Boolean retrieval, which cannot deeply understand the semantic relationship between the user's search intent and reference material knowledge, resulting in low relevance and accuracy of the retrieval results. Traditional methods mainly adopt static sorting strategies based on specific rules or single indicators, such as sorting by relevance score or user rating;

[0003] Static sorting ignores the complex interactions among multiple influencing factors such as user preferences, query characteristics, and reference material properties, and cannot dynamically adapt to changes in user needs, resulting in low relevance and satisfaction of sorting results.

[0004] Therefore, a solution is urgently needed to solve the problems existing in the prior art. Summary of the invention

[0005] The embodiment of the present invention provides a method and system for searching and sorting reference materials and reference products based on a search engine, which can at least solve some of the problems existing in the prior art.

[0006] A first aspect of an embodiment of the present invention provides a method for searching and sorting reference materials and reference standards based on a search engine, comprising:

[0007] Obtain search keywords input by the user, build a general language model based on a bidirectional encoder, fine-tune the general language model in combination with a small amount of reference material domain corpus to obtain a domain adaptive language model, build a reference material and standard product knowledge base based on ontology and knowledge graph, map entities, relationships and attributes in the reference material and standard product knowledge base to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, generate a distributed representation of reference material knowledge and combine it with the domain adaptive language model, calculate the semantic relevance between the search keywords and reference material knowledge, build a heterogeneous information network with multi-source heterogeneous information corresponding to reference material and standard products in combination with a heterogeneous information network representation learning model, determine semantic associations in combination with a meta-path method and generate a comprehensive relevance metric in combination with the semantic relevance, perform multi-granular semantic matching based on the comprehensive relevance metric, and obtain candidate search results;

[0008] For the structured features corresponding to the candidate search results, a fuzzy rule base is constructed through evidence theory and fuzzy ensemble learning, and corresponding adaptive feature weights are assigned to the structured features of each dimension. The subspaces corresponding to the structured features are independently clustered based on the adaptive feature weights. The density peak search algorithm and the spectral clustering algorithm are combined to obtain the subspace clustering results corresponding to each subspace. The association dependency is modeled by Markov random field and the subspace clustering results are optimized by the variational inference algorithm to obtain the global optimal clustering results. For the global optimal clustering results, high-order semantic association features are mined through the attribute-weighted heterogeneous information network representation learning model and cross-category semantic associations are performed through collaborative filtering. The cross-category semantic associations are combined to generate high-quality search results;

[0009] Based on the high-quality search results, the search result influencing factors are constructed as a high-dimensional scoring tensor through a hierarchical factor decomposition machine model, and low-rank decomposition is performed in combination with a high-order singular value decomposition algorithm to generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results. Based on the user's personal preferences, a ranking model is constructed in combination with the novelty of the search keywords and the initial comprehensive score. The ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, and the action space and state space are determined in combination with a dual-tower architecture. A reward function is constructed and the ranking position of each standard substance is determined based on the reward function value to obtain the optimal ranking result.

[0010] In an optional embodiment,

[0011] The search keywords input by the user are obtained, a general language model is constructed based on a bidirectional encoder, and the general language model is fine-tuned in combination with a small amount of reference material domain corpus to obtain a domain adaptive language model, a reference material and standard product knowledge base is constructed based on ontology and knowledge graph, and entities, relationships and attributes in the reference material and standard product knowledge base are mapped to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, a distributed representation of reference material knowledge is generated and combined with the domain adaptive language model, the semantic relevance between the search keywords and reference material knowledge is calculated, and the multi-source heterogeneous information corresponding to reference material and standard products is constructed into a heterogeneous information network in combination with a heterogeneous information network representation learning model, the semantic association is determined in combination with a meta-path method, and a comprehensive relevance metric is generated in combination with the semantic relevance, and multi-granular semantic matching is performed based on the comprehensive relevance metric to obtain candidate search results including:

[0012] Acquire search keywords input by a user, perform standardization on the search keywords and convert the search keywords into preliminary semantic representations through a word embedding algorithm, build a general language model based on a bidirectional encoder, collect professional corpora corresponding to the field of standard substances, build a corpus in the field of a small amount of standard substances, use the corpus in the field of a small amount of standard substances to train the general language model, fine-tune the general language model according to the training results, perform word segmentation and part-of-speech tagging on the corpus in the field of a small amount of standard substances, build a domain vocabulary and use the domain vocabulary to replace the original vocabulary corresponding to the general language model to obtain the domain adaptive language model;

[0013] Use an ontology editor to define a reference material and standard product ontology, define core concepts, attributes and relationships, extract reference material and standard product information from an open source reference material database through a dictionary matching method, perform named entity recognition on the reference material and standard product information, perform relationship extraction through a dependency analysis algorithm, generate a reference material and standard product knowledge base, perform rule reasoning on the reference material and standard product knowledge base for quality assessment, identify erroneous relationships and correct the erroneous relationships through manual review, and obtain a corrected reference material and standard product knowledge base;

[0014] The knowledge in the reference material and standard product knowledge base is converted into a tensor form to obtain a knowledge tensor, and the knowledge tensor is decomposed into an entity embedding matrix and a relationship embedding matrix by a knowledge representation learning algorithm based on tensor decomposition, and the reconstruction loss is minimized according to a stochastic gradient descent algorithm to obtain a low-dimensional embedding representation corresponding to the entity, relationship and attribute, and the low-dimensional embedding representation is standardized by a regularization term to obtain reference material knowledge, and the distributed representation of the reference material knowledge is spliced ​​with the domain adaptive language model to obtain a joint embedding space;

[0015] The search keywords are encoded by the domain adaptive language model to obtain an embedding vector representation in the joint embedding space, the semantic relevance score between the embedding vector representation and the standard material knowledge embedding is determined by calculating the vector dot product and the vector modulus, the entities and relationships corresponding to the standard material knowledge embedding are extracted from the multi-source heterogeneous information corresponding to the standard material knowledge embedding, a heterogeneous information network is constructed in combination with a heterogeneous information network representation learning model, a random walk sequence is generated according to a predefined meta-path and mapped to a low-dimensional embedding space, for node pairs in the heterogeneous information network, the semantic association is determined by calculating the meta-path feature vector between each node pair, the meta-path feature vector and the semantic relevance score are weightedly combined to obtain the comprehensive relevance measure, the entities and relationships in the standard material and standard product knowledge base are divided into multiple granularities, multiple granularity levels are determined, and the search keywords are matched and clustered at each granularity level to obtain the candidate search results.

[0016] In an optional embodiment,

[0017] Minimizing the reconstruction loss according to the stochastic gradient descent algorithm is shown in the following formula:

[0018]

[0019] Among them, L represents reconstruction loss, h represents head entity, t represents tail entity, r represents relationship, and Real() represents real part operation. Represents the head entity embedding vector e h The transpose of , T represents the transpose, d r represents the complex diagonal matrix corresponding to the relation r, ⊙ represents element-by-element multiplication, e t represents the embedding vector of the tail entity, λ represents the regularization coefficient, |E| 2 represents the square of the norm of the entity embedding matrix E, |d| 2 represents the complex diagonal matrix d r The square of the norm of .

[0020] In an optional embodiment,

[0021] For the structured features corresponding to the candidate search results, a fuzzy rule base is constructed through evidence theory and fuzzy ensemble learning, and corresponding adaptive feature weights are assigned to the structured features of each dimension. The subspaces corresponding to the structured features are independently clustered based on the adaptive feature weights. The density peak search algorithm and the spectral clustering algorithm are combined to obtain the subspace clustering results corresponding to each subspace. The association dependency is modeled by Markov random field and the subspace clustering results are optimized by the variational inference algorithm to obtain the global optimal clustering results. For the global optimal clustering results, high-order semantic association features are mined through the attribute-weighted heterogeneous information network representation learning model and cross-category semantic associations are performed through collaborative filtering. The generation of high-quality search results combined with the cross-category semantic associations includes:

[0022] For the candidate search results, the corresponding structured features are extracted by a feature extraction algorithm, the structured features are modeled by a quality function in evidence theory, the quality function and importance coefficient corresponding to the structured features of each dimension are defined, the quality functions corresponding to each dimension are combined by a Dempster combination rule to generate a comprehensive quality function, the structured features are clustered by a fuzzy C-means clustering algorithm to obtain a plurality of fuzzy rules and the comprehensive quality function is updated by using the fuzzy rules as a source of evidence, and the adaptive feature weight is assigned to the structured features corresponding to each dimension based on the updated comprehensive quality function;

[0023] For the structured features, a subspace corresponding to each structured feature is constructed based on the adaptive feature weights. For each subspace, the local density corresponding to each data point and the minimum distance between the current data point and the data point with the maximum local density are calculated by a density peak search algorithm, and the cluster center is automatically determined to obtain a first clustering result. A Laplace matrix is ​​constructed according to the similarity matrix between different data points by a spectral clustering algorithm, and feature decomposition and clustering are performed to obtain a second clustering result. The first clustering result and the second clustering result are combined to obtain a subspace clustering result corresponding to each subspace. The subspace clustering result is used as a random variable by a Markov random field algorithm and an energy function is defined. The associated dependency is determined by minimizing the energy function and approximately inferred by a variational inference algorithm. The subspace clustering result is optimized according to the variational distribution to obtain the global optimal clustering result.

[0024] The candidate retrieval results are regarded as nodes in a heterogeneous information network and the structured features are used as attributes of the nodes in the heterogeneous information network. Node embedding learning is performed through an attribute-weighted heterogeneous information network representation learning model, and the network structure preservation items and attribute reconstruction items are jointly optimized to generate high-order semantic association features. The global optimal clustering result is used as the user's preference category in combination with the collaborative filtering idea to determine the cross-category semantic association between the candidate retrieval results. Based on the cross-category semantic association and the global optimal clustering result, the high-quality retrieval result is solved.

[0025] In an optional embodiment,

[0026] The objective function of the attribute-weighted heterogeneous information network representation learning model is shown in the following formula:

[0027]

[0028] Among them, H represents the objective function value, w ij represents the edge weight between the i-th node and the j-th node, v i represents the embedding vector of the i-th node, v j represents the embedding vector of the jth node, ρ represents the balance parameter, and a i represents the attribute feature of the i-th node, μ represents the regularization coefficient, and ||θ|| 2 represents the norm of the model parameters θ.

[0029] In an optional embodiment,

[0030] Based on the high-quality search results, the search result influencing factors are constructed as a high-dimensional scoring tensor through a hierarchical factor decomposition machine model, and low-rank decomposition is performed in combination with a high-order singular value decomposition algorithm to generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results. Based on the user's personal preferences, a ranking model is constructed in combination with the novelty of the search keywords and the initial comprehensive score. The ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, and the action space and state space are determined in combination with a dual-tower architecture. A reward function is constructed and the ranking position of each standard substance is determined based on the reward function value. The optimal ranking results include:

[0031] For the high-quality search results, define the search result influencing factors, determine the corresponding causal decomposition machine loss function through the hierarchical causal decomposition machine model, repeatedly optimize the causal decomposition machine loss function through the alternating least squares method and gradually update the corresponding factor matrix, calculate the high-dimensional score tensor corresponding to the search result influencing factors according to the updated factor matrix, perform high-order singular value decomposition on the high-dimensional score tensor, generate a core tensor and a factor matrix, extract high-order interaction information in different search result influencing factors from the core tensor, and for each high-quality search result, calculate the initial comprehensive score through tensor operation based on the high-order interaction information and the factor matrix;

[0032] Based on the initial comprehensive score, based on the pre-acquired personal preferences corresponding to the current user, and combined with the novelty corresponding to the search keyword, a ranking model is constructed, the standard substance and standard product ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, a state tower and an action tower are constructed through a dual-tower strategy network, the state tower models the input sequence composed of the high-quality search results through a bidirectional long short-term memory network, determines the hidden state and generates a state embedding representation through a multi-layer perceptron, uses the personal preference, novelty and the initial comprehensive score as inputs to the action tower, generates an input feature vector and generates a corresponding action embedding representation through a multi-layer perceptron, constructs a reward function, and calculates the function value of the reward function by calculating the inner product of the state embedding representation and the action embedding representation;

[0033] The maximum reward function value is selected and the corresponding action embedding representation and state embedding representation are determined, an optimal action is generated and the optimal action is parsed into a sorting result to obtain the optimal sorting result.

[0034] In an optional embodiment,

[0035] The high-order singular value decomposition of the high-dimensional score tensor is shown in the following formula:

[0036]

[0037] Among them, X represents a high-dimensional rating tensor, represents the elements of the core tensor, represents the singular vectors in the first dimension, represents the singular vectors in the second dimension, represents the singular vector in the Mth dimension, R M represents the rank of dimension m, where m represents the total number of dimensions. Represents the outer product operator.

[0038] A second aspect of an embodiment of the present invention provides a reference material and standard product retrieval and sorting system based on a search engine, comprising:

[0039] The first unit is used to obtain the search keywords input by the user, build a general language model based on a bidirectional encoder, fine-tune the general language model in combination with a small amount of standard material domain corpus to obtain a domain adaptive language model, build a standard material and standard product knowledge base based on ontology and knowledge graph, map the entities, relationships and attributes in the standard material and standard product knowledge base to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, generate a distributed representation of standard material knowledge and combine it with the domain adaptive language model, calculate the semantic relevance between the search keywords and the standard material knowledge, build the multi-source heterogeneous information corresponding to the standard material and standard product into a heterogeneous information network in combination with a heterogeneous information network representation learning model, determine the semantic association in combination with a meta-path method and generate a comprehensive relevance measure in combination with the semantic relevance, perform multi-granular semantic matching based on the comprehensive relevance measure, and obtain candidate search results;

[0040] The second unit is used to construct a fuzzy rule base for the structured features corresponding to the candidate search results through evidence theory and fuzzy ensemble learning, and assign corresponding adaptive feature weights to the structured features of each dimension, independently cluster the subspaces corresponding to the structured features based on the adaptive feature weights, obtain the subspace clustering results corresponding to each subspace by combining the density peak search algorithm and the spectral clustering algorithm, model the association dependency through Markov random field and optimize the subspace clustering results through the variational inference algorithm to obtain the global optimal clustering results, and for the global optimal clustering results, mine high-order semantic association features through the attribute-weighted heterogeneous information network representation learning model and perform cross-category semantic association through collaborative filtering, and generate high-quality search results in combination with the cross-category semantic association;

[0041] The third unit is used to construct the search result influencing factors into a high-dimensional scoring tensor based on the high-quality search results through a hierarchical factor decomposition machine model, perform low-rank decomposition in combination with a high-order singular value decomposition algorithm, generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results, construct a ranking model based on the user's personal preferences in combination with the novelty of the search keywords and the initial comprehensive score, model the ranking problem as a Markov decision process based on a deep reinforcement learning algorithm, determine the action space and state space in combination with a dual-tower architecture, construct a reward function and determine the ranking position of each standard substance based on the reward function value to obtain the optimal ranking result.

[0042] According to a third aspect of the embodiments of the present invention,

[0043] An electronic device is provided, comprising:

[0044] processor;

[0045] a memory for storing processor-executable instructions;

[0046] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0047] According to a fourth aspect of the embodiments of the present invention,

[0048] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0049] In the present invention, the enhanced semantic understanding and domain adaptation capabilities enable the retrieval system to accurately analyze the user's retrieval needs and provide more accurate and comprehensive retrieval results in professional fields. The introduction of knowledge representation learning enables the retrieval system to deeply explore the potential semantic associations between reference material knowledge, and through the combination with the domain adaptive language model, the semantic correlation calculation and matching capabilities between retrieval keywords and reference material knowledge are significantly improved. The fusion of heterogeneous information networks and the mining of multi-granular semantic associations enable the retrieval system to capture the complex associations between reference materials from multiple dimensions and granularity levels, and generate more comprehensive and accurate candidate retrieval results. The optimization and aggregation of structured features This type of method can effectively discover the potential similarities and association patterns between standard substances, and generate higher-quality and more relevant search results. The modeling of high-order interactions of influencing factors and the dynamically optimized sorting strategy can comprehensively consider the complex influence of multiple factors, and dynamically adjust the sorting strategy according to the real-time feedback of users to generate more accurate, personalized and satisfactory sorting results. In summary, the present invention has comprehensively optimized and innovated the retrieval and sorting problems of standard substances, greatly improved the retrieval quality, user experience and satisfaction in the field of standard substances, not only promoted technological progress and service upgrades in the field of standard substances, but also provided valuable ideas and experience for intelligent information retrieval and knowledge services in other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 The present invention is a flowchart of a method for searching and sorting reference materials and reference products based on a search engine according to an embodiment of the present invention;

[0051] Figure 2 The present invention is a schematic diagram of the structure of a search engine-based reference material and standard product retrieval and ranking system. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0054] Figure 1 FIG. 1 is a flow chart of a method for searching and sorting reference materials and reference products based on a search engine according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0055] S1. Obtain search keywords input by the user, build a general language model based on a bidirectional encoder, fine-tune the general language model in combination with a small amount of standard material domain corpus to obtain a domain adaptive language model, build a standard material and standard product knowledge base based on ontology and knowledge graph, map the entities, relationships and attributes in the standard material and standard product knowledge base to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, generate a distributed representation of standard material knowledge and combine it with the domain adaptive language model, calculate the semantic relevance between the search keywords and standard material knowledge, build the multi-source heterogeneous information corresponding to the standard material and standard product into a heterogeneous information network in combination with a heterogeneous information network representation learning model, determine the semantic association in combination with a meta-path method and generate a comprehensive relevance metric in combination with the semantic relevance, perform multi-granular semantic matching based on the comprehensive relevance metric, and obtain candidate search results;

[0056] The bidirectional encoder is a deep learning model architecture, which is usually used for natural language processing tasks. The general language model is a model for generating and understanding natural language, which aims to capture general laws and patterns in language. The small amount of standard substance domain corpus is a small amount of corpus specially built for a specific field (such as the field of standard substances), which is used to train and test the performance of the model with a small amount of samples. The domain adaptive language model is a model that makes domain-specific adjustments to the general language model to improve its performance in specific fields (such as medicine, law, etc.). The ontology is a structured framework for representing knowledge, which defines concepts, entities and their relationships in the field. The knowledge representation learning algorithm based on tensor decomposition uses tensor decomposition technology. to carry out knowledge representation learning. The standard material and standard product knowledge base is a knowledge base specially used to store and manage information related to standard materials and standards. The low-dimensional semantic space maps high-dimensional data (such as word embedding, text representation) to a lower-dimensional space through dimensionality reduction technology. The heterogeneous information network representation learning model is used to learn the representation of nodes and edges in heterogeneous information networks. The heterogeneous information network is a data structure containing multiple types of nodes and edges. The meta-path method is a method for knowledge reasoning and information retrieval in heterogeneous information networks. The comprehensive relevance measure is a method for evaluating the comprehensive impact of multiple factors or features. The granular semantic matching is a semantic matching technology that aims to match at different semantic levels or granularities.

[0057] In an optional embodiment,

[0058] The search keywords input by the user are obtained, a general language model is constructed based on a bidirectional encoder, and the general language model is fine-tuned in combination with a small amount of reference material domain corpus to obtain a domain adaptive language model, a reference material and standard product knowledge base is constructed based on ontology and knowledge graph, and entities, relationships and attributes in the reference material and standard product knowledge base are mapped to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, a distributed representation of reference material knowledge is generated and combined with the domain adaptive language model, the semantic relevance between the search keywords and reference material knowledge is calculated, and the multi-source heterogeneous information corresponding to reference material and standard products is constructed into a heterogeneous information network in combination with a heterogeneous information network representation learning model, the semantic association is determined in combination with a meta-path method, and a comprehensive relevance metric is generated in combination with the semantic relevance, and multi-granular semantic matching is performed based on the comprehensive relevance metric to obtain candidate search results including:

[0059] Acquire search keywords input by a user, perform standardization on the search keywords and convert the search keywords into preliminary semantic representations through a word embedding algorithm, build a general language model based on a bidirectional encoder, collect professional corpora corresponding to the field of standard substances, build a corpus in the field of a small amount of standard substances, use the corpus in the field of a small amount of standard substances to train the general language model, fine-tune the general language model according to the training results, perform word segmentation and part-of-speech tagging on the corpus in the field of a small amount of standard substances, build a domain vocabulary and use the domain vocabulary to replace the original vocabulary corresponding to the general language model to obtain the domain adaptive language model;

[0060] Use an ontology editor to define a reference material and standard product ontology, define core concepts, attributes and relationships, extract reference material and standard product information from an open source reference material database through a dictionary matching method, perform named entity recognition on the reference material and standard product information, perform relationship extraction through a dependency analysis algorithm, generate a reference material and standard product knowledge base, perform rule reasoning on the reference material and standard product knowledge base for quality assessment, identify erroneous relationships and correct the erroneous relationships through manual review, and obtain a corrected reference material and standard product knowledge base;

[0061] The knowledge in the reference material and standard product knowledge base is converted into a tensor form to obtain a knowledge tensor, and the knowledge tensor is decomposed into an entity embedding matrix and a relationship embedding matrix by a knowledge representation learning algorithm based on tensor decomposition, and the reconstruction loss is minimized according to a stochastic gradient descent algorithm to obtain a low-dimensional embedding representation corresponding to the entity, relationship and attribute, and the low-dimensional embedding representation is standardized by a regularization term to obtain reference material knowledge, and the distributed representation of the reference material knowledge is spliced ​​with the domain adaptive language model to obtain a joint embedding space;

[0062] The search keywords are encoded by the domain adaptive language model to obtain an embedding vector representation in the joint embedding space, the semantic relevance score between the embedding vector representation and the standard material knowledge embedding is determined by calculating the vector dot product and the vector modulus, the entities and relationships corresponding to the standard material knowledge embedding are extracted from the multi-source heterogeneous information corresponding to the standard material knowledge embedding, a heterogeneous information network is constructed in combination with a heterogeneous information network representation learning model, a random walk sequence is generated according to a predefined meta-path and mapped to a low-dimensional embedding space, for node pairs in the heterogeneous information network, the semantic association is determined by calculating the meta-path feature vector between each node pair, the meta-path feature vector and the semantic relevance score are weightedly combined to obtain the comprehensive relevance measure, the entities and relationships in the standard material and standard product knowledge base are divided into multiple granularities, multiple granularity levels are determined, and the search keywords are matched and clustered at each granularity level to obtain the candidate search results.

[0063] The preliminary semantic representation refers to the preliminary understanding and representation of text or data, which is usually achieved through word embedding or word vector. The part-of-speech tagging is a natural language processing task that aims to assign a part-of-speech label to each word in the text, such as noun, verb, adjective, etc. The domain vocabulary is a vocabulary built for a specific field or industry (such as medicine, law, finance, etc.). The named entity recognition is a natural language processing task that aims to identify specific types of entities (such as names of people, places, organization names, time, etc.) from text. The dependency analysis algorithm is used to analyze the syntactic structure of a sentence and understand the grammatical structure of a sentence by determining the dependency relationship between words in the sentence. The entity embedding matrix is ​​a matrix that represents entities (such as names of people, places, organizations, etc.) as vectors. The relationship embedding matrix is ​​a matrix that represents relationships. The reconstruction loss is used to measure the error of the model when reconstructing the input data. It is usually used in unsupervised learning or autoencoder models. The goal is to minimize the difference between the input data and the reconstructed data. The vector dot product is the inner product operation of two vectors. The calculation formula is the sum of the products of the corresponding elements of the two vectors. The vector modulus is the length of a vector. The calculation formula is the square root of the sum of the squares of each component of the vector. The semantic relevance score is used to measure the semantic similarity between two texts or words. The random walk sequence refers to a node sequence generated by a random walk algorithm in a graph structure. The multi-granularity partitioning refers to the segmentation or division of data or information at different granularity levels. The granularity level refers to the division details or levels of data or information.

[0064] Obtain the search keywords input by the user, and perform standardization on the search keywords, including word segmentation, removal of stop words, word form restoration, etc., to obtain the standardized keywords, and convert the standardized keywords into preliminary semantic representations through word embedding algorithms (such as Word2Vec) to obtain the word vectors corresponding to the keywords. A general language model is constructed based on a bidirectional encoder (such as BERT), and professional corpora corresponding to the field of standard materials are collected, such as standard material manuals, research papers, etc., to construct a small amount of standard material field corpora, and the corpora in the standard material field corpora are used to train the general language model. The general language model is fine-tuned according to the training results, and the corpora in the standard material field corpora are segmented and POS tagged to construct a domain vocabulary, and the domain vocabulary is used to replace the original vocabulary corresponding to the general language model to obtain a domain adaptive language model.

[0065] Use an ontology editor (such as Protégé) to define the reference material ontology, including core concepts, attributes, and relationships, and use a dictionary matching method to obtain reference material databases (such as NIST) from open source databases. extracting standard material and standard product information from a standard material and standard product information repository, performing named entity recognition on the extracted standard material and standard product information, identifying entities such as material, purity, and purpose, performing relationship extraction through a dependency analysis algorithm, such as material-purity, material-purpose, and other relationships, generating a standard material and standard product knowledge base based on the extracted entities and relationships, performing rule reasoning on the standard material and standard product knowledge base for quality assessment, identifying erroneous relationships, correcting erroneous relationships through manual review, obtaining a corrected standard material and standard product knowledge base, converting the knowledge in the standard material and standard product knowledge base into a tensor form, obtaining a knowledge tensor, decomposing the knowledge tensor into an entity embedding matrix and a relationship embedding matrix through a knowledge representation learning algorithm based on tensor decomposition (such as RESCAL), minimizing the reconstruction loss according to a stochastic gradient descent algorithm, obtaining a low-dimensional embedding representation corresponding to entities, relationships, and attributes, normalizing the low-dimensional embedding representation through a regularization term, obtaining standard material knowledge, and splicing the distributed representation of the standard material knowledge with a domain adaptive language model to obtain a joint embedding space;

[0066] The search keywords are encoded through a domain adaptive language model to obtain an embedded vector representation in a joint embedding space. The semantic relevance score between the embedded vector representation and the reference material knowledge embedding is determined by calculating the vector dot product and the vector modulus. Entities and relationships are extracted from the multi-source heterogeneous information corresponding to the reference material knowledge embedding. A heterogeneous information network is constructed by combining a heterogeneous information network representation learning model (such as metapath2vec). A random walk sequence is generated according to a predefined metapath and mapped to a low-dimensional embedding space. For node pairs in the heterogeneous information network, the semantic association is determined by calculating the metapath feature vector between each node pair. The metapath feature vector and the semantic relevance score are weighted and combined to obtain a comprehensive relevance measure. The entities and relationships in the reference material knowledge base are divided into multiple granularities, and multiple granularity levels are determined. The search keywords are matched and clustered at each granularity level to obtain candidate search results.

[0067] Exemplarily, the user enters the search keyword: high-purity copper, the system standardizes it to obtain: high purity, copper, and encodes it into a vector representation in the joint embedding space through the domain adaptive language model. The system calculates the semantic relevance score between the vector and each entity in the standard material knowledge base (such as: high-purity copper SRM 683, high-purity copper: SRM 885, etc.), and at the same time constructs a heterogeneous information network and calculates the semantic association based on the meta-path to comprehensively obtain the relevance measure. The system matches and clusters at different granularities such as material and purity, and returns the top-k most relevant standard substances as candidate search results, such as: SRM 683 high-purity copper (99.95%), SRM 885 high-purity copper (99.999%), etc.

[0068] In this embodiment, a domain-adaptive language model is obtained by fine-tuning the general language model in combination with the standard material domain corpus and replacing the original vocabulary with the domain vocabulary, so that the system can better understand and represent the professional terms and semantic information in the field of standard materials, thereby improving the accuracy of retrieval. Mapping it to a low-dimensional embedding space through knowledge representation learning not only enriches the semantic representation of standard materials, but also reveals the implicit connection between entities and relationships, and expands the breadth and depth of retrieval. Multi-angle semantic understanding helps to capture the diversity of user query intentions and improve the relevance of retrieval results. By performing multi-granularity division and clustering on the standard material knowledge base, the system can match at different granularity levels such as materials and purity, meet the multi-level retrieval needs of users, improve the matching efficiency, and take into account the comprehensiveness of the results. In summary, this embodiment significantly improves the effect of standard material retrieval, provides convenience for practical applications, and is expected to greatly improve work efficiency and decision-making quality.

[0069] In an optional embodiment,

[0070] Minimizing the reconstruction loss according to the stochastic gradient descent algorithm is shown in the following formula:

[0071]

[0072] Among them, L represents reconstruction loss, h represents head entity, t represents tail entity, r represents relationship, and Real() represents real part operation. Represents the head entity embedding vector e h The transpose of , T represents the transpose, d r represents the complex diagonal matrix corresponding to the relation r, ⊙ represents element-by-element multiplication, e t represents the embedding vector of the tail entity, λ represents the regularization coefficient, |E| 2 represents the square of the norm of the entity embedding matrix E, |d| 2 represents the complex diagonal matrix d r The square of the norm of .

[0073] In this embodiment, complex embedding can represent more diverse interaction modes between entities and relationships, provides more degrees of freedom, can characterize different semantic roles of entities under different relationships, and capture more fine-grained semantic information. By introducing a complex diagonal matrix, each relationship is given unique rotation and scaling characteristics, so that the model can flexibly adjust the representation of entity embedding under different relationships, thereby improving the expressive power of knowledge representation. Knowledge representation learning based on complex embedding and diagonal matrices can explicitly model the complex interactions between entities and relationships, help reveal implicit patterns and laws in knowledge graphs, and enhance the reasoning and generalization capabilities of the model. Stochastic gradient descent iteratively updates parameters in the form of small batches, reduces memory overhead, and is suitable for industrial-level application scenarios. In summary, this embodiment has significant technical effects in improving expressive power, enhancing reasoning generalization, and optimizing computing efficiency. It provides a new idea for building high-quality, explainable, and efficient knowledge graph embedding, and is of great value for application scenarios such as intelligent retrieval, knowledge reasoning, and decision support.

[0074] S2. For the structured features corresponding to the candidate search results, a fuzzy rule base is constructed through evidence theory and fuzzy ensemble learning, and corresponding adaptive feature weights are assigned to the structured features of each dimension. The subspaces corresponding to the structured features are independently clustered based on the adaptive feature weights. The density peak search algorithm and the spectral clustering algorithm are combined to obtain the subspace clustering results corresponding to each subspace. The association dependency is modeled by Markov random field and the subspace clustering results are optimized by the variational inference algorithm to obtain the global optimal clustering results. For the global optimal clustering results, high-order semantic association features are mined through the attribute-weighted heterogeneous information network representation learning model and cross-category semantic associations are performed through collaborative filtering. The cross-category semantic associations are combined to generate high-quality search results;

[0075] The structured features refer to features with clear organization and hierarchy in the data. The evidence theory is a method for dealing with uncertainty and evidence fusion. The fuzzy ensemble learning combines the advantages of fuzzy logic and ensemble learning, and processes and fuses the prediction results of multiple models by using fuzzy rules. The fuzzy rule base is a collection of fuzzy rules used for reasoning and decision-making in fuzzy logic systems. The independent clustering refers to clustering data assuming that each cluster is independent and has no dependencies. The density peak search algorithm is a density-based clustering method that determines the cluster center by finding the density peak in the data space. The spectral clustering algorithm uses the spectral features of the data (such as the eigenvalues ​​and eigenvectors of the Laplace matrix) to perform clustering. The subspace refers to a low-dimensional space defined by certain specific dimensions or features in a high-dimensional space. The Markov random field is A probability model for describing a set of random variables, in which the state of each variable depends only on the states of variables in its neighborhood. The association dependency describes the association relationship and mutual dependence between different variables or data. The variational inference algorithm is a technology for approximately inferring complex probability distributions, which approaches the true posterior distribution by introducing variational distributions and minimizes the KL divergence between the variational distribution and the true posterior distribution. The attribute-weighted heterogeneous information network refers to assigning weights to different types of nodes and edges in a heterogeneous information network to reflect their importance or influence. The high-order semantic association features refer to data features captured by complex semantic association models. These features not only consider direct semantic relationships, but also include high-order, indirect semantic connections. The collaborative filtering is a recommendation system technology that predicts user preferences for unknown items by utilizing historical interaction data between users and items.

[0076] In an optional embodiment,

[0077] For the structured features corresponding to the candidate search results, a fuzzy rule base is constructed through evidence theory and fuzzy ensemble learning, and corresponding adaptive feature weights are assigned to the structured features of each dimension. The subspaces corresponding to the structured features are independently clustered based on the adaptive feature weights. The density peak search algorithm and the spectral clustering algorithm are combined to obtain the subspace clustering results corresponding to each subspace. The association dependency is modeled by Markov random field and the subspace clustering results are optimized by the variational inference algorithm to obtain the global optimal clustering results. For the global optimal clustering results, high-order semantic association features are mined through the attribute-weighted heterogeneous information network representation learning model and cross-category semantic associations are performed through collaborative filtering. The generation of high-quality search results combined with the cross-category semantic associations includes:

[0078] For the candidate search results, the corresponding structured features are extracted by a feature extraction algorithm, the structured features are modeled by a quality function in evidence theory, the quality function and importance coefficient corresponding to the structured features of each dimension are defined, the quality functions corresponding to each dimension are combined by a Dempster combination rule to generate a comprehensive quality function, the structured features are clustered by a fuzzy C-means clustering algorithm to obtain a plurality of fuzzy rules and the comprehensive quality function is updated by using the fuzzy rules as a source of evidence, and the adaptive feature weight is assigned to the structured features corresponding to each dimension based on the updated comprehensive quality function;

[0079] For the structured features, a subspace corresponding to each structured feature is constructed based on the adaptive feature weights. For each subspace, the local density corresponding to each data point and the minimum distance between the current data point and the data point with the maximum local density are calculated by a density peak search algorithm, and the cluster center is automatically determined to obtain a first clustering result. A Laplace matrix is ​​constructed according to the similarity matrix between different data points by a spectral clustering algorithm, and feature decomposition and clustering are performed to obtain a second clustering result. The first clustering result and the second clustering result are combined to obtain a subspace clustering result corresponding to each subspace. The subspace clustering result is used as a random variable by a Markov random field algorithm and an energy function is defined. The associated dependency is determined by minimizing the energy function and approximately inferred by a variational inference algorithm. The subspace clustering result is optimized according to the variational distribution to obtain the global optimal clustering result.

[0080] The candidate retrieval results are regarded as nodes in a heterogeneous information network and the structured features are used as attributes of the nodes in the heterogeneous information network. Node embedding learning is performed through an attribute-weighted heterogeneous information network representation learning model, and the network structure preservation items and attribute reconstruction items are jointly optimized to generate high-order semantic association features. The global optimal clustering result is used as the user's preference category in combination with the collaborative filtering idea to determine the cross-category semantic association between the candidate retrieval results. Based on the cross-category semantic association and the global optimal clustering result, the high-quality retrieval result is solved.

[0081] The quality function is used to evaluate the effect of the model or algorithm. In cluster analysis, the quality function (or objective function) is used to measure the quality of the clustering results, such as the compactness within the cluster and the separation between clusters. The importance coefficient is used to quantify the relative importance of each feature, variable or element in a model or algorithm. The fuzzy C-means clustering algorithm is a fuzzy clustering method that allows a data point to belong to multiple clusters. The membership of each cluster is represented by a membership function. The evidence source refers to the source of data or information for decision-making, reasoning or model training. The local density refers to the density of data points around a point in the data space. The cluster center refers to the representative point of each cluster in the clustering algorithm. The Laplace matrix is ​​a matrix in graph theory used to describe the structural characteristics of the graph. The energy function is an objective function used to quantify the state of the model. In statistical physics, the energy function is used to describe the energy state of the system. In machine learning, the energy function is often used in optimization problems as the optimization target. The heterogeneous information network is a network structure containing multiple types of nodes and edges.

[0082] The structured features corresponding to the candidate search results, such as material type, purity, and purpose, are extracted through feature extraction algorithms (such as TF-IDF, Word2Vec, etc.). The structured features are modeled through the quality function in the evidence theory, and the quality function and importance coefficient corresponding to the structured features of each dimension are defined. The quality functions corresponding to each dimension are combined through the Dempster combination rule to generate a comprehensive quality function. The structured features are clustered through the fuzzy C-means clustering algorithm to obtain multiple fuzzy rules. The fuzzy rules are used as evidence sources to update the comprehensive quality function. Based on the updated comprehensive quality function, adaptive feature weights are assigned to the structured features corresponding to each dimension.

[0083] For structured features, a subspace corresponding to each structured feature is constructed based on adaptive feature weights. For each subspace, the local density corresponding to each data point and the minimum distance between the current data point and the data point with the maximum local density are calculated through the density peak search algorithm, and the cluster center is automatically determined to obtain the first clustering result. The Laplace matrix is ​​constructed according to the similarity matrix between different data points through the spectral clustering algorithm, and feature decomposition and clustering are performed to obtain the second clustering result. The first clustering result and the second clustering result are combined to obtain the subspace clustering result corresponding to each subspace. The subspace clustering result is used as a random variable through the Markov random field algorithm and an energy function is defined. The associated dependency is determined by minimizing the energy function. Approximate inference is performed through the variational inference algorithm, and the subspace clustering result is optimized according to the variational distribution to obtain the global optimal clustering result.

[0084] The candidate search results are regarded as nodes in a heterogeneous information network, and the structural features are used as the attributes of the nodes. The node embedding learning is performed through the attribute-weighted heterogeneous information network representation learning model. The network structure preservation item and the attribute reconstruction item are jointly optimized to generate high-order semantic association features. The global optimal clustering result is used as the user's preference category in combination with the collaborative filtering idea to determine the cross-category semantic association between the candidate search results. Based on the cross-category semantic association and the global optimal clustering result, the high-quality search results are obtained.

[0085] For example, suppose the user enters the search keyword: high-purity copper, and the system returns a set of candidate search results, including standard substances of different materials, purities, and uses. By extracting the structured features of the candidate results, such as: material: copper, purity: 99.99%, use: electronics industry, etc., the quality function of each feature dimension is defined based on the evidence theory, and the comprehensive quality function is generated by the Dempster combination rule. The fuzzy rules are obtained by using fuzzy C-means clustering, and the comprehensive quality function is updated. Finally, the adaptive feature weights are obtained. Based on the adaptive feature weights, three subspaces of material, purity, and use are constructed. For each subspace, the density peak search is first used. The search algorithm is used to obtain the first clustering result, and then the spectral clustering algorithm is used to obtain the second clustering result. The two clustering results are combined to obtain the subspace clustering result. The subspace clustering is optimized by Markov random field and variational inference to obtain the global optimal clustering result, such as: high-purity copper (above 99.99%) → semiconductor materials, high-purity copper (99.95%-99.99%) → wires and cables, etc. The candidate results and structured features are used to construct a heterogeneous information network, and high-order semantic association features are generated through the attribute-weighted representation learning model. The global optimal clustering result is used as the user preference, and the cross-category semantic association between the candidate results is mined, such as: the association between high-purity copper SRM 683 and 5N high-purity copper wire in the dimensions of materials and uses. The semantic association and clustering results are combined to obtain high-quality retrieval results, such as: high-purity copper SRM 683, 5N high-purity copper target, etc.

[0086] In this embodiment, by extracting the structured features of the candidate results and using the quality function in the evidence theory for modeling, the system can comprehensively characterize the multi-dimensional properties of the standard material, making the feature representation more accurate and robust. The adaptive feature weight allocation mechanism enables the system to flexibly adjust the importance of each feature dimension according to different retrieval requirements and data distribution, thereby improving the adaptability and generalization ability of the feature representation. By optimizing the subspace clustering results through Markov random fields and variational inference, the system obtains the globally optimal clustering results, providing high-quality semantic category information for subsequent semantic association mining. The introduction of heterogeneous information networks expands the dimension and depth of semantic associations, enabling the system to discover more implicit, cross-category correlations. In summary, this embodiment significantly improves the recall, precision and practicality of standard material retrieval, provides strong technical support for the management and application of standard materials, and is expected to play an important role in materials science, chemical analysis, quality control and other fields.

[0087] In an optional embodiment,

[0088] The objective function of the attribute-weighted heterogeneous information network representation learning model is shown in the following formula:

[0089]

[0090] Among them, H represents the objective function value, w ij represents the edge weight between the i-th node and the j-th node, v i represents the embedding vector of the i-th node, v j represents the embedding vector of the jth node, ρ represents the balance parameter, and a i represents the attribute feature of the i-th node, μ represents the regularization coefficient, and ||θ|| 2 represents the norm of the model parameters θ.

[0091] In this embodiment, the first term of the objective function is the network structure preservation term, which is used to minimize the difference between the distance between the node embedding vectors and the edge weights, which helps to capture the high-order interactions and transitive relationships between nodes and improves the quality of the embedded representation. The attribute-aware embedding not only retains the network topology, but also incorporates the semantic information of the nodes, so that semantically similar nodes are closer in the embedding space, enhancing the expressive power of the embedded representation and improving the performance of downstream tasks. The adjustable balancing mechanism enables the model to adapt to different types of heterogeneous information networks and task requirements, improving the applicability and robustness of the method. Through regularization, the model can learn a more generalized and stable embedded representation, improving the performance on unknown data. In summary, this embodiment comprehensively utilizes network topology and node semantics, has the advantages of strong expressive power, good generalization, and wide applicability, and provides a powerful tool for the analysis and mining of complex heterogeneous networks.

[0092] S3. Based on the high-quality search results, the factors affecting the search results are constructed as a high-dimensional scoring tensor through a hierarchical factor decomposition machine model, and low-rank decomposition is performed in combination with a high-order singular value decomposition algorithm to generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results. Based on the user's personal preferences, a ranking model is constructed in combination with the novelty of the search keywords and the initial comprehensive score. The ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, and the action space and state space are determined in combination with a dual-tower architecture. A reward function is constructed and the ranking position of each standard substance is determined based on the reward function value to obtain the optimal ranking result.

[0093] The hierarchical factorization machine is an extended factorization machine model, which models complex interactions in high-dimensional data by introducing a hierarchical structure. The high-dimensional scoring tensor is a tensor structure that represents multi-dimensional interactions of data and is used to capture the interactions of multiple feature dimensions in the data. The high-order singular value decomposition is a tensor decomposition method that is used to decompose a high-dimensional tensor into the product of a core tensor and a set of factor matrices. The low-rank decomposition is a process of representing high-dimensional data as a low-rank structure, which aims to reduce computational complexity and storage requirements by simplifying the model while retaining the main features of the data. The high-order interaction refers to the complex interactions between multiple features in the data. The Markov decision process is a framework for modeling decision problems with randomness. The dual-tower architecture is a model architecture commonly used in recommendation systems and deep learning, and mainly includes two independent network towers (or branches), which are used to process different types of input data respectively.

[0094] In an optional embodiment,

[0095] Based on the high-quality search results, the search result influencing factors are constructed as a high-dimensional scoring tensor through a hierarchical factor decomposition machine model, and low-rank decomposition is performed in combination with a high-order singular value decomposition algorithm to generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results. Based on the user's personal preferences, a ranking model is constructed in combination with the novelty of the search keywords and the initial comprehensive score. The ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, and the action space and state space are determined in combination with a dual-tower architecture. A reward function is constructed and the ranking position of each standard substance is determined based on the reward function value. The optimal ranking results include:

[0096] For the high-quality search results, define the search result influencing factors, determine the corresponding causal decomposition machine loss function through the hierarchical causal decomposition machine model, repeatedly optimize the causal decomposition machine loss function through the alternating least squares method and gradually update the corresponding factor matrix, calculate the high-dimensional score tensor corresponding to the search result influencing factors according to the updated factor matrix, perform high-order singular value decomposition on the high-dimensional score tensor, generate a core tensor and a factor matrix, extract high-order interaction information in different search result influencing factors from the core tensor, and for each high-quality search result, calculate the initial comprehensive score through tensor operation based on the high-order interaction information and the factor matrix;

[0097] Based on the initial comprehensive score, based on the pre-acquired personal preferences corresponding to the current user, and combined with the novelty corresponding to the search keyword, a ranking model is constructed, the standard substance and standard product ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, a state tower and an action tower are constructed through a dual-tower strategy network, the state tower models the input sequence composed of the high-quality search results through a bidirectional long short-term memory network, determines the hidden state and generates a state embedding representation through a multi-layer perceptron, uses the personal preference, novelty and the initial comprehensive score as inputs to the action tower, generates an input feature vector and generates a corresponding action embedding representation through a multi-layer perceptron, constructs a reward function, and calculates the function value of the reward function by calculating the inner product of the state embedding representation and the action embedding representation;

[0098] The maximum reward function value is selected and the corresponding action embedding representation and state embedding representation are determined, an optimal action is generated and the optimal action is parsed into a sorting result to obtain the optimal sorting result.

[0099] The alternating least squares method is an optimization algorithm for matrix decomposition, the core tensor is a key component in high-order tensor decomposition (such as high-order singular value decomposition or tensor decomposition), capturing the main interactive information in the original high-dimensional tensor, the factor matrix is ​​a part of the tensor decomposition, representing the implicit features of each dimension in the data, the state tower is a part of the dual-tower architecture, used to process and encode state information (such as user behavior or features), the action tower is the other part of the dual-tower architecture, used to process and encode action information (such as recommended items or decisions), the state embedding representation is a vector representation generated by the state tower, used to capture and encode the characteristics of the state, the action embedding representation is a vector representation generated by the action tower, used to capture and encode the characteristics of the action, the multi-layer perceptron is a feedforward neural network composed of multiple fully connected layers, and the inner product is a basic operation in vector algebra, used to calculate the similarity or correlation between two vectors.

[0100] Define the factors that affect the ranking of high-quality search results, such as material properties, user preferences, novelty, etc., determine the corresponding causal decomposition machine loss function through the hierarchical causal decomposition machine model, repeatedly optimize the causal decomposition machine loss function through the alternating least squares method, and gradually update the corresponding factor matrix. According to the updated factor matrix, calculate the high-dimensional score tensor corresponding to the factors affecting the search results, perform high-order singular value decomposition on the high-dimensional score tensor, generate a core tensor and a factor matrix, extract high-order interaction information from different search result influencing factors from the core tensor, and for each high-quality search result, calculate the initial comprehensive score through tensor operations based on the high-order interaction information and the factor matrix;

[0101] Based on the initial comprehensive score, combined with the pre-acquired current user personal preferences and the novelty of the search keywords, a ranking model is constructed. Based on the deep reinforcement learning algorithm, the standard material ranking problem is modeled as a Markov decision process. The state tower and action tower are constructed through a dual-tower strategy network. The state tower models the input sequence composed of high-quality search results through a bidirectional long short-term memory network, determines the hidden state, and converts the hidden state into a state embedding representation through a multi-layer perceptron. Personal preferences, novelty and initial comprehensive scores are used as inputs to the action tower to generate input feature vectors. The input feature vectors are converted into action embedding representations through a multi-layer perceptron. A reward function is constructed. The function value of the reward function is obtained by calculating the inner product of the state embedding representation and the action embedding representation. The largest reward function value is selected, and the corresponding action embedding representation and state embedding representation are determined. The optimal action is generated, and the optimal action is parsed into a ranking result to obtain the optimal ranking result.

[0102] For example, suppose the user enters the search keyword: high-purity copper. The system obtains a set of high-quality search results through the previous steps, including standard substances with different material properties, manufacturers, prices and other information, defines influencing factors such as material purity, user preferences, and prices, and determines the causal relationship and weight between factors through the causal decomposition machine model, and constructs the corresponding loss function. By alternately optimizing the loss function and updating the factor matrix, the system obtains the high-dimensional score tensor of each influencing factor. After high-order singular value decomposition, the high-order interaction information between the influencing factors is extracted, and the initial comprehensive score of each search result is calculated, such as: high-purity copper SRM 683 is rated 4.2, 5N high-purity copper target is rated 4.5, etc. Based on the initial comprehensive score, the system builds a ranking model in combination with the user's preference for material purity and the novelty of the search keyword high-purity copper. By modeling the ranking problem as a Markov decision process, the system uses a dual-tower strategy network for modeling. The state tower encodes the sequence of high-quality retrieval results through a bidirectional long short-term memory network to obtain a hidden state, and generates a state embedding through a multi-layer perceptron. The action tower takes user preferences, novelty and initial scores as inputs, and generates action embedding through a multi-layer perceptron. The system builds a reward function and obtains the reward value by calculating the inner product of the state embedding and the action embedding. Through continuous interaction and learning with the environment, the system selects the action with the largest reward value and parses it into the final ranking result. For example, based on the user's preference for high-purity materials, the system ranks 5N high-purity copper target before high-purity copper SRM 683. At the same time, because 6N high-purity copper powder has a higher novelty for the keyword high-purity copper, it is also ranked in front. Finally, the system generates a ranking result that comprehensively considers multiple influencing factors and is dynamically optimized.

[0103] In this embodiment, the modeling of high-order interactions makes the sorting results more accurate and in line with actual needs. The introduction of reinforcement learning breaks through the static and fixed limitations of traditional sorting methods, enabling the model to dynamically adjust the sorting results according to the user's immediate feedback and preference changes, providing a real-time optimization and personalized sorting mechanism. The intelligent sorting method maximizes the mining of effective information in the data and provides users with high-quality, high-relevance and high-satisfaction retrieval services. Deep reinforcement learning provides an end-to-end optimization framework that can seamlessly integrate different feature extractors, policy networks and reward functions, making the method applicable to various types of sorting problems. In summary, this embodiment greatly improves the sorting quality and user experience, and provides key technical support and innovative ideas for the intelligent development of the field of standard materials and even other related fields.

[0104] In an optional embodiment,

[0105] The high-order singular value decomposition of the high-dimensional score tensor is shown in the following formula:

[0106]

[0107] Among them, X represents a high-dimensional rating tensor, represents the elements of the core tensor, represents the singular vectors in the first dimension, represents the singular vectors in the second dimension, represents the singular vector in the Mth dimension, R M represents the rank of dimension m, where m represents the total number of dimensions. Represents the outer product operator.

[0108] In this embodiment, the characterization of high-order interactions enables the model to better understand and utilize complex dependencies in the data, improves the ability of feature representation and pattern mining, and by retaining only the most important singular values ​​and corresponding singular vectors, the dimension and scale of the data can be greatly reduced while retaining the main information, which not only improves the computing efficiency but also saves storage space, enabling the model to process large-scale high-dimensional tensor data more efficiently. The extraction of potential features provides a new perspective for in-depth understanding of the data, helps to improve feature engineering and model design, and by drawing heat maps or scatter plots of core tensors and factor matrices, potential patterns and features can be intuitively displayed, which is convenient for manual analysis and interpretation. In summary, this embodiment provides a powerful tool for processing and analyzing high-dimensional complex data, and shows broad application prospects in multiple fields.

[0109] Exemplarily, another embodiment of the present invention includes: the search module searches for a product list according to the search content and filter conditions input by the user, and then obtains a new filter condition according to the filter condition and the order of selection. Each time a filter condition is selected, its content is fixed, and the content of the remaining filter conditions is refreshed according to the search results. For example, the user first selects the product type as a single-label solution, then selects the brand as jar ink quality inspection, and finally selects the specification as 2mL. The search module does not set any filter conditions, searches for a product list that meets the search content, and aggregates the product type field to obtain a new product type filter condition, the content of which is all product types included in the product list. Set the filter condition to "product type is single-label solution", query the product list that meets the search content and aggregate the brand field to get a new brand filter condition, the content is all brands included in the product list with product type of single-label solution, set the filter condition to "product type is single-label solution and brand is Tanmo Quality Inspection", query all product lists that meet the search content and aggregate the specification field to get a new specification filter condition, the content is all specifications included in the product list with product type of single-label solution and brand of Tanmo Quality Inspection, the search module sends the product list and the new filter condition to the data completion module , the data completion module will query all the fields of product information in Mysql according to the product ID in the product list and add the fields required by the user to the product list. Finally, the product list and the new filtering conditions are displayed to the user. So far, multi-dimensional dynamic condition search has been realized. The now_sign field is added. If the delivery type is spot, the value is 1, otherwise the value is 0. The tmrm_sign field is added. If the product type is Tanmo self-operated, the value is 1, otherwise the value is 0. The gbw_sign field is added. If the product is a national standard material, the value is 1, otherwise the value is 0. When the user searches, the product list is sorted. The product delivery period is spot Priority is given to products with the highest matching degree from high to low, and the matching score is obtained by the score calculation formula and sorted in reverse order. The standard value is from high to low, and is sorted in reverse order according to the standard value numeric field in the product information. The results of the associated word query need to be retained and placed at the end, and this is achieved using the painless script. The script will return different values ​​based on whether the associated word contains the search condition. If it does, it returns 1, and if it does not, it returns 0, and the results are sorted in reverse order based on the return value.

[0110] Figure 2 FIG. 1 is a schematic diagram of a search engine-based reference material and standard product retrieval and ranking system according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0111] The first unit is used to obtain the search keywords input by the user, build a general language model based on a bidirectional encoder, fine-tune the general language model in combination with a small amount of standard material domain corpus to obtain a domain adaptive language model, build a standard material and standard product knowledge base based on ontology and knowledge graph, map the entities, relationships and attributes in the standard material and standard product knowledge base to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, generate a distributed representation of standard material knowledge and combine it with the domain adaptive language model, calculate the semantic relevance between the search keywords and the standard material knowledge, build the multi-source heterogeneous information corresponding to the standard material and standard product into a heterogeneous information network in combination with a heterogeneous information network representation learning model, determine the semantic association in combination with a meta-path method and generate a comprehensive relevance measure in combination with the semantic relevance, perform multi-granular semantic matching based on the comprehensive relevance measure, and obtain candidate search results;

[0112] The second unit is used to construct a fuzzy rule base for the structured features corresponding to the candidate search results through evidence theory and fuzzy ensemble learning, and assign corresponding adaptive feature weights to the structured features of each dimension, independently cluster the subspaces corresponding to the structured features based on the adaptive feature weights, obtain the subspace clustering results corresponding to each subspace by combining the density peak search algorithm and the spectral clustering algorithm, model the association dependency through Markov random field and optimize the subspace clustering results through the variational inference algorithm to obtain the global optimal clustering results, and for the global optimal clustering results, mine high-order semantic association features through the attribute-weighted heterogeneous information network representation learning model and perform cross-category semantic association through collaborative filtering, and generate high-quality search results in combination with the cross-category semantic association;

[0113] The third unit is used to construct the search result influencing factors into a high-dimensional scoring tensor based on the high-quality search results through a hierarchical factor decomposition machine model, perform low-rank decomposition in combination with a high-order singular value decomposition algorithm, generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results, construct a ranking model based on the user's personal preferences in combination with the novelty of the search keywords and the initial comprehensive score, model the ranking problem as a Markov decision process based on a deep reinforcement learning algorithm, determine the action space and state space in combination with a dual-tower architecture, construct a reward function and determine the ranking position of each standard substance based on the reward function value to obtain the optimal ranking result.

[0114] According to a third aspect of the embodiments of the present invention,

[0115] An electronic device is provided, comprising:

[0116] processor;

[0117] a memory for storing processor-executable instructions;

[0118] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0119] According to a fourth aspect of the embodiments of the present invention,

[0120] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0121] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for searching and sorting reference materials and reference standards based on a search engine, characterized in that: include: Obtain search keywords input by the user, build a general language model based on a bidirectional encoder, fine-tune the general language model in combination with a small amount of reference material domain corpus to obtain a domain adaptive language model, build a reference material and standard product knowledge base based on ontology and knowledge graph, map entities, relationships and attributes in the reference material and standard product knowledge base to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, generate a distributed representation of reference material knowledge and combine it with the domain adaptive language model, calculate the semantic relevance between the search keywords and reference material knowledge, build a heterogeneous information network with multi-source heterogeneous information corresponding to reference material and standard products in combination with a heterogeneous information network representation learning model, determine semantic associations in combination with a meta-path method and generate a comprehensive relevance metric in combination with the semantic relevance, perform multi-granular semantic matching based on the comprehensive relevance metric, and obtain candidate search results; For the structured features corresponding to the candidate search results, a fuzzy rule base is constructed through evidence theory and fuzzy ensemble learning, and corresponding adaptive feature weights are assigned to the structured features of each dimension. The subspaces corresponding to the structured features are independently clustered based on the adaptive feature weights. The density peak search algorithm and the spectral clustering algorithm are combined to obtain the subspace clustering results corresponding to each subspace. The association dependency is modeled by Markov random field and the subspace clustering results are optimized by the variational inference algorithm to obtain the global optimal clustering results. For the global optimal clustering results, high-order semantic association features are mined through the attribute-weighted heterogeneous information network representation learning model and cross-category semantic associations are performed through collaborative filtering. The cross-category semantic associations are combined to generate high-quality search results; Based on the high-quality search results, the search result influencing factors are constructed as a high-dimensional scoring tensor through a hierarchical factor decomposition machine model, and low-rank decomposition is performed in combination with a high-order singular value decomposition algorithm to generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results. Based on the user's personal preferences, a ranking model is constructed in combination with the novelty of the search keywords and the initial comprehensive score. The ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, and the action space and state space are determined in combination with a dual-tower architecture. A reward function is constructed and the ranking position of each standard substance is determined based on the reward function value to obtain the optimal ranking result.

2. The method according to claim 1, characterized in that The search keywords input by the user are obtained, a general language model is constructed based on a bidirectional encoder, and the general language model is fine-tuned in combination with a small amount of reference material domain corpus to obtain a domain adaptive language model, a reference material and standard product knowledge base is constructed based on ontology and knowledge graph, and entities, relationships and attributes in the reference material and standard product knowledge base are mapped to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, a distributed representation of reference material knowledge is generated and combined with the domain adaptive language model, the semantic relevance between the search keywords and reference material knowledge is calculated, and the multi-source heterogeneous information corresponding to reference material and standard products is constructed into a heterogeneous information network in combination with a heterogeneous information network representation learning model, the semantic association is determined in combination with a meta-path method, and a comprehensive relevance metric is generated in combination with the semantic relevance, and multi-granular semantic matching is performed based on the comprehensive relevance metric to obtain candidate search results including: Acquire search keywords input by a user, perform standardization on the search keywords and convert the search keywords into preliminary semantic representations through a word embedding algorithm, build a general language model based on a bidirectional encoder, collect professional corpora corresponding to the field of standard substances, build a corpus in the field of a small amount of standard substances, use the corpus in the field of a small amount of standard substances to train the general language model, fine-tune the general language model according to the training results, perform word segmentation and part-of-speech tagging on the corpus in the field of a small amount of standard substances, build a domain vocabulary and use the domain vocabulary to replace the original vocabulary corresponding to the general language model to obtain the domain adaptive language model; Use an ontology editor to define a reference material and standard product ontology, define core concepts, attributes and relationships, extract reference material and standard product information from an open source reference material database through a dictionary matching method, perform named entity recognition on the reference material and standard product information, perform relationship extraction through a dependency analysis algorithm, generate a reference material and standard product knowledge base, perform rule reasoning on the reference material and standard product knowledge base for quality assessment, identify erroneous relationships and correct the erroneous relationships through manual review, and obtain a corrected reference material and standard product knowledge base; The knowledge in the reference material and standard product knowledge base is converted into a tensor form to obtain a knowledge tensor, and the knowledge tensor is decomposed into an entity embedding matrix and a relationship embedding matrix by a knowledge representation learning algorithm based on tensor decomposition, and the reconstruction loss is minimized according to a stochastic gradient descent algorithm to obtain a low-dimensional embedding representation corresponding to the entity, relationship and attribute, and the low-dimensional embedding representation is standardized by a regularization term to obtain reference material knowledge, and the distributed representation of the reference material knowledge is spliced ​​with the domain adaptive language model to obtain a joint embedding space; The search keywords are encoded by the domain adaptive language model to obtain an embedding vector representation in the joint embedding space, the semantic relevance score between the embedding vector representation and the standard material knowledge embedding is determined by calculating the vector dot product and the vector modulus, the entities and relationships corresponding to the standard material knowledge embedding are extracted from the multi-source heterogeneous information corresponding to the standard material knowledge embedding, a heterogeneous information network is constructed in combination with a heterogeneous information network representation learning model, a random walk sequence is generated according to a predefined meta-path and mapped to a low-dimensional embedding space, for node pairs in the heterogeneous information network, the semantic association is determined by calculating the meta-path feature vector between each node pair, the meta-path feature vector and the semantic relevance score are weightedly combined to obtain the comprehensive relevance measure, the entities and relationships in the standard material and standard product knowledge base are divided into multiple granularities, multiple granularity levels are determined, and the search keywords are matched and clustered at each granularity level to obtain the candidate search results.

3. The method according to claim 2, characterized in that Minimizing the reconstruction loss according to the stochastic gradient descent algorithm is shown in the following formula: Among them, L represents reconstruction loss, h represents head entity, t represents tail entity, r represents relationship, and Real() represents real part operation. Represents the head entity embedding vector e h The transpose of , T represents the transpose, d r represents the complex diagonal matrix corresponding to the relation r, ⊙ represents element-by-element multiplication, e t represents the embedding vector of the tail entity, λ represents the regularization coefficient, |E| 2 represents the square of the norm of the entity embedding matrix E, |d r | 2 represents the complex diagonal matrix d r The square of the norm of .

4. The method according to claim 1, characterized in that: For the structured features corresponding to the candidate search results, a fuzzy rule base is constructed through evidence theory and fuzzy ensemble learning, and corresponding adaptive feature weights are assigned to the structured features of each dimension. The subspaces corresponding to the structured features are independently clustered based on the adaptive feature weights. The density peak search algorithm and the spectral clustering algorithm are combined to obtain the subspace clustering results corresponding to each subspace. The association dependency is modeled by Markov random field and the subspace clustering results are optimized by the variational inference algorithm to obtain the global optimal clustering results. For the global optimal clustering results, high-order semantic association features are mined through the attribute-weighted heterogeneous information network representation learning model and cross-category semantic associations are performed through collaborative filtering. The generation of high-quality search results combined with the cross-category semantic associations includes: For the candidate search results, the corresponding structured features are extracted by a feature extraction algorithm, the structured features are modeled by a quality function in evidence theory, the quality function and importance coefficient corresponding to the structured features of each dimension are defined, the quality functions corresponding to each dimension are combined by a Dempster combination rule to generate a comprehensive quality function, the structured features are clustered by a fuzzy C-means clustering algorithm to obtain a plurality of fuzzy rules and the comprehensive quality function is updated by using the fuzzy rules as a source of evidence, and the adaptive feature weight is assigned to the structured features corresponding to each dimension based on the updated comprehensive quality function; For the structured features, a subspace corresponding to each structured feature is constructed based on the adaptive feature weights. For each subspace, the local density corresponding to each data point and the minimum distance between the current data point and the data point with the maximum local density are calculated by a density peak search algorithm, and the cluster center is automatically determined to obtain a first clustering result. A Laplace matrix is ​​constructed according to the similarity matrix between different data points by a spectral clustering algorithm, and feature decomposition and clustering are performed to obtain a second clustering result. The first clustering result and the second clustering result are combined to obtain a subspace clustering result corresponding to each subspace. The subspace clustering result is used as a random variable by a Markov random field algorithm and an energy function is defined. The associated dependency is determined by minimizing the energy function and approximately inferred by a variational inference algorithm. The subspace clustering result is optimized according to the variational distribution to obtain the global optimal clustering result. The candidate retrieval results are regarded as nodes in a heterogeneous information network and the structured features are used as attributes of the nodes in the heterogeneous information network. Node embedding learning is performed through an attribute-weighted heterogeneous information network representation learning model, and the network structure preservation items and attribute reconstruction items are jointly optimized to generate high-order semantic association features. The global optimal clustering result is used as the user's preference category in combination with the collaborative filtering idea to determine the cross-category semantic association between the candidate retrieval results. Based on the cross-category semantic association and the global optimal clustering result, the high-quality retrieval result is solved.

5. The method according to claim 4, characterized in that The objective function of the attribute-weighted heterogeneous information network representation learning model is shown in the following formula: Among them, H represents the objective function value, w ij represents the edge weight between the i-th node and the j-th node, v i represents the embedding vector of the i-th node, v j represents the embedding vector of the jth node, ρ represents the balance parameter, and a i represents the attribute feature of the i-th node, μ represents the regularization coefficient, and ||θ|| 2 represents the norm of the model parameters θ.

6. The method according to claim 1, characterized in that Based on the high-quality search results, the search result influencing factors are constructed as a high-dimensional scoring tensor through a hierarchical factor decomposition machine model, and low-rank decomposition is performed in combination with a high-order singular value decomposition algorithm to generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results. Based on the user's personal preferences, a ranking model is constructed in combination with the novelty of the search keywords and the initial comprehensive score. The ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, and the action space and state space are determined in combination with a dual-tower architecture. A reward function is constructed and the ranking position of each standard substance is determined based on the reward function value. The optimal ranking results include: For the high-quality search results, define the search result influencing factors, determine the corresponding causal decomposition machine loss function through the hierarchical causal decomposition machine model, repeatedly optimize the causal decomposition machine loss function through the alternating least squares method and gradually update the corresponding factor matrix, calculate the high-dimensional score tensor corresponding to the search result influencing factors according to the updated factor matrix, perform high-order singular value decomposition on the high-dimensional score tensor, generate a core tensor and a factor matrix, extract high-order interaction information in different search result influencing factors from the core tensor, and for each high-quality search result, calculate the initial comprehensive score through tensor operation based on the high-order interaction information and the factor matrix; Based on the initial comprehensive score, based on the pre-acquired personal preferences corresponding to the current user, and combined with the novelty corresponding to the search keyword, a ranking model is constructed, the standard substance and standard product ranking problem is modeled as a Markov decision process based on a deep reinforcement learning algorithm, a state tower and an action tower are constructed through a dual-tower strategy network, the state tower models the input sequence composed of the high-quality search results through a bidirectional long short-term memory network, determines the hidden state and generates a state embedding representation through a multi-layer perceptron, uses the personal preference, novelty and the initial comprehensive score as inputs to the action tower, generates an input feature vector and generates a corresponding action embedding representation through a multi-layer perceptron, constructs a reward function, and calculates the function value of the reward function by calculating the inner product of the state embedding representation and the action embedding representation; The maximum reward function value is selected and the corresponding action embedding representation and state embedding representation are determined, an optimal action is generated and the optimal action is parsed into a sorting result to obtain the optimal sorting result.

7. The method according to claim 6, characterized in that The high-order singular value decomposition of the high-dimensional score tensor is shown in the following formula: Among them, X represents the high-dimensional rating tensor, represents the elements of the core tensor, represents the singular vectors in the first dimension, represents the singular vectors in the second dimension, represents the singular vector in the Mth dimension, R M represents the rank of dimension M, M represents the total number of dimensions, and ° represents the outer product operator.

8. A reference material and standard product retrieval and ranking system based on a search engine, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the search keywords input by the user, build a general language model based on a bidirectional encoder, fine-tune the general language model in combination with a small amount of standard material domain corpus to obtain a domain adaptive language model, build a standard material and standard product knowledge base based on ontology and knowledge graph, map the entities, relationships and attributes in the standard material and standard product knowledge base to a low-dimensional semantic space in combination with a knowledge representation learning algorithm based on tensor decomposition, generate a distributed representation of standard material knowledge and combine it with the domain adaptive language model, calculate the semantic relevance between the search keywords and the standard material knowledge, build the multi-source heterogeneous information corresponding to the standard material and standard product into a heterogeneous information network in combination with a heterogeneous information network representation learning model, determine the semantic association in combination with a meta-path method and generate a comprehensive relevance measure in combination with the semantic relevance, perform multi-granular semantic matching based on the comprehensive relevance measure, and obtain candidate search results; The second unit is used to construct a fuzzy rule base for the structured features corresponding to the candidate search results through evidence theory and fuzzy ensemble learning, and assign corresponding adaptive feature weights to the structured features of each dimension, independently cluster the subspaces corresponding to the structured features based on the adaptive feature weights, obtain the subspace clustering results corresponding to each subspace by combining the density peak search algorithm and the spectral clustering algorithm, model the association dependency through Markov random field and optimize the subspace clustering results through the variational inference algorithm to obtain the global optimal clustering results, and for the global optimal clustering results, mine high-order semantic association features through the attribute-weighted heterogeneous information network representation learning model and perform cross-category semantic association through collaborative filtering, and generate high-quality search results in combination with the cross-category semantic association; The third unit is used to construct the search result influencing factors into a high-dimensional scoring tensor based on the high-quality search results through a hierarchical factor decomposition machine model, perform low-rank decomposition in combination with a high-order singular value decomposition algorithm, generate high-order interactions between different influencing factors and obtain an initial comprehensive score corresponding to the high-quality search results, construct a ranking model based on the user's personal preferences in combination with the novelty of the search keywords and the initial comprehensive score, model the ranking problem as a Markov decision process based on a deep reinforcement learning algorithm, determine the action space and state space in combination with a dual-tower architecture, construct a reward function and determine the ranking position of each standard substance based on the reward function value to obtain the optimal ranking result.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data retrieval / intelligent question and answer method and device and storage medium

    CN112463926A

  • Artificial intelligence data search and distribution method and system

    CN118467851A