Cultural relic knowledge processing method and related device
By improving the quantum spin optimization and inter-layer nonlinear coupling constraint loss function of the Transformer model, the problem that traditional methods are difficult to deal with large-scale multi-dimensional cultural relics knowledge is solved, and efficient automatic processing and analysis of cultural relics knowledge is achieved.
Patent Information
- Application Number
- CN202510907795.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Traditional cultural relics knowledge processing methods are difficult to cope with large-scale and multi-dimensional cultural relics knowledge, and are inefficient in processing and difficult to meet the needs of modern research.
The pre-trained word vector is used to embed cultural relics text data, and the improved Transformer model is used to process cultural relics knowledge. The weight of the feedforward network is initialized through the quantum spin optimization algorithm, and the inter-layer nonlinear coupling constraint loss function and quantum acceleration gradient update mechanism is combined to improve processing speed and accuracy.
It realizes efficient and automatic processing of large-scale and multi-dimensional cultural relics knowledge, improves the classification, semantic analysis and feature extraction capabilities of cultural relics knowledge, and ensures the stability and generalization capabilities of the model.
Smart Images

Figure CN120409475A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and more particularly, to a method for processing cultural relic knowledge and related devices. Background Art
[0002] The comprehensiveness and complexity of cultural relic knowledge make its processing and analysis extremely difficult. Traditional processing of cultural relic knowledge usually relies on manual analysis. However, with the increasing demand for cultural relic research and the development of big data technology, traditional methods for processing cultural relic knowledge are difficult to cope with large-scale and multi-dimensional cultural relic knowledge. Summary of the Invention
[0003] In view of this, the present invention discloses a method for processing cultural relic knowledge and related devices to solve the problem that traditional methods for processing cultural relic knowledge are difficult to cope with large-scale and multi-dimensional cultural relic knowledge.
[0004] A method for processing cultural relic knowledge includes:
[0005] Obtaining cultural relic text data corresponding to the cultural relic knowledge to be processed;
[0006] Performing text embedding on the cultural relic text data by using pre-trained word vectors to convert the cultural relic text data into cultural relic word vectors;
[0007] Inputting the cultural relic word vectors into the encoder of a pre-trained improved Transformer model to map them into a set of context-aware target word vectors, wherein, during the training of the improved Transformer model, the weights of the feed-forward network are initialized by using a quantum spin optimization algorithm, a layer-intermediate non-linear coupling constraint loss function is adopted, and a quantum acceleration gradient update mechanism is adopted to accelerate the processing speed;
[0008] Inputting the target word vectors into the decoder of the improved Transformer model for decoding to obtain the result of processing cultural relic knowledge.
[0009] Optionally, during the training process of the improved Transformer model, the process of initializing the weights of the feed-forward network by using a quantum spin optimization algorithm includes:
[0010] Calculating the product of a quantum optimization operator and the original weight matrix of the feed-forward network to obtain the weight matrix of the feed-forward network, wherein the weight matrix is the matrix after the original weight matrix is initialized, and the quantum optimization operator represents a quantum spin rotation operation and is determined based on the quantum spin rotation angle.
[0011] Optionally, the process of determining the quantum spin rotation angle includes:
[0012] Calculate the product of the artifact word vector and the original weight matrix in the artifact text data to obtain an inner product representing the initial association strength between the artifact word vector and the original weight matrix;
[0013] Based on the inner product, the initial offset compensation of the sparse features in the artifact text data, and the adjustment factor, obtain the quantum spin rotation angle.
[0014] Optionally, the training process of the improved Transformer model further includes:
[0015] Based on the weight matrix of the feed-forward network, determine the inter-layer non-linear coupling constraint loss function.
[0016] Optionally, determining the inter-layer non-linear coupling constraint loss function based on the weight matrix of the feed-forward network includes:
[0017] Determine the first non-linear function corresponding to the weight matrix of the i-th layer of the feed-forward network, the second non-linear function corresponding to the weight matrix of the (i + 1)-th layer of the feed-forward network, and the coupling degree coefficient of the i-th layer;
[0018] Determine the product of the first non-linear function, the second non-linear function, and the coupling degree coefficient of the i-th layer as the inter-layer non-linear coupling constraint loss function.
[0019] Optionally, the determination process of the coupling degree coefficient of the i-th layer includes:
[0020] According to the correlation between the activation value of the i-th layer and the activation value of the (i + 1)-th layer of the feed-forward network, as well as the artifact word vector and the number of training samples used as training samples, obtain the coupling degree coefficient of the i-th layer.
[0021] Optionally, the training process of the improved Transformer model further includes:
[0022] Based on the inter-layer non-linear coupling constraint loss function, dynamically update the weight matrix of the feed-forward network using a quantum feedback mechanism.
[0023] Optionally, dynamically updating the weight matrix of the feed-forward network using a quantum feedback mechanism based on the inter-layer non-linear coupling constraint loss function includes:
[0024] Based on the learning rate representing the step size of each iterative update, the weight gradient, the quantum acceleration gradient, the gradient of the inter-layer non-linear coupling constraint loss function with respect to the weight, the quantum spin optimization feedback adjustment coefficient, and the weight adjustment brought by the quantum spin feedback, update the weight matrix of the feed-forward network at the t-th iteration to obtain the weight matrix of the feed-forward network at the (t + 1)-th iteration.
[0025] Optionally, the process of determining the learning rate includes:
[0026] Based on the weight gradient, feedback gain factor, and quantum spin feedback amount, the learning rate is obtained.
[0027] Optionally, the process of determining the quantum acceleration gradient includes:
[0028] Based on the quantum acceleration factor, the quantum optimization operator representing the quantum spin rotation operation, the cultural relic word vector as a training sample, and the gradient of the cross-entropy loss function with respect to the nth training sample, the quantum acceleration gradient is obtained.
[0029] A cultural relic knowledge processing device includes:
[0030] An acquisition unit, configured to acquire cultural relic text data corresponding to the cultural relic knowledge to be processed;
[0031] A vector conversion unit, configured to perform text embedding on the cultural relic text data by using pre-trained word vectors, and convert the cultural relic text data into cultural relic word vectors;
[0032] An encoding unit, configured to input the cultural relic word vectors into the encoder of a pre-trained improved Transformer model, and map them to a set of context-aware target word vectors, where, when training the improved Transformer model, the weights of the feed-forward network are initialized by using a quantum spin optimization algorithm, a layer-wise non-linear coupling constraint loss function is adopted, and a quantum acceleration gradient update mechanism is adopted to accelerate the processing speed;
[0033] A decoding unit, configured to input the target word vectors into the decoder of the improved Transformer model for decoding, and obtain the cultural relic knowledge processing result.
[0034] A computer storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, any one of the cultural relic knowledge processing methods is implemented.
[0035] An electronic device includes: a memory and a processor;
[0036] The memory is used to store at least one instruction;
[0037] The processor is used to execute the at least one instruction to implement any one of the cultural relic knowledge processing methods.
[0038] As can be seen from the above technical solutions, the present invention discloses a method for processing cultural relic knowledge and related devices. The method includes obtaining cultural relic text data corresponding to the cultural relic knowledge to be processed, performing text embedding on the cultural relic text data using pre-trained word vectors to convert the cultural relic text data into cultural relic word vectors, inputting the cultural relic word vectors into the encoder of a pre-trained improved Transformer model to map them into a set of context-aware target word vectors, and inputting the target word vectors into the decoder of the improved Transformer model for decoding to obtain the processing result of the cultural relic knowledge. The present invention realizes the automatic processing of cultural relic knowledge through the improved Transformer model. The improved Transformer model initializes the weights of the feed-forward network using the quantum spin optimization algorithm to ensure the stability of weight initialization, uses an inter-layer non-linear coupling constraint loss function to better process multi-dimensional features in the cultural relic text data, and uses a quantum-accelerated gradient update mechanism to speed up the processing speed. Therefore, it can efficiently process large-scale and multi-dimensional cultural relic knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the disclosed drawings without creative efforts.
[0040] Figure 1 It is a schematic diagram of the architecture of a traditional Transformer model;
[0041] Figure 2 It is a flowchart of a method for processing cultural relic knowledge disclosed in an embodiment of the present invention;
[0042] Figure 3 It is a comparison diagram of the weight distributions of the quantum initialization of the feed-forward network weights and the traditional random initialization of the feed-forward network weights disclosed in an embodiment of the present invention;
[0043] Figure 4 It is a comparison diagram of the gradient flow field with inter-layer coupling constraints and the gradient flow field without coupling constraints disclosed in an embodiment of the present invention;
[0044] Figure 5 It is a comparison schematic diagram of a quantum feedback path and a traditional optimization path disclosed in an embodiment of the present invention;
[0045] Figure 6 It is a schematic diagram of the structure of a cultural relic knowledge processing device disclosed in an embodiment of the present invention;
[0046] Figure 7 It is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. Detailed implementation manners
[0047] The improved Transformer model used in the cultural relic knowledge processing method disclosed by the present invention is obtained by improving the traditional Transformer model. For the convenience of understanding the working principle of the improved Transformer model, the structure and working principle of the traditional Transformer model are described below as follows:
[0048] See Figure 1 The schematic architecture diagram of the traditional Transformer model shown. The Transformer model mainly consists of two parts: an encoder and a decoder. The encoder maps the input into a series of context-aware vector representations, and the decoder gradually decodes these vector representations into target features.
[0049] Specifically, the encoder part is composed of N trans identical encoding layers, where N trans represents the number of encoding layers. For example, the value of N trans is 5. Each encoding layer consists of two sub-layers: a multi-head attention mechanism sub-layer and a feed-forward network sub-layer. The encoder adds a residual connection ( Figure 1 not shown in the figure) and a normalization operation after each sub-layer.
[0050] The multi-head attention mechanism sub-layer is implemented through the multi-head self-attention mechanism. The multi-head self-attention mechanism is the core module of the encoder, responsible for learning the importance of each position from the input features, and combining the information of all positions to obtain a global context representation, thereby improving the performance and generalization ability of the model. Specifically, the implementation method of the multi-head self-attention mechanism sub-layer is as follows: The input features pass through three different linear transformation layers to generate "query", "key", and "value" vectors respectively. The linear transformation is realized by multiplying the weight matrix. Further, the generated query, key, and value vectors are divided into multiple smaller parts, that is, "heads", so that the model can process multiple different representation sub-spaces in parallel, thereby capturing the diversity information in the input data. Further, for each head, the dot product of the query and key vectors is calculated respectively to obtain the attention score. The attention score reflects the relative importance between different positions in the sequence. Moreover, in order to make the numerical value of the attention score stable, the score is divided by a scaling factor, and the scaling factor can be the square root of the dimension of the key vector. Further, the Softmax function is used to process the attention score to obtain the attention weight of each position. The Softmax function ensures that all weights add up to 1, so that it can be processed as a probability distribution, expressed as:
[0051] (1);
[0052] where head i is the attention weight of the i-th head, Q i is the query vector of the i-th head, K i is the key vector of the i-th head, K i T is the transpose of the key vector of the i-th head, d ks is the dimension of the input features, h ks is the number of attention heads, V i is the value vector of the i-th head, and Sof() is the Softmax function.
[0053] Furthermore, the obtained attention weights are used to perform a weighted sum on the value vectors to obtain the output vector. The output vectors of all heads are concatenated together, and then the output of the final multi-head self-attention mechanism sub-layer is generated through a linear transformation layer.
[0054] The feed-forward network sub-layer consists of two fully-connected layers and an activation function, and is used to perform non-linear transformation on the vectors at each position. In the feed-forward network sub-layer, the input vector first undergoes a linear transformation through the first fully-connected layer, then undergoes a non-linear transformation through a ReLU activation function, and finally undergoes a linear transformation through the second fully-connected layer to improve the expression ability and generalization ability of the model, so that the model can better process complex input data.
[0055] The decoder consists of N trans identical decoding layers. Different from the encoding layer, each decoding layer has one more multi-head attention mechanism sub-layer, that is, the decoding layer contains two multi-head attention mechanism sub-layers and one feed-forward network sub-layer. Moreover, the second multi-head attention mechanism sub-layer of the decoding layer (i.e., Figure 1 the masked multi-head attention layer in
[0056] is implemented through the multi-head cross-attention mechanism. The multi-head cross-attention mechanism is a mechanism for information interaction between the encoder and the decoder. It takes the output features of the encoder and the input features of the decoder as inputs, and simultaneously uses the attention mechanism to perform cross-attention calculation on the inputs, so as to obtain a set of weighted decoder output features.
[0057] It should be noted that in the encoder and decoder of the Transformer model, the output of each sub-layer is the result of layer normalization and residual connection, which is expressed as:
[0058] (2);
[0059] where y cgyis the feature output after layer normalization, LNc() is the layer normalization function, and x trans is the input of the sublayer, and Sublayer() is the output of the sublayer with residual connection.
[0060] In the present invention, the improved Transformer model is based on the traditional Transformer model. The weights of the feed-forward network (including the feed-forward network sublayer in the encoder and the feed-forward network sublayer in the decoder) are initialized using the quantum spin optimization algorithm, reducing the weight initialization instability and error caused by the random initialization method of the traditional feed-forward network, as well as the phenomenon of gradient disappearance or gradient explosion. The improved Transformer model in the present invention uses an inter-layer non-linear coupling constraint loss function, enabling the improved Transformer model to better handle multi-dimensional features (such as vocabulary, syntax, and semantics) in cultural relic text data, enhancing the model's expressive ability and generalization ability. The improved Transformer model in the present invention uses a quantum-accelerated gradient update mechanism to accelerate the processing speed, significantly improving the gradient calculation speed, and thus can significantly improve the model training efficiency when processing cultural relic text data.
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0062] The embodiments of the present invention disclose a method and related device for processing cultural relic knowledge. The cultural relic text data corresponding to the cultural relic knowledge to be processed is obtained, and the cultural relic text data is text-embedded using pre-trained word vectors to convert the cultural relic text data into cultural relic word vectors. The cultural relic word vectors are input into the encoder of the pre-trained improved Transformer model and mapped into a set of context-aware target word vectors. The target word vectors are input into the decoder of the improved Transformer model for decoding to obtain the cultural relic knowledge processing result. The present invention realizes the automatic processing of cultural relic knowledge through the improved Transformer model. The improved Transformer model uses the quantum spin optimization algorithm to initialize the weights of the feed-forward network to ensure the stability of weight initialization, uses an inter-layer non-linear coupling constraint loss function to better handle multi-dimensional features in cultural relic text data, and uses a quantum-accelerated gradient update mechanism to accelerate the processing speed. Therefore, it can efficiently process large-scale and multi-dimensional cultural relic knowledge.
[0063] See Figure 2 , the embodiments of the present invention disclose a flowchart of a method for processing cultural relic knowledge, and the method includes:
[0064] Step S101: Obtain the cultural relic text data corresponding to the cultural relic knowledge to be processed.
[0065] The cultural relic knowledge to be processed includes, but is not limited to, multi-dimensional contents such as the historical background, cultural connotation, text description, unearthed information, etc. of cultural relics, and the cultural relic knowledge to be processed has high precision and high reliability.
[0066] Step S102: Perform text embedding on the cultural relic text data by using pre-trained word vectors, and convert the cultural relic text data into cultural relic word vectors.
[0067] Since the input of the improved Transformer model is an embedded vector, it is necessary to convert the cultural relic text data into a vector representation suitable for the improved Transformer model to ensure that the cultural relic text data can be effectively processed by the model. In the present invention, pre-trained word vectors are used to perform text embedding on the cultural relic text data to obtain cultural relic word vectors suitable for the improved Transformer model.
[0068] Examples of pre-trained word vectors include BERT (Bidirectional Encoder Representations from Transformers, a bidirectional encoder representation model based on Transformers), GPT (Generative Pre-trained Transformer, a generative pre-trained Transformer model), etc.
[0069] In practical applications, before performing text embedding on the cultural relic text data by using pre-trained word vectors, preprocessing operations can be first performed on the cultural relic text data, including: denoising, standardization, normalization, etc. operations, to provide clear and consistent data for subsequent data analysis. Then, pre-trained word vectors are used to perform text embedding on the preprocessed cultural relic text data, and the preprocessed cultural relic text data is converted into cultural relic word vectors.
[0070] Step S103: Input the cultural relic word vectors into the encoder of the pre-trained improved Transformer model, and map them into a set of context-aware target word vectors.
[0071] The encoder of the pre-trained improved Transformer model maps the input cultural relic word vectors into a set of context-aware vector representations, that is, target word vectors, according to the learned cultural relic features and context information.
[0072] It should be noted that the pre-trained improved Transformer model in the present invention is actually a large cultural relic knowledge model for realizing the processing of cultural relic knowledge.
[0073] Step S104: Input the target word vector into the decoder of the improved Transformer model for decoding to obtain the processing result of cultural relic knowledge.
[0074] Among them, the processing result of cultural relic knowledge obtained by decoding can be a text classification label, a feature extraction result, a speech analysis result, etc.
[0075] The improved Transformer model in the present invention is a knowledge - answering model obtained by pre - training. The output result of the improved Transformer model is related to the input question.
[0076] It should be noted that during the inference process, the improved Transformer model uses the multi - head attention mechanism between each layer to capture important information in the text and processes the input data according to the optimized weight matrix. At the same time, the inference process does not involve new parameter updates, but uses the cultural relic knowledge obtained by the improved Transformer model during the training stage to generate the processing result of the input data through inference.
[0077] In summary, the present invention discloses a method for processing cultural relic knowledge. It acquires the cultural relic text data corresponding to the cultural relic knowledge to be processed, performs text embedding on the cultural relic text data using pre - trained word vectors, converts the cultural relic text data into cultural relic word vectors, inputs the cultural relic word vectors into the encoder of the pre - trained improved Transformer model to be mapped into a set of context - aware target word vectors, and inputs the target word vectors into the decoder of the improved Transformer model for decoding to obtain the processing result of cultural relic knowledge. The present invention realizes the automatic processing of cultural relic knowledge through the improved Transformer model. The improved Transformer model uses the quantum spin optimization algorithm to initialize the weights of the feed - forward network to ensure the stability of weight initialization, uses the inter - layer non - linear coupling constraint loss function to better process the multi - dimensional features in the cultural relic text data, and uses the quantum - accelerated gradient update mechanism to speed up the processing speed. Therefore, it can efficiently process large - scale and multi - dimensional cultural relic knowledge.
[0078] The training process of the improved Transformer model in the present invention mainly includes the following units:
[0079] (1) Data acquisition unit
[0080] The data acquisition unit is used to acquire and organize the cultural relic text data of cultural relic knowledge to ensure the comprehensiveness and quality of the cultural relic text data.
[0081] The functions of the data acquisition unit include collecting cultural relic text data, covering multi - dimensional contents such as the historical background of cultural relics, cultural connotations, text descriptions, unearthed information, etc., and the cultural relic text data should have high precision and high reliability.
[0082] The data acquisition unit also includes a data preprocessing function, which performs operations such as denoising, standardization, and normalization on the collected original cultural relic text data, providing clear and consistent input data for subsequent data analysis and modeling.
[0083] Furthermore, the collected cultural relic text data is constructed into a training set. Taking the improved Transformer model as the basic architecture of the large model as an example, in order to train the improved Transformer model, it is necessary to prepare the cultural relic text data of cultural relic knowledge. Specifically, the cultural relic text data needs to be labeled and classified according to dimensions such as theme, time, and region (labeled according to different classification dimensions) to ensure data structuring so that the model can understand different categories and context relationships. In addition, the collected cultural relic text data should have sufficient diversity and scale to ensure that it can support the deep learning requirements of the improved Transformer model.
[0084] Since the input of the improved Transformer model is an embedded vector, it is necessary to encode the cultural relic text data. The present invention uses pre-trained word vectors for text embedding, converting the cultural relic text data into a vector representation suitable for the improved Transformer model, that is, cultural relic word vectors, ensuring that cultural relic knowledge can be effectively understood by the model.
[0085] Furthermore, the cultural relic text data is divided into a training set, a validation set, and a test set according to a certain proportion to ensure that the model performance can be effectively evaluated during the training process.
[0086] The training set is used for model learning, the validation set is used for model tuning, and the test set is used to evaluate the generalization ability of the final model.
[0087] Furthermore, since the improved Transformer model depends on the input sequence data when performing self-attention mechanism calculations, the cultural relic text data needs to be segmented by sequence. According to the characteristics of cultural relic knowledge, the data acquisition unit can set an appropriate maximum sequence length to ensure that each input sequence can contain sufficient context information and avoid information loss.
[0088] (2) Data storage unit
[0089] The data storage unit is used to store and manage large-scale cultural relic text data, ensuring the efficient access and security of cultural relic text data.
[0090] The functions of the data storage unit include the long-term storage and fast query of cultural relic text data, and the data is stored through a database or a distributed storage system.
[0091] The data storage unit should also support high-concurrency data access and be able to ensure data backup, recovery, and data consistency to guarantee the integrity and security of cultural relic text data.
[0092] (III) Machine learning modeling unit
[0093] The machine learning modeling unit is used to analyze and model the collected cultural relic text data through machine learning algorithms to discover potential laws and knowledge in the data.
[0094] The functions of the machine learning modeling unit include selecting appropriate natural language processing algorithms for modeling, such as word vector models (Word2Vec), long short-term memory networks (LSTM), transformers, etc., and optimizing the model using the training set to enable it to perform tasks such as automatic classification, feature extraction, and semantic analysis of cultural relic knowledge.
[0095] The machine learning modeling unit also includes model verification and evaluation functions, and comprehensively evaluates the accuracy, robustness, and generalization ability of the model using evaluation methods such as cross-validation, accuracy, and recall rate.
[0096] Taking the improved Transformer model as the basic architecture of the large model, the machine learning modeling unit constructs and trains the improved Transformer model, and uses the trained improved Transformer model to process cultural relic knowledge.
[0097] (IV) Model deployment and application unit
[0098] The model deployment and application unit is used to apply the trained improved Transformer model to actual tasks and provide intelligent services related to cultural relic knowledge for users.
[0099] The functions of the model deployment and application unit include deploying the trained model to the production environment and interacting with users through APIs (Application Programming Interfaces), front-end interfaces, etc. The model deployment and application unit also includes real-time inference functions, which can perform real-time analysis on the cultural relic text information submitted by users and provide corresponding classification, recognition, or recommendation results.
[0100] (V) Knowledge update and management unit
[0101] Knowledge update and management is used to continuously update the cultural relic knowledge base to maintain the timeliness and accuracy of cultural relic-related knowledge.
[0102] The functions of the knowledge update and management unit include fine-tuning the existing model according to new cultural relic knowledge text data or knowledge reasoning results to ensure that the model can adapt to new data changes.
[0103] The knowledge update and management unit also includes a knowledge management function, which classifies, stores, and retrieves cultural relic knowledge, supports intelligent knowledge update and version management, and provides continuous and effective cultural relic information support.
[0104] (6) Evaluation and feedback unit
[0105] The evaluation and feedback unit is used to monitor and evaluate the effect of the trained improved Transformer model in real time to ensure the continuous optimization and efficient operation of the model.
[0106] The functions of the evaluation and feedback unit include evaluating the model through indicators such as user feedback and prediction accuracy, analyzing the performance of the model in actual applications, and timely discovering and fixing potential problems.
[0107] The evaluation and feedback unit also includes a model retraining mechanism, which retrains the model based on the feedback results to continuously improve the performance and adaptability of the model.
[0108] It should be noted that the feed-forward network of the Transformer model is a neural network structure composed of neurons, which plays a role in enhancing the non-linear processing ability of the model. The neural network structure includes weight and bias parameters. To avoid the gradient disappearance and gradient descent phenomena brought about by using the gradient descent method to update parameters in traditional neural networks, and to avoid the phenomenon that parameters fall into local optimal solutions after training, in the training process of the improved Transformer model of the present invention, a training method based on a quantum feedback mechanism is adopted for training the feed-forward network, and the training process is as follows:
[0109] (1) Initialize the weights of the feed-forward network using the quantum spin optimization algorithm.
[0110] In the training of cultural relic text data, the optimization of weight initialization can significantly reduce the gradient disappearance or gradient explosion phenomena caused by random initialization. The present invention optimizes the initial weights of the feed-forward network to avoid the instability and error problems brought about by random initialization of traditional feed-forward networks, and at the same time improve the convergence speed of feed-forward network training. By initializing the weights of the feed-forward network using the quantum spin optimization algorithm, it can more accurately approach the global optimal solution and avoid falling into local minima.
[0111] Therefore, in the training process of the improved Transformer model, the process of initializing the weights of the feed-forward network using the quantum spin optimization algorithm includes:
[0112] Calculate the product of the quantum optimization operator and the original weight matrix of the feedforward network to obtain the weight matrix of the feedforward network.
[0113] Among them, the weight matrix is the matrix after initializing the original weight matrix, and the quantum optimization operator represents the quantum spin rotation operation, which is determined based on the quantum spin rotation angle.
[0114] The expression of the weight matrix of the feedforward network is as follows:
[0115] (3);
[0116] In the formula, W is the weight matrix of the feedforward network, which represents the connection strength between neurons and is used to adjust the influence between each neuron in the network model, directly affecting the feature learning and representation ability of cultural relic text data; is the quantum optimization operator, which represents the quantum spin rotation operation, [[ID=1�]] is the quantum spin rotation angle, which represents the adjustment parameter of quantum optimization; R is the original weight matrix in the feedforward network, which represents the initial weight of the traditional feedforward network.
[0117] The quantum spin optimization algorithm adjusts the weight matrix of the feedforward network by adopting the quantum spin rotation angle, and the quantum spin rotation angle is calculated through the inner product of the input data and the weight matrix. Because of the diversity and complexity of cultural relic text data, more accurate initialization is required to avoid the instability of model training, so it is particularly important for feature learning in cultural relic text data.
[0118] The determination process of the quantum spin rotation angle in the present invention includes:
[0119] Calculate the product of the cultural relic word vector in the cultural relic text data and the original weight matrix to obtain the inner product representing the initial correlation strength between the cultural relic word vector and the original weight matrix;
[0120] Based on the inner product, the initial offset compensation of the sparse features in the cultural relic text data, and the adjustment factor, obtain the quantum spin rotation angle.
[0121] Quantum spin rotation angle The expression is as follows:
[0122] (4);
[0123] In the formula, is X i is the i-th training sample, representing the input data. Among them, in the cultural relic text data processing task, N is the number of training samples, and X i is the cultural relic word vector; R i is the i-th weight matrix element, representing the value at the corresponding position in the original weight matrix, and X i ·R iis the inner product of the training sample and the original weight matrix, representing the initial correlation strength between the cultural relic word vector and the original weight matrix; is a constant bias term, which characterizes the initial offset in the quantum optimization process, specifically the initial offset compensation of sparse features in cultural relic text data; To adjust the factor, control the scale of the quantum rotation angle, and prevent the inner product value caused by the high-dimensional word vector of the artifact from being too large (to avoid gradient explosion). Preferably, Set to 0.2.
[0124] In one embodiment, the parameters Figure 3 This paper presents a comparison of the weight distributions of a quantum-initialized feedforward network weights and a traditional randomly initialized feedforward network weights. To verify the effectiveness of the quantum spin optimization algorithm initialization method (referred to as "quantum initialization") in improving the stability of the weight matrix, by comparing it with the traditional random initialization method (referred to as "traditional initialization"), it is shown that the weight values of the present invention exhibit a multi-peak clustered distribution characteristic, rather than the dispersed unimodal Gaussian distribution used in the traditional method. Experimental results demonstrate that the quantum rotation operation precisely controls the initial position of the weights, bringing the weight matrix closer to the neighborhood of multiple potential global optimal solutions. This avoids the problem of excessively wide weight distribution caused by the randomness of the traditional random initialization method, confirming the optimization effect of quantum initialization on gradient stability and convergence direction from a mathematical distribution perspective.
[0125] (2) Calculate the inter-layer nonlinear coupling constraint loss function.
[0126] In practical applications, the training process of the improved Transformer model also includes:
[0127] Based on the weight matrix of the feedforward network, the inter-layer nonlinear coupling constraint loss function is determined.
[0128] The training of traditional feedforward networks usually assumes that the weight adjustments between each layer are performed independently. However, in complex data scenarios, the coupling relationship between parameters is often ignored, which may lead to insufficient model fitting ability for the data. In order to improve the performance of feedforward networks when processing complex and nonlinear data, the present invention adopts nonlinear coupling constraints to ensure that the parameter adjustments between the layers of the feedforward network can affect each other, avoiding the degradation of model performance caused by ignoring the interaction between layers. The nonlinear coupling constraints are constrained during the training process through the loss function. In the processing of cultural relics text data, it helps to improve the feedforward network's ability to express multidimensional features such as vocabulary, syntax and semantics.
[0129] In the present invention, based on the weight matrix of the feedforward network, the inter-layer nonlinear coupling constraint loss function is determined, including:
[0130] Determine the first non - linear function corresponding to the weight matrix of the \(i\) - th layer of the feed - forward network, the second non - linear function corresponding to the weight matrix of the \((i + 1)\) - th layer of the feed - forward network, and the coupling degree coefficient of the \(i\) - th layer;
[0131] Determine the product of the first non - linear function, the second non - linear function, and the coupling degree coefficient of the \(i\) - th layer as the inter - layer non - linear coupling constraint loss function.
[0132] The expression of the inter - layer non - linear coupling constraint loss function is as follows:
[0133] (5);
[0134] In the formula, \(L\) cou is the inter - layer non - linear coupling constraint loss function, which is used to control the coupling relationship between layers. In the task of cultural relic text data processing, cultural relic text data usually has a multi - level structure (such as vocabulary, syntax, and semantics), and inter - layer coupling helps to improve the model's ability to express these multi - dimensional features; \(W\) i is the weight matrix of the \(i\) - th layer of the feed - forward network, \(W\) i+1 is the weight matrix of the \((i + 1)\) - th layer of the feed - forward network, \(W\) i and \(W\) i+1 represent the connection strength between two adjacent layers in the feed - forward network. \(f(W)\) is a non - linear function, which characterizes the non - linear transformation of the weights of the feed - forward network. \(f(W)\) can adopt the Sigmoid non - linear activation function. \(f(W\) i ) represents the first non - linear function corresponding to the weight matrix of the \(i\) - th layer of the feed - forward network, \(f(W\) i+1 ) is the second non - linear function corresponding to the weight matrix of the \((i + 1)\) - th layer of the feed - forward network. is the coupling degree coefficient of the \(i\) - th layer.
[0135] Furthermore, the determination process of the coupling degree coefficient \(\lambda_i\) includes:
[0136] Obtain the coupling degree coefficient \(\lambda_i\) of the \(i\) - th layer according to the correlation between the activation value of the \(i\) - th layer and the activation value of the \((i + 1)\) - th layer of the feed - forward network, as well as the cultural relic word vectors and the number of training samples used as training samples.
[0137] In order to accurately reflect the non - linear relationship between layers, adjusting the coefficient of the coupling degree of each layer can effectively capture the non - linear relationship between input features and activations, and improve the model's processing ability on complex and high - dimensional cultural relic knowledge text data. The calculation expression of the coupling degree coefficient \(\lambda_i\) of the \(i\) - th layer is as follows:
[0138] (6);
[0139] In the formula, \(A\) i is the activation value of the \(i\) - th layer of the feed - forward network, \(A\)i+1 is the activation value of the (i + 1)-th layer of the feedforward network, A i and A i+1 both represent the layer output of the feedforward network, A T i+1 is the transpose of A i+1 A i ·A T i+1 characterizes the correlation between the activation value of the i-th layer and the activation value of the (i + 1)-th layer of the feedforward network, capturing the hierarchical features of vocabulary → syntax → semantics in the cultural relic text, X j is the j-th training sample, representing the input features, specifically the cultural relic word vectors in the cultural relic text data, X T j is the transpose of X j M is the number of training samples.
[0140] In one embodiment, to analyze the enhancement mechanism of the inter-layer nonlinear coupling constraint loss function on inter-layer information transmission, see Figure 4 , a comparison diagram of the gradient flow field with inter-layer coupling constraints and the gradient flow field without coupling constraints provided by the embodiments of the present invention. Among them, (a) is a schematic diagram of the gradient flow field with inter-layer coupling constraints, and (b) is a schematic diagram of the gradient flow field without coupling constraints. Figure 4 The abscissa in is the weight dimension 1, the ordinate is the weight dimension 2, and the gradient intensity is shown. By comparing the propagation directions and intensities of the gradient flow field with inter-layer coupling constraints and the gradient flow field without coupling constraints, it can be seen that the gradient flow field with inter-layer coupling constraints presents an obvious vortex structure and direction consistency, while the gradient flow direction shown by the traditional gradient flow field without coupling constraints has disorder and divergence phenomena. The experimental results prove that the inter-layer coupling constraint guides the gradient flow field to propagate along the hierarchical direction of feature expression by establishing a non-linear association between inter-layer parameters, thereby significantly improving the collaborative extraction ability of the Transformer model for vocabulary, syntax, and semantic features in the text, and solving the problem of feature transmission attenuation caused by the isolation between traditional feedforward network layers.
[0141] (3) Dynamically update the weights using a quantum feedback mechanism.
[0142] To improve the optimization accuracy and avoid local minima, the present invention adopts the feedback mechanism in quantum spin optimization. During the quantum spin optimization process, the adjustment of the optimization path is not static, but dynamically adjusted through the feedback mechanism to ensure that in the training process, each optimization can gradually approach the optimal solution, avoiding possible local minima in the traditional gradient descent method. In the cultural relic text data, the quantum feedback mechanism enables the model to adjust the weights in real time during the training process, more precisely fitting the complex patterns in the cultural relic text data.
[0143] Therefore, the improvement of the training process of the Transformer model further includes:
[0144] Based on the inter-layer non-linear coupling constraint loss function, the weight matrix of the feed-forward network is dynamically updated by using a quantum feedback mechanism.
[0145] Specifically: Based on the learning rate representing the update step size of each iteration, the weight gradient, the quantum acceleration gradient, the gradient of the inter-layer non-linear coupling constraint loss function with respect to the weight, the quantum spin optimization feedback adjustment coefficient, and the weight adjustment brought by the quantum spin feedback, the weight matrix of the feed-forward network at the t-th iteration is updated to obtain the weight matrix of the feed-forward network at the (t + 1)-th iteration.
[0146] In quantum spin optimization, the implementation method of the feedback mechanism is expressed as:
[0147] (7);
[0148] In the formula, W t+1 is the weight matrix of the feed-forward network at the (t + 1)-th iteration, W t is the weight matrix of the feed-forward network at the t-th iteration, is the learning rate representing the update step size of each iteration, is the weight gradient, representing the partial derivative of the loss function with respect to the weight, is the quantum acceleration gradient, is the gradient of the inter-layer non-linear coupling constraint loss function with respect to the weight, is the quantum spin optimization feedback adjustment coefficient, is the weight adjustment brought by the quantum spin feedback,
[0149] Preferably, takes a value of 0.2.
[0150] In one embodiment, the determination process of the learning rate includes:
[0151] Based on the weight gradient, the feedback gain factor, and the quantum spin feedback amount, the learning rate is obtained.
[0152] It should be noted that the quantum feedback mechanism enhances the optimization accuracy by using the learning rate. The learning rate depends on the difference between the current gradient and the feedback information. The learning rate has the following expression:
[0153] (8);
[0154] In the formula, is the feedback gain factor, which controls the intensity of the feedback effect. For example, the training error of the current iteration can be selected as the value of the feedback gain factor to adapt to the gradient fluctuations caused by noisy annotations (such as fuzzy descriptions) in cultural relic text data. L is the cross-entropy loss function. is the weight gradient, representing the partial derivative of the loss function with respect to the weight. is the quantum spin feedback amount, representing the feedback adjustment given by quantum computing. The non-steadiness of text features (such as the difference in gradient directions of different cultural relic categories) is captured by using the change amount of the loss. For example, the change amount of the loss function between the current iteration and the previous iteration can be selected as the value of the quantum spin feedback amount.
[0155] In one embodiment, refer to Figure 5 , a comparison schematic diagram of a quantum feedback path and a traditional optimization path disclosed in an embodiment of the present invention. Among them, to analyze the dynamic regulation ability of the quantum feedback mechanism on the path optimization process, on a complex non-convex loss surface, the traditional optimization path shows irregular oscillations and falls into a local minimum region, while the quantum feedback path precisely approaches the global optimal point in a spiral descent form, indicating that the feedback mechanism dynamically adjusts the parameter update step size and direction by fusing quantum gradients and classical gradient information in real time, effectively overcoming the "overshoot" or "stagnation" phenomena caused by the fixed learning rate of the traditional gradient descent method, and proving the adaptive advantage of quantum feedback in the feature learning of complex cultural relic text data.
[0156] (4) Determine the quantum accelerated gradient.
[0157] To solve the problem of large computational complexity caused by the traditional gradient descent method, the present invention adopts a quantum accelerated gradient update mechanism, which significantly accelerates the gradient calculation speed by using the parallel computing ability of quantum computing, and can significantly improve the training efficiency when processing cultural relic text data.
[0158] The determination process of the quantum accelerated gradient includes: obtaining the quantum accelerated gradient based on the quantum acceleration factor, the quantum optimization operator representing the quantum spin rotation operation, the cultural relic word vector as the training sample, and the gradient of the cross-entropy loss function with respect to the nth training sample.
[0159] The expression of the quantum accelerated gradient is as follows:
[0160] (9);
[0161] In the formula, is the quantum accelerated gradient, is the quantum acceleration factor, which down-weights high-gradient samples (such as cultural relic named entities) to suppress the interference of abnormal samples (such as incomplete texts) on training, X n is the nth input sample, specifically the cultural relic word vector as the training sample, is Xn The transpose of is a quantum optimization operator, representing a quantum spin rotation operation. is used to accelerate the interaction between word vectors and weights in quantum parallel computing, and is suitable for large-scale feature dimensions in cultural relic text data, such as BERT (Bidirectional Encoder Representations from Transformers) word vectors. is the gradient of the cross-entropy loss function with respect to the nth training sample.
[0162] Quantum acceleration factor The expression of
[0163] is as follows: (10);
[0164] In the formula, is a preset quantum acceleration factor influence parameter. For example, takes the value of 2.
[0165] (5) Iteration termination determination.
[0166] Repeat iterations (2) to (4) until the preset iteration stop condition is met, which indicates that the training of the improved Transformer model is completed.
[0167] In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0168] In summary, through the multi-dimensional acquisition, preprocessing, encoding of cultural relic text data and the reasonable division of the training data set, the present invention ensures that the improved Transformer model trained can effectively learn and understand the potential laws and context relationships in cultural relic knowledge, and improves the automatic classification, semantic analysis and feature extraction capabilities of cultural relic knowledge. By using the quantum spin optimization algorithm, the inter-layer non-linear coupling constraint loss function and the quantum accelerated gradient, the stability and efficiency of model training are improved, the problems of gradient disappearance, gradient explosion and local minimum in the traditional training process are avoided, and at the same time, the processing of large-scale data sets is accelerated. The training method of quantum spin optimization, the inter-layer non-linear coupling constraint loss and the feedback mechanism enable the improved Transformer model to better capture the multi-level features in the cultural relic text data when processing complex and high-dimensional cultural relic knowledge data, and improve the generalization ability and robustness of the model. In addition, the knowledge update and management mechanism ensures the continuous update and optimization of the cultural relic knowledge base, and the evaluation and feedback mechanism supports the real-time monitoring and retraining of the model, enhancing the adaptability of the improved Transformer model and ensuring the long-term efficient operation of the improved Transformer model.
[0169] Corresponding to the above method embodiments, the present invention also discloses a cultural relic knowledge processing device.
[0170] See Figure 6 , a schematic structural diagram of a cultural relic knowledge processing device disclosed in an embodiment of the present invention. The device may include:
[0171] An acquisition unit 201, configured to acquire cultural relic text data corresponding to the cultural relic knowledge to be processed.
[0172] A vector conversion unit 202, configured to perform text embedding on the cultural relic text data by using pre-trained word vectors, and convert the cultural relic text data into cultural relic word vectors.
[0173] An encoding unit 203, configured to input the cultural relic word vectors into the encoder of the pre-trained improved Transformer model, and map them into a set of context-aware target word vectors.
[0174] Wherein, when the improved Transformer model is trained, the weights of the feed-forward network are initialized by using the quantum spin optimization algorithm, the inter-layer non-linear coupling constraint loss function is adopted, and the quantum accelerated gradient update mechanism is adopted to accelerate the processing speed.
[0175] A decoding unit 204, configured to input the target word vectors into the decoder of the improved Transformer model for decoding to obtain the cultural relic knowledge processing result.
[0176] In summary, the present invention discloses a cultural relic knowledge processing device, which obtains the cultural relic text data corresponding to the cultural relic knowledge to be processed, performs text embedding on the cultural relic text data using pre-trained word vectors, converts the cultural relic text data into cultural relic word vectors, inputs the cultural relic word vectors into the encoder of the pre-trained improved Transformer model, maps them into a set of context-aware target word vectors, and inputs the target word vectors into the decoder of the improved Transformer model for decoding to obtain the cultural relic knowledge processing result. The present invention realizes the automatic processing of cultural relic knowledge through the improved Transformer model. The improved Transformer model initializes the weights of the feed-forward network using the quantum spin optimization algorithm to ensure the stability of weight initialization, uses the inter-layer non-linear coupling constraint loss function to better process the multi-dimensional features in the cultural relic text data, and uses the quantum acceleration gradient update mechanism to accelerate the processing speed. Therefore, it can efficiently process large-scale and multi-dimensional cultural relic knowledge.
[0177] The cultural relic knowledge processing device may further include a model training unit:
[0178] The model training unit is used to train the improved Transformer model.
[0179] In one embodiment, the process of the model training unit initializing the weights of the feed-forward network using the quantum spin optimization algorithm includes:
[0180] Calculate the product of the quantum optimization operator and the original weight matrix of the feed-forward network to obtain the weight matrix of the feed-forward network, where the weight matrix is the matrix after the original weight matrix is initialized, and the quantum optimization operator represents the quantum spin rotation operation and is determined based on the quantum spin rotation angle.
[0181] Among them, the determination process of the quantum spin rotation angle includes:
[0182] Calculate the product of the cultural relic word vector in the cultural relic text data and the original weight matrix to obtain the inner product representing the initial correlation strength between the cultural relic word vector and the original weight matrix;
[0183] Based on the inner product, the initial offset compensation of the sparse features in the cultural relic text data, and the adjustment factor, obtain the quantum spin rotation angle.
[0184] In one embodiment, the model training unit may specifically be used for:
[0185] Based on the weight matrix of the feed-forward network, determine the inter-layer non-linear coupling constraint loss function.
[0186] In one embodiment, the process by which the model training unit determines the inter-layer non-linear coupling constraint loss function based on the weight matrix of the feed-forward network includes:
[0187] Determine the first non-linear function corresponding to the weight matrix of the i-th layer of the feed-forward network, the second non-linear function corresponding to the weight matrix of the (i + 1)-th layer of the feed-forward network, and the coupling degree coefficient of the i-th layer;
[0188] Determine the product of the first non-linear function, the second non-linear function, and the coupling degree coefficient of the i-th layer as the inter-layer non-linear coupling constraint loss function.
[0189] Among them, the process of determining the coupling degree coefficient of the i-th layer includes:
[0190] Obtain the coupling degree coefficient of the i-th layer according to the correlation between the activation values of the i-th layer and the (i + 1)-th layer of the feed-forward network, the cultural relic word vectors as training samples, and the number of training samples.
[0191] In one embodiment, the model training unit can specifically be used to:
[0192] Dynamically update the weight matrix of the feed-forward network based on the inter-layer non-linear coupling constraint loss function by using a quantum feedback mechanism.
[0193] In one embodiment, the process by which the model training unit dynamically updates the weight matrix of the feed-forward network based on the inter-layer non-linear coupling constraint loss function by using a quantum feedback mechanism includes:
[0194] Update the weight matrix of the feed-forward network at the t-th iteration based on the learning rate representing the step size of each iterative update, the weight gradient, the quantum acceleration gradient, the gradient of the inter-layer non-linear coupling constraint loss function with respect to the weight, the quantum spin optimization feedback adjustment coefficient, and the weight adjustment brought by the quantum spin feedback, to obtain the weight matrix of the feed-forward network at the (t + 1)-th iteration.
[0195] Among them, the process of determining the learning rate includes:
[0196] Obtain the learning rate based on the weight gradient, the feedback gain factor, and the quantum spin feedback amount.
[0197] The process of determining the quantum acceleration gradient includes:
[0198] Obtain the quantum acceleration gradient based on the quantum acceleration factor, the quantum optimization operator representing the quantum spin rotation operation, the cultural relic word vectors as training samples, and the gradient of the cross-entropy loss function with respect to the n-th training sample.
[0199] It should be noted that for the specific working principles of the components in the device embodiments, please refer to the corresponding parts of the method embodiments, which will not be elaborated here.
[0200] Corresponding to the above embodiments, the present invention also discloses a computer storage medium, which stores at least one instruction, and when the at least one instruction is executed by a processor, the steps shown in the method embodiment of cultural relic knowledge processing are implemented.
[0201] The computer storage medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium can be a machine-readable signal medium or a machine-readable storage medium. The computer storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0202] Corresponding to the above embodiments, as Figure 7 shown, the present invention also provides a schematic structural diagram of an electronic device, which may include: a processor 1 and a memory 2;
[0203] Among them, the processor 1 and the memory 2 communicate with each other through a communication bus 3;
[0204] The processor 1 is used to execute at least one instruction;
[0205] The memory 2 is used to store at least one instruction;
[0206] The processor 1 may be a central processing unit (CPU), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0207] The memory 2 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0208] Among them, the processor executes at least one instruction to implement the steps shown in the embodiments of the cultural relic knowledge processing method.
[0209] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0210] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other.
[0211] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing cultural relic knowledge, characterized in that, Including: Obtain the cultural relic text data corresponding to the cultural relic knowledge to be processed; Perform text embedding on the cultural relic text data using pre-trained word vectors, and convert the cultural relic text data into cultural relic word vectors; Input the cultural relic word vectors into the encoder of a pre-trained improved Transformer model, and map them into a set of context-aware target word vectors. Among them, when training the improved Transformer model, the weights of the feed-forward network are initialized using a quantum spin optimization algorithm, a layer-interlayer non-linear coupling constraint loss function is adopted, and a quantum-accelerated gradient update mechanism is used to speed up the processing speed; Input the target word vectors into the decoder of the improved Transformer model for decoding to obtain the cultural relic knowledge processing result.
2. The method for processing cultural relic knowledge according to claim 1, wherein During the training process of the improved Transformer model, the process of initializing the weights of the feed-forward network using the quantum spin optimization algorithm includes: Calculate the product of the quantum optimization operator and the original weight matrix of the feed-forward network to obtain the weight matrix of the feed-forward network. Among them, the weight matrix is the matrix after the original weight matrix is initialized, and the quantum optimization operator represents a quantum spin rotation operation, which is determined based on the quantum spin rotation angle.
3. The method for processing cultural relic knowledge according to claim 2, characterized in that, The determination process of the quantum spin rotation angle includes: Calculate the product of the cultural relic word vectors in the cultural relic text data and the original weight matrix to obtain the inner product representing the initial correlation strength between the cultural relic word vectors and the original weight matrix; Based on the inner product, the initial offset compensation of the sparse features in the cultural relic text data, and the adjustment factor, obtain the quantum spin rotation angle.
4. The method for processing cultural relic knowledge according to claim 2, wherein, The training process of the improved Transformer model further includes: Determine the layer-interlayer non-linear coupling constraint loss function based on the weight matrix of the feed-forward network.
5. The method for processing cultural relics knowledge according to claim 4, wherein, Determining the layer-interlayer non-linear coupling constraint loss function based on the weight matrix of the feed-forward network includes: Determine the first non-linear function corresponding to the weight matrix of the i-th layer of the feed-forward network, the second non-linear function corresponding to the weight matrix of the (i + 1)-th layer of the feed-forward network, and the coupling degree coefficient of the i-th layer; Determine the product of the first non-linear function, the second non-linear function, and the coupling degree coefficient of the i-th layer as the layer-interlayer non-linear coupling constraint loss function.
6. The method for processing cultural relic knowledge according to claim 5, wherein The determination process of the coupling degree coefficient of the i-th layer includes: Obtain the coupling degree coefficient of the i-th layer according to the correlation between the activation values of the i-th layer and the (i + 1)-th layer of the feed-forward network, the cultural relic word vectors used as training samples, and the number of training samples.
7. The method for processing cultural relics knowledge according to any one of claims 4 to 6, characterized in that, The training process of the improved Transformer model further includes: Dynamically update the weight matrix of the feed-forward network using a quantum feedback mechanism based on the layer-interlayer non-linear coupling constraint loss function.
8. The method for processing cultural relics knowledge according to claim 7, wherein Dynamically updating the weight matrix of the feed-forward network using a quantum feedback mechanism based on the layer-interlayer non-linear coupling constraint loss function includes: Update the weight matrix of the feedforward network at the t-th iteration based on the learning rate, weight gradient, quantum acceleration gradient, gradient of the inter-layer non-linear coupling constraint loss function with respect to the weight, quantum spin optimization feedback adjustment coefficient, and weight adjustment brought by quantum spin feedback, to obtain the weight matrix of the feedforward network at the (t + 1)-th iteration.
9. The method for processing cultural relic knowledge according to claim 8, wherein The determination process of the learning rate includes: Obtain the learning rate based on the weight gradient, feedback gain factor, and quantum spin feedback amount.
10. The method for processing cultural relic knowledge according to claim 8 or 9, characterized in that, The determination process of the quantum acceleration gradient includes: Obtain the quantum acceleration gradient based on the quantum acceleration factor, quantum optimization operator representing the quantum spin rotation operation, cultural relic word vector as a training sample, and gradient of the cross-entropy loss function with respect to the n-th training sample.
11. A cultural relic knowledge processing device, characterized in that, Includes: An acquisition unit for acquiring cultural relic text data corresponding to the cultural relic knowledge to be processed; A vector conversion unit for performing text embedding on the cultural relic text data using pre-trained word vectors to convert the cultural relic text data into cultural relic word vectors; An encoding unit for inputting the cultural relic word vectors into the encoder of a pre-trained improved Transformer model to map them into a set of context-aware target word vectors, where the improved Transformer model uses the quantum spin optimization algorithm to initialize the weights of the feedforward network during training, uses the inter-layer non-linear coupling constraint loss function, and uses the quantum acceleration gradient update mechanism to accelerate the processing speed; A decoding unit for inputting the target word vectors into the decoder of the improved Transformer model for decoding to obtain the cultural relic knowledge processing result.
12. A computer storage medium, characterized in that, The computer storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, it implements the cultural relic knowledge processing method according to any one of claims 1 to 10.
13. An electronic device, characterized in that, The electronic device includes: a memory and a processor; The memory is used to store at least one instruction; The processor is used to execute the at least one instruction to implement the cultural relic knowledge processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Task processing method and device based on model quantification, equipment and storage medium
CN115860068A
Low-resource legal information extraction method and system based on large-scale language model
CN117763103A
Intelligent auxiliary inquiry medical large model training method and system
CN119831054A
Parameter fine tuning method and device of pre-training model, equipment and medium
CN119849576A
Sea surface temperature prediction method and system
CN120146253A
Cited By
Power battery diagnosis method and device, electronic equipment and storage medium
CN122471262A