Text information processing method based on artificial intelligence

By integrating multimodal information and optimizing the iterative update of semantic vectors, the problem of insufficient recognition ability of existing pre-trained models in text processing in the railway industry has been solved, and the accuracy and robustness of keyword extraction have been improved, especially in identifying key elements such as key equipment models, fault locations, and maintenance steps in railway documents.

CN120633663APending Publication Date: 2025-09-12CHINA ACADEMY OF RAILWAY SCI CORP LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510562084.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing pre-trained models cannot fully capture industry-specific terminology and expressions in text processing in the railway industry, resulting in inaccurate keyword extraction or omission of key terms, especially when processing specialized texts involving vehicle computers, industrial electronics, etc.

Method used

By introducing multimodal information fusion and optimizing the iterative update mechanism of semantic vectors, using pre-trained language models for word segmentation and context perception, combining sparse values ​​and key entity mapping, the keyword extraction process is optimized, including multimodal feature fusion of non-text information and the application of knowledge graphs.

Benefits of technology

It significantly improves the accuracy and robustness of railway document processing, enabling more accurate identification of key information and avoiding misjudgments or omissions, especially in professional fields such as complex equipment and fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633663A_ABST
    Figure CN120633663A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text processing, and discloses a text information processing method based on artificial intelligence, which remarkably improves the accuracy and robustness of railway official document text processing by introducing multi-modal information fusion and an iterative updating mechanism for optimizing semantic vectors. Firstly, an initial semantic vector obtained by an existing pre-training model serves as a first initial semantic vector, and it is ensured that basic semantic representation of vocabularies has high accuracy. On this basis, non-text information in the railway official document text is further considered, a multi-modal feature vector of the non-text information is extracted, and deep fusion is performed on the multi-modal feature and the first initial semantic vector through a multi-layer perceptron to generate a second initial semantic vector. The process not only enriches semantic representation of vocabularies, but also provides additional context information for the model, so that the model can more comprehensively understand text content, and the accuracy of keyword extraction is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text processing technology, and in particular to a text information processing method based on artificial intelligence. Background Art

[0002] In modern railway operations, management, and maintenance, a significant amount of information exchange relies on formal written documents, known as railway official documents. These documents cover a wide range of content, from daily dispatch orders and safety advisories to equipment maintenance records. They incorporate specialized terminology and technical details in areas such as locomotives (trains and their operation) and electrical engineering (engineering and power systems). With the increasing complexity and volume of data in railway systems, the efficient management and utilization of these documents has become crucial. Therefore, automated text information processing methods, particularly keyword extraction technologies targeted at specific fields, are of significant value in improving work efficiency and assisting in decision support. Keywords are words or phrases that can highly summarize the main idea of ​​an article and play a vital role in understanding the core content of a document. In the railway industry, accurately identifying keywords related to "locomotives" and "electrical engineering" can help quickly locate key information, optimize retrieval efficiency, and enhance the intelligence of document management. Furthermore, it provides a solid foundation for subsequent data analysis, trend forecasting, and risk assessment. Existing technologies use artificial intelligence to extract keywords from railway official documents. Although pre-trained models such as BERT have demonstrated powerful language understanding capabilities, they may not be able to fully capture industry-specific terminology and expressions when applied in specific fields (such as the railway industry), resulting in inaccurate keyword extraction or missing key terms, especially when dealing with highly specialized texts involving vehicle computers, industrial electronics, etc. The railway industry has its own unique vocabulary system, including a large number of professional terms and technical terms. For example, terms such as "vehicle-computer joint control" and "contact network maintenance" appear very rarely in general corpora, so the pre-trained model has limited understanding and recognition capabilities for these terms by default. In summary, the problem of limited understanding and recognition capabilities of pre-trained models for railway terminology by default in existing technologies needs to be solved urgently. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides a text information processing method based on artificial intelligence.

[0004] The present invention proposes an artificial intelligence-based text information processing method, comprising:

[0005] Step 100: receiving a railway official document to be processed and determining the text category of the railway official document;

[0006] Step 200: Segment each word in the railway official document using a pre-trained language model to generate a first initial semantic vector and a context vector for each word.

[0007] Step 300: Optimize the first initial semantic vector to obtain a second initial semantic vector;

[0008] Step 400, optimizing the second initial semantic vector to obtain an optimized semantic vector, includes:

[0009] Step 410 , calculating the importance weight of each word relative to other words in the railway official document text based on a context-aware mechanism;

[0010] Step 420, obtaining a sparse value for each word according to the text category;

[0011] Step 430: construct an optimization loss function based on the importance weight, the sparsity value, and the second initial semantic vector;

[0012] Step 440: Calculate the gradient of the second initial semantic vector according to the optimization loss function;

[0013] Step 450: Preset an adjustment step size and a learning rate; iteratively update the second initial semantic vector based on the second initial semantic vector, the adjustment step size, the learning rate, the context vector, and the gradient of the second initial semantic vector until the optimization loss function converges or the maximum number of iterations is reached, and then stop updating. The latest second initial semantic vector is used as the optimized semantic vector.

[0014] Step 500, constructing a knowledge graph; selecting a structured information template according to the text category, extracting key elements from the railway official document text as key entities, and assigning label values ​​to the key elements; mapping the key entities to the mapping nodes of the knowledge graph, and obtaining the vocabulary association according to the mapping nodes; calculating the score items according to the optimized semantic vector, calculating the comprehensive score of the vocabulary according to the importance weight, sparsity value, score item, vocabulary association, and key element label value, taking the vocabulary with the highest comprehensive score as the keyword, and outputting the keyword list.

[0015] Furthermore, the sparseness value of a word is calculated by inverse document frequency of word frequency.

[0016] Furthermore, the first initial semantic vector is optimized to obtain a second initial semantic vector, including obtaining non-text information of the railway official document text, extracting a multimodal feature vector from the non-text information through a pre-trained visual model; and using a multilayer perceptron to fuse the multimodal feature vector with the first initial semantic vector to obtain a second initial semantic vector.

[0017] Furthermore, the optimized loss function is:

[0018]

[0019] Where L is the optimization loss function, CE is the cross entropy loss term, R is the regularization term, D is the penalty term, and s i is the importance weight between the i-th word and other words, p(w i , C) is the sparse value of the i-th word in the text category C, is the second initial semantic vector of the t-th iteration of the i-th word, is the optimized semantic vector of the i-th word, n is the total number of words, and α, β, γ, and δ are adjustment coefficients respectively.

[0020] Furthermore, the gradient of the second initial semantic vector is:

[0021]

[0022] Where, To optimize the gradient of the loss function for the second initial semantic vector at the tth iteration, is the partial derivative, is the gradient.

[0023] Furthermore, the optimized semantic vector is:

[0024]

[0025] Where, is the second initial semantic vector of the i-th word at the t-th iteration, is the adjustment step size, τ is the learning rate, x i is the context vector of the i-th word.

[0026] Furthermore, the score items are calculated based on the optimized semantic vector, including:

[0027] The mean of the optimized semantic vectors of all words in the railway official document is calculated, and the similarity between the optimized semantic vector and the mean of the optimized semantic vector is calculated, and the similarity is used as the score item.

[0028] Furthermore, the key entities are mapped to the mapping nodes of the knowledge graph, and the vocabulary relevance is obtained according to the mapping nodes, including:

[0029] For each key entity, search for candidate entities from the knowledge graph, calculate the key similarity between the key entity and the candidate entity, use the candidate entity with the largest key similarity as the mapping node of the key entity, and map the key entity to the mapping node; obtain the neighboring entities of the key entity, determine the connection weight based on the relationship type between the key entity and the neighboring entities, and use the connection weight as the vocabulary association.

[0030] Furthermore, the comprehensive score is:

[0031]

[0032] Where C i is the comprehensive score of the i-th word, is the score item calculated based on the optimized semantic vector, KG(w i ) is the vocabulary association of the i-th word, KE is the label value of whether the i-th word is a key element, if the i-th word is a key element, KE = 1; if the i-th word is not a key element, KE = 0.

[0033] The embodiments of the present invention have the following technical effects:

[0034] The present invention significantly improves the accuracy and robustness of railway official document processing by introducing an iterative update mechanism that integrates multimodal information and optimizes semantic vectors. First, the present invention uses the initial semantic vector obtained by the existing pre-training model as the first initial semantic vector, ensuring that the basic semantic representation of the vocabulary already has a high degree of accuracy. On this basis, the non-text information in the railway official document text is further considered, the multimodal feature vector of the non-text information is extracted, and the multimodal features are deeply fused with the first initial semantic vector through a multi-layer perceptron to generate a second initial semantic vector. This process not only enriches the semantic representation of the vocabulary, but also provides additional contextual information for the model, enabling the model to understand the text content more comprehensively, especially in railway official document texts involving professional fields such as complex equipment, fault diagnosis, and maintenance procedures, which greatly improves the accuracy of keyword extraction.

[0035] In addition, based on the generation of the second initial semantic vector, the present invention further introduces an optimized loss function and a gradient descent mechanism. By iteratively updating the second initial semantic vector, an optimized semantic vector is ultimately obtained. This optimization process not only takes into account factors such as the importance weight, sparsity value, and context vector of the vocabulary, but also incorporates the influence of multimodal features, ensuring that the optimized semantic vector can better reflect the actual meaning of the vocabulary in a specific context. In this way, the model can more accurately identify key information in complex railway official documents, avoiding misjudgments or omissions that may result from relying solely on text information.

[0036] Finally, the present invention calculates the score items based on the optimized semantic vector, and comprehensively evaluates the score of each word in combination with the importance weight, sparsity value, word association and key element label value, and finally outputs the word with the highest comprehensive score as the keyword list. This comprehensive evaluation mechanism not only takes into account the semantic similarity of the words, but also fully combines the importance of the words in the text and the relationship with other words, ensuring that the selection of keywords is more reasonable and accurate. In particular, for key elements such as key equipment models, fault locations, maintenance steps, etc. contained in railway official documents, the present invention can ensure that these key elements are prioritized and included in the keyword list through knowledge graph mapping and word association calculation, thereby providing strong support for subsequent document analysis, fault diagnosis, maintenance decision-making and other tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 This is a flowchart of an artificial intelligence-based text information processing method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0039] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0040] Figure 1 This is a flowchart of an artificial intelligence-based text information processing method provided by an embodiment of the present invention.

[0041] See also Figure 1 , a text information processing method based on artificial intelligence, comprising the following steps:

[0042] Step 100: receiving a railway official document to be processed and determining the text category of the railway official document.

[0043] Received railway official documents are cleaned to remove irrelevant symbols, punctuation, and spaces to ensure text integrity. The text is segmented into paragraphs and sentences to facilitate subsequent processing and analysis. If the text contains multiple character encodings, unified encoding conversion is required to ensure compatibility across different systems.

[0044] Based on railway documents' titles, keywords, and content, we design a series of rules to automatically categorize text. For example, if a title contains keywords like "fault report," "maintenance record," "equipment inspection," or "vehicle power consumption report," it can be assigned to the corresponding category.

[0045] Step 200 : Segment each word in the railway official document using a pre-trained language model to generate a first initial semantic vector for each word and a context vector for each word.

[0046] For each sentence, construct an input sequence consisting of tokenIDs and assign corresponding attention masks to them to indicate which words are valid and which are fillers. For example, the text of railway official documents is: CRH380A EMU had a brake system failure during operation. After inspection, it was found that it was caused by damage to the brake valve. The word segmentation result is: ["CRH380A","type","EMU","in","operation","process","in","appeared","brake","system","fault",","after","inspection","found","is","due to","brake valve","damage","cause","of","."]. Convert the words after word segmentation into corresponding tokenIDs to form an input sequence. Pass the input sequence to the pre-trained language model to generate the hidden layer representation of each word. Extract the hidden layer representation of each word from the output of the model as its first initial semantic vector. For example, the semantic vector of "CRH380A" is represented as [y1, y2, y3,...,y d ], d is the dimension of the vector. The context vector refers to the representation of a word in its context, reflecting the meaning of the word in a specific context. Unlike the first initial semantic vector, the context vector not only contains the semantic information of the word itself, but also takes into account factors such as its position in the sentence and its relationship with other words. The self-attention mechanism is used to calculate the context vector of a word. The self-attention mechanism calculates the similarity between a word and other words in its context to generate a weight matrix that represents the importance of each word in the current context. Then, the hidden layer representations of all words are combined by weighted summation to obtain the context vector of each word.

[0047] Step 300: Optimize the first initial semantic vector to obtain a second initial semantic vector.

[0048] Railway official documents include both text and non-text information. Types of non-text information:

[0049] Charts: equipment working principle diagram, fault diagnosis flow chart, maintenance steps diagram, etc.

[0050] Images: photos of the equipment, photos of the failure site, detailed pictures of parts, etc.

[0051] Forms: Equipment technical parameter table, maintenance record table, fault statistics table, etc.

[0052] Formulas: Mathematical or physical formulas that may be included in technical documents.

[0053] Symbol: Professional symbol or logo in the railway field.

[0054] The non-text information of railway official documents is obtained, and the multimodal feature vector in the non-text information is extracted through a pre-trained visual model. The multimodal feature vector is fused with the first initial semantic vector using a multi-layer perceptron to obtain a second initial semantic vector.

[0055] Step 400 , optimizing the second initial semantic vector to obtain an optimized semantic vector. Specific step 400 includes steps 410 to 450 .

[0056] Step 410 : Calculate the importance weight of each word relative to other words in the railway official document text based on the context-aware mechanism.

[0057] This example uses a self-attention mechanism to calculate the similarity between each word and other words in its context, generating a weight matrix. The self-attention mechanism calculates the dot product or scaled dot product between pairs of words to obtain attention weights for each word. These weights represent the importance of a word in the current context.

[0058]

[0059] Where s i is the importance weight of the i-th word relative to other words, q is the query vector, k is the key vector, and d k is the dimension of the key vector, x i is the context vector of the i-th word, and T is the transposition operation.

[0060] Step 420, obtaining a sparse value for each word according to the text category;

[0061] The sparse value of a word is calculated by using the term frequency inverse document frequency (TF-IDF). Specifically,

[0062]

[0063] Where, p(w i , C) is the sparse value of the i-th word in the text category C, TF-IDF(w i , C) is the word frequency inverse document frequency of the i-th word in text category C, ∑ i TF-IDF(w i , C) is the word frequency inverse document frequency of all words in text category C.

[0064] Step 430: construct an optimization loss function based on the importance weight, the sparsity value, and the second initial semantic vector.

[0065]

[0066] Where L is the optimization loss function, CE is the cross entropy loss term, R is the regularization term, D is the penalty term, and s i is the importance weight between the i-th word and other words, p(w i , C) is the sparse value of the i-th word in the text category C, is the second initial semantic vector of the i-th word, is the optimized semantic vector of the i-th word, α, β, γ, and δ are adjustment coefficients, and n is the total number of words.

[0067] CE is used to measure the difference between the distribution predicted by the model and the true label. Specifically, the cross-entropy loss reflects the difference between the semantic vector generated by the model and the true semantics. The smaller the cross-entropy loss, the closer the model's prediction is to the true label, which is a conventional design for those skilled in the art. R encourages model parameters to maintain small values ​​or sparseness by penalizing the sum of squares or absolute values ​​of model parameters, which is a conventional design for those skilled in the art. D is used to introduce additional constraints to ensure that the optimized semantic vector meets certain specific requirements. For example, sparsity constraints can be introduced to encourage certain dimensions in the semantic vector to remain zero; or smoothness constraints can be introduced to ensure that the semantic vector maintains a smooth transition between adjacent words:

[0068]

[0069] Where μ is the penalty coefficient, y i is the first initial semantic vector of the i-th word, y i-1 is the first initial semantic vector of the i-1th word, and n is the total number of words.

[0070] This is a semantic vector optimization term used to measure the difference between the second initial semantic vector and the optimized semantic vector. This term combines the importance weight and sparsity value of the vocabulary to ensure that the optimized semantic vector can better reflect the actual meaning of the vocabulary in the context. By minimizing this optimization term, we can ensure that the optimized semantic vector not only captures the semantic information of the vocabulary itself, but also fully considers its dependencies and sparsity in the context.

[0071] By optimizing the loss function described above, the present invention can effectively optimize the semantic vector of each word in railway official documents. Specifically, the cross-entropy loss term ensures that the model's predictions are as close as possible to the true label; the regularization term prevents model overfitting; the penalty term introduces additional constraints to ensure the rationality of the semantic vector; and the semantic vector optimization term combines the importance weight and sparsity value of the word to ensure that the optimized semantic vector can better reflect the actual meaning of the word in context. This process not only captures the semantic information of the word itself, but also fully considers its dependencies and sparsity in the context, significantly improving the accuracy and robustness of railway official document processing.

[0072] Step 440: Calculate the gradient of the second initial semantic vector according to the optimization loss function.

[0073] The gradient is calculated by solving the partial derivative of the optimization loss function with respect to the second initial semantic vector. The direction of the gradient indicates how to adjust the semantic vector to minimize the optimization loss function.

[0074]

[0075] Where, To optimize the gradient of the loss function for the second initial semantic vector at the tth iteration, is the partial derivative, is the gradient.

[0076] Step 450, preset the adjustment step size and learning rate; iteratively update the second initial semantic vector according to the second initial semantic vector, the adjustment step size, the learning rate, the context vector, and the gradient of the second initial semantic vector until the optimization loss function converges or the maximum number of iterations is reached and the update is stopped, and the latest second initial semantic vector is used as the optimized semantic vector.

[0077] First iteration: Based on the current second initial semantic vector context vector x i , optimize the loss function L and calculate the gradient And do the first update:

[0078]

[0079] Multiple iterations: In each iteration, repeat the following steps:

[0080] According to the current second initial semantic vector context vector x i And optimize the loss function L and calculate the gradient Use the gradient descent method to update the second initial semantic vector:

[0081]

[0082] Where, is the second initial semantic vector of the i-th word at the t-th iteration, is the second initial semantic vector of the i-th word at the t+1-th iteration, is the adjustment step size, τ is the learning rate, x i is the context vector of the i-th word.

[0083] Update the number of iterations t and check again whether the optimization loss function converges or reaches the maximum number of iterations. If the convergence condition is met, stop the iteration and go to the next step, using the latest second initial semantic vector as the optimized semantic vector Output.

[0084]

[0085] Through the above steps, the present invention can effectively optimize the semantic vector of each word in the railway official document text. Specifically, the importance weight of each word is first calculated through the context-aware mechanism, and then the sparse value of each word is obtained according to the text category. Then, based on the importance weight, sparse value and the second initial semantic vector, an optimization loss function is constructed, and the second initial semantic vector is iteratively updated by the gradient descent method. Finally, when the optimization loss function converges or reaches the maximum number of iterations, the optimized semantic vector is output. This process not only captures the semantic information of the word itself, but also fully considers its dependency and sparsity in the context, significantly improving the accuracy and robustness of railway official document text processing.

[0086] Step 500, according to the text category (such as fault report, maintenance record, vehicle power report, etc.), select a suitable structured information template to ensure that the key elements in the railway official document text can be accurately extracted as key entities, map the key entities to the mapping nodes of the knowledge graph, and obtain the vocabulary association according to the mapping nodes; calculate the score items based on the optimized semantic vector, calculate the comprehensive score of the vocabulary based on the importance weight, sparsity value, score item, vocabulary association, and key element label value, use the vocabulary with the highest comprehensive score as the keyword, and output the keyword list.

[0087] Design a dedicated structured information template for each text category. The template contains specific fields to extract key information from the text. For example, for a vehicle power consumption report, the template can include the following fields:

[0088] Equipment model: record the specific model of vehicles, locomotives, and industrial and electrical equipment.

[0089] Operating parameters: including equipment operating time, speed, temperature, pressure and other parameters.

[0090] Energy consumption data: record the equipment's power consumption, fuel consumption and other energy consumption information.

[0091] Maintenance status: record the maintenance status of the equipment, such as the last maintenance time, next maintenance plan, replaced parts, etc.

[0092] Fault description: record the faults that occur during the operation of the equipment and their treatment measures.

[0093] For example, the key entities extracted from the above vehicle mechanic power report are:

[0094] Equipment model: CRH380A.

[0095] Operating parameters: operating time (10 hours), average speed (250 km / h).

[0096] Energy consumption data: electricity consumption (500kWh), fuel consumption (0 liters).

[0097] Maintenance status: last maintenance time (November 25, 2024), next maintenance plan (December 20, 2024).

[0098] The key entities are mapped to the mapping nodes of the knowledge graph, and the vocabulary association is obtained based on the mapping nodes, including: for each key entity, searching for candidate entities from the knowledge graph, calculating the key similarity between the key entity and the candidate entity, taking the candidate entity with the largest key similarity as the mapping node of the key entity, and mapping the key entity to the mapping node; obtaining the neighboring entities of the key entity, determining the connection weight according to the relationship type between the key entity and the neighboring entity, and taking the connection weight as the vocabulary association.

[0099] For each key entity, candidate entities are searched from the knowledge graph. Candidate entities are nodes with similar names or attributes to the key entity. For example, for the equipment model "CRH380A," multiple candidate entities can be found in the knowledge graph (e.g., CRH380A EMUs produced by different manufacturers). The similarity between the key entity and the candidate entities is calculated. Similarity can be based on various factors, such as name matching, attribute matching, and context matching. For example, methods such as cosine similarity and Jaccard similarity can be used to calculate similarity. The candidate entity with the highest similarity is selected as the mapping node for the key entity. For example, if "CRH380A" has the highest similarity with a candidate entity, that candidate entity is selected as the mapping node for "CRH380A." Neighboring entities of the key entity (i.e., entities with a direct relationship to it) are obtained from the knowledge graph. For example, for "CRH380A," its neighboring entities may include "brake system," "traction system," "air conditioning system," and so on. Based on the relationship type between the key entity and the neighboring entities, a connection weight is determined. The connection weight reflects the strength of the relationship between the two. For example, if there is a "contains" relationship between "CRH380A" and "brake system," the connection weight is higher; if there is a "auxiliary" relationship between "air conditioning system," the connection weight is lower. A connection weight of 0-1 is assigned based on the key type. The connection weight is used as the word relevance, indicating the strength of the relationship between the word and its related entities. For example, "CRH380A" is mapped to a node in the knowledge graph, and its neighboring entities are obtained:

[0100] Adjacent entities: braking system, traction system, air conditioning system.

[0101] Connection weight: "CRH380A" and "Braking System": 0.9 (strong relationship).

[0102] CRH380A” and “traction system”: 0.8 (strong relationship).

[0103] CRH380A” and “air conditioning system”: 0.6 (weak relationship).

[0104] Therefore, the word relevance of “CRH380A” is high because it has strong connections with multiple important neighboring entities.

[0105] Scoring items are calculated based on the optimized semantic vectors, including: calculating the mean of the optimized semantic vectors of all words in the railway official document text, calculating the similarity between the optimized semantic vectors and the mean of the optimized semantic vectors, and using the similarity as the scoring item.

[0106]

[0107] Where C iis the comprehensive score of the i-th word, is the score item calculated based on the optimized semantic vector, KG(w i ) is the vocabulary association of the i-th word, KE is the label value of whether the i-th word is a key element, if the i-th word is a key element, KE = 1; if the i-th word is not a key element, KE = 0.

[0108] Sort all words by their comprehensive scores from high to low, select the top k words with the highest scores as keywords, k is preferably 10, select the top 10 words as keywords, and output the selected keywords as a list.

[0109] It is worth noting that the score item, relevance, and label value in the comprehensive score formula are scalars, but the importance weight is a matrix. The matrix needs to be converted to a scalar before calculation:

[0110]

[0111] In the formula, Z is the importance weight value, which is a scalar. IK is the element in the I-th row and K-th column of the matrix Z, where M and N are the number of rows and columns of the matrix.

[0112] The present invention can effectively extract key entities from railway official documents and map them to mapping nodes of the knowledge graph to obtain vocabulary associations. At the same time, score items are calculated based on the optimized semantic vectors, and a comprehensive score for each vocabulary word is calculated by combining importance weights, sparsity values, vocabulary associations, and key element label values. Finally, the vocabulary with the highest comprehensive score is selected as the keyword, and a keyword list is output. This process not only captures the semantic information of the vocabulary itself, but also fully considers its dependencies in the context and its associations with other entities, significantly improving the accuracy and robustness of keyword extraction.

[0113] It should be noted that the terms used in the present invention are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method or device comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.

[0114] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A text information processing method based on artificial intelligence, characterized in that: include: Step 100: receiving a railway official document to be processed and determining the text category of the railway official document; Step 200: Segment each word in the railway official document using a pre-trained language model to generate a first initial semantic vector and a context vector for each word. Step 300: Optimize the first initial semantic vector to obtain a second initial semantic vector; Step 400, optimizing the second initial semantic vector to obtain an optimized semantic vector, includes: Step 410 , calculating the importance weight of each word relative to other words in the railway official document text based on a context-aware mechanism; Step 420, obtaining a sparse value for each word according to the text category; Step 430: construct an optimization loss function based on the importance weight, the sparsity value, and the second initial semantic vector; Step 440: Calculate the gradient of the second initial semantic vector according to the optimization loss function; Step 450: Preset an adjustment step size and a learning rate; iteratively update the second initial semantic vector based on the second initial semantic vector, the adjustment step size, the learning rate, the context vector, and the gradient of the second initial semantic vector until the optimization loss function converges or the maximum number of iterations is reached, and then stop updating. The latest second initial semantic vector is used as the optimized semantic vector. Step 500, constructing a knowledge graph; selecting a structured information template according to the text category, extracting key elements from the railway official document text as key entities, and assigning label values ​​to the key elements; mapping the key entities to the mapping nodes of the knowledge graph, and obtaining the vocabulary association according to the mapping nodes; calculating the score items according to the optimized semantic vector, calculating the comprehensive score of the vocabulary according to the importance weight, sparsity value, score item, vocabulary association, and key element label value, taking the vocabulary with the highest comprehensive score as the keyword, and outputting the keyword list.

2. The text information processing method based on artificial intelligence according to claim 1, characterized in that: The sparseness value of a word is calculated by the inverse document frequency of the word frequency.

3. The text information processing method based on artificial intelligence according to claim 1, characterized in that: Optimizing the first initial semantic vector to obtain a second initial semantic vector includes: The non-text information of railway official documents is obtained, and the multimodal feature vector in the non-text information is extracted through a pre-trained visual model. The multimodal feature vector is fused with the first initial semantic vector using a multi-layer perceptron to obtain a second initial semantic vector.

4. The text information processing method based on artificial intelligence according to claim 1, characterized in that: The optimized loss function is: Where L is the optimization loss function, CE is the cross entropy loss term, R is the regularization term, D is the penalty term, and s i is the importance weight between the i-th word and other words, p(w i , C) is the sparse value of the i-th word in the text category C, is the second initial semantic vector of the t-th iteration of the i-th word, is the optimized semantic vector of the i-th word, n is the total number of words, and α, β, γ, and δ are adjustment coefficients respectively.

5. The text information processing method based on artificial intelligence according to claim 4, characterized in that: The gradient of the second initial semantic vector is: Where, To optimize the gradient of the loss function for the second initial semantic vector at the tth iteration, is the partial derivative, is the gradient.

6. The text information processing method based on artificial intelligence according to claim 5, characterized in that: The optimized semantic vector is: Where, is the second initial semantic vector of the i-th word at the t-th iteration, is the adjustment step size, τ is the learning rate, x i is the context vector of the i-th word.

7. The text information processing method based on artificial intelligence according to claim 1, characterized in that: Scoring items are calculated based on the optimized semantic vector, including: The mean of the optimized semantic vectors of all words in the railway official document is calculated, and the similarity between the optimized semantic vector and the mean of the optimized semantic vector is calculated, and the similarity is used as the score item.

8. The text information processing method based on artificial intelligence according to claim 1, characterized in that: Map key entities to mapping nodes of the knowledge graph and obtain vocabulary relevance based on the mapping nodes, including: For each key entity, search for candidate entities from the knowledge graph, calculate the key similarity between the key entity and the candidate entity, use the candidate entity with the largest key similarity as the mapping node of the key entity, and map the key entity to the mapping node; obtain the neighboring entities of the key entity, determine the connection weight based on the relationship type between the key entity and the neighboring entities, and use the connection weight as the vocabulary association.

9. The text information processing method based on artificial intelligence according to claim 4, characterized in that: The comprehensive score is: Where C i is the comprehensive score of the i-th word, is the score item calculated based on the optimized semantic vector, KG(w i ) is the vocabulary association of the i-th word, KE is the label value of whether the i-th word is a key element, if the i-th word is a key element, KE = 1; if the i-th word is not a key element, KE = 0.