A logistics policy document analysis method based on multimodal features

Through the analysis method of logistics policy document based on multimodal characteristics, the problem of difficult to determine the theme of logistics policy document in the existing technology is solved, and more accurate theme extraction and timing characteristics fusion are achieved, which improves the evaluation efficiency of policy documents.

CN119557428BActive Publication Date: 2025-05-06EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510122282.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-06
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The prior art is difficult to accurately determine the theme of logistics policy documents, resulting in inconsistent evaluation standards, lack of clear theme guidance, difficulty in determining the focus and dimensions of evaluation, and failure to fully utilize the timing characteristics and format characteristics.

Method used

A logistics policy document analysis method based on multimodal features is proposed, through the construction of the thesaurus, the use of the LDA2Vec model to extract keywords, the timing enhancement module fusion timing characteristics, and the coordination of superior and subordinate policy documents is evaluated through the similarity function.

Benefits of technology

More accurate theme keyword extraction and sorting is achieved, keyword boundaries are clarified, timing characteristics are integrated, the accuracy of topic classification and analysis is improved, and the implementation efficiency of policy documents is improved through synergistic evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557428B_ABST
    Figure CN119557428B_ABST
Patent Text Reader

Abstract

The present invention provides a logistics policy document analysis method based on multimodal features, which extracts keywords of the subject from the logistics policy document, and then obtains the sorted keywords by sorting the keywords by importance score, and clarifies the boundaries of the keywords; further, by marking timestamps, multimodal keywords integrating temporal features are obtained, which has a more accurate prediction effect, and further lays the foundation for the subsequent evaluation of logistics policy documents according to time periods; in addition, the present invention evaluates the similarity of the superior logistics policy document and the subordinate logistics policy document through a similarity function, and obtains the superior-subordinate coordination degree to evaluate the implementation efficiency of the logistics policy document. In addition, in the process of obtaining the superior policy document, the present invention obtains more accurate keywords of the subject by adjusting the preset weight of the subject at the subheading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of text processing, and in particular to a logistics policy document analysis method based on multimodal features. Background Art

[0002] In the research and application of logistics policy documents, determining the document theme is a basic and key task. However, the current problem of difficulty in determining the theme of logistics policy documents is more prominent, which brings many challenges to the evaluation of policy documents. On the one hand, the logistics industry involves multiple fields and links, such as transportation, warehousing, distribution, supply chain management, etc. Policy documents often cover multiple aspects of content, resulting in blurred boundaries of keywords in their themes.

[0003] On the other hand, the formulation of logistics policy documents involves multiple departments in various regions, among which the policy objectives and emphases of the upper and lower departments are different, which further increases the complexity of the subject determination. In addition, with the rapid development of the logistics industry and the continuous emergence of new technologies, policy documents need to be constantly updated and adjusted, which also makes the evaluation of logistics policy documents face the problem of dynamic changes.

[0004] In this case, it is difficult to evaluate logistics policy documents. Specifically, it is difficult to unify the evaluation criteria, lack clear subject guidance, and it is difficult to determine the focus and dimensions of the evaluation. At the same time, the choice of evaluation methods is also limited, making it difficult to accurately measure the implementation effect and influence of policy documents.

[0005] In addition, in the existing process of analyzing logistics policy documents, the temporal characteristics of logistics policy documents are not fully utilized, and it is impossible to accurately and flexibly classify and analyze logistics policies according to time.

[0006] In addition, in the process of obtaining the topics of logistics policy documents, the format of the policy documents was not weighted, resulting in inaccurate keywords for the obtained topics. Summary of the invention

[0007] In view of the shortcomings of the prior art, the present invention proposes a logistics policy document analysis method based on multimodal features, comprising the following steps:

[0008] Step S1: construct a logistics policy document word library, import logistics policy documents for word segmentation preprocessing, and obtain preprocessed documents; wherein, keywords of the topics of some logistics policy documents are manually annotated; the annotated tags include keywords of the manually selected topics and the release time of the corresponding logistics policy documents;

[0009] Step S2: construct an analysis model, which includes a keyword extraction module based on LDA2Vec, a sorting module, a time series enhancement module and an evaluation module, import the annotated preprocessing file in step S1 into the keyword extraction module based on LDA2Vec, and obtain several keywords of the topic in the annotated preprocessing file;

[0010] Step S3: inputting several keywords obtained in step S2 into the sorting module, sorting the keywords by calculating the importance score of each keyword through TF-IDF, and obtaining the determined number of keywords of the topic by calculating the perplexity, retaining the determined number of keywords ranked first, and indexing and marking the determined number of keywords according to the sorting;

[0011] Step S4: inputting the determined number of keywords in step S3 into the time series enhancement module, marking a timestamp on each keyword for time series indexing, and obtaining keywords of the topic integrating the time series features;

[0012] Step S5: construct a loss function, and optimize the parameters of the analysis model according to the label minimization loss function preset in step S1;

[0013] Step S6: Import the subordinate logistics policy documents of a specific time period into the trained analysis model, and obtain the keywords of the theme of the subordinate logistics policy documents according to the timestamp; import the superior logistics policy documents of the same time period into the trained analysis model, and obtain the keywords of the theme of the superior logistics policy documents according to the timestamp. In this process, the full text of the superior logistics policy documents and the keywords of the theme in the subheadings are obtained respectively, and dynamically adjusted by setting weight parameters;

[0014] Step S7: The evaluation module processes the keywords of the lower-level logistics policy document theme and the keywords of the upper-level logistics policy document theme through a similarity function to obtain the upper-lower and lower-level coordination degree and the similarity of keywords of different themes in different time periods.

[0015] Furthermore, in step S1, a logistics policy document word library consisting of logistics keywords is constructed; specifically:

[0016] Step S11: extracting potential words from the flow policy file as candidate words according to the window size based on the N-Gram model;

[0017] Step S12: Rank the candidate words based on the scores of self-information and mutual information, and import the candidate words into the logistics policy document word library according to the preset ranking threshold.

[0018] The ranking process is expressed as:

[0019] ;

[0020] ;

[0021] ;

[0022] ;

[0023] ;

[0024] ;

[0025] in, represents the entropy of the left neighbor string set, represents the entropy of the right neighbor string set, is the set of left neighbor strings, is the set of right neighbor strings, Indicates that in a given string Candidate words under the condition The probability of occurrence, is a combination of candidate word strings, Indicates the score function used to calculate candidate words The novelty, Represents a string combination And string combination The mutual information of is used to calculate the mutual dependence between two string combinations, Represents the normalization information function, which is used to eliminate the influence of the string combination length. Is a string combination And string combination The joint probability distribution function of , They are string combinations And string combination Marginal probability distribution function, To form candidate words The marginal probability distribution function of each sub-word of represents the length of the candidate word W, is the score of candidate word W, and are the preset weights respectively.

[0026] Furthermore, step S2 is specifically as follows:

[0027] Step S21: Obtain the probability distribution of the subject of the logistics policy document, specifically:

[0028] Define the total number of logistics policy documents as ,The total number of words in the logistics policy document is expressed as , An index of logistics policy documents, for one of a number of topics; specifically:

[0029] According to Dirichlet distribution Sampling generates Logistics Policy Documents The topic distribution , Indicates logistics policy document The proportion of each topic in; according to the Dirichlet distribution Generate each topic Word distribution , word distribution Indicates the subject The proportion of each word in is expressed as:

[0030] ;

[0031] ;

[0032] in, and Represents the preset parameters used to control the sparsity of the Dirichlet distribution;

[0033] Distribution by topic Sampling generates the Article Theme of the word Probability , where; from the multinomial distribution of words Sampling generates the In the document The probability of a word ;

[0034] Step S22: Obtaining word probability distribution , expressed as:

[0035] ;

[0036] For logistics policy documents , the probability distribution of its words After the text segmentation processing, it has been determined; the parameters are , Make an estimate and adjust and optimize according to the formula in step S22 to obtain the logistics policy document - theme and subject-word relationship;

[0037] Step S23: Obtain logistics policy documents Corresponding topic , The definition is as follows:

[0038] ;

[0039] in, ;

[0040] Step S24: Obtain the topic based on the Skip-gram word embedding model The context words of the words of the word, the central word and the optimal context word are concatenated to obtain the keywords of the topic; the details are as follows:

[0041] The size of the default logistics policy document vocabulary is (that is, the number of neurons in the input layer), the word vector dimension is , Indicates the first Candidate words;

[0042] Step S241: The topic in step S24 is embedded in the input layer of the Skip-gram word embedding model. The corresponding central word Perform one-hot encoding for initialization to obtain the input vector , represents the set of real numbers, The central word index in the vocabulary;

[0043] Step S242: The hidden layer of the Skip-gram word embedding model processes the input vector, which is expressed as:

[0044] ;

[0045] in, represents the output of the hidden layer, represents the weight matrix of the hidden layer;

[0046] Step S243: The output layer of the Skip-gram word embedding model processes the output of the hidden layer, which is expressed as:

[0047] ;

[0048] in, represents the output of the output layer, Represents the weight matrix of the output layer, that is, the central word The context word score of ,

[0049] The context word score is further converted into a score probability distribution through the activation function ; expressed as:

[0050] ;

[0051] in, represents the predicted context word, and c represents the central word Index in the Logistics Policy Documents Dictionary, Indicates a given central word The context word appears in the case of The probability of represents the score of the context word, Indicates the score of the central word;

[0052] Constructing the loss function , expressed as:

[0053] ;

[0054] in, is the context window size, Represents context words to the central word The offset of

[0055] Minimize the loss function to obtain the optimal context word, and concatenate the center word with the optimal context word to obtain the keywords of the topic;

[0056] Repeat step S24 to obtain several keywords of the topic according to different central words.

[0057] Furthermore, step S3 is specifically as follows:

[0058] The importance score of the keywords in the topic is calculated using TF-IDF, expressed as:

[0059] ;

[0060] ;

[0061] ;

[0062] in, Indicates word frequency, The keyword t representing the topic is in the document The number of times it appears in Representation Document The sum of the number of occurrences of keywords in all topics; represents the inverse document frequency, N represents the total number of documents in the logistics policy document set D, represents the number of logistics policy documents containing the keyword t of the topic, The importance score of the keywords representing the topic;

[0063] The importance score of each keyword is calculated by TF-IDF to sort the keywords, and a certain number of keywords are obtained by minimizing the perplexity, and a certain number of keywords ranked first are retained, and the certain number of keywords are indexed and marked according to the ranking; the perplexity calculation process is expressed as:

[0064] ;

[0065] in, represents the perplexity, D represents the set of logistics policy documents, Represents the topic distribution.

[0066] Furthermore, step S4 is specifically as follows:

[0067] Input the determined number of keywords in step S3 into the timing enhancement module, mark the timestamp according to the release time of the logistics policy document where the keywords are located for timing indexing, convert the keywords and the corresponding timestamps into feature dimensions respectively, and then concatenate the keywords and the corresponding timestamps after feature conversion to obtain the keywords of the topic that integrates the timing features.

[0068] Furthermore, in step S6, the subordinate logistics policy documents of a specific period are imported into the trained analysis model, and keywords of the subject of the subordinate logistics policy documents are obtained according to the timestamp; it is expressed as:

[0069] ;

[0070] ;

[0071] in, For the lower level logistics policy documents Time period theme Keywords in the Words, For the lower level logistics policy documents Time period theme The total number of keywords in For the lower level logistics policy documents Time period theme The probability weight of the topic-keyword, For the lower level logistics policy documents Time period theme Keywords, that is, keywords of the subject of the subordinate logistics policy document, For the lower level logistics policy documents The number of keyword categories for the topic of the time period, for The number of paragraphs corresponding to the topic, For the lower level logistics policy documents Keywords for all topics in the time period;

[0072] Import the superior logistics policy documents of the same period into the trained analysis model, and obtain the keywords of the theme of the superior logistics policy documents according to the timestamp. In this process, obtain the keywords of the theme of the full text and subtitle of the superior logistics policy documents respectively, and dynamically adjust them by setting weight parameters; expressed as:

[0073] ;

[0074] ;

[0075] ;

[0076] in, For The keywords of the topic of the superior logistics policy document in the time period Words, For The total number of keywords in the theme of the superior logistics policy documents in the time period, For The first keyword in the theme of the superior logistics policy document in the time period The TF-IDF importance score corresponding to each word, For the superior logistics policy documents Full text of the time period; For The first keyword in the subtitle of the superior logistics policy document of the time period Words, For The number of keywords in the subheadings of the superior logistics policy documents during the time period, For The first keyword in the subheading of the topic in the superior logistics policy document of the time period The frequency of occurrence of the word, For The subtitle of the superior logistics policy document for the time period; is the weight of the full-text keyword, For Superior logistics policy documents for the time period.

[0077] Furthermore, the similarity function in step S7 is specifically:

[0078] ;

[0079] ;

[0080] in, is the i-th subordinate logistics policy document in Keywords for all topics in the time period, is the i-th superior logistics policy document in Full text of the time period; For the coordination between superiors and subordinates, is the similarity of keywords of different topics in different time periods, is the i-th subordinate logistics policy document in Time period theme Keywords, is the i-th subordinate logistics policy document in Time period theme Keywords.

[0081] The beneficial effects of the present invention are:

[0082] The present invention provides a logistics policy document analysis method based on multimodal features, which extracts keywords of the subject from the logistics policy document, and then obtains the sorted keywords by sorting the keywords by importance score, and clarifies the boundaries of the keywords; further, by marking timestamps, multimodal keywords integrating temporal features are obtained, which has a more accurate prediction effect, and further lays the foundation for the subsequent evaluation of logistics policy documents according to time periods; in addition, the present invention evaluates the similarity of the superior logistics policy document and the subordinate logistics policy document through a similarity function, and obtains the superior-subordinate coordination degree to evaluate the implementation efficiency of the logistics policy document. In addition, in the process of obtaining the superior policy document, the present invention obtains more accurate keywords of the subject by adjusting the preset weight of the subject at the subheading. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 A flowchart of the steps of a logistics policy document analysis method based on multimodal features. DETAILED DESCRIPTION

[0084] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, in the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate the orientation or position relationship based on the orientation or position relationship shown in the accompanying drawings, which is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "No. 1", "No. 2" and "No. 3" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. The present invention is further explained below in conjunction with specific embodiments.

[0085] Reference Figure 1 ,A logistics policy document analysis method based on multimodal features includes the following steps:

[0086] Step S1: construct a logistics policy document word library, import logistics policy documents for word segmentation preprocessing, and obtain preprocessed documents; wherein, keywords of the topics of some logistics policy documents are manually annotated; the annotated tags include keywords of the manually selected topics and the release time of the corresponding logistics policy documents;

[0087] Step S2: construct an analysis model, which includes a keyword extraction module based on LDA2Vec, a sorting module, a time series enhancement module and an evaluation module, import the annotated preprocessing file in step S1 into the keyword extraction module based on LDA2Vec, and obtain several keywords of the topic in the annotated preprocessing file;

[0088] Step S3: inputting several keywords obtained in step S2 into the sorting module, sorting the keywords by calculating the importance score of each keyword through TF-IDF, and obtaining the determined number of keywords of the topic by calculating the perplexity, retaining the determined number of keywords ranked first, and indexing and marking the determined number of keywords according to the sorting;

[0089] Step S4: inputting the determined number of keywords in step S3 into the time series enhancement module, marking a timestamp on each keyword for time series indexing, and obtaining keywords of the topic integrating the time series features;

[0090] Step S5: construct a loss function, and optimize the parameters of the analysis model according to the label minimization loss function preset in step S1;

[0091] Step S6: Import the subordinate logistics policy documents of a specific time period into the trained analysis model, and obtain the keywords of the theme of the subordinate logistics policy documents according to the timestamp; import the superior logistics policy documents of the same time period into the trained analysis model, and obtain the keywords of the theme of the superior logistics policy documents according to the timestamp. In this process, the full text of the superior logistics policy documents and the keywords of the theme in the subheadings are obtained respectively, and dynamically adjusted by setting weight parameters;

[0092] Step S7: The evaluation module processes the keywords of the lower-level logistics policy document theme and the keywords of the upper-level logistics policy document theme through a similarity function to obtain the upper-lower and lower-level coordination degree and the similarity of keywords of different themes in different time periods.

[0093] Furthermore, in step S1, a logistics policy document word library consisting of logistics keywords is constructed; specifically:

[0094] Step S11: extracting potential words from the flow policy file as candidate words according to the window size based on the N-Gram model;

[0095] Step S12: Rank the candidate words based on the scores of self-information and mutual information, and import the candidate words into the logistics policy document word library according to the preset ranking threshold.

[0096] The ranking process is expressed as:

[0097] ;

[0098] ;

[0099] ;

[0100] ;

[0101] ;

[0102] ;

[0103] in, represents the entropy of the left neighbor string set, represents the entropy of the right neighbor string set, is the set of left neighbor strings, is the set of right neighbor strings, Indicates that in a given string Candidate words under the condition The probability of occurrence, is a combination of candidate word strings, Indicates the score function used to calculate candidate words The novelty, Represents a string combination And string combination The mutual information of is used to calculate the mutual dependence between two string combinations, Represents the normalization information function, which is used to eliminate the influence of the string combination length. Is a string combination And string combination The joint probability distribution function of , They are string combinations And string combination Marginal probability distribution function, To form candidate words The marginal probability distribution function of each sub-word of represents the length of the candidate word W, is the score of candidate word W, and are the preset weights respectively.

[0104] Furthermore, step S2 is specifically as follows:

[0105] Step S21: Obtain the probability distribution of the subject of the logistics policy document, specifically:

[0106] Define the total number of logistics policy documents as ,The total number of words in the logistics policy document is expressed as , An index of logistics policy documents, for one of a number of topics; specifically:

[0107] According to Dirichlet distribution Sampling generates Logistics Policy Documents The topic distribution , Indicates logistics policy document The proportion of each topic in; according to the Dirichlet distribution Generate each topic Word distribution , word distribution Indicates the subject The proportion of each word in is expressed as:

[0108] ;

[0109] ;

[0110] in, and Represents the preset parameters used to control the sparsity of the Dirichlet distribution;

[0111] Distribution by topic Sampling generates the Article Theme of the word Probability , where; from the multinomial distribution of words Sampling generates the In the document The probability of a word ;

[0112] Step S22: Obtaining word probability distribution , expressed as:

[0113] ;

[0114] For logistics policy documents , the probability distribution of its words After the text segmentation processing, it has been determined; the parameters are , Make an estimate and adjust and optimize according to the formula in step S22 to obtain the logistics policy document - theme and subject-word relationship;

[0115] Step S23: Obtain logistics policy documents Corresponding topic , The definition is as follows:

[0116] ;

[0117] in, ;

[0118] Step S24: Obtain the topic based on the Skip-gram word embedding model The context words of the words of the word, the central word and the optimal context word are concatenated to obtain the keywords of the topic; the details are as follows:

[0119] The size of the default logistics policy document vocabulary is (that is, the number of neurons in the input layer), the word vector dimension is , Indicates the first Candidate words;

[0120] Step S241: The topic in step S24 is embedded in the input layer of the Skip-gram word embedding model. The corresponding central word Perform one-hot encoding for initialization to obtain the input vector , represents the set of real numbers, The central word index in the vocabulary;

[0121] Step S242: The hidden layer of the Skip-gram word embedding model processes the input vector, which is expressed as:

[0122] ;

[0123] in, represents the output of the hidden layer, represents the weight matrix of the hidden layer;

[0124] Step S243: The output layer of the Skip-gram word embedding model processes the output of the hidden layer, which is expressed as:

[0125] ;

[0126] in, represents the output of the output layer, Represents the weight matrix of the output layer, that is, the central word The context word score of ,

[0127] The context word score is further converted into a score probability distribution through the activation function ; expressed as:

[0128] ;

[0129] in, represents the predicted context word, and c represents the central word Index in the Logistics Policy Documents Dictionary, Indicates a given central word The context word appears in the case of The probability of represents the score of the context word, Indicates the score of the central word;

[0130] Constructing the loss function , expressed as:

[0131] ;

[0132] in, is the context window size, Represents context words to the central word The offset of

[0133] Minimize the loss function to obtain the optimal context word, and concatenate the center word with the optimal context word to obtain the keywords of the topic;

[0134] Repeat step S24 to obtain several keywords of the topic according to different central words.

[0135] Furthermore, step S3 is specifically as follows:

[0136] The importance score of the keywords in the topic is calculated using TF-IDF, expressed as:

[0137] ;

[0138] ;

[0139] ;

[0140] in, Indicates word frequency, The keyword t representing the topic is in the document The number of times it appears in Representation Document The sum of the number of occurrences of keywords in all topics; represents the inverse document frequency, N represents the total number of documents in the logistics policy document set D, represents the number of logistics policy documents containing the keyword t of the topic, The importance score of the keywords representing the topic;

[0141] The importance score of each keyword is calculated by TF-IDF to sort the keywords, and a certain number of keywords are obtained by minimizing the perplexity, and a certain number of keywords ranked first are retained, and the certain number of keywords are indexed and marked according to the ranking; the perplexity calculation process is expressed as:

[0142] ;

[0143] in, represents the perplexity, D represents the set of logistics policy documents, Represents the topic distribution.

[0144] Furthermore, step S4 is specifically as follows:

[0145] Input the determined number of keywords in step S3 into the timing enhancement module, mark the timestamp according to the release time of the logistics policy document where the keywords are located for timing indexing, convert the keywords and the corresponding timestamps into feature dimensions respectively, and then concatenate the keywords and the corresponding timestamps after feature conversion to obtain the keywords of the topic that integrates the timing features.

[0146] Furthermore, in step S6, the subordinate logistics policy documents of a specific period are imported into the trained analysis model, and keywords of the subject of the subordinate logistics policy documents are obtained according to the timestamp; it is expressed as:

[0147] ;

[0148] ;

[0149] in, For the lower level logistics policy documents Time period theme The first keyword Words, For the lower level logistics policy documents Time period theme The total number of keywords in For the lower level logistics policy documents Time period theme The probability weight of the topic-keyword, For the lower level logistics policy documents Time period theme Keywords, that is, keywords of the subject of the subordinate logistics policy document, For the lower level logistics policy documents The number of keyword categories for the topic of the time period, for The number of paragraphs corresponding to the topic, For the lower level logistics policy documents Keywords for all topics in the time period;

[0150] Import the superior logistics policy documents of the same period into the trained analysis model, and obtain the keywords of the theme of the superior logistics policy documents according to the timestamp. In this process, obtain the keywords of the theme of the full text and subtitle of the superior logistics policy documents respectively, and dynamically adjust them by setting weight parameters; expressed as:

[0151] ;

[0152] ;

[0153] ;

[0154] in, For The keywords of the topic of the superior logistics policy document in the time period Words, For The total number of keywords in the theme of the superior logistics policy documents in the time period, For The first keyword in the theme of the superior logistics policy document in the time period The TF-IDF importance score corresponding to each word, For the superior logistics policy documents Full text of the time period; For The first keyword in the subtitle of the superior logistics policy document of the time period Words, For The number of keywords in the subheadings of the superior logistics policy documents during the time period, For The first keyword in the subheading of the topic in the superior logistics policy document of the time period The frequency of occurrence of the word, For The subtitle of the superior logistics policy document for the time period; is the weight of the full-text keyword, For Superior logistics policy documents for the time period.

[0155] Furthermore, the similarity function in step S7 is specifically:

[0156] ;

[0157] ;

[0158] in, is the i-th subordinate logistics policy document in Keywords for all topics in the time period, is the i-th superior logistics policy document in Full text of the time period; For the coordination between superiors and subordinates, is the similarity of keywords of different topics in different time periods, is the i-th subordinate logistics policy document in Time period theme Keywords, is the i-th subordinate logistics policy document in Time period theme Keywords.

[0159] The specific implementation modes as described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation mode of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A logistics policy document analysis method based on multimodal features, characterized in that: The following steps are involved: Step S1: construct a logistics policy document word library, import logistics policy documents for word segmentation preprocessing, and obtain preprocessed documents; wherein, keywords of the topics of some logistics policy documents are manually annotated; the annotated tags include keywords of the manually selected topics and the release time of the corresponding logistics policy documents; Step S2: construct an analysis model, which includes a keyword extraction module based on LDA2Vec, a sorting module, a time series enhancement module and an evaluation module, import the annotated preprocessing file in step S1 into the keyword extraction module based on LDA2Vec, and obtain several keywords of the topic in the annotated preprocessing file; Step S3: inputting several keywords obtained in step S2 into the sorting module, sorting the keywords by calculating the importance score of each keyword through TF-IDF, and obtaining the determined number of keywords of the topic by calculating the perplexity, retaining the determined number of keywords ranked first, and indexing and marking the determined number of keywords according to the sorting; Step S4: inputting the determined number of keywords in step S3 into the time series enhancement module, marking a timestamp on each keyword for time series indexing, and obtaining keywords of the topic integrating the time series features; Step S5: construct a loss function, and optimize the parameters of the analysis model according to the label minimization loss function preset in step S1; Step S6: Import the subordinate logistics policy documents of a specific time period into the trained analysis model, and obtain the keywords of the theme of the subordinate logistics policy documents according to the timestamp; import the superior logistics policy documents of the same time period into the trained analysis model, and obtain the keywords of the theme of the superior logistics policy documents according to the timestamp. In this process, the full text of the superior logistics policy documents and the keywords of the theme in the subheadings are obtained respectively, and dynamically adjusted by setting weight parameters; Step S7: The evaluation module processes the keywords of the lower-level logistics policy document theme and the keywords of the upper-level logistics policy document theme through a similarity function to obtain the upper-lower and lower-level coordination degree and the similarity of keywords of different themes in different time periods.

2. The method for analyzing logistics policy documents based on multimodal features according to claim 1 is characterized in that: In step S1, a logistics policy document word library consisting of logistics keywords is constructed; specifically: Step S11: extracting potential words from the flow policy file as candidate words according to the window size based on the N-Gram model; Step S12: Rank the candidate words based on the scores of self-information and mutual information, and import the candidate words into the logistics policy document word library according to the preset ranking threshold. The ranking process is expressed as: ; ; ; ; ; ; in, represents the entropy of the left neighbor string set, represents the entropy of the right neighbor string set, is the set of left neighbor strings, is the set of right neighbor strings, Indicates that in a given string Candidate words under the condition The probability of occurrence, is a combination of candidate word strings, Indicates the score function used to calculate candidate words The novelty, Represents a string combination And string combination The mutual information of is used to calculate the mutual dependence between two string combinations, Represents the normalization information function, which is used to eliminate the influence of the string combination length. Is a string combination And string combination The joint probability distribution function of , They are string combinations And string combination Marginal probability distribution function, To form candidate words The marginal probability distribution function of each sub-word of represents the length of the candidate word W, is the score of candidate word W, and are the preset weights respectively.

3. The method for analyzing logistics policy documents based on multimodal features according to claim 2 is characterized in that: Step S2 is specifically as follows: Step S21: Obtain the probability distribution of the subject of the logistics policy document, specifically: Define the total number of logistics policy documents as ,The total number of words in the logistics policy document is expressed as , An index of logistics policy documents, for one of a number of topics; specifically: According to Dirichlet distribution Sampling generates Logistics Policy Documents The topic distribution , Indicates logistics policy document The proportion of each topic in; according to the Dirichlet distribution Generate each topic Word distribution , word distribution Indicates the subject The proportion of each word in is expressed as: ; ; in, and Represents the preset parameters used to control the sparsity of the Dirichlet distribution; Distribution by topic Sampling generates the Article Theme of the word Probability , where; from the multinomial distribution of words Sampling generates the In the document The probability of a word ; Step S22: Obtaining word probability distribution , expressed as: ; For logistics policy documents , the probability distribution of its words After the text segmentation processing, it has been determined; the parameters are , Make an estimate and adjust and optimize according to the formula in step S22 to obtain the logistics policy document - theme and subject-word relationship; Step S23: Obtain logistics policy documents Corresponding topic , The definition is as follows: ; in, ; Step S24: Obtain the topic based on the Skip-gram word embedding model The context words of the words of the word, the central word and the optimal context word are concatenated to obtain the keywords of the topic; the details are as follows: The size of the default logistics policy document vocabulary is , the word vector dimension is , Indicates the first Candidate words; Step S241: The topic in step S24 is embedded in the input layer of the Skip-gram word embedding model. The corresponding central word Perform one-hot encoding for initialization to obtain the input vector , represents the set of real numbers, The central word index in the vocabulary; Step S242: The hidden layer of the Skip-gram word embedding model processes the input vector, which is expressed as: ; in, represents the output of the hidden layer, represents the weight matrix of the hidden layer; Step S243: The output layer of the Skip-gram word embedding model processes the output of the hidden layer, which is expressed as: ; in, represents the output of the output layer, Represents the weight matrix of the output layer, that is, the central word The context word score of , Further, the context word score is converted into a score probability distribution through the activation function ; expressed as: ; in, represents the predicted context word, and c represents the central word Index in the Logistics Policy Documents Dictionary, Indicates a given central word The context word appears in the case of The probability of represents the score of the context word, Indicates the score of the central word; Constructing the loss function , expressed as: ; in, is the context window size, Represents context words to the central word The offset of Minimize the loss function to obtain the optimal context word, and concatenate the center word with the optimal context word to obtain the keywords of the topic; Repeat step S24 to obtain several keywords of the topic according to different central words.

4. The method for analyzing logistics policy documents based on multimodal features according to claim 3 is characterized in that: Step S3 is specifically as follows: The importance score of the keywords in the topic is calculated using TF-IDF, expressed as: ; ; ; in, Indicates word frequency, The keyword t representing the topic is in the document The number of times it appears in Representation Document The sum of the occurrences of keywords in all topics; represents the inverse document frequency, N represents the total number of documents in the logistics policy document set D, represents the number of logistics policy documents containing the keyword t of the topic, The importance score of the keywords representing the topic; The importance score of each keyword is calculated by TF-IDF to sort the keywords, and a certain number of keywords are obtained by minimizing the perplexity, and a certain number of keywords ranked first are retained, and the certain number of keywords are indexed and marked according to the ranking; the perplexity calculation process is expressed as: ; in, represents the perplexity, D represents the set of logistics policy documents, Represents the topic distribution.

5. The method for analyzing logistics policy documents based on multimodal features according to claim 4 is characterized in that: Step S4 is specifically as follows: Input the determined number of keywords in step S3 into the timing enhancement module, mark the timestamp according to the release time of the logistics policy document where the keywords are located for timing indexing, convert the keywords and the corresponding timestamps into feature dimensions respectively, and then concatenate the keywords and the corresponding timestamps after feature conversion to obtain the keywords of the topic that integrates the timing features.

6. The method for analyzing logistics policy documents based on multimodal features according to claim 5 is characterized in that: In step S6, the subordinate logistics policy documents of a specific period are imported into the trained analysis model, and the keywords of the subject of the subordinate logistics policy documents are obtained according to the timestamp; it is expressed as: ; ; in, For the lower level logistics policy documents Time period theme Keywords in the Words, For the lower level logistics policy documents Time period theme The total number of keywords in For the lower level logistics policy documents Time period theme The probability weight of the topic-keyword, For the lower level logistics policy documents Time period theme Keywords, that is, keywords of the subject of the subordinate logistics policy document, For the lower level logistics policy documents The number of keyword categories for the topic of the time period, for The number of paragraphs corresponding to the topic, For the lower level logistics policy documents Keywords for all topics in the time period; Import the superior logistics policy documents of the same period into the trained analysis model, and obtain the keywords of the theme of the superior logistics policy documents according to the timestamp. In this process, obtain the keywords of the theme of the full text and subtitle of the superior logistics policy documents respectively, and dynamically adjust them by setting weight parameters; expressed as: ; ; ; in, For The keywords of the topic of the superior logistics policy document in the time period Words, For The total number of keywords in the theme of the superior logistics policy documents in the time period, For The first keyword in the theme of the superior logistics policy document in the time period The TF-IDF importance score corresponding to each word, For the superior logistics policy documents Full text of the time period; For The first keyword in the subtitle of the superior logistics policy document of the time period Words, For The number of keywords in the subheadings of the superior logistics policy documents during the time period, For The first keyword in the subheading of the topic in the superior logistics policy document of the time period The frequency of occurrence of the word, For The subtitle of the superior logistics policy document for the time period; is the weight of the full-text keyword, For Superior logistics policy documents for the time period.

7. The method for analyzing logistics policy documents based on multimodal features according to claim 6 is characterized in that: The similarity function in step S7 is specifically: ; ; in, is the i-th subordinate logistics policy document in Keywords for all topics in the time period, is the i-th superior logistics policy document in Full text of the time period; For the coordination between superiors and subordinates, is the similarity of keywords of different topics in different time periods, is the i-th subordinate logistics policy document in Time period theme Keywords, is the i-th subordinate logistics policy document in Time period theme Keywords.

Citation Information

Patent Citations

  • Sales prediction method based on policy-public opinion-purchase double-order deep learning

    CN114511345A

  • Resource classification screening method and system based on science and technology policies

    CN116644174A