A government affairs information management system based on data matching

By designing a government information management system based on data matching, the integration and accurate matching of multimodal data is achieved, and the problems of cross-modal information separation and keyword extraction efficiency in the existing system are solved, and the efficiency and accuracy of government information management are improved.

CN119831812BActive Publication Date: 2025-05-27四川省大数据技术服务中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510324031.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-05-27
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The existing government information management system lacks the ability to integrate multi-source data in data processing, resulting in cross-modal information separation, low keyword extraction efficiency, and difficult to meet the needs of digital government construction.

Method used

A government information management system based on data matching is designed, and the fusion processing and precise matching of multimodal data is achieved through the government information collection and processing module, keyword group extraction classification module, keyword group matching module, identity verification module, user keyword group search module and storage module.

Benefits of technology

Through multimodal fusion and precise extraction, the accuracy of multimodal matching is improved, dynamic weight optimization improves the recall rate of key information, intelligent classification and time similarity factors reduce the error correlation rate, and layered encrypted storage and dynamic permission control improve data security and user retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831812B_ABST
    Figure CN119831812B_ABST
Patent Text Reader

Abstract

The present invention discloses a government information management system based on data matching, which relates to the technical field of data management. The present invention realizes cross-modal keyword association through multimodal fusion and precise extraction, improves the accuracy of multimodal matching; dynamic weight optimization improves the recall rate of key information; combines BERT semantic vectors with source department weights to achieve intelligent classification; introduces time similarity factors to improve the matching priority of recent policy documents and reduce the misassociation rate. The use of layered encrypted storage and dynamic permission management improves the decryption throughput and unauthorized access interception rate, and intelligent association retrieval and multimodal result presentation improve user retrieval efficiency and cross-device compatibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and particularly to a government information management system based on data matching. Background Art

[0002] Current government information management mainly relies on traditional technologies, presenting significant bottlenecks. In data processing, the system mostly adopts single-modal analysis (such as pure text), lacking the ability to fuse multi-source data such as images and videos, resulting in cross-modal information fragmentation. For example, the flood control notice text and the material dispatching image cannot be automatically associated and require manual matching. Keyword extraction relies on traditional algorithms such as TF-IDF, which is difficult to identify compound terms, ignores the weight of file names, and does not combine the authority of data sources. The classification results are easily interfered by high-frequency common words. These problems lead to low efficiency in government information management, difficult cross-departmental collaboration, and are difficult to meet the needs of digital government construction. Summary of the Invention

[0003] The purpose of the present invention is to provide a government information management system based on data matching to improve the above technical problems.

[0004] To achieve the above invention purpose, the embodiments of the present invention provide the following technical solutions:

[0005] A government information management system based on data matching includes:

[0006] A government information collection and processing module, configured to obtain the government information of the current batch and perform preprocessing to obtain the preprocessed government information;

[0007] A keyword group extraction and classification module, configured to extract keyword groups from the preprocessed government information and classify them to obtain corresponding classification results;

[0008] A keyword group matching module, configured to match the government information of the current batch and the government information of the historical batch according to the keyword groups and classification results to obtain corresponding matching results;

[0009] An identity verification module, configured to perform identity verification on the user; obtain user keyword groups and user classification results;

[0010] A user keyword group retrieval module, configured to retrieve and decrypt the user keyword groups based on the user classification results to obtain retrieval results, thereby completing the management of government information;

[0011] A storage module, configured to encrypt / decrypt and store the government information, keyword groups, classification results, and matching results.

[0012] Further, the preprocessed government information includes encoded text data, RGB-image data, and RGB-video data.

[0013] Furthermore, the processing process of the keyword group extraction and classification module includes:

[0014] Using the TF-IDF algorithm and the BERT-TextRank hybrid model to extract keywords from the encoded text data, RGB-image data, and RGB-video data, obtaining the corresponding first keyword group, second keyword group, and third keyword group;

[0015] Calculating the relevance between the first keyword group and the second keyword group, and the third keyword group respectively;

[0016] Merging the first keyword group, second keyword group, or / and third keyword group whose relevance meets the fusion condition, and retaining the remaining first keyword group, second keyword group, and third keyword group to obtain the keyword group;

[0017] Using the Bert-LightGBM model to classify the keyword group to obtain the corresponding classification result; the classification result includes policy, administration, social service, and others.

[0018] Furthermore, the processing process of the keyword extraction includes:

[0019] Concatenating, segmenting, and removing the encoded text data and its file name to obtain the processed encoded text; setting the word weight, using the BERT-TextRank hybrid model to perform semantic analysis on the processed encoded text, selecting the top 5 high-frequency words in terms of frequency, and determining the first keyword group;

[0020] Using the OCR-ResNet model to extract the first text data of the RGB-image data; concatenating the first text data and its corresponding file name; processing the concatenated first text data in the same way as extracting the first keyword group of the encoded text data to determine the second keyword group;

[0021] Using the speech converter to convert the RGB-video data into the second text data; randomly intercepting N video frames, using the OCR-ResNet model to extract the corresponding third text data; concatenating the second text data, third text data, and their file names to obtain the fourth text data; processing the concatenated fourth text data in the same way as extracting the first keyword group of the encoded text data to determine the third keyword group.

[0022] Furthermore, the calculation of the relevance between the first keyword group and the second keyword group, and the third keyword group respectively includes:

[0023] Calculating the corresponding cosine similarity based on the vectors corresponding to the first keyword group and the second keyword group;

[0024] Calculate the corresponding semantic vector similarity based on the keywords corresponding to the first keyword group and the second keyword group;

[0025] Set the relevance weight coefficient and the rule weight;

[0026] Calculate the corresponding relevance based on the cosine similarity, semantic vector similarity, relevance weight coefficient, and rule weight;

[0027] The calculation process of the relevance between the first keyword group and the third keyword group is the same as that between the first keyword group and the second keyword group.

[0028] Further, classifying the keyword groups using the Bert-LightGBM model to obtain the corresponding classification results, including:

[0029] Obtain the corresponding source department and set the corresponding source weight;

[0030] Use the Bert model to extract the semantic vectors of the keyword groups and combine them with the source weights to obtain the feature vectors;

[0031] Input the feature vectors into the LightGBM model and output the classification results.

[0032] Further, the processing process of the keyword group matching module includes:

[0033] Based on the classification results, call the historical batches of government affairs information and their corresponding keyword groups in the corresponding database; calculate the relevance between the keyword groups of the current batch and the keyword groups of the historical batches; select the historical batch of government affairs information with the highest relevance and construct a link with the government affairs information of the current batch to obtain the corresponding matching results.

[0034] Further, calculating the relevance between the keyword groups of the current batch and the keyword groups of the historical batches includes:

[0035] Calculate the initial semantic similarity between each keyword in the keyword group of the current batch and each keyword in the keyword group of the historical batch respectively;

[0036] Calculate the average value of each initial semantic similarity and use it as the semantic similarity;

[0037] Calculate the corresponding word frequency similarity based on the intersection and union of the keyword groups of the current batch and the keyword groups of the historical batch;

[0038] Calculate the corresponding time similarity based on the storage times corresponding to the keyword groups of the current batch and the keyword groups of the historical batch;

[0039] Calculate the corresponding correlation based on semantic similarity, word frequency similarity, and time similarity.

[0040] Further, the authentication module includes:

[0041] Obtain the user's identity information or work permit information and conduct authentication; if the authentication is successful, input the user keyword group through the client, and classify the user keyword group through the keyword group extraction and classification module to obtain the user classification result; otherwise, the authentication fails.

[0042] Further, the storage module includes:

[0043] An encryption / decryption unit for encrypting / decrypting government affairs information, its keyword group, classification result, and matching result;

[0044] A keyword storage unit for storing keyword groups;

[0045] A political and government affairs information storage unit for all batches of political and government affairs information, classification results, and matching results;

[0046] An administrative government affairs information storage unit for all batches of administrative government affairs information, classification results, and matching results;

[0047] A social service government affairs information storage unit for all batches of social service government affairs information, classification results, and matching results;

[0048] An other government affairs information storage unit for all batches of other government affairs information, classification results, and matching results.

[0049] The beneficial effects of the present invention are as follows:

[0050] Through multi-modal fusion and precise extraction, the present invention realizes cross-modal keyword association, improves the multi-modal matching accuracy; dynamic weight optimization improves the recall rate of key information; combines BERT semantic vectors and source department weights to achieve intelligent classification; introduces a time similarity factor to improve the matching priority of recent policy documents and reduce the mis-association rate; adopts hierarchical encrypted storage and dynamic permission control to improve the decryption throughput and unauthorized access interception rate, and intelligent associated retrieval and multi-modal result presentation improve the user retrieval efficiency and cross-device compatibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0052] Figure 1 This is the system structure diagram in the embodiment of the present invention;

[0053] Figure 2 This is the storage module structure diagram in the embodiment of the present invention;

[0054] Figure 3 This is the method flow chart in the embodiment of the present invention. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0056] Please refer to Figure 1 , a government information management system based on data matching provided in this embodiment, which includes:

[0057] A government information collection and processing module, configured to obtain the government information of the current batch and perform preprocessing, encode unstructured text data such as pdf, txt, and word, and perform format conversion on image data and video data to obtain preprocessed government information, that is, obtain encoded text data, RGB-image data, and RGB-video data;

[0058] A keyword group extraction and classification module, configured to extract keyword groups from the preprocessed government information and perform classification to obtain corresponding classification results;

[0059] The processing process of the keyword group extraction and classification module includes:

[0060] S21. Use the TF-IDF algorithm and the BERT-TextRank hybrid model to extract keywords from the encoded text data, RGB-image data, and RGB-video data to obtain corresponding first keyword groups, second keyword groups, and third keyword groups;

[0061] The S21 includes:

[0062] Concatenate, segment, and remove the encoded text data and its file name, removing common but meaningless words such as "de", "shi", "zai", etc. to reduce the text complexity and obtain the processed encoded text; set the word weights, and use the BERT-TextRank hybrid model to perform semantic analysis on the processed encoded text to determine the first keyword group;

[0063] Use the BERT-TextRank hybrid model to perform semantic analysis on the processed encoded text, select the top 5 high-frequency words in terms of frequency, and determine the first keyword group, including:

[0064] Use the TextRank algorithm to process the processed encoded text and extract the corresponding key sentences; input the key sentences and the processed encoded text into the BERT model, and output the fused text representation; use the TF-IDF algorithm to count the high-frequency words in the fused text representation, multiply the frequency of the keywords extracted from the name by the word weight of 1.5, and multiply the frequency of the keywords extracted from the text data by the word weight of 1.0, select the top 5 high-frequency words in terms of frequency, and determine the first keyword group.

[0065] Use the OCR-ResNet model to extract the first text data of the RGB-image data; concatenate the first text data and its corresponding file name; use the same method as extracting the first keyword group of the encoded text data to process the concatenated first text data to determine the second keyword group.

[0066] Use the speech converter to convert the RGB-video data into the second text data; randomly intercept N video frames, and use the OCR-ResNet model to extract the corresponding third text data; concatenate the second text data, the third text data and their file names to obtain the fourth text data; use the same method as extracting the first keyword group of the encoded text data to process the concatenated fourth text data to determine the third keyword group.

[0067] Precisely extract high-frequency keywords from three aspects of text, image, and video, set corresponding weights, which can prevent important keywords from being ignored, improve the accuracy and comprehensiveness of subsequent retrieval, and is beneficial to the accuracy of subsequent classification.

[0068] S22. Calculate the relevance between the first keyword group and the second keyword group and the third keyword group respectively;

[0069] The S22 includes:

[0070] Based on the vectors corresponding to the first keyword group and the second keyword group, calculate the corresponding cosine similarity;

[0071] Based on the keywords corresponding to the first keyword group and the second keyword group, calculate the corresponding semantic vector similarity;

[0072] Set the relevance weight coefficient and the rule weight;

[0073] Calculate the corresponding relevance based on the cosine similarity, semantic vector similarity, relevance weight coefficient, and rule weight;

[0074] The calculation process of the relevance between the first keyword group and the third keyword group is the same as that between the first keyword group and the second keyword group.

[0075] Furthermore, the relevance of this embodiment has the formula:

[0076] ;

[0077] ;

[0078] ;

[0079] Among them, , represent the relevance weight coefficient, , represent the cosine similarity function and the sine function respectively, , represent the first keyword group and the keyword group respectively, , represent the first keyword group and the keyword group corresponding vectors, , represent the th keyword in the first keyword group and the keyword group and the th keyword respectively, represents the maximum value function, represents the normalized first keyword group, represents the rule weight, represents the first keyword group and the keyword group corresponding cosine similarity, represents the first keyword group and the keyword group corresponding semantic vector similarity.

[0080] When the keyword involves emergency measures, the weight rule The value is 0.2; in other cases, the weight rule The value is 0.

[0081] S23. Merge the first keyword group, the second keyword group, and / or the third keyword group whose relevance meets the fusion condition, and retain the remaining first keyword group, second keyword group, and third keyword group to obtain a keyword group; the fusion condition is: if the relevance between the first keyword group and the second keyword group is greater than the first fusion threshold and the relevance between the first keyword group and the third keyword group is greater than the second fusion threshold, then merge the corresponding first keyword group, second keyword group, and third keyword group to obtain an updated keyword group; if the relevance is greater than the first fusion threshold or the relevance between the first keyword group and the third keyword group is greater than the second fusion threshold, then merge the corresponding first keyword group and the second keyword group or merge the corresponding first keyword group and the third keyword group to obtain an updated keyword group; in other cases, retain the corresponding first keyword group, second keyword group, and third keyword group.

[0082] The keyword group includes the updated keyword group and the first keyword group, second keyword group, and third keyword group whose relevance does not exceed the first fusion threshold or the second fusion threshold.

[0083] S24. Classify the keyword group using the Bert-LightGBM model to obtain the corresponding classification result; among them, the classification result includes policy, administration, social service, and others.

[0084] The S24 includes:

[0085] Obtain the corresponding source department, set the corresponding source weight. For example, the source weight of the document issued by the State Council is 1.5, the source weight of the document issued by the department or bureau unit is 1.0, and the source weight of the document issued by the district or county unit is 0.8; use the Bert model to extract the semantic vector of the keyword group and combine it with the source weight to obtain a feature vector, and the corresponding format is [semantic vector, source weight, number of keywords]; input the feature vector into the LightGBM model, and output the classification result.

[0086] The training process of the LightGBM model is:

[0087] Set the initial model parameters of the LightGBM model, the number of leaf nodes is 31, the learning rate is 0.05, and the number of base learners (trees) is 200;

[0088] Obtain the training feature vector and the test feature vector; input the training feature vector into the LightGBM model, use the incremental learning mechanism to train the LightGBM model, and adjust the model parameters of the LightGBM model in real time to obtain the initially trained LightGBM model;

[0089] Input the test feature vector into the initially trained LightGBM model to obtain the corresponding test results; based on the test results, formulate the corresponding confusion matrix according to the number of correct classifications and incorrect classifications; in this embodiment, the confusion matrix is shown in Table 1.

[0090] Table 1

[0091]

[0092] Based on the confusion matrix, analyze the initially trained LightGBM model to obtain the corresponding analysis results, and calculate the corresponding evaluation indicators, such as F2 score, recall rate, etc.

[0093] Based on the analysis results, adjust the model parameters of the initially trained LightGBM model to obtain the trained LightGBM model.

[0094] The keyword group matching module is used to match the current batch of government affairs information and the historical batch of government affairs information according to the keyword group and the classification result to obtain the corresponding matching result.

[0095] The processing process of the keyword group matching module includes:

[0096] Based on the classification result, call the historical batch of government affairs information and its corresponding keyword group in the corresponding database; calculate the correlation between the keyword group of the current batch and the keyword group of the historical batch; select the historical batch of government affairs information with the highest correlation and construct a link with the current batch of government affairs information, that is, obtain the corresponding matching result.

[0097] The calculation of the correlation between the keyword group of the current batch and the keyword group of the historical batch includes:

[0098] Calculate the initial semantic similarity between each keyword in the keyword group of the current batch and each keyword in the keyword group of the historical batch respectively.

[0099] Calculate the average value of each initial semantic similarity and use it as the semantic similarity.

[0100] Based on the union and intersection of the keyword group of the current batch and the keyword group of the historical batch, calculate the corresponding word frequency similarity.

[0101] Based on the storage time corresponding to the keyword group of the current batch and the keyword group of the historical batch, calculate the corresponding time similarity.

[0102] Based on the semantic similarity, word frequency similarity and time similarity, calculate the corresponding correlation.

[0103] Thus, the relevance of this embodiment is given by the formula:

[0104] ;

[0105] ;

[0106] ;

[0107] ;

[0108] where , , represent the relevance weight coefficients, , , represent semantic similarity, word frequency similarity, and time similarity respectively, , represent the th keyword group of the current batch , the th keyword group of the historical batch corresponding vectors, , represent the size of the intersection and the size of the union of vectors and vector respectively, represents cosine similarity, , represent the first keywords in keyword group and keyword group respectively, represents the summation function, , represent the total number of keywords in keyword group and keyword group respectively, , represent the th keyword of keyword group , the th keyword in keyword group respectively, represents the time difference corresponding to keyword group and keyword group , represents the maximum time difference.

[0109] By calculating the correlation between the keyword groups of the current batch and those of the historical batches, the historical information most relevant to the current government affairs information can be quickly found, improving the data matching efficiency. The introduction of time similarity enhances the matching priority of recent policy documents and reduces the mis-association rate between old policies and new documents. By constructing the links between the government affairs information of the current batch and the historical batches, the relevance and traceability of the data are enhanced.

[0110] An authentication module for authenticating users; obtaining user keyword groups and user classification results;

[0111] The authentication module includes:

[0112] Obtain the user's identity information or work permit information and conduct authentication; if the authentication is successful, input the user keyword groups through the client, and classify the user keyword groups through the keyword group extraction and classification module to obtain the user classification results; otherwise, the authentication fails.

[0113] The user keyword group retrieval module is used to retrieve and decrypt the user keyword groups based on the user classification results to obtain the retrieval results and complete the management of government affairs information;

[0114] The retrieval results include the government affairs information with the highest correlation with the user keyword groups and the government affairs information connected to this government affairs information. Connection refers to other government affairs information with the highest correlation with the keywords of this government affairs information and the closest time, which can provide additional information to the user. The user can obtain similar government affairs information in the previous time period without additional retrieval, reducing the retrieval workload and working time and obtaining relevant documents in the shortest time.

[0115] Based on the user classification results, determine the corresponding storage unit; in this storage unit, search for the government affairs information corresponding to the keyword group with the highest correlation with the user keyword groups; according to the connection, determine other government affairs information related to this government affairs information; after decrypting the retrieved government affairs information, package it into a retrieval report and output it to the client.

[0116] The storage module is used to encrypt / decrypt and store government affairs information, keyword groups, classification results, and matching results.

[0117] As Figure 2 shown, store the keyword groups into the corresponding keyword storage units; based on the classification results, encrypt the corresponding government affairs information and matching results and input them into the corresponding storage units, that is, store them into the political government affairs information storage unit, administrative government affairs information storage unit, social service government affairs information storage unit, and other government affairs information storage units respectively. Thus, the storage module includes:

[0118] An encryption / decryption unit for encrypting / decrypting government affairs information, its keyword groups, classification results, and matching results;

[0119] A keyword storage unit for storing keyword groups;

[0120] A political government affairs information storage unit for all batches of political government affairs information, classification results, and matching results;

[0121] An administrative government affairs information storage unit for all batches of administrative government affairs information, classification results, and matching results;

[0122] A social service government affairs information storage unit for all batches of social service government affairs information, classification results, and matching results;

[0123] An other government affairs information storage unit for all batches of other government affairs information, classification results, and matching results.

[0124] Adopt hierarchical encrypted storage to encrypt different types of government affairs information (text, image, video), improving data security. Store government affairs information, its keyword groups, classification results, and matching results in different storage units respectively, facilitating management and retrieval. By optimizing the encryption algorithm and storage structure, improve the decryption throughput and reduce data access latency.

[0125] As Figure 3 shown, the government affairs information management method corresponding to the government affairs information management system includes:

[0126] S1. Obtain the government affairs information of the current batch and perform preprocessing to obtain the preprocessed government affairs information; wherein, the preprocessed government affairs information includes encoded text data, RGB-image data, and RGB-video data;

[0127] S2. Extract, associate, and classify keywords from the preprocessed government affairs information to obtain the corresponding keyword groups and classification results;

[0128] S3. Based on the keyword groups and classification results, match the government affairs information with the government affairs information of historical batches to obtain the corresponding matching results;

[0129] S4. Encrypt the government affairs information, keyword groups, classification results, and matching results and store them in the database (storage module); store the keyword groups in the corresponding keyword storage unit; based on the classification results, encrypt and input the corresponding government affairs information and matching results into the corresponding storage units, that is, store them in the political government affairs information storage unit, administrative government affairs information storage unit, social service government affairs information storage unit, and other government affairs information storage unit respectively.

[0130] S5. Authenticate the user; obtain the user keyword group and the user classification result;

[0131] S6. Based on the user classification result, retrieve and decrypt the user keyword group to obtain the retrieval result, and complete the management of government affairs information.

[0132] In summary, through multi-modal fusion and precise extraction, the present invention realizes cross-modal keyword association, improves the multi-modal matching accuracy; dynamic weight optimization improves the recall rate of key information; combines BERT semantic vectors with the weights of source departments to achieve intelligent classification; introduces a time similarity factor to improve the matching priority of recent policy documents and reduce the mis-association rate; adopts hierarchical encrypted storage and dynamic permission control to improve the decryption throughput and unauthorized access interception rate, and intelligent association retrieval and multi-modal result presentation improve the user retrieval efficiency and cross-device compatibility. The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A government information management system based on data matching, characterized in that: include: The government information collection and processing module is used to obtain the current batch of government information and pre-process it to obtain the pre-processed government information; A keyword group extraction and classification module is used to extract and classify keyword groups in the pre-processed government information to obtain corresponding classification results; A keyword group matching module is used to match the current batch of government affairs information with the historical batch of government affairs information according to the keyword group and the classification result to obtain the corresponding matching result; The identity verification module is used to verify the identity of the user and obtain the user's keyword phrases and user classification results; A user keyword group search module is used to search and decrypt user keyword groups based on user classification results, obtain search results, and complete the management of government information; A storage module, used to encrypt / decrypt and store government information and keyword groups, classification results and matching results; The processing process of the keyword group matching module includes: Based on the classification results, the government affairs information of the historical batches and their corresponding keyword groups in the corresponding database are called; the correlation between the keyword groups of the current batch and the keyword groups of the historical batches is calculated; the government affairs information of the historical batch with the highest correlation is selected, and a link is constructed between the government affairs information of the current batch, so as to obtain the corresponding matching results; The calculation of the correlation between the keyword groups of the current batch and the keyword groups of the historical batches includes: Calculate the initial semantic similarity between each keyword in the keyword group of the current batch and each keyword in the keyword group of the historical batch respectively; Calculate the average of each initial semantic similarity and use it as the semantic similarity; Based on the intersection and union of the keyword groups of the current batch and the keyword groups of the historical batches, the corresponding word frequency similarity is calculated; Based on the storage time corresponding to the keyword groups of the current batch and the keyword groups of the historical batches, the corresponding time similarity is calculated; Based on semantic similarity, word frequency similarity and time similarity, the corresponding correlation is calculated.

2. The government affairs information management system based on data matching according to claim 1 is characterized in that: The preprocessed government information includes encoded text data, RGB-image data and RGB-video data.

3. The government affairs information management system based on data matching according to claim 2 is characterized in that: The processing process of the keyword group extraction classification module includes: Using the TF-IDF algorithm and the BERT-TextRank hybrid model, keywords are extracted from the encoded text data, the RGB-image data, and the RGB-video data to obtain the corresponding first keyword group, the second keyword group, and the third keyword group; Calculate the relevance of the first keyword group with the second keyword group and the third keyword group respectively; Merge the first keyword group, the second keyword group, or / and the third keyword group whose relevance meets the fusion condition, and retain the remaining first keyword group, second keyword group, and third keyword group to obtain a keyword group; The Bert-LightGBM model is used to classify the keyword groups to obtain corresponding classification results; the classification results include policy, administration, social services and others.

4. The government affairs information management system based on data matching according to claim 3 is characterized in that: The keyword extraction process includes: The coded text data and its file name are concatenated, segmented and removed to obtain the processed coded text; the word weight is set, and the BERT-TextRank hybrid model is used to perform semantic analysis on the processed coded text, and the high-frequency words with the top 5 frequencies are selected to determine the first keyword group; Extracting first text data from the RGB-image data using the OCR-ResNet model; splicing the first text data and its corresponding file name; processing the spliced ​​first text data using the same method as the first keyword group extracted from the encoded text data to determine a second keyword group; The RGB-video data is converted into second text data by using a voice converter; N video frames are randomly intercepted, and the corresponding third text data are extracted by using an OCR-ResNet model; the second text data, the third text data and their file names are spliced ​​to obtain fourth text data; the spliced ​​fourth text data is processed by the same method as the first keyword group of the encoded text data to determine the third keyword group.

5. The government affairs information management system based on data matching according to claim 3 is characterized in that: The calculating the relevance of the first keyword group with the second keyword group and the third keyword group respectively includes: Based on the vectors corresponding to the first keyword group and the second keyword group, the corresponding cosine similarity is calculated; Based on the keywords corresponding to the first keyword group and the second keyword group, calculating the corresponding semantic vector similarity; Set the relevance weight coefficient and rule weight; Based on cosine similarity, semantic vector similarity, relevance weight coefficient, and rule weight, the corresponding relevance is calculated; The calculation process of the correlation between the first keyword group and the third keyword group is the same as the calculation process of the correlation between the first keyword group and the second keyword group.

6. The government affairs information management system based on data matching according to claim 3 is characterized in that: The Bert-LightGBM model is used to classify the keyword groups to obtain corresponding classification results, including: Get the corresponding source department and set the corresponding source weight; The semantic vector of the keyword group is extracted using the Bert model and combined with the source weight to obtain the feature vector; The feature vector is input into the LightGBM model and the classification result is output.

7. The government affairs information management system based on data matching according to claim 1 is characterized in that: The identity verification module comprises: Obtain the user's identity information or work badge information and perform authentication; if the authentication is successful, input the user's keyword group through the client, classify the user's keyword group through the keyword group extraction and classification module, and obtain the user classification result; otherwise, the authentication fails.

8. The government affairs information management system based on data matching according to claim 3 is characterized in that: The storage module comprises: An encryption / decryption unit, used to encrypt / decrypt government information and its keyword groups, classification results and matching results; A keyword storage unit, used for storing keyword groups; The political and government affairs information storage unit is used for all batches of political and government affairs information and classification results and matching results; Administrative government information storage unit, used for all batches of administrative government information and classification results and matching results; The social service government affairs information storage unit is used for all batches of social service government affairs information and classification results and matching results; Other types of government affairs information storage unit, used for all batches of other types of government affairs information and classification results and matching results.

Citation Information

Patent Citations

  • Policy portrait AI modeling system and method based on big data

    CN111813890A

  • Big data operation supervision platform based on smart government affairs

    CN119048311A