Regional digital economic data management method and system
By summarizing and matching regional digital economy data through large language models and semantic coding technology, the problem of low data integration and retrieval efficiency in traditional solutions has been solved, enabling fast and accurate data management and retrieval, and promoting economic development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 杭州市萧山区区域经济促进会
- Filing Date
- 2023-11-07
- Publication Date
- 2026-04-21
Smart Images

Figure CN121901154A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent data management technology, and in particular to a regional digital economy data management method and system. Background Technology
[0002] Regional digital economy data is a crucial indicator reflecting the state of regional economic development and holds significant reference value for regional economic planning and decision-making. Regional digital economy data typically comprises a large amount of structured and unstructured data, scattered across various data sources, including government databases, enterprise data warehouses, and internet data. With the rapid development of the digital economy, a robust regional digital economy data storage and management system is needed to effectively utilize, store, manage, and analyze this data.
[0003] However, due to the diverse sources, complex formats, and massive scale of regional digital economy data, traditional data management solutions are insufficient to meet the needs of data storage, management, and retrieval. Furthermore, traditional relational databases are inefficient when processing large-scale unstructured data. Meanwhile, traditional file systems lack semantic understanding and intelligent search capabilities.
[0004] Therefore, there is a need for an optimized regional digital economy data management solution that can effectively integrate and utilize data from different data sources and provide efficient and accurate data retrieval services. Summary of the Invention
[0005] This invention provides a regional digital economy data management method and system. It uses a large language model to summarize various files in a database to obtain a general description of multiple database files. Then, semantic understanding technology is introduced at the backend to perform semantic analysis and semantic matching measurement of the general descriptions of the multiple database files and the retrieval input. This achieves fast and accurate data retrieval, providing efficient data storage and management functions, ensuring data security and reliability, and helping users quickly find the data they need for analysis and decision-making, thus promoting economic growth and innovation.
[0006] This invention also provides a regional digital economy data management method, which includes: A large language model is used to summarize each file in the database to obtain a summary description of multiple database files; Semantic encoding is performed on the general descriptions of the multiple database files respectively to obtain semantic encoded feature vectors of the general descriptions of the multiple database files. The context semantic association encoding is performed on the semantic encoding feature vectors of the general description of the multiple database files to obtain multiple context database file general description semantic encoding feature vectors. Get search input; The retrieval input is semantically encoded to obtain a retrieval input semantic encoding feature vector; The semantic encoded feature vectors of the retrieval input and the semantic encoded feature vectors of the summary descriptions of each context database file are respectively subjected to association metric analysis to obtain multiple matching degree metrics; and Based on the sorting of the multiple matching metrics, the top K files are returned.
[0007] This invention also provides a regional digital economy data management system, which includes: The file summary description generation module is used to summarize each file in the database using a large language model to obtain summary descriptions of multiple database files. The first semantic encoding module is used to perform semantic encoding on the general descriptions of the multiple database files respectively to obtain semantic encoding feature vectors of the general descriptions of the multiple database files. The context association encoding module is used to perform context semantic association encoding on the semantic encoding feature vectors of the general description of the multiple database files to obtain multiple context database file general description semantic encoding feature vectors. The search input acquisition module is used to acquire search input. The second semantic encoding module is used to perform semantic encoding on the retrieval input to obtain a semantic encoding feature vector of the retrieval input; The association measurement analysis module is used to perform association measurement analysis on the semantic encoded feature vector of the retrieval input and the semantic encoded feature vector of the summary description of each context database file to obtain multiple matching degree measurement values; and The control return module is used to return the top K files based on the sorting of the multiple matching degree metrics. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a flowchart of a regional digital economy data management method provided in an embodiment of the present invention.
[0009] Figure 2 This is a schematic diagram of the system architecture of a regional digital economy data management method provided in an embodiment of the present invention.
[0010] Figure 3This is a block diagram of a regional digital economy data management system provided in an embodiment of the present invention.
[0011] Figure 4 This is an application scenario diagram of a regional digital economy data management method provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0013] Unless otherwise stated, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application.
[0014] In the embodiments described in this application, it should be noted that, unless otherwise stated and limited, the term "connection" should be interpreted broadly. For example, it can be an electrical connection, or a connection between two internal components. It can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above term according to the specific circumstances.
[0015] It should be noted that the terms "first," "second," and "third" used in the embodiments of this application are merely used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first," "second," and "third" can be interchanged in a specific order or sequence where permitted. It should be understood that the objects distinguished by "first," "second," and "third" can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein.
[0016] Regional digital economy data refers to a collection of data reflecting the development status of the digital economy in a specific region. The digital economy refers to economic activities driven by information and communication technologies (ICT), including e-commerce, digital payments, internet finance, cloud computing, big data analytics, etc. Regional digital economy data can help understand and assess a region's development level and potential in the field of digital economy.
[0017] Regional digital economy data typically contains a large amount of structured and unstructured data. Structured data refers to data stored in tabular form, such as records and fields in a database, which can be easily stored and processed. Unstructured data refers to data without a specific format or organization, such as text, images, audio, and video, which require special technologies for processing and analysis.
[0018] These data can come from various data sources, including government databases, enterprise data warehouses, and internet data. Government departments, enterprises, and research institutions typically collect and maintain this data to support regional economic planning, policy making, and decision analysis. The importance of regional digital economy data lies in its provision of a comprehensive understanding of a region's digital economy ecosystem. This data may include, but is not limited to, the following: economic indicators, including GDP growth rate, employment rate, and industrial structure, used to assess the contribution and impact of the digital economy on the regional economy; e-commerce data, including online transaction volume, number of e-commerce platform users, and number of e-commerce enterprises, used to assess the scale and trends of e-commerce development; internet penetration rate, including the number of internet users and mobile internet penetration rate, used to measure the degree of penetration of digital technology in society; digital payment data, including the transaction volume and amount of mobile payments and electronic payments, used to assess the popularity of digital payments and changes in payment habits; innovation indicators, including the number of research institutions and patent applications, used to assess the role of the digital economy in promoting innovation capabilities; and talent data, including the quantity and structure of talent in digital economy-related fields, used to assess the region's talent pool and development potential.
[0019] With the rapid development of the digital economy, the storage, management, and analysis of regional digital economy data have become a significant challenge and opportunity. To effectively utilize this data, a robust regional digital economy data storage and management system is needed to meet the demands of data storage, management, and analysis. Due to the massive scale of regional digital economy data, the system requires large-scale data storage and processing capabilities. Utilizing distributed storage and computing technologies, such as distributed file systems and distributed databases, can achieve scalable and high-performance data storage and processing.
[0020] Regional digital economy data originates from diverse data sources, including government databases, enterprise data warehouses, and internet data. The system needs to support access to and integration of these diverse data sources to consolidate data from different sources into a unified data storage facility, enabling comprehensive data analysis. Regional digital economy data comes in various formats, including structured and unstructured data. The system needs to support flexible data models capable of storing and managing different types of data. Simultaneously, the system should provide flexible query and analysis functions so users can retrieve and analyze data according to their needs. The quality of regional digital economy data is crucial for subsequent analysis and decision-making. The system should provide data quality management functions, including data cleaning, data verification, and data completion, to ensure data accuracy and integrity. Regional digital economy data involves a large amount of sensitive information. The system needs advanced data security and privacy protection mechanisms, including data encryption, access control, and identity authentication, to ensure data security and compliance. The system should possess intelligent search and analysis functions, enabling rapid data retrieval and analysis based on user needs. This includes the application of technologies such as keyword-based search, data visualization, data mining, and machine learning to help users discover patterns and correlations within the data. With the rapid development of the digital economy, real-time data processing has become increasingly important. Systems need to support real-time data stream processing and real-time analysis in order to acquire and process the latest data in a timely manner and support real-time decision-making and applications.
[0021] However, due to the diverse sources, complex formats, and massive scale of regional digital economy data, traditional data management solutions are insufficient to meet the needs of data storage, management, and retrieval. Furthermore, traditional relational databases are inefficient when processing large-scale unstructured data. Meanwhile, traditional file systems lack semantic understanding and intelligent search capabilities.
[0022] Therefore, this application provides an optimized regional digital economy data management solution.
[0023] In one embodiment of the present invention, Figure 1 This is a flowchart of a regional digital economy data management method provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the system architecture of a regional digital economy data management method provided in an embodiment of the present invention. Figure 1 and Figure 2As shown, the regional digital economy data management method according to an embodiment of the present invention includes: 110, summarizing each file in the database using a large language model to obtain multiple database file summary descriptions; 120, semantically encoding the multiple database file summary descriptions to obtain multiple database file summary description semantic encoding feature vectors; 130, performing contextual semantic association encoding on the multiple database file summary description semantic encoding feature vectors to obtain multiple contextual database file summary description semantic encoding feature vectors; 140, obtaining retrieval input; 150, semantically encoding the retrieval input to obtain retrieval input semantic encoding feature vectors; 160, performing association metric analysis on the retrieval input semantic encoding feature vectors and the various contextual database file summary description semantic encoding feature vectors to obtain multiple matching degree metrics; and 170, returning the top K files based on the sorting of the multiple matching degree metrics.
[0024] In step 110, a large language model is used to summarize each file in the database to obtain multiple database file summary descriptions. When using a large language model (such as GPT-3) to summarize the database files, it is necessary to ensure that the model has sufficient understanding of the database files. Furthermore, it is essential to ensure that the structure and content of the database files can be correctly understood and processed by the model. By using a large language model, summary descriptions of database files can be automatically generated, reducing the workload of manually writing descriptions. This improves work efficiency and allows for a quick understanding of the content and characteristics of the database files.
[0025] In step 120, semantic encoding is performed on the summaries of the multiple database files to obtain semantically encoded feature vectors for each summary. Semantic encoding converts text into a vector representation that represents its semantic meaning. When semantically encoding the summaries of the database files, an appropriate encoding method is selected, such as using a pre-trained word vector model (e.g., Word2Vec, GloVe) or a deep learning model (e.g., BERT). By semantically encoding the summaries of the database files, text can be converted into a computer-processable vector representation, facilitating subsequent semantic association and matching analysis.
[0026] In step 130, contextual semantic association encoding is performed on the semantically encoded feature vectors of the multiple database file summary descriptions to obtain multiple contextual database file summary description semantically encoded feature vectors. Contextual semantic association encoding refers to combining multiple semantically encoded feature vectors to capture the contextual relationships and semantic associations between them. Techniques such as recurrent neural networks (RNNs), convolutional neural networks (CNNs), or attention mechanisms can be used to implement contextual semantic association encoding. Through contextual semantic association encoding, the semantic associations and contextual information between the summary descriptions of multiple database files can be captured, thereby better understanding the relationships between database files and providing more accurate feature representations for subsequent retrieval and matching.
[0027] In step 140, the search input is obtained. The search input is the query information provided by the user for searching database files. Ensuring the accuracy and completeness of the user input is crucial for subsequent semantic encoding and matching analysis. Obtaining accurate and complete search input ensures the accuracy and effectiveness of subsequent search and matching results.
[0028] In step 150, the retrieval input is semantically encoded to obtain a semantically encoded feature vector. The purpose of semantically encoding the retrieval input is to convert the user's query information into a vector representation that can be processed by a computer. Similarly, pre-trained word vector models or deep learning models can be used to implement semantic encoding. By semantically encoding the retrieval input, the user's query information can be matched with the semantics of the database file, improving the accuracy and effectiveness of the retrieval.
[0029] In step 160, association metric analysis is performed on the semantic encoded feature vector of the retrieval input and the semantic encoded feature vectors of the summary descriptions of each context database file to obtain multiple matching degree metrics. Association metric analysis assesses the degree of matching between the semantic encoded feature vector of the retrieval input and the semantic encoded feature vectors of the database files by calculating the similarity or distance between them. Metrics such as cosine similarity and Euclidean distance can be used for this analysis. Through association metric analysis, the matching degree metrics between the retrieval input and the database files can be calculated, thereby determining which database files are most relevant to the retrieval input and providing more accurate search results.
[0030] In step 170, based on the sorting of the multiple matching metrics, the top K files are returned. The database files are sorted according to the matching metrics, and the top K files most relevant to the search input are returned. The sorting can be in ascending or descending order based on the matching metrics. By sorting and returning the files most relevant to the search input, the most relevant and valuable database files can be provided to the user, improving the quality of search results and user satisfaction.
[0031] In recent years, the development of large language models has provided new possibilities for regional digital economy data management. Large language models can summarize various files in a database, generating concise descriptions of their content. This helps users quickly understand the content and characteristics of files, improving data searchability and discoverability.
[0032] Based on this, the technical concept of this application is to obtain a summary description of multiple database files by summarizing each file in the database using a large language model, and then introduce semantic understanding technology in the backend to perform semantic analysis and semantic matching measurement of the summary description of the multiple database files and the retrieval input, thereby achieving fast and accurate data retrieval, providing efficient data storage and management functions, ensuring data security and reliability, and helping users quickly find the data they need for analysis and decision-making, thus promoting economic growth and innovation.
[0033] Specifically, in the technical solution of this application, firstly, a large language model is used to summarize each file in the database to obtain a summary description of multiple database files. It should be understood that the large language model can summarize the content of files, generating concise and accurate descriptions. These descriptions can help users quickly understand the content and characteristics of files without having to read the entire file in detail. At the same time, the summarization method provides a summary of key information, enabling users to find files of interest more quickly.
[0034] Then, in order to perform semantic understanding on the summary descriptions of each database file separately, the technical solution of this application further performs semantic encoding on the summary descriptions of the multiple database files to obtain multiple semantic encoding feature vectors of database file summary descriptions. In this way, the summary descriptions of each file can be transformed into vector representations in semantic space, thereby capturing the semantic feature information of the content of each file. Through semantic encoding, the meaning of the files and the semantic relationships within them can be better understood, thereby improving the semantic matching and retrieval performance of the data.
[0035] In one specific embodiment of this application, semantic encoding is performed on the general descriptions of the plurality of database files to obtain semantic encoding feature vectors of the general descriptions of the plurality of database files, including: passing the general descriptions of the plurality of database files through a semantic encoder of general descriptions of database files to obtain semantic encoding feature vectors of the general descriptions of the plurality of database files.
[0036] Next, considering the correlation and dependency between the semantic features of the various files in the database, it is necessary to further process the semantic feature vectors summarizing the multiple database files through a converter-based context encoder to obtain multiple context database file summary description semantic feature vectors. This not only considers the independent semantic information of each file but also the associations and dependencies between them, thus providing a more comprehensive and accurate semantic representation and better capturing the semantic relationships between files. It should be understood that the context database file summary description semantic feature vectors generated by the context encoder can be used for context-aware search. During the search process, not only the user's search input and the matching degree of a single file are considered, but also the contextual relationship between the file and other relevant files. This can provide more accurate and relevant search results, meeting the specific needs of users.
[0037] In one specific embodiment of this application, performing context semantic association encoding on the plurality of database file summary description semantic encoding feature vectors to obtain a plurality of context database file summary description semantic encoding feature vectors includes: passing the plurality of database file summary description semantic encoding feature vectors through a converter-based context encoder to obtain the plurality of context database file summary description semantic encoding feature vectors.
[0038] In the process of user input retrieval, firstly, the retrieval input is obtained and semantically encoded to capture the semantic understanding feature information of the retrieval input, thereby obtaining the semantic encoding feature vector of the retrieval input.
[0039] In one specific embodiment of this application, semantically encoding the retrieval input to obtain a retrieval input semantic encoded feature vector includes: passing the retrieval input through a retrieval input semantic encoder to obtain the retrieval input semantic encoded feature vector.
[0040] Furthermore, to match the user's search input with various files in the database, the semantic features of these two elements need to be correlated and encoded. Specifically, the semantic encoding feature vector of the search input and the semantic encoding feature vectors summarizing the context database files are correlated and encoded separately to obtain multiple semantic matching expression feature matrices. Through correlation encoding, the semantic information of the search input and each database file can be associated, thereby quantifying their similarity or relevance. Each semantic matching expression feature matrix represents the degree of semantic matching between the search input and each database file. This provides multiple different semantic matching expressions, helping users to more comprehensively understand the degree of correlation between the search results and the search input, facilitating the subsequent generation of a semantic matching degree metric related to the search input for each database file.
[0041] In one specific embodiment of this application, association metric analysis is performed on the retrieval input semantic encoding feature vector and the semantic encoding feature vector of the summary description of each context database file to obtain multiple matching degree metrics, including: performing association encoding on the retrieval input semantic encoding feature vector and the semantic encoding feature vector of the summary description of each context database file to obtain multiple semantic matching expression feature matrices; and passing the multiple semantic matching expression feature matrices through a metric network based on a deep neural network model to obtain the multiple matching degree metrics.
[0042] Subsequently, the multiple semantic matching expression feature matrices are processed through a metric network based on a deep neural network model to obtain multiple matching degree metrics. It should be understood that the metric network can learn the functional relationship mapping each semantic matching expression feature matrix to the matching degree metric value. Through the learning capability of the deep neural network, the metric network can capture the complex semantic relationships and patterns in each semantic matching expression feature matrix. In this way, multiple semantic matching expression feature matrices can be transformed into corresponding matching degree metric values, quantifying the degree of matching between files and search inputs, improving the accuracy and expressive power of the matching degree metric, thereby helping users find files relevant to their needs more quickly and providing personalized recommendation services. Furthermore, based on the ranking of the multiple matching degree metric values, the top K files are returned.
[0043] In one embodiment of this application, the regional digital economy data management method further includes a training step: training the database file summary description semantic encoder, the converter-based context encoder, the retrieval input semantic encoder, and the metric network based on a deep neural network model. The training step includes: summarizing each training file in the database using a large language model to obtain multiple training database file summary descriptions; semantically encoding the multiple training database file summary descriptions to obtain multiple training database file summary description semantic encoding feature vectors; performing context semantic association encoding on the multiple training database file summary description semantic encoding feature vectors to obtain multiple training context database file summary description semantic encoding feature vectors; obtaining a training retrieval input; semantically encoding the training retrieval input to obtain a training retrieval input semantic encoding feature vector; and separately encoding the training retrieval input semantic encoding feature vectors and the respective training context database file summary descriptions. The semantic encoding feature vectors are correlated and encoded to obtain multiple training semantic matching expression feature matrices; the multiple training semantic matching expression feature matrices are arranged along the channel dimension to form a training semantic matching expression feature map; each feature value of the training semantic matching expression feature map is optimized to obtain an optimized training semantic matching expression feature map; the optimized training semantic matching expression feature map is passed through a metric network based on a deep neural network model to obtain multiple training matching degree metrics; and the database file summary description semantic encoder, the converter-based context encoder, the retrieval input semantic encoder, and the metric network based on the multiple training matching degree metrics are trained.
[0044] Specifically, in the technical solution of this application, the training retrieval input semantic encoding feature vector and the semantic encoding feature vector of each training context database file summary description respectively express the encoded text semantic features of the training retrieval input and the context association semantic encoding features of the multiple training database files summary description based on the source text semantic context. Thus, after associating and encoding the training retrieval input semantic encoding feature vector and the semantic encoding feature vector of each training context database file summary description, the resulting multiple training semantic matching expression feature matrices express the cross-semantic space association between text semantic features.
[0045] However, considering that although the text semantic features of the semantically encoded feature vectors described in the various training context database files have undergone context association based on the source text semantic context, the differences in the source text semantics still lead to differences in the explicit feature distribution of the encoded text semantic features. This results in different cross-semantic space association distribution properties among the multiple training semantic matching expression feature matrices, leading to sparsity in the cross-semantic space association feature distribution of the overall feature distribution of the multiple training semantic matching expression feature matrices. Consequently, when the multiple training semantic matching expression feature matrices are regressed and mapped to the matching degree metric value through a metric network based on a deep neural network model, the convergence of the probability density distribution of the regression probability of each feature value of the multiple training semantic matching expression feature matrices is poor, affecting the accuracy of the obtained matching degree metric value.
[0046] Therefore, preferably, the plurality of training semantic matching expression feature matrices are treated as a whole, i.e., the channels are arranged into a training semantic matching expression feature map, and the feature values of the training semantic matching expression feature map are optimized. Specifically, the feature values of the training semantic matching expression feature map, which is treated as a whole, are optimized using the following optimization formula; wherein, the optimization formula is: in, and The training semantic matching representation feature map The and the 1 eigenvalue, and The training semantic matching representation feature map The global feature mean, It is the first step in optimizing the training of semantic matching representation feature maps. 1 eigenvalue, This indicates the calculation of the natural exponential function value raised to the power of the numerical value.
[0047] Specifically, for the trained semantic matching representation feature map The sparse distribution in the high-dimensional feature space leads to local probability density mismatches in the probability density distribution within the probability space. This is addressed by using a regularized global self-consistent class encoding to mimic the trained semantic matching representation feature map. The global self-consistency relationship of the encoding behavior of high-dimensional features in the probability space is used to adjust the error landscape of the feature manifold in the high-dimensional open space domain, thereby realizing the training semantic matching representation feature map. The high-dimensional features are used for self-consistent matching class encoding of explicit probability space embedding, thereby improving the trained semantic matching representation feature map. The convergence of the probability density distribution of the regression probability is improved to enhance the accuracy of the obtained multiple matching degree measures. This enables fast and accurate data retrieval based on semantic matching between general descriptions of multiple database files and user search input, providing efficient data storage and management functions to ensure data security and reliability. Simultaneously, it helps users quickly find the data they need for analysis and decision-making, promoting economic growth and innovation.
[0048] In summary, the regional digital economy data management method based on the embodiments of the present invention has been clarified. It can achieve fast and accurate data retrieval, thereby providing efficient data storage and management functions, ensuring data security and reliability, and helping users quickly find the data they need for analysis and decision-making, thus promoting economic growth and innovation.
[0049] Figure 3 This is a block diagram of a regional digital economy data management system provided in an embodiment of the present invention. Figure 3 As shown, the regional digital economy data management system 200 includes: a file summary description generation module 210, used to summarize each file in the database using a large language model to obtain multiple database file summary descriptions; a first semantic encoding module 220, used to perform semantic encoding on the multiple database file summary descriptions to obtain multiple database file summary description semantic encoding feature vectors; a context association encoding module 230, used to perform context semantic association encoding on the multiple database file summary description semantic encoding feature vectors to obtain multiple context database file summary description semantic encoding feature vectors; a retrieval input acquisition module 240, used to acquire retrieval input; a second semantic encoding module 250, used to perform semantic encoding on the retrieval input to obtain retrieval input semantic encoding feature vectors; an association measurement analysis module 260, used to perform association measurement analysis on the retrieval input semantic encoding feature vectors and the semantic encoding feature vectors of each context database file summary description to obtain multiple matching degree measurement values; and a control return module 270, used to return the top K files based on the sorting of the multiple matching degree measurement values.
[0050] In the regional digital economy data management system, the first semantic encoding module is used to: obtain the semantic encoding feature vector of the multiple database file summary descriptions by passing the database file summary description semantic encoder.
[0051] In the regional digital economy data management system, the context association encoding module is used to: obtain the semantic encoding feature vector of the summary description of the multiple database files by passing the context encoder based on a converter.
[0052] Those skilled in the art will understand that the specific operations of each step in the aforementioned regional digital economy data management system have been referenced above. Figures 1 to 2 The description of the regional digital economy data management method is detailed here, and therefore, its repeated description will be omitted.
[0053] As described above, the regional digital economy data management system 200 according to embodiments of the present invention can be implemented in various terminal devices, such as servers for regional digital economy data management. In one example, the regional digital economy data management system 200 according to embodiments of the present invention can be integrated into a terminal device as a software module and / or a hardware module. For example, the regional digital economy data management system 200 can be a software module in the operating system of the terminal device, or it can be an application developed for the terminal device; of course, the regional digital economy data management system 200 can also be one of many hardware modules of the terminal device.
[0054] Alternatively, in another example, the regional digital economy data management system 200 and the terminal device can also be separate devices, and the regional digital economy data management system 200 can connect to the terminal device via wired and / or wireless networks and transmit interactive information in accordance with an agreed data format.
[0055] Figure 4 This is an application scenario diagram of a regional digital economy data management method provided in an embodiment of the present invention. For example... Figure 4 As shown, in this application scenario, firstly, a large language model is used to summarize each file in the database to obtain a summary description of multiple database files (e.g., ...). Figure 4 As shown in C1), and, obtaining the retrieval input (e.g., such as Figure 4 (as shown in C2); then, the acquired multiple database files are summarized and the retrieval input is fed into a server deployed with regional digital economy data management algorithms (e.g., such as...). Figure 4 In the S shown, the server is able to process the general description of the plurality of database files and the search input based on the regional digital economy data management algorithm, and return the top K files based on the sorting of the plurality of matching degree metrics.
[0056] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for managing regional digital economy data, characterized in that, include: A large language model is used to summarize each file in the database to obtain a summary description of multiple database files; Semantic encoding is performed on the general descriptions of the multiple database files respectively to obtain semantic encoded feature vectors of the general descriptions of the multiple database files. The context semantic association encoding is performed on the semantic encoding feature vectors of the general description of the multiple database files to obtain multiple context database file general description semantic encoding feature vectors. Get search input; The retrieval input is semantically encoded to obtain a retrieval input semantic encoding feature vector; The semantic encoded feature vectors of the retrieval input and the semantic encoded feature vectors of the summary descriptions of each context database file are respectively subjected to association metric analysis to obtain multiple matching degree metrics; and Based on the sorting of the multiple matching metrics, the top K files are returned.
2. The regional digital economy data management method according to claim 1, characterized in that, Semantic encoding is performed on the general descriptions of the multiple database files to obtain semantic encoding feature vectors for the general descriptions of the multiple database files, including: passing the general descriptions of the multiple database files through a database file general description semantic encoder to obtain semantic encoding feature vectors for the general descriptions of the multiple database files.
3. The regional digital economy data management method according to claim 2, characterized in that, Performing context semantic association encoding on the multiple database file summary description semantic encoding feature vectors to obtain multiple context database file summary description semantic encoding feature vectors includes: passing the multiple database file summary description semantic encoding feature vectors through a converter-based context encoder to obtain the multiple context database file summary description semantic encoding feature vectors.
4. The regional digital economy data management method according to claim 3, characterized in that, Semantically encoding the retrieval input to obtain a retrieval input semantic encoding feature vector includes: passing the retrieval input through a retrieval input semantic encoder to obtain the retrieval input semantic encoding feature vector.
5. The regional digital economy data management method according to claim 4, characterized in that, The semantic encoding feature vector of the retrieval input and the semantic encoding feature vector of each context database file are subjected to association metric analysis to obtain multiple matching degree metrics, including: The semantic encoding feature vector of the retrieval input and the semantic encoding feature vector of the summary description of each context database file are respectively correlated and encoded to obtain multiple semantic matching expression feature matrices; and The multiple semantic matching expression feature matrices are passed through a metric network based on a deep neural network model to obtain the multiple matching degree metric values.
6. The regional digital economy data management method according to claim 5, characterized in that, It also includes a training step: training the database file general description semantic encoder, the converter-based context encoder, the retrieval input semantic encoder, and the metric network based on the deep neural network model.
7. The regional digital economy data management method according to claim 6, characterized in that, The training steps include: A large language model is used to summarize each training file in the database to obtain a summary description of multiple training database files. Semantic encoding is performed on the general descriptions of the multiple training database files respectively to obtain semantic encoded feature vectors of the general descriptions of the multiple training database files. The semantic encoding feature vectors of the general description of the multiple training database files are subjected to context semantic association encoding to obtain multiple semantic encoding feature vectors of the general description of the training context database files. Obtain the training retrieval input; The training retrieval input is semantically encoded to obtain the training retrieval input semantically encoded feature vector; The semantic encoded feature vectors of the training retrieval input and the semantic encoded feature vectors of the summary descriptions of each training context database file are respectively correlated and encoded to obtain multiple training semantic matching expression feature matrices; Arrange the multiple training semantic matching expression feature matrices along the channel dimension to form a training semantic matching expression feature map; The feature values of the trained semantic matching representation feature map are optimized to obtain an optimized trained semantic matching representation feature map; The optimized training semantic matching representation feature map is passed through a metric network based on a deep neural network model to obtain multiple training matching degree metrics; and The database file general description semantic encoder, the converter-based context encoder, the retrieval input semantic encoder, and the metric network based on the multiple training matching degree metrics are trained.
8. A regional digital economy data management system, characterized in that, include: The file summary description generation module is used to summarize each file in the database using a large language model to obtain summary descriptions of multiple database files. The first semantic encoding module is used to perform semantic encoding on the general descriptions of the multiple database files respectively to obtain semantic encoding feature vectors of the general descriptions of the multiple database files. The context association encoding module is used to perform context semantic association encoding on the semantic encoding feature vectors of the general description of the multiple database files to obtain multiple context database file general description semantic encoding feature vectors. The search input acquisition module is used to acquire search input. The second semantic encoding module is used to perform semantic encoding on the retrieval input to obtain a semantic encoding feature vector of the retrieval input; The association measurement analysis module is used to perform association measurement analysis on the semantic encoded feature vector of the retrieval input and the semantic encoded feature vector of the summary description of each context database file to obtain multiple matching degree measurement values. as well as The control return module is used to return the top K files based on the sorting of the multiple matching degree metrics.
9. The regional digital economy data management system according to claim 8, characterized in that, The first semantic encoding module is used to: obtain the semantic encoding feature vector of the multiple database file summary descriptions by passing the database file summary description semantic encoder.
10. The regional digital economy data management system according to claim 9, characterized in that, The context association encoding module is used to: pass the semantic encoding feature vectors of the summaries of the multiple database files through a converter-based context encoder to obtain the semantic encoding feature vectors of the summaries of the multiple context database files.