File arrangement system and method capable of realizing multilingual conversion

By building a multilingual archive sorting system, using semantic depth mining and cross-language semantic association algorithms, the problem of inaccurate multilingual archive retrieval is solved, efficient archive storage and query is achieved, and user experience and retrieval efficiency is improved.

CN120278119AInactive Publication Date: 2025-07-08山东聚鑫科技服务有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510435323.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the existing archive sorting system to accurately locate semantic similar content in a multilingual environment, the search results are incomplete and inaccurate, the manual sorting efficiency is low, and the translation accuracy and conversion efficiency cannot meet the actual needs.

Method used

Semantic deep mining algorithm, cross-language semantic association algorithm and semantic correlation algorithm are adopted, combined with hierarchical convolutional neural network and cross-language semantic association algorithm, a multilingual archive database is built to accurately match semantic similar contents of archives of different languages, and to support multiple interaction methods through the user interaction interface.

Benefits of technology

It improves the accuracy and efficiency of multilingual file organization, reduces the cumbersomeness of index updates and resource waste, and improves the accuracy and efficiency of user operation experience and search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278119A_ABST
    Figure CN120278119A_ABST
Patent Text Reader

Abstract

The invention discloses an archive arrangement system and method capable of achieving multilingual conversion, and relates to the technical field of archive arrangement. The system comprises a multilingual archive database, a semantic analysis module, a retrieval matching module and a user interaction interface; according to the method, a multi-language association retrieval system based on semantic understanding is constructed, a semantic deep mining algorithm, a cross-language semantic association algorithm and a semantic association degree algorithm are applied, the effect of accurately matching semantic similar contents in files of different languages is achieved, and automatic index updating and incremental updating algorithms are designed for a multi-language file database, so that the accuracy of the multi-language file database is improved. According to the method, the effect of efficient file storage and query is achieved, complexity and resource waste of index updating during new file input or information change are avoided, the effect of extracting semantic features more accurately is achieved by customizing preprocessing processes for different language texts and combining a word vector generation algorithm based on semantic context, and the efficiency of file storage and query is improved. And the accuracy and efficiency of multi-language archive arrangement are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of file arrangement, and specifically provides a file arrangement system and method capable of multi-language conversion. Background Art

[0002] A file arrangement system is a comprehensive system that uses information technology to collect, classify, store, retrieve, and maintain various types of files. It arranges a large number of files in an orderly manner. By establishing a standardized process and a digital storage architecture, it facilitates users to quickly locate and use the required files. In the past single-language environment, it has been able to meet the basic file management needs of many institutions, assisting in daily office work, business promotion, and data access. With the deep development of globalization and the all-round expansion of international exchanges, in the operation of multinational enterprises, files such as contracts and reports from branch companies in different countries use multiple languages. In order to effectively manage and utilize these multi-language file resources, a file arrangement system capable of multi-language conversion has emerged, realizing the unified management and efficient utilization of multi-language files.

[0003] However, the existing technologies have certain defects. Traditional systems mostly rely on literal keyword matching. Due to the huge differences in vocabulary, grammar, and expression methods between different languages, it is difficult to accurately locate semantically similar content when retrieving files in other languages with a keyword in one language, resulting in incomplete and inaccurate retrieval results. There is a lack of an intelligent classification and processing mechanism for the characteristics of different languages. The file formats and structures of different languages vary greatly, and manual arrangement is inefficient and prone to errors. The translation accuracy and conversion efficiency of the existing technologies cannot meet the actual needs. For files in professional fields, due to the complexity of terms, the converted content often cannot accurately convey the original meaning, seriously affecting the utilization value of the files. Therefore, it is of great significance to develop a file arrangement system and method capable of multi-language conversion. Summary of the Invention

[0004] The purpose of the present invention is to make up for the deficiencies of the existing technologies, and provides a file arrangement system and method capable of multi-language conversion. It can use semantic deep mining algorithms, cross-language semantic association algorithms, and semantic association degree algorithms to achieve the effect of accurately matching semantically similar content in files of different languages, achieve the effect of high efficiency in file storage and query, avoid the cumbersome process and resource waste of index update when new files are entered or information is changed, achieve the effect of more accurately extracting semantic features, and improve the accuracy and efficiency of multi-language file arrangement.

[0005] To solve the above technical problems, the present invention provides the following technical solution: A file arrangement system capable of multi-language conversion, the system includes: a multi-language file database, a semantic analysis module, a retrieval and matching module, and a user interface;

[0006] A multilingual archive database for storing various multilingual archives, indexing them in multiple dimensions according to language, archive type, and creation time;

[0007] A semantic parsing module that uses a semantic deep mining algorithm to perform in-depth semantic analysis on archive texts in different languages, extracts semantic features. Given the input text as T, it processes through a hierarchical convolutional neural network to obtain a semantic feature matrix where W i and b i are network parameters, σ is the activation function. Using a cross-lingual semantic association algorithm, given the semantic feature matrices of two languages as M1 and M2 respectively, it constructs a cross-lingual semantic association relationship by calculating the association matrix ;

[0008] A retrieval matching module that receives keywords input by the user, generates keyword semantic features through the semantic parsing module, and matches them with the semantic features of archives in the database according to the semantic association degree algorithm, screening out relevant multilingual archives. The semantic association degree R = α·dot(K,A)+β·cosine(K,A), where K is the keyword semantic feature, A is the archive semantic feature, α and β are weight coefficients, dot is the dot product operation, and cosine is the cosine similarity calculation;

[0009] A user interaction interface for receiving user retrieval instructions, displaying retrieval results, and supporting users to sort and filter the results.

[0010] Furthermore, when new archives are entered or the information of existing archives is changed, the multilingual archive database automatically recalculates and updates the index according to the preset multi-dimensional indexing rules. During the index update process, an incremental update algorithm is used to only adjust the indexes of the changed parts.

[0011] Furthermore, the hierarchical convolutional neural network structure in the semantic parsing module is alternately composed of multiple convolutional layers and pooling layers. The convolutional layers perform sliding convolutional operations through convolutional kernels of different sizes to extract semantic features at different levels of the text. The pooling layers reduce the feature dimensions, and each layer is connected through a non-linear activation function. When training this neural network, an adaptive learning rate adjustment algorithm is used to adjust the learning rate according to the change of the loss function during the training process.

[0012] Furthermore, when processing texts in different languages, the semantic parsing module adopts a customized preprocessing process according to the grammar and vocabulary characteristics of each language. For Chinese texts, word segmentation is first performed, the part-of-speech of each word is marked using a part-of-speech tagging tool, and then it is converted into a word vector representation. For English texts, after stemming and stop-word filtering operations, they are converted into word vectors. In the process of converting texts into word vectors, a word vector generation algorithm based on semantic context is adopted, taking into account the semantic relationship before and after the words in the sentence to generate word vectors.

[0013] Furthermore, when the retrieval and matching module calculates the semantic relevance, the weight coefficients α and β are dynamically adjusted according to the user's historical retrieval behavior data. The system records the keywords of each user retrieval and the click information of the retrieval results. By analyzing these data, it judges the user's preference for the retrieval results based on dot product operation and cosine similarity calculation, so as to adjust the values of α and β.

[0014] Furthermore, the user interface supports text input and voice input of retrieval instructions. The system integrates a speech recognition module, which converts the user's voice into text and then passes it to the retrieval and matching module for retrieval operations. The user interface provides a visual way to display retrieval results, showing the retrieval quantity distribution of different language archives in the form of charts.

[0015] Furthermore, a method for file arrangement that can be converted between multiple languages is applicable to the above-mentioned file arrangement system that can be converted between multiple languages. This method includes the following steps:

[0016] File storage: Store various multi-language files in a multi-language file database, and establish indexes according to multi-dimensional attributes such as language, file type, and creation time;

[0017] Semantic parsing: Use specific semantic depth mining methods and cross-language semantic association construction methods to deeply analyze the semantics of file texts, extract semantic features, and realize the construction of cross-language semantic associations;

[0018] Retrieval and matching: Receive the keywords input by the user, generate the semantic features of the keywords, match them with the semantic features of the files in the database, and screen out relevant multi-language files;

[0019] Result display: Display the retrieval results to the user through the user interface, and support the user to sort and filter the results.

[0020] Further, in the step of file storage, when the amount of file data reaches the set threshold, the system automatically starts the data compression program. For text files, a compression algorithm combining dictionary coding and Huffman coding is adopted. First, a dictionary of common words in the file text is constructed, and the words in the text are replaced with dictionary indexes. Then, Huffman coding is performed on the index sequence, and the corresponding relationship between the data before and after compression is recorded.

[0021] Furthermore, in the step of semantic parsing, to improve the efficiency and accuracy of semantic parsing, parallel computing technology is adopted. The text data is segmented into multiple subtasks according to paragraphs and sentences and assigned to multiple computing cores to perform semantic depth mining and cross - language semantic association construction simultaneously. Finally, the computing results of each core are integrated to obtain complete semantic features and cross - language semantic association relationships. During the parallel computing process, a task scheduling algorithm is used to allocate computing resources.

[0022] Compared with the prior art, the file sorting system and method capable of multi - language conversion have the following beneficial effects:

[0023] First, by constructing a multi - language association retrieval system based on semantic understanding and applying semantic depth mining algorithms, cross - language semantic association algorithms, and semantic association degree algorithms, the present invention realizes the effect of accurately matching semantically similar content in files of different languages. By designing an automatic update index and incremental update algorithm for the multi - language file database, the effect of high - efficiency file storage and query is achieved, avoiding the cumbersome process and resource waste of index update when new files are entered or information is changed. By customizing the pre - processing process for different language texts and combining the word vector generation algorithm based on semantic context, the effect of more accurately extracting semantic features is realized, improving the accuracy and efficiency of multi - language file sorting.

[0024] Second, by integrating various interaction methods into the user interaction interface, such as voice - input retrieval instructions and cooperating with a multi - language speech recognition model based on deep learning, the present invention realizes the convenience and efficiency of retrieval instruction input, breaking through the limitation of traditional single - text input. At the same time, visual charts are used to display the distribution of the retrieval quantities of files in different languages, enabling users to quickly grasp the overall picture of retrieval results, greatly improving the users' understanding and screening efficiency of multi - language file retrieval results, comprehensively optimizing the operation experience of users in the multi - language file sorting system, and meeting the diverse needs of users for multi - language file management in different scenarios.

[0025] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1 It is a schematic structural diagram of an archive sorting system capable of multi-language conversion;

[0028] Figure 2 It is a flowchart of an archive sorting method capable of multi-language conversion. Specific Embodiments

[0029] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in conjunction with the drawings and preferred embodiments, details the specific embodiments, structures, features, and their effects of the present invention as follows.

[0030] Embodiment 1

[0031] Refer to Figure 1 - Figure 2 , a multinational financial group conducts business in more than 30 countries around the world. It has multiple financial subsidiaries such as banks, securities, and insurance under its umbrella. In daily operations, a large amount of multi-language archives are generated every day, such as loan contracts of customers from different countries, financial statements of various regions, and local market research reports. To achieve the unified management and efficient utilization of multi-language archives, the group introduced this archive sorting system capable of multi-language conversion.

[0032] Each subsidiary transmits the archives to the multi-language archive database through the group's exclusive secure upload channel. The system is built-in with an advanced language recognition engine that can accurately identify the language of the archives within milliseconds, and then determine the archive type according to the preset complex classification system. For example, the files are classified into loan contract categories, financial statement categories, research report categories, etc., and the creation time is accurately recorded.

[0033] When the storage capacity of the database reaches the pre-set 80% threshold, the system automatically triggers the data compression program. For text-based archives, a compression algorithm combining dictionary coding and Huffman coding is used. Taking a French loan contract as an example, the system first scans the text, constructs a dictionary containing the common words in the contract, replaces each word in the contract text with the dictionary index one by one, and then performs Huffman coding on the index sequence for compressed storage. During the compression process, the system details the corresponding relationship between the data before and after compression for subsequent decompression and restoration. At the same time, the multi-language archive database establishes indexes according to multiple dimensions such as language, archive type, and creation time.

[0034] When a new file is entered or the information of an existing file is changed, the system automatically recalculates and updates the index according to the preset multi-dimensional indexing rules. During the index update process, an incremental update algorithm is adopted to only adjust the indexes of the changed parts, ensuring the efficiency of file storage and query. For example, if a part of the data in an English financial statement is updated, the system only recalculates and updates the indexes involved in the changed data block, rather than completely reconstructing the indexes of the entire statement.

[0035] Semantic parsing: The semantic parsing module starts to work. Taking a Chinese market research report as an example, first, a professional Chinese word segmentation tool is used to accurately segment the text into individual words. Then, a part-of-speech tagging tool is used to tag the part of speech of each word, such as nouns, verbs, adjectives, etc. Then, it is converted into a word vector representation. Here, a self-created semantic depth mining algorithm is used. Let the input text be T. Through a hierarchical convolutional neural network processing, a semantic feature matrix is obtained. Among them, W i , b i are network parameters, σ is the activation function, so as to deeply analyze the text and extract semantic features.

[0036] The hierarchical convolutional neural network structure in the semantic parsing module is composed of multiple convolutional layers and pooling layers alternating. In the convolutional layer, sliding convolutional operations are performed through convolutional kernels of different sizes. Smaller convolutional kernels are used to extract word-level semantic features, such as identifying keywords like "financial market investment return rate", and larger convolutional kernels are used to extract phrase- and sentence-level semantic features to understand the overall meaning expressed by the sentence.

[0037] The pooling layer is used to reduce the feature dimension, reduce the computational amount, and at the same time retain the key semantic information. Each layer is connected through a non-linear activation function to enhance the model's ability to express complex semantics. When training this neural network, an adaptive learning rate adjustment algorithm is adopted. According to the change of the loss function during the training process, the learning rate is dynamically adjusted to accelerate the model convergence speed and improve the accuracy of semantic parsing. For texts in different languages, parallel computing technology is adopted. The text data is segmented into multiple sub-tasks according to paragraphs or sentences and assigned to multiple computing cores to perform semantic depth mining and cross-language semantic association construction simultaneously.

[0038] Through the cross-language semantic association algorithm, let the semantic feature matrices of two languages be M1 and M2 respectively. By calculating the association matrix cross-language semantic association relationships are constructed. For example, for a Chinese and an English financial market analysis report, through this algorithm, the semantic associations in the part of market trend analysis in both are found.

[0039] Retrieval Matching: Analysts of the group input retrieval keywords, such as "Risk Assessment of Emerging Market Investments", in the user interface. After the system receives the keywords, the semantic parsing module immediately processes them to generate keyword semantic features. The retrieval matching module matches the keyword semantic features with the archive semantic features in the database according to the semantic correlation algorithm (semantic correlation R = α·dot(K,A)+β·cosine(K,A), where K is the keyword semantic feature, A is the archive semantic feature, α and β are weight coefficients, dot is the dot product operation, and cosine is the cosine similarity calculation).

[0040] The weight coefficients α and β are not fixed values but are dynamically adjusted according to the historical retrieval behavior data of analysts. The system continuously records information such as the keywords retrieved by analysts each time and the click situation of retrieval results. By analyzing these data, it judges whether analysts are more inclined to retrieval results based on the dot product operation or the cosine similarity calculation, so as to adjust the values of α and β. For example, if analysts have clicked on results with a high matching degree based on the dot product operation many times in the past, then appropriately increase the value of α, so that in subsequent retrievals, more emphasis is placed on the semantic correlation calculation based on the dot product operation to screen out relevant multi-language archives.

[0041] The retrieval results are displayed to analysts through the user interface. The user interface supports multiple interaction methods. Analysts can either input retrieval instructions through traditional text or through voice input. The system integrates a voice recognition module, which uses a deep learning-based voice recognition model and is trained with a large amount of multi-language voice data to accurately recognize voice instructions in different languages. At the same time, the interface displays the retrieval quantity distribution of archives in different languages in the form of visual charts, which is convenient for analysts to intuitively understand the overall situation of retrieval results. It also supports operations such as sorting and filtering the results, such as sorting according to dimensions such as relevance, creation time, and language, or filtering out archives of specific languages and specific types.

[0042] In summary, through this embodiment, the multinational financial group has achieved a qualitative leap in multi-language archive management. In terms of archive storage, data compression has greatly saved storage space, and multi-dimensional indexing and incremental update algorithms have ensured the efficiency of archive storage and query. During the semantic parsing process, parallel computing has significantly improved the processing efficiency. Precise semantic extraction and cross-language semantic association construction have greatly improved the accuracy of archive semantic understanding. During retrieval matching, according to the algorithm of dynamically adjusting weight coefficients, more retrieval results that meet the needs of analysts are provided.

[0043] Embodiment 2

[0044] See Figure 1 - Figure 2, a large international logistics enterprise with business coverage across all continents of the world, has frequent interactions with customers, suppliers and partners around the world in its daily operations. This has enabled the enterprise to accumulate a vast amount of multi-language files, including customer shipping orders, supplier contract agreements, logistics transportation route planning documents and customs clearance documents. In order to optimize the internal management process and improve the logistics operation efficiency, the enterprise has deployed this file sorting system that can perform multi-language conversion.

[0045] Each department of the enterprise transfers the multi-language files generated in daily work through the upload interface integrated in the internal office system to the multi-language file database. The system uses efficient language recognition technology to quickly and accurately determine the language of each file. For example, for an Arabic shipping order from a customer in the Middle East, the system can identify it in a short time. Then, according to the file classification standard customized by the enterprise, it is classified into the shipping order category and the creation time is accurately recorded.

[0046] When the storage capacity of the database reaches the preset 85%, the data compression program is automatically triggered. For text-based files, such as a German supplier contract agreement, the system first constructs a dictionary of common words in the contract, replaces the words in the contract text with dictionary indexes, and then performs Huffman coding on the index sequence, thus greatly reducing the data storage space. During the compression process, the system rigorously records the corresponding relationship between the data before and after compression so that the file data can be quickly decompressed and restored when needed later.

[0047] At the same time, the multi-language file database establishes indexes in multiple dimensions such as language, file type, creation time, etc. When a new file is entered or the information of an existing file changes, according to the preset multi-dimensional index rules, the indexes are automatically recalculated and updated. During the index update process, an incremental update algorithm is adopted, and only the changed parts are adjusted for indexing, avoiding large-scale recalculation of the entire database index and effectively saving computing resources and time.

[0048] Taking a Chinese logistics transportation route planning document as an example, the semantic analysis module first uses a professional Chinese word segmentation tool to segment the text, splitting the continuous text into meaningful words one by one. Then, it uses a part-of-speech tagging tool to accurately tag the part of speech of each word, such as noun, verb, preposition, etc. After that, it uses its own self-developed deep semantic mining algorithm. Let the input text be T, and through processing by a hierarchical convolutional neural network, a semantic feature matrix is obtained where W i , b i are network parameters, σ is the activation function, and the text is deeply semantically analyzed to extract key semantic features.

[0049] The hierarchical convolutional neural network structure in the semantic parsing module is composed of multiple alternating convolutional layers and pooling layers. In the convolutional layer, sliding convolutional operations are performed using convolutional kernels of different sizes. Smaller convolutional kernels are used to extract lexical-level semantic features, such as identifying professional terms like "logistics hub" and "transportation route", while larger convolutional kernels are used to extract phrase- and sentence-level semantic features to understand the overall transportation planning intention expressed by the sentence. The pooling layer is used to reduce the feature dimension, retaining key semantic information while reducing the computational amount. The layers are connected through non-linear activation functions to enhance the model's ability to express complex semantics. When training this neural network, an adaptive learning rate adjustment algorithm is adopted. According to the changes in the loss function during the training process, the learning rate is dynamically adjusted to accelerate the model's convergence speed and improve the accuracy of semantic parsing.

[0050] For texts in different languages, parallel computing technology is adopted. The text data is segmented into multiple subtasks according to paragraphs or sentences and distributed to multiple computing cores to simultaneously perform in-depth semantic mining and cross-language semantic association construction. Through the cross-language semantic association algorithm, assuming the semantic feature matrices of two languages are M1 and M2 respectively, the association matrix is calculated to construct cross-language semantic association relationships. For example, for a file about a logistics distribution plan in English and a Chinese one, the semantic associations of key information such as delivery time and location in both are found through this algorithm.

[0051] The enterprise's logistics dispatcher enters a retrieval keyword, such as "Measures for Responding to Winter Transportation Risks in the European Region", in the user interface. After the system receives the keyword, the semantic parsing module immediately processes it to generate keyword semantic features. The retrieval matching module matches the keyword semantic features with the file semantic features in the database according to the semantic association degree algorithm (semantic association degree R = α·dot(K,A)+β·cosine(K,A), where K is the keyword semantic feature, A is the file semantic feature, α and β are weight coefficients, dot is the dot product operation, and cosine is the cosine similarity calculation).

[0052] The weight coefficients α and β are not fixed values but are dynamically adjusted according to the historical retrieval behavior data of the logistics dispatcher. The system continuously records information such as the keywords retrieved by the dispatcher each time and the click situation of the retrieval results. By analyzing this data, it is judged whether the dispatcher is more inclined to the retrieval results based on the dot product operation or the cosine similarity calculation, and then the values of α and β are adjusted. For example, if the dispatcher has clicked on the results with a high matching degree based on the cosine similarity calculation many times in the past, the value of β is appropriately increased, so that in subsequent retrievals, the semantic association degree calculation based on the cosine similarity calculation is more emphasized to screen out relevant multi-language files.

[0053] The retrieval results are presented to the logistics dispatcher through the user interaction interface, which supports multiple interaction methods. The dispatcher can either input retrieval instructions through traditional text or by voice. The system integrates a speech recognition module that uses a deep learning-based speech recognition model and is trained with a large amount of multi-language speech data to accurately recognize speech instructions in different languages.

[0054] Meanwhile, the interface displays the retrieval quantity distribution of archives in different languages in the form of visual charts, facilitating the dispatcher to intuitively understand the overall situation of the retrieval results. It also supports operations such as sorting and filtering the results. For example, sorting can be performed according to dimensions such as relevance, creation time, and language, or specific languages and types of archives can be filtered out. For instance, the dispatcher can quickly filter out all English archives related to transportation in the European region to obtain the required information more efficiently.

[0055] In summary, through this embodiment, the international logistics enterprise has achieved a comprehensive optimization of multi-language archive management. During the semantic parsing process, parallel computing significantly improves the processing efficiency. The accurate semantic extraction and cross-language semantic association construction greatly enhance the accuracy of semantic understanding of multi-language archives. During retrieval matching, an algorithm that dynamically adjusts the weight coefficients provides retrieval results that better meet the actual needs of the logistics dispatcher.

[0056] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. An archive sorting system capable of multi-language conversion, characterized in that, The system includes: a multilingual archive database, a semantic parsing module, a retrieval matching module, and a user interface; The multilingual archive database is used to store various multilingual archives, and indexes are established in multiple dimensions according to language, archive type, and creation time; The semantic analysis module uses a semantic depth mining algorithm to deeply analyze the text of archives in different languages, extract semantic features. Set the input text as T, and through processing by a hierarchical convolutional neural network, obtain a semantic feature matrix where W i , b i are network parameters, σ is an activation function. Adopt a cross-lingual semantic association algorithm. Set the semantic feature matrices of two languages as M1 and M2 respectively. By calculating the association matrix A = softmax(M T 1·M2), construct a cross-lingual semantic association relationship; The retrieval matching module receives keywords input by the user, generates keyword semantic features through the semantic parsing module, and matches them with the archive semantic features in the database according to the semantic correlation degree algorithm. The relevant multilingual archives are filtered out. The semantic correlation degree R = α·dot(K,A)+β·cosine(K,A), where K is the keyword semantic feature, A is the archive semantic feature, α and β are weight coefficients, dot is the dot product operation, and cosine is the cosine similarity calculation; The user interface is used to receive the user's retrieval instruction, display the retrieval result, and support the user to sort and filter the result.

2. The archival arrangement system capable of multi-language conversion according to claim 1, wherein When new archives are entered or the information of existing archives is changed in the multilingual archive database, the index is automatically recalculated and updated according to the preset multi-dimensional index rule. During the index update process, an incremental update algorithm is adopted to only adjust the index of the changed part.

3. A file sorting system capable of multi-language conversion according to claim 1, characterized in that, The hierarchical convolutional neural network structure in the semantic parsing module is alternately composed of multiple convolutional layers and pooling layers. The convolutional layer performs sliding convolutional operations through convolutional kernels of different sizes to extract semantic features at different levels of the text. The pooling layer reduces the feature dimension, and each layer is connected through a non-linear activation function. When training this neural network, an adaptive learning rate adjustment algorithm is adopted to adjust the learning rate according to the change of the loss function during the training process.

4. The archival arrangement system capable of multi-language conversion according to claim 1, wherein, When processing texts in different languages, the semantic parsing module adopts a customized preprocessing process according to the grammar and vocabulary characteristics of each language. For Chinese texts, word segmentation is first performed, the part-of-speech of each word is marked using a part-of-speech tagging tool, and then it is converted into a word vector representation. For English texts, after stemming and stop word filtering operations, they are converted into word vectors. During the process of converting texts into word vectors, a word vector generation algorithm based on semantic context is adopted to generate word vectors considering the front and back semantic relationships of words in sentences.

5. A file sorting system capable of multi-language conversion according to claim 1, characterized in that, When calculating the semantic correlation degree, the weight coefficients α and β of the retrieval matching module are dynamically adjusted according to the user's historical retrieval behavior data. The system records the keywords retrieved by the user each time and the click information of the retrieval results. By analyzing these data, the system judges the user's preference for the retrieval results based on the dot product operation and the cosine similarity calculation, so as to adjust the values of α and β.

6. The archival arrangement system capable of multi-language conversion according to claim 1, characterized in that The user interface supports text input and voice input retrieval instructions. The system integrates a speech recognition module, which converts the user's voice into text and then passes it to the retrieval matching module for retrieval operations. The user interface provides a visual retrieval result display method, and displays the retrieval quantity distribution of different language archives in the form of a chart.

7. A method for file sorting with multi - language conversion, applicable to a file sorting system with multi - language conversion described in claims 1 - 6, characterized in that, The method includes the following steps: Archive storage: Store various multilingual archives in the multilingual archive database, and establish indexes with multi-dimensional attributes according to language, archive type, and creation time; Semantic parsing: Using specific semantic depth mining methods and cross-language semantic association construction methods, deeply analyze the semantic content of archival texts, extract semantic features, and realize the construction of cross-language semantic associations; Retrieval and matching: Receive keywords input by users, generate semantic features of keywords, match them with the semantic features of archives in the database, and filter out relevant multi-language archives; Result display: Display the retrieval results to users through the user interface, and support users to sort and filter the results.

8. A method for file arrangement capable of multi-language conversion according to claim 7, characterized in that In the archival storage step, when the amount of archival data reaches the set threshold, the system automatically starts the data compression program. For text-based archives, a compression algorithm combining dictionary coding and Huffman coding is used. First, construct a dictionary of common words in the archival text, replace the words in the text with dictionary indexes, and then perform Huffman coding on the index sequence, recording the corresponding relationship between the data before and after compression.

9. A method for file sorting that can be converted into multiple languages according to claim 7, characterized in that, In the semantic parsing step, parallel computing technology is adopted. The text data is segmented into multiple subtasks according to paragraphs and sentences, and distributed to multiple computing cores to simultaneously perform semantic depth mining and cross-language semantic association construction. Finally, the calculation results of each core are integrated to obtain complete semantic features and cross-language semantic association relationships. During the parallel computing process, a task scheduling algorithm is used to allocate computing resources.

Citation Information

Cited By

  • Multilingual semantic analysis and decision support system for international climate negotiation scene

    CN121328552A

  • RAG mixed retrieval method and device

    CN121501944A

  • Rag mixed search method and device

    CN121501944B