Processing method for improving collection efficiency of localized translation demand data

Through collaborative filtering recommendation strategy, the text information similarity and professionalism of translation requirements data are analyzed, and database recommendation collection is constructed, which solves the problem of low efficiency in data collection of translation requirements and achieves efficient and accurate data collection.

CN120561284AActive Publication Date: 2025-08-29EC INNOVATIONS (SHENYANG) INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511052857.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-08-29
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

In the prior art, the collection efficiency of translation required data is low. With the increase of data sources and data types, the cost and difficulty of the work during the processing process increase, resulting in inefficient data collection.

Method used

A collaborative filtering recommendation strategy based on real-time translation requirements and historical translation requirements is adopted to analyze the text information similarity and professionalism of the translation requirements data, a database recommendation collection is constructed, and a collaborative filtering recommendation is carried out to optimize the data collection process.

Benefits of technology

It improves the efficiency of data collection required for translation, ensures the accuracy and comprehensiveness of the data, reduces labor costs, enhances system scalability, and promotes data analysis and decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561284A_ABST
    Figure CN120561284A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a processing method for improving the collection efficiency of localized translation demand data, which comprises the following steps: acquiring translation demand data at each moment; obtaining a database recommendation set of the translation demand data at each moment by combining the professional degree of the translation demand data at each moment according to the text information similarity condition between the translation demand data at each moment and the translation demand data in each classification database; according to the intersection condition of the database recommendation set corresponding to the current moment and the database recommendation set corresponding to the historical moment, the database recommendation set at the current moment is updated in combination with the recommendation priority degree of the classification database in the database recommendation set at the historical moment; and carrying out collaborative filtering recommendation on the translation demand data at the current moment based on the updated database recommendation set at the current moment. According to the invention, the collection efficiency of the localized translation demand data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a processing method for improving the efficiency of collecting localized translation demand data. Background Art

[0002] Translation demand data determines the translation content and workload of localization translation projects and serves as the raw data for localization projects. Project team members need to manually classify, collect, organize, and standardize this data, integrating it into foundational data that meets localization project execution specifications. Therefore, improving the efficiency of collecting translation demand data not only improves customer satisfaction, reduces labor costs, and lowers error rates, but also speeds up task allocation, enhances system scalability, facilitates data analysis, and supports decision-making.

[0003] In the localization field, translation requirement data is often scattered across various data sources. Project team members must manually classify, collect, organize, and standardize this data, integrating it into foundational data that meets localization project execution specifications. However, as the number of data sources and types increases, the processing cost and difficulty increase, resulting in lower efficiency in collecting translation requirement data. Summary of the Invention

[0004] To address the technical problem of low efficiency in collecting translation demand data in existing methods, the present invention provides a method for improving the efficiency of collecting localized translation demand data. The method aims to propose a collaborative filtering recommendation strategy based on intelligent data classification based on real-time translation requirements and translation results. The technical solution adopted is as follows: Obtaining translation demand data at each moment; wherein the moments include the current moment and historical moments, and the translation demand data at the historical moments belong to different types of classified databases; Based on the textual similarity between the translation demand data at each moment and the translation demand data in each classification database, and combined with the professional level of the translation demand data at each moment, a database recommendation set of the translation demand data at each moment is obtained; Update the current database recommendation set based on the intersection of the current database recommendation set and the historical database recommendation set, combined with the recommendation priority of the classified database in the historical database recommendation set; Based on the currently updated database recommendation set, collaborative filtering recommendation is performed on the current translation demand data.

[0005] Preferably, the method of obtaining a database recommendation set of translation demand data at each moment based on the similarity of text information between the translation demand data at each moment and the translation demand data in each classification database, combined with the professional level of the translation demand data at each moment, specifically includes: Obtain similarity feature indicators between the translation demand data at each moment and each classification database based on text similarity and language style similarity between the translation demand data at each moment and each translation demand data in each classification database; Based on the professional distribution of translation demand data at each moment and the diversity of translation results of the translation demand data in each classification database, a translation efficiency index between the translation demand data at each moment and each classification database is obtained; Determining the matching priority corresponding to each moment and each classification database based on the similarity feature index and the translation effectiveness index; According to the size distribution of the matching priority corresponding to each moment and each classification database, the classification database is screened to obtain a database recommendation set of translation demand data at each moment.

[0006] Preferably, obtaining similarity feature indicators between the translation demand data at each moment and each classification database based on text similarity and language style similarity between the translation demand data at each moment and each translation demand data in each classification database specifically includes: The Jaccard similarity between the translation demand data at each moment and each classification database is used as the first similarity coefficient between the translation demand data at each moment and each classification database; Based on the context information of each translation requirement data, a language style vector of each translation requirement data is obtained; For any moment and any classification database, determining a second similarity coefficient between the any moment and the any classification database based on a cosine similarity between a language style vector of each translation requirement data at the moment and a language style vector of each translation requirement data in the classification database; The product of the first similarity coefficient and the second similarity coefficient is used as a similarity feature index between any moment and any classification database.

[0007] Preferably, the translation efficiency index between the translation demand data at each moment and each classification database is obtained based on the professional distribution of the translation demand data at each moment and the diversity of translation results of the translation demand data in each classification database, specifically including: The total number of professional terms contained in the translation demand data at each moment is obtained as the first feature number; the total number of translation results of different types of all translation demand data in each classification database is obtained as the second feature number; and the ratio of the first feature number to the second feature number is used as the translation efficiency indicator between the translation demand data at each moment and each classification database.

[0008] Preferably, the method of screening the classification databases according to the size distribution of the matching priorities corresponding to each classification database at each moment to obtain a database recommendation set of translation demand data at each moment specifically includes: At any moment, the classification database corresponding to each classification database at that moment whose matching priority is greater than or equal to a preset priority threshold is obtained, and a database recommendation set of the translation requirement data at the said any moment is constructed in descending order of matching priority.

[0009] Preferably, the updating of the database recommendation set at the current moment according to the intersection of the database recommendation set corresponding to the current moment and the database recommendation set corresponding to the historical moment, combined with the recommendation priority of the classified database in the database recommendation set at the historical moment, specifically includes: The difference between the database recommendation set corresponding to the previous historical moment adjacent to the current moment and the database recommendation set corresponding to the current moment is used as the reference recommendation set for the current moment; According to the correlation between the translation demand data at the historical moment and the translation demand data at the current moment, the reference priority of each classification database in the reference recommendation set at the current moment is obtained by combining the recommendation priority of each classification database in the reference recommendation set at the historical moment with the translation demand data at the corresponding historical moment; According to the reference priority, the classification databases in the reference recommendation set are screened, and combined with the database recommendation set at the current moment to obtain an updated database recommendation set.

[0010] Preferably, the step of obtaining the reference priority of each classification database in the reference recommendation set based on the correlation between the translation demand data at the historical moment and the translation demand data at the current moment, in combination with the recommendation priority of each classification database in the reference recommendation set at the historical moment to the translation demand data at the current moment, specifically includes: Take any classification database in the current reference recommendation set as the target classification database; For each historical moment corresponding to the target classification database in the database recommendation set, calculate the similarity coefficient between the translation demand data at each historical moment and the translation demand data at the current moment to obtain the text reference weight of each historical moment; The text reference weight is used to perform weighted averaging on the matching priorities corresponding to the target classification database at each historical moment to obtain the reference priority of the target classification database at the current moment.

[0011] Preferably, the filtering of the classification databases in the reference recommendation set according to the reference priority and combining the database recommendation set at the current moment to obtain an updated database recommendation set specifically includes: The classification databases corresponding to the reference priority levels in the reference recommendation set that are greater than a preset reference threshold are added to the database recommendation set at the current moment for updating, thereby obtaining an updated database recommendation set.

[0012] Preferably, the method for obtaining the different types of classification databases specifically includes: For translation demand data at all historical moments, obtain all types of translation results for each type of translation demand data; Based on the similarities of all types of translation results between each type of translation demand data and other types of translation demand data, combined with the differences in the number of types of translation results between each type of translation demand data and other types of translation demand data, a classification characteristic index between each pair of translation demand data is obtained; When the classification characteristic index is greater than or equal to a preset classification threshold, the corresponding types of translation demand data are classified into the same classification database.

[0013] Preferably, the classification characteristic index between each pair of translation demand data is obtained based on the similarities of all types of translation results between each type of translation demand data and other types of translation demand data, combined with the difference in the number of types of translation results between each type of translation demand data and other types of translation demand data, and specifically includes: Obtaining all translation results for each translation requirement data to form a global translation text for each translation requirement data; Calculating the semantic similarity between the global translation data of any two translation demand texts to obtain a similarity feature factor; determining a difference feature factor based on the difference in the number of types of translation results in the global translation texts of any two translation demand data; A normalized result of the ratio of the similarity feature factor to the difference feature factor is used as a classification feature index between any two translation requirement data.

[0014] The embodiments of the present invention have at least the following beneficial effects: The present invention first collects translation demand data at different times, including real-time translation demand at the current moment and translation demand at historical moments. Furthermore, all translation demand data at historical moments belong to different types of classified databases, that is, data classification management is achieved by constructing different types of databases, which further improves the efficiency of the coordinated recommendation process. Then, the database recommendation set is determined by analyzing the matching degree between the real-time translation demand and the translation data classification result, and the feature association between the classified database and the translation demand is fully analyzed. Furthermore, based on the correlation and priority between the translation demand at different historical moments and the real-time translation demand at the current moment, the database recommendation set is further updated, and finally collaborative filtering recommendation is performed, which improves the collection efficiency of localized translation demand data and ensures the accuracy and comprehensiveness of the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 This is a flowchart of the steps of a processing method for improving the efficiency of collecting localized translation demand data provided by the present invention; Figure 2 This is a schematic diagram of the overall concept of establishing a collaborative filtering recommendation strategy provided by the present invention; Figure 3 It is a flowchart of the steps of the method for obtaining a database recommendation set provided by the present invention; Figure 4 This is a flowchart of the steps of the method for updating the database recommendation set at the current moment provided by the present invention. DETAILED DESCRIPTION

[0017] To further illustrate the technical means and effectiveness of the present invention in achieving its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effectiveness of a method for improving the efficiency of collecting localized translation demand data. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0018] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0019] The following describes in detail a specific solution of a processing method for improving the efficiency of collecting localized translation demand data provided by the present invention with reference to the accompanying drawings.

[0020] See also Figure 1 , which shows a flowchart of a method for improving the efficiency of collecting localization translation demand data provided by an embodiment of the present invention. The method includes the following steps: Step S100 , obtaining translation demand data at each moment; wherein the moments include the current moment and historical moments, and the translation demand data at historical moments belong to different types of classification databases.

[0021] First, translation requirements data determines the translation content and workload of localization translation projects and serves as the raw data for localization projects. Translation requirements data generally exists in two forms: data and files. Data types include plain strings and structured data like JSON, while file types include Excel, JSON, and Word file formats. After collecting this translation requirements data, project team members conduct workload analysis and content processing to generate project quotations and the underlying data for subsequent project execution.

[0022] There are many different ways to obtain translation requirement data, depending on the scenario. In the localization field, this data can often be scattered across various data sources. In this embodiment, depending on the product form, it may exist in various resources such as local resource files, game development tools, and databases. In most cases, the translation requirement data for game engines is provided by game developers.

[0023] Therefore, as a specific example, in this embodiment, first, when data is provided, all translation requirement data are marked and translation results of all possible types of corresponding data are obtained to obtain the translation requirement data at a historical moment.

[0024] At the same time, because the real-time data requirements in the game engine cannot be directly obtained and converted, developers need to provide them in real time as needed during their work process. Therefore, this embodiment provides componentized processing and develops a data collection component. This component has a built-in data structure of the content database and can be installed by game developers in the development IDE tool. When there is an actual translation requirement, the developer will directly use this component and append the translation results of the translation requirement data to the component interface according to the structured requirements in the component. At this point, the real-time translation requirement data at the current moment can be obtained.

[0025] Since developers need to manually process real-time translation requirements, the workload and time required will increase as the number of data sources and data types involved in the project increases. Therefore, this solution establishes a data collaborative filtering recommendation strategy by analyzing the correlation classification results between translation requirement data and real-time translation requirements. Developers can then review the recommended results based on the translation data and append them to the component interface, thereby improving the quality and efficiency of the game engine translation requirement data collection process and establishing an overall conceptual diagram of the collaborative filtering recommendation strategy, as shown below. Figure 2 shown.

[0026] In order to make the data at each historical moment more organized and more convenient for the formulation of collaborative recommendation strategies to improve data collection efficiency, it is first necessary to construct a database of historical big data. The purpose is to group the data with relatively similar translation results in the historical big data into the same database, so as to facilitate the comparison of similarities between requirements and provide a reference for real-time translation demand data. Based on this, the translation demand data collected at each historical moment in this embodiment belongs to different types of classified databases, with each classified database representing a category. The translation demand data of all historical moments contained in the same classified database are relatively similar, and the corresponding translation results are also relatively similar.

[0027] As a first step, it should be noted that the same translation requirement data needs to be classified into the same category, that is, the terms corresponding to the translation requirement data of the same category are exactly the same.

[0028] More specifically, the translation requirement data marked at different data source locations are traversed to perform term comparisons. If the terms corresponding to the translation requirement data at certain marked locations are exactly the same, they are classified into the same category.

[0029] The second step is to construct a classification metric to measure the similarity between each two different translation demand data.

[0030] Specifically, for the translation demand data at all historical moments, all types of translation results for each type of translation demand data are obtained; based on the similarities of all types of translation results between each type of translation demand data and other types of translation demand data, combined with the differences in the number of types of translation results between each type of translation demand data and other types of translation demand data, classification feature indicators between every two types of translation demand data are obtained.

[0031] More specifically, all translation results of each translation requirement data are obtained to form a global translation text of each translation requirement data; the semantic similarity between the global translation data of any two translation requirement texts is calculated to obtain a similarity feature factor; based on the difference in the number of types of translation results in the global translation texts of any two translation requirement data, a difference feature factor is determined; and a normalized processing result of the ratio of the similarity feature factor to the difference feature factor is used as a classification feature indicator between any two translation requirement data.

[0032] The semantic similarity between texts can be calculated by the BERT algorithm. The BERT model is a well-known technology. BERT embedding can be used to measure the semantic similarity between sentences or documents, which will not be introduced in detail here.

[0033] As a specific example, the calculation method of the classification feature index between the a-th type of translation demand data and the b-th type of translation demand data can be expressed as: in, Represents the classification feature index between the a-th type of translation demand data and the b-th type of translation demand data, Represents the semantic similarity between the global translation texts corresponding to the a-th translation requirement data and the b-th translation requirement data, that is, the similarity feature factor; Indicates the number of different types of translation results contained in the global translation text corresponding to the a-th type of translation requirement data, It represents the number of different types of translation results contained in the global translation text corresponding to the a-th translation requirement data; Norm is the normalization function, It is a preset hyperparameter, and its value is a very small positive number, such as 0.01, to prevent the denominator from being 0 and affecting the calculation results.

[0034] The greater the semantic similarity between the translation results corresponding to two types of translation demand data, the smaller the quantitative difference in the translation results generated by the two types of translation demand data. This indicates that the two types of translation demand data have similar performance in terms of translation results, and the larger the value of the corresponding classification feature index. The classification feature index reflects the similarity of the translation results between any two types of translation demand data and can indicate the likelihood of classifying two types of translation demand data into the same category.

[0035] In the third step, the translation demand data of corresponding categories when the classification characteristic index is greater than or equal to a preset classification threshold are classified into the same classification database.

[0036] The larger the value of the classification feature index between each two types of translation demand data, the more similar the feature performance of the translation results of the corresponding two types of translation demand data are. They can be classified into the same category to classify the translation demand data.

[0037] In this embodiment, the classification threshold is set to 0.85, which can be customized by the implementer based on the specific implementation scenario. At this point, different types of translation demand data are categorized and placed into separate classification databases. It should be understood that the primary purpose of this step is to construct a classification database of historical big data, analyze all translation demand data at all historical moments, implement data classification, and provide a data foundation for subsequent data recommendation processes based on real-time translation demand data.

[0038] Step S200 , obtaining a database recommendation set of translation demand data at each moment based on the text information similarity between the translation demand data at each moment and the translation demand data in each classification database and the professional level of the translation demand data at each moment.

[0039] In the translation demand scenario of the game engine, different translation requirements will be generated in real time during operation. The real-time translation requirements are essentially the game text data that needs to be translated. At this time, the data has not yet been translated, so it is necessary to compare the real-time translation requirements with the translation requirements that have been classified in the historical big data, match the real-time translation requirements with the historical big data, and provide a reference for the real-time translation requirements to improve data collection and translation efficiency.

[0040] Furthermore, for any translation requirement, if the similarity between its text and the text in different classification databases in historical big data is greater, the matching effect of the corresponding more similar classification database will be better. At the same time, real-time translation needs may come from the translation of current game plots or game dialogue data. Therefore, when analyzing similar situations, language style characteristics can also be considered to ensure that the translation style between the matched classification database and the implemented translation requirement is consistent. Based on this feature, the degree of reference value of each classification database to the current translation requirement can be constructed based on the feature similarity between the current translation requirement data and each classification database.

[0041] In some embodiments, as Figure 3 As shown, the method for obtaining the database recommendation set can be implemented by steps S201 to S204.

[0042] Step S201 : obtaining similarity feature indices between the translation demand data at each moment and each classification database based on text similarity and language style similarity between the translation demand data at each moment and each translation demand data in each classification database.

[0043] Specifically, the Jaccard similarity between the translation demand data at each moment and each classification database is used as the first similarity coefficient between the translation demand data at each moment and each classification database; based on the context information of each translation demand data, the language style vector of each translation demand data is obtained; for any moment and any classification database, based on the cosine similarity between the language style vector of each translation demand data at that moment and the language style vector of each translation demand data in the classification database, the second similarity coefficient between the any moment and any classification database is determined; and the product of the first similarity coefficient and the second similarity coefficient is used as the similarity feature index between any moment and any classification database.

[0044] In some embodiments, the BERT model is used to obtain contextual information for each translation request data and construct a language style vector for the translation request data. The BERT model can be used to extract vectors from text, and this is a well-known technique and will not be further described here. It is understood that the language style vector of the translation request data represents the contextual information of the corresponding translation request data. In other embodiments, the implementer may select other appropriate methods to obtain a vector that can express the data context information based on the specific implementation scenario.

[0045] As a specific example, taking the similarity feature index between the i-th moment and its corresponding K-th classification database as an example, the calculation formula of the similarity feature index can be expressed as: in, Represents the similarity feature index between the translation demand data at the i-th moment and the corresponding K-th classification database; represents the Jaccard similarity between the translation demand data at the i-th moment and the corresponding K-th classification database, which is also the first similarity coefficient; Represents the language style vector of the translation demand data at the i-th moment, Represents the language style vector of the k-th translation requirement data in the K-th classification database, represents the cosine similarity between two language style vectors, Represents the mean function. The mean value of the cosine similarity between the language style vector of the translation demand data at the i-th moment and the language style vector of each translation demand data in the K-th classification database is the second similarity coefficient.

[0046] The first similarity coefficient reflects the similarity between the translation demand data at the i-th moment and the K-th classification database in terms of text keywords or phrases. The second similarity coefficient reflects the similarity between the translation demand data at the i-th moment and all the translation demand data in the K-th classification data in terms of language style and context information. The mean is used as the balanced similarity to measure the features. The larger the value of the first similarity coefficient and the larger the value of the second similarity coefficient, the larger the value of the corresponding similarity feature index, indicating that there is a greater similarity feature between the translation demand data and the corresponding classification database. The similarity feature index characterizes the degree of information similarity between the translation demand data and the classification database. At the same time, the similarity feature index combines the similarity of keywords or phrases in the data, as well as the similarity between single text styles, to more comprehensively and accurately measure the degree of similarity features.

[0047] It should be understood that the analysis process of the similarity feature indicators between the translation demand data at each moment and each corresponding classification database is the same. Here, any moment is used as an example for explanation. The similarity feature indicators between the current moment and each historical moment corresponding to each classification database can be obtained using the same method.

[0048] Step S202 : obtaining a translation performance index between the translation demand data at each moment and each classification database based on the professional distribution of the translation demand data at each moment and the diversity of translation results of the translation demand data in each classification database.

[0049] Considering that real-time translation needs are more likely to contain more professional terms, and in the game translation scenario, the translation results of professional terms are unique, based on this feature, if the translation results corresponding to the translation demand data in certain classified databases are more concentrated and have fewer types, when these classified databases are more similar to the real-time translation needs, the corresponding matching relationship between the two is stronger.

[0050] Specifically, the total number of professional terms contained in the translation demand data at each moment is obtained as the first feature number; the total number of translation results of different types of all translation demand data in each classification database is obtained as the second feature number; and the ratio of the first feature number to the second feature number is used as the translation efficiency indicator between the translation demand data at each moment and each classification database.

[0051] The translation efficiency index represents the corresponding translation convenience when the translation demand data at the i-th moment is matched with the K-th classification database. When the number of professional terms contained in the translation demand data at the i-th moment is large, the fewer the types of all translation results contained in all translation demands in the K-th classification database, the higher the effect and priority of matching the translation demand data at the i-th moment with the K-th classification database. After the two are matched, the efficiency and convenience of using the K-th classification database to refer to the translation results will be greatly improved.

[0052] Step S203 : determining the matching priority corresponding to each moment and each classification database based on the similarity feature index and the translation efficiency index.

[0053] As a specific example, the similarity feature index between the translation demand data at the i-th moment and the corresponding K-th classification database is calculated. The translation performance index between the translation demand data at the i-th moment and the K-th classification database The translation requirement data at the i-th moment and the matching priority of the corresponding K-th classification database are obtained by multiplying the product of and performing normalization processing. The normalization processing method is a well-known technology and will not be described in detail here.

[0054] The matching priority represents the priority of matching the translation requirement data at each moment with the classification database, and reflects the efficiency and convenience of translating the translation requirement data using the classification database with reference to the matching relationship.

[0055] Step S204 : Screening the classification databases based on the size distribution of the matching priorities corresponding to each moment and each classification database to obtain a database recommendation set for the translation demand data at each moment.

[0056] The larger the value of the matching priority, the greater the correlation between the translation demand data at each moment and the corresponding classification database. Therefore, in the scenario of real-time translation, the translation results in the classification database with a higher matching priority should be referred to.

[0057] Specifically, at any given moment, the databases corresponding to each categorized database whose matching priority at that moment is greater than or equal to a preset priority threshold are retrieved. The databases are then ranked from highest to lowest matching priority to form a recommended set of databases for translation demand data at that moment. In this embodiment, the priority threshold is set to 0.3; implementers can adjust this threshold based on their specific implementation scenario.

[0058] The recommended database set corresponding to each moment represents the set of categorized databases that should be referenced, in order, when translating the translation demand data at that moment. Thus, a historical translation demand data set has been constructed for reference before translation at each moment. It should be understood that the recommended database set is constructed by analyzing the characteristics of all historical data corresponding to each moment.

[0059] Step S300 , updating the database recommendation set at the current moment according to the intersection of the database recommendation set corresponding to the current moment and the database recommendation set corresponding to the historical moment, combined with the recommendation priority of the classified database in the database recommendation set at the historical moment.

[0060] When constructing a historical recommendation database for translation demand data at the current moment, only the correlation between the real-time translation demand and the categorized database with similar characteristics constructed at all historical moments was considered, lacking temporal contextual information. To improve the efficiency of real-time translation requests, collaborative filtering recommendations were further performed on the categorized database, targeting the textual correlation between the real-time translation demand data and adjacent historical moments.

[0061] Based on this feature, the database recommendation sets of the most recent historical moment adjacent to the current moment are compared. When some classified databases do not exist in the database recommendation set of the current moment, the text correlation between the translation requirements of the historical moment and the translation requirements of the current moment, as well as the matching relationship between the historical moment and the classified database, can be further combined to coordinate filtering and update the database recommendation set of the current moment.

[0062] In some embodiments, as Figure 4 As shown, the method for updating the database recommendation set at the current moment can be implemented by steps S301 to S303.

[0063] Step S301 : taking the difference between the database recommendation set corresponding to the previous historical moment adjacent to the current moment and the database recommendation set corresponding to the current moment as the reference recommendation set for the current moment.

[0064] The classification databases in the current moment's reference recommendation set are present in the database recommendation set of the adjacent historical moment, but not in the current moment's database recommendation set. Based on this, using the coordinated filtering concept, classification databases that are present in the database recommendation set of the closest historical moment to the current moment, but not in the current moment's database recommendation set, can be recommended to the current moment.

[0065] Step S302, based on the correlation between the translation demand data at the historical moment and the translation demand data at the current moment, combined with the recommendation priority of each classification database in the reference recommendation set at the historical moment to the translation demand data at the current moment and the corresponding historical moment, obtains the reference priority of each classification database in the reference recommendation set.

[0066] Specifically, any classification database in the current reference recommendation set is used as the target classification database. For each historical moment in the database recommendation set corresponding to the target classification database, the similarity coefficient between the translation requirement data at each historical moment and the translation requirement data at the current moment is calculated to obtain a text reference weight for each historical moment. Using this text reference weight, the matching priorities corresponding to the target classification database at each historical moment are weighted and averaged to obtain the reference priority of the target classification database at the current moment.

[0067] As a specific example, taking the classification database A in the reference recommendation set at the current moment as the target classification database, the calculation formula for the reference priority of the classification database A at the current moment can be expressed as: in, represents the reference priority of the classification database A in the reference recommendation set at the current moment, t represents the current moment, Indicates the total number of historical moments in which classification database A exists in the database recommendation set. Indicates the text reference weight between the translation demand data of the nth historical moment in the classification database A and the translation demand data at the current moment in the database recommendation set, Indicates the matching priority between the nth historical moment and the classification database A.

[0068] The greater the similarity between the translation demand data corresponding to the historical moment and the current moment, the larger the value of the corresponding text reference weight, indicating that the features of the translation demand data at the historical moment and the translation demand data at the current moment are more similar, that is, the classification database that recommends translations for the translation demand data at the historical moment can also be recommended to the translation demand data at the current moment, that is, the text reference weight is used to weight the matching priority between the historical moment and the classification database A, and the reference priority between the classification database A and the current moment can be constructed by comprehensively considering all historical moments.

[0069] It should be noted that the specific method for calculating the similarity coefficient between the translation demand data at each historical moment and the translation demand data at the current moment is to calculate the Jaccard similarity between the translation demand data at each historical moment and the translation demand data at the current moment as the similarity coefficient. The BERT algorithm can also be used to calculate the semantic similarity between two texts as the similarity coefficient. It is also possible to extract the word vector of the translation demand data and then calculate the cosine similarity between the word vector of each historical moment and the current moment as the similarity coefficient. Implementers can make choices based on the specific implementation scenario, aiming to quantify the degree of text similarity between the translation demand data at each historical moment and the translation demand data at the current moment.

[0070] Step S303 : screening the classification databases in the reference recommendation set according to the reference priority, and obtaining an updated database recommendation set by combining the database recommendation set at the current moment.

[0071] The classification database in the reference recommendation combination represents the textual association between the translation needs at historical moments and the real-time translation needs. By using the coordinated filtering recommendation results corresponding to the translation needs at historical moments and quantifying the coordinated filtering recommendation results of the current real-time translation needs, we can further screen out the classification databases with greater correlation to be recommended.

[0072] Specifically, the classification databases corresponding to the reference priority levels in the reference recommendation set that are greater than a preset reference threshold are added to the database recommendation set at the current moment for updating, thereby obtaining an updated database recommendation set.

[0073] It should be noted that in this embodiment, the reference threshold is set to 0.7, and the reference priority of each classification database in the reference recommendation set is normalized. The normalization method is well known and will not be further described here. When the normalized reference priority is greater than the reference threshold, the corresponding classification database is added to the current database recommendation set, completing the update operation for the current database recommendation set.

[0074] Step S400 : performing collaborative filtering recommendation on the translation requirement data at the current moment based on the database recommendation set updated at the current moment.

[0075] It should be understood that the currently updated database recommendation set is the coordinated filtered recommendation result for the current translation request. The translation results of the real-time translation request can be translated and aggregated into the database with reference to the results of the corresponding database, and finally displayed after manual review.

[0076] In some other embodiments, since the final translation of the translation demand data requires manual effort from project team members, the satisfaction and evaluation of the translation results represent the quality of the translation results. Furthermore, after each translation is completed, the feedback from project team members can indicate whether the recommended content meets expectations, the relevance of the recommendations, the timeliness of the recommendations, etc., thereby converting the feedback from project team members into quantifiable data. For example, a rating system can use 1-5 stars for quantification, and a satisfaction survey can be converted into a rating of 0-10.

[0077] By processing and analyzing feedback results in real time, a feedback loop is established, allowing the recommendation system to continuously receive feedback and make adjustments, forming a continuous improvement process, thereby improving the efficiency of collecting translation demand data and the accuracy of translation results.

[0078] In summary, the embodiments of the present invention address the high processing costs and low accuracy associated with the increasing number of data sources and data types associated with manual collection of translation demand data. Data is classified by analyzing the similarities in translation results for different translation demand data. Data recommendation priorities are then determined based on the degree of match between real-time translation demand and the classification results. Collaborative filtering recommendations are then performed based on the relevance and priority of translation demand at different times. Finally, the data recommendation strategy is optimized based on feedback from processing personnel. This further improves the efficiency of collecting localized translation demand data while ensuring its accuracy and comprehensiveness.

[0079] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A processing method for improving the efficiency of collecting localization translation demand data, characterized in that: The method comprises the following steps: Obtaining translation demand data at each moment; wherein the moments include the current moment and historical moments, and the translation demand data at the historical moments belong to different types of classified databases; Based on the textual similarity between the translation demand data at each moment and the translation demand data in each classification database, and combined with the professional level of the translation demand data at each moment, a database recommendation set of the translation demand data at each moment is obtained; Update the current database recommendation set based on the intersection of the current database recommendation set and the historical database recommendation set, combined with the recommendation priority of the classified database in the historical database recommendation set; Perform collaborative filtering recommendations on the current translation demand data based on the currently updated database recommendation set; The method for obtaining the database recommendation set specifically includes: Obtain similarity feature indicators between the translation demand data at each moment and each classification database based on text similarity and language style similarity between the translation demand data at each moment and each translation demand data in each classification database; Based on the professional distribution of translation demand data at each moment and the diversity of translation results of the translation demand data in each classification database, we obtain the translation efficiency indicators between the translation demand data at each moment and each classification database, including: The total number of professional terms contained in the translation demand data at each moment is obtained as a first feature number; the total number of translation results of different types of all translation demand data in each classification database is obtained as a second feature number; and the ratio of the first feature number to the second feature number is used as a translation performance indicator between the translation demand data at each moment and each classification database; Determining the matching priority corresponding to each moment and each classification database based on the similarity feature index and the translation effectiveness index; According to the size distribution of the matching priority corresponding to each moment and each classification database, the classification database is screened to obtain a database recommendation set of translation demand data at each moment.

2. A processing method for improving the efficiency of collecting localization translation demand data according to claim 1, characterized in that: The similarity feature index between the translation demand data at each moment and each classification database is obtained based on the text similarity and language style similarity between the translation demand data at each moment and each translation demand data in each classification database, specifically including: The Jaccard similarity between the translation demand data at each moment and each classification database is used as the first similarity coefficient between the translation demand data at each moment and each classification database; Based on the context information of each translation requirement data, a language style vector of each translation requirement data is obtained; For any moment and any classification database, determining a second similarity coefficient between the any moment and the any classification database based on a cosine similarity between a language style vector of each translation requirement data at the moment and a language style vector of each translation requirement data in the classification database; The product of the first similarity coefficient and the second similarity coefficient is used as a similarity feature index between any moment and any classification database.

3. The method for improving the efficiency of collecting localization translation demand data according to claim 1, characterized in that: The classification database is screened based on the size distribution of the matching priorities corresponding to each moment and each classification database to obtain a database recommendation set for the translation demand data at each moment, specifically including: At any moment, the classification database corresponding to each classification database at that moment whose matching priority is greater than or equal to a preset priority threshold is obtained, and a database recommendation set of the translation requirement data at the said any moment is constructed in descending order of matching priority.

4. The method for improving the efficiency of collecting localization translation demand data according to claim 1, characterized in that: The updating of the database recommendation set at the current moment according to the intersection of the database recommendation set corresponding to the current moment and the database recommendation set corresponding to the historical moment, combined with the recommendation priority of the classified database in the database recommendation set at the historical moment, specifically includes: The difference between the database recommendation set corresponding to the previous historical moment adjacent to the current moment and the database recommendation set corresponding to the current moment is used as the reference recommendation set for the current moment; According to the correlation between the translation demand data at the historical moment and the translation demand data at the current moment, the reference priority of each classification database in the reference recommendation set at the current moment is obtained by combining the recommendation priority of each classification database in the reference recommendation set at the historical moment with the translation demand data at the corresponding historical moment; According to the reference priority, the classification databases in the reference recommendation set are screened, and combined with the database recommendation set at the current moment to obtain an updated database recommendation set.

5. The method for improving the efficiency of collecting localization translation demand data according to claim 4, characterized in that: The reference priority of each classification database in the reference recommendation set is obtained based on the correlation between the translation demand data at the historical moment and the translation demand data at the current moment, in combination with the recommendation priority of each classification database in the reference recommendation set at the current moment and the translation demand data at the corresponding historical moment, specifically including: Take any classification database in the current reference recommendation set as the target classification database; For each historical moment corresponding to the target classification database in the database recommendation set, calculate the similarity coefficient between the translation demand data at each historical moment and the translation demand data at the current moment to obtain the text reference weight of each historical moment; The text reference weight is used to perform weighted averaging on the matching priorities corresponding to the target classification database at each historical moment to obtain the reference priority of the target classification database at the current moment.

6. The method for improving the efficiency of collecting localization translation demand data according to claim 4, characterized in that: The step of screening the classified databases in the reference recommendation set according to the reference priority and combining the database recommendation set at the current moment to obtain an updated database recommendation set specifically includes: The classification databases corresponding to the reference priority levels in the reference recommendation set that are greater than a preset reference threshold are added to the database recommendation set at the current moment for updating, thereby obtaining an updated database recommendation set.

7. The method for improving the efficiency of collecting localization translation demand data according to claim 1, characterized in that: The method for obtaining the different types of classification databases specifically includes: For translation demand data at all historical moments, obtain all types of translation results for each type of translation demand data; Based on the similarities of all types of translation results between each type of translation demand data and other types of translation demand data, combined with the differences in the number of types of translation results between each type of translation demand data and other types of translation demand data, a classification characteristic index between each pair of translation demand data is obtained; When the classification characteristic index is greater than or equal to a preset classification threshold, the corresponding types of translation demand data are classified into the same classification database.

8. The method for improving the efficiency of collecting localization translation demand data according to claim 7, characterized in that: The classification characteristic indicators between each pair of translation demand data are obtained based on the similarities of all types of translation results between each type of translation demand data and other types of translation demand data, combined with the differences in the number of types of translation results between each type of translation demand data and other types of translation demand data, specifically including: Obtaining all translation results for each translation requirement data to form a global translation text for each translation requirement data; Calculating the semantic similarity between the global translation data of any two translation demand texts to obtain a similarity feature factor; determining a difference feature factor based on the difference in the number of types of translation results in the global translation texts of any two translation demand data; A normalized result of the ratio of the similarity feature factor to the difference feature factor is used as a classification feature index between any two translation requirement data.

Citation Information

Patent Citations

  • Article recommendation method and device and computer storage medium

    CN113449200A

  • Auxiliary translation method based on computer-aided translation system

    CN114139555A

  • Medical machine translation method based on reinforcement learning

    CN117688952A

  • Term translation recommendation method and device, electronic equipment and storage medium

    CN117952128A

  • Document Translation Feasibility Analysis Systems and Methods

    US20250148210A1