Method and system for quickly checking trados word number report
By analyzing and generating user interface tables and applying exception checking rules, multiple trados word count report checking and comparison are solved, and fast and simple inspection operations are achieved, suitable for multilingual projects.
Patent Information
- Application Number
- CN202510045870.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to efficiently check and compare multiple trados word count reports, especially in multilingual projects, resulting in high workload and inefficiency.
By extracting and parsing trados word count reports in XML format, generating user interface tables, and using exception check rules to mark the tables, achieving quick checking and comparison.
It greatly simplifies the translation package inspection operation, improves efficiency, and reduces the possibility of errors. It is suitable for non-professional localization engineers and has a wide range of application prospects.
Smart Images

Figure CN120106050A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of localization engineering processing, and in particular relates to a method and a system for quickly checking a TRADOS word count report. Background Art
[0002] The content of the present invention is to develop a word count report generated by the translation assistance software Trados, specifically a word count report in XML format, which is a text format that covers the word count, language, and other project information and complies with the XML general parsing standard. The above-mentioned information can be easily imported into other systems using this XML format. This XML-formatted Trados word count report is generally prepared by a technically strong and experienced localization engineer from the client's original file into a Trados project package before the translator officially uses the Trados software to carry out translation work. The original intention of the present invention is also to check whether there is an abnormality in the XML-formatted word count report generated after preparing the Trados project package, and then to remind the localization engineer that there may be a problem with the preparation of the Trados project package based on this abnormality.
[0003] Word count reports are often available in two formats, XML and Excel. The purpose of the XML format has been introduced above, which is to be compatible with other systems to read information. The Excel format is convenient for users to directly open the report for preview. In order to ensure the correctness of the translation package, localization engineers often manually open the Excel format word count report to check multiple things after they are ready, such as whether the number of translation memories, the total number of words in the project, the number of new words, and other information are consistent with expectations. In multilingual projects, it is often necessary to open the word count report for each language. For example, if a source file of this project is translated into 10 target languages, it means that the localization engineer needs to open these 10 word count reports for viewing. The workload is too large, and it is not convenient to compare the word count reports of each language with the naked eye. For such multilingual projects, comparing the reports is still very helpful for discovering anomalies. For example, if the source files are the same, then the total number of words in each language should be the same. If there is a difference, it is likely to be an anomaly. There is a problem with the file preparation or setting of this language, and the localization engineer needs to confirm it again.
[0004] From the description in the previous paragraph, we can understand two pain points. One is that multiple word count reports cannot be efficiently viewed at the same time. The other is that the information comparison between multi-language word count reports is very inefficient and difficult to guarantee accuracy when done manually. Therefore, in order to solve the above technical problems, there is an urgent need for a method and system to quickly read the word count report information and display it to the user in a convenient form, and then make automatic judgments based on some relatively fixed abnormality inspection rules, and then feed back the results to the user, that is, a method and system for quickly checking the trados word count report. Summary of the invention
[0005] To solve the above problems, the present invention provides a method and system for quickly checking a trados word count report. A method for quickly checking a trados word count report comprises:
[0006] Step S1, extracting and parsing the word count report in XML format to obtain word count report information in XML format;
[0007] Step S2: Based on the word count report information in XML format, generate a user interface table according to user needs;
[0008] Step S3, applying the exception check rule, checking and marking the user interface form, adding the marked result to the user interface form, and completing the check of the TRADOS word count report.
[0009] Optionally, the process of extracting the word count report in XML format specifically includes:
[0010] Filter out files that are not word count reports in XML format and obtain the user's input path;
[0011] Traversing each XML file under the input path of the user, and judging whether each XML file belongs to a word count report file according to the basic format structure characteristics of the Trados XML word count report;
[0012] If it is a word count report file, in the user-defined information item acquisition list, each information item is taken as a unit to traverse each word count report file in XML format input by the user to find the node where the current information item is located.
[0013] Optionally, the user-defined information items specifically include:
[0014] The word count report contains the file name, total number of segments, total word count, number of files included, number of translation memory files used when analyzing the word count, name of the memory, number of locked segments and words, number of segments and words of repeated segments, and number of new segments and words.
[0015] Optionally, the method for finding the node where the current information item is located specifically includes:
[0016] Traverse the nearest nodes of the current information item node;
[0017] Compare the name of the current information item node with the name of the nearest node. If the names are consistent, it is considered to be found.
[0018] If no consistent node is found at the end of the traversal, the default value is returned.
[0019] The present invention also discloses a system for quickly checking a TRADOS word count report, the system comprising:
[0020] An initial data extraction module is used to extract and parse the word count report in XML format to obtain the word count report information in XML format;
[0021] A table mapping module, used to generate a user interface table according to user needs based on the word count report information in XML format;
[0022] The word count checking module is used to apply the exception checking rules, check and mark the user interface form, add the marking results to the user interface form, and complete the check of the TRADOS word count report.
[0023] Optionally, in the initial data extraction module, the process of extracting the word count report in XML format specifically includes:
[0024] Filter out files that are not word count reports in XML format and obtain the user's input path;
[0025] Traversing each XML file under the input path of the user, and judging whether each XML file belongs to a word count report file according to the basic format structure characteristics of the Trados XML word count report;
[0026] If it is a word count report file, in the user-defined information item acquisition list, each information item is taken as a unit to traverse each word count report file in XML format input by the user to find the node where the current information item is located.
[0027] Optionally, the user-defined information items specifically include:
[0028] The word count report contains the file name, total number of segments, total word count, number of files included, number of translation memory files used when analyzing the word count, name of the memory, number of locked segments and words, number of segments and words of repeated segments, and number of new segments and words.
[0029] Optionally, the method for finding the node where the current information item is located specifically includes:
[0030] Traverse the nearest nodes of the current information item node;
[0031] Compare the name of the current information item node with the name of the nearest node. If the names are consistent, it is considered to be found.
[0032] If no consistent node is found at the end of the traversal, the default value is returned.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] (1) The present invention can make the operation of checking the prepared translation package by the user simple and efficient. Even non-professional localization engineers can easily and independently perform the checking operation, which greatly reduces the possibility of known problems occurring when preparing the translation package;
[0035] (2) The method of the present invention has abundant room for expansion. For example, it is useful to extract which information items from the TRADOS word count report in XML format and how to set the abnormality checking rules to be efficient and more targeted in discovering problems. With the support of these two features, the method and system of the present invention for quickly checking the TRADOS word count report will undoubtedly become a potential stock with broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0037] Figure 1 A method step diagram of a method for quickly checking a TRADOS word count report according to an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, such as the adjustment of the reading and abnormality checking rules of the TRADOS word count report information items, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0039] Embodiment 1
[0040] like Figure 1 As shown, this embodiment provides a method for quickly checking the word count report of Trados, where the word count report specifically refers to the XML format exported by Trados, including:
[0041] Step S1, extract and parse the word count report in XML format to obtain word count report information in XML format.
[0042] Parse the word count report in XML format to obtain information; parse each XML file in the user input path, and use the XML parsing function of the programming language (such as XMLDocument of vb.net) to obtain the root node of the XML file. If the root node is not "<task name="analyse"> ", it is considered that this XML file is not a word count report file in the XML format of Trados, and it is ignored and not processed further, because these files are irrelevant to this method.
[0043] Furthermore, it is learned that the information items of the word count report in XML format that this method needs to obtain often include: the file name of the word count report, the total number of segments, the total number of words, the number of files included, the number of translation memory files used when analyzing the word count and the names of the memories, the number of locked segments and words, the number of repeated segments and words, the number of new segments and words, etc. Only some commonly used information items are listed here. The actual information items that can be obtained are far more than these, but they can all be obtained by applying this method.
[0044] Furthermore, the method for obtaining information items includes: filtering out files whose suffix is not .xml, obtaining the user's input path, and then traversing each XML file under the path, judging whether the XML file belongs to a word count report file according to the basic format structure characteristics of the Trados XML word count report, and if it is a word count report file, traversing each XML format word count report file input by the user in the user-defined information item acquisition list with each information item as a unit, and trying to find the information item. To parse these XML files one by one, you can use the XML parsing method that is conveniently supported by the familiar programming language, such as the XMLDocument of the .NET platform.
[0045] There are two methods of parsing:
[0046] If you are familiar with the location of information items and can clearly obtain the xpath information of a specified information item (for example, to obtain the total number of segments), you can directly use the method supported by xmldocument to select nodes based on the xpath path to obtain the information item. At the same time, you must also make fault-tolerant judgments to prevent unexpected situations where the node does not exist and cause acquisition errors.
[0047] You can also manually traverse the nearest nodes of the XML word count report level by level, and then find the information item to be obtained based on the node name. This method also requires fault tolerance, that is, when all nodes are traversed and the information item to be obtained is still not found, a preset value of "False" of this method is returned to end the acquisition of this information item.
[0048] According to the actual situation, two XML word count report parsing methods can be selected to complete the acquisition of each information item. At this point, the work of step S1 "parsing the XML format word count report to obtain information" is completed.
[0049] Step S2: Based on the word count report information in XML format, generate a user interface table according to user needs.
[0050] After breaking the ice of "parsing the XML format word count report to obtain information", the information items needed at this time are already obvious and can be obtained at any time.
[0051] First, the information of the information item is obtained and placed in the interface table shown to the user in the order of the first letters of the pinyin of the first letter of the information item name; further, the interface table shown to the user is not a specific hard-coded table, it can be flexible and changeable, and the principle is to facilitate users to view, search, analyze and compare; although the design of the user interface table information items is flexible and diverse, it must also adhere to a principle, that is, the set information items cannot be the kind of information items that do not exist or cannot be obtained in the XML word count report; further, the information items used in the user interface table are available, and the information can also be parsed in the XML format word count report. At this time, a mapping relationship between the XML word count report information items and the user interface table information items is formed, and when needed during the method process, the information can be directly mapped over.
[0052] Step S3, applying the exception check rule, checking and marking the user interface form, adding the marked result to the user interface form, and completing the check of the TRADOS word count report.
[0053] The working scope of this method is each data item of each data in the user interface table, that is, it often appears after the "information item output to interface table" method; in addition to the data entries of the user interface data table, this method also requires a list of exception check item rules preset by the program. The examples of exception check item rules given by this technology are: the mounted tm is empty, the check option value is not 1, the number of new words is 0, and the number of files is 0. It can also be flexibly set according to actual needs, such as checking whether the total number of words in each word count report is the same, and treating the few that appear as exceptions and marking them with colors to ensure that users can clearly and quickly see the abnormal location or no abnormality based on the color marking.
[0054] Use the exception check rules set by the program or user one by one, apply the exception check rules to all the data in the interface table, and then obtain the abnormal data items, and then mark the abnormal data items with eye-catching colors. When all the exception check rules are run, the abnormal points in the interface table are also marked, and the user is reminded to check them one by one. It is also possible that all the inspection items are completely passed and there is no problem with the interface data.
[0055] Embodiment 2
[0056] This embodiment also provides a system for quickly checking a TRADOS word count report, comprising:
[0057] The initial data extraction module is used to extract and parse the word count report in XML format to obtain the word count report information in XML format.
[0058] Parse the word count report in XML format to obtain information; parse each XML file in the user input path, and use the XML parsing function of the programming language (such as XMLDocument of vb.net) to obtain the root node of the XML file. If the root node is not "<task name="analyse"> ", it is considered that this XML file is not a word count report file in the XML format of Trados, and it is ignored and not processed further, because these files are irrelevant to this method.
[0059] Furthermore, it is learned that the information items of the word count report in XML format that this method needs to obtain often include: the file name of the word count report, the total number of segments, the total number of words, the number of files included, the number of translation memory files used when analyzing the word count and the names of the memories, the number of locked segments and words, the number of repeated segments and words, the number of new segments and words, etc. Only some commonly used information items are listed here. The actual information items that can be obtained are far more than these, but they can all be obtained by applying this method.
[0060] Furthermore, the method for obtaining information items includes: filtering out files whose suffix is not .xml, obtaining the user's input path, and then traversing each XML file under the path, judging whether the XML file belongs to a word count report file according to the basic format structure characteristics of the Trados XML word count report, and if it is a word count report file, traversing each XML format word count report file input by the user in the user-defined information item acquisition list with each information item as a unit, and trying to find the information item. To parse these XML files one by one, you can use the XML parsing method that is conveniently supported by the familiar programming language, such as the XMLDocument of the .NET platform.
[0061] There are two methods of parsing:
[0062] If you are familiar with the location of information items and can clearly obtain the xpath information of a specified information item (for example, to obtain the total number of segments), you can directly use the method supported by xmldocument to select nodes based on the xpath path to obtain the information item. At the same time, you must also make fault-tolerant judgments to prevent unexpected situations where the node does not exist and cause acquisition errors.
[0063] You can also manually traverse the nearest nodes of the XML word count report level by level, and then find the information item to be obtained based on the node name. This method also requires fault tolerance, that is, when all nodes are traversed and the information item to be obtained is still not found, a preset value of "False" of this method is returned to end the acquisition of this information item.
[0064] According to the actual situation, two XML word count report parsing methods can be selected to obtain each information item.
[0065] The table mapping module is used to generate a user interface table according to user needs based on the word count report information in XML format.
[0066] After breaking the ice with the module "Parsing XML-format word count report to obtain information", the information items needed at this time are already obvious and can be obtained at any time.
[0067] The focus of this module is to obtain the information of the information items and place them in the interface table shown to the user in the order of the first letters of the pinyin of the first letters of the information item names; further, the interface table shown to the user is not a specific hard-coded table, it can be flexible and changeable, and the principle is to facilitate user viewing, retrieval, analysis and comparison; the design of the table is not the focus of this method, and too many examples and explanations will not be given here; although the design of the user interface table information items is flexible and diverse, it must also adhere to a principle, that is, the set information items cannot be the kind of information items that do not exist or cannot be obtained in the XML word count report; further, the information items used in the user interface table are available, and the information can also be parsed in the XML format word count report. At this time, a mapping relationship between the XML word count report information items and the user interface table information items is formed. When needed during the method, the information can be directly mapped over.
[0068] The word count checking module is used to apply the exception checking rules, check and mark the user interface form, add the marking results to the user interface form, and complete the check of the TRADOS word count report.
[0069] The working scope of this method is each data item of each data in the user interface table, that is, it often appears after the "information item output to interface table" method; in addition to the data entries of the user interface data table, this method also requires a list of exception check item rules preset by the program. The examples of exception check item rules given by this technology are: the mounted tm is empty, the check option value is not 1, the number of new words is 0, and the number of files is 0. It can also be flexibly set according to actual needs, such as checking whether the total number of words in each word count report is the same, and treating the few that appear as exceptions and marking them with colors to ensure that users can clearly and quickly see the abnormal location or no abnormality based on the color marking.
[0070] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A method for quickly checking the word count report of Trados, characterized in that: The method specifically comprises: Step S1, extracting and parsing the word count report in XML format in Trados to obtain word count report information in XML format; Step S2: Based on the word count report information in XML format, generate a user interface table according to user needs; Step S3, applying the exception check rule, checking and marking the user interface form, adding the marked result to the user interface form, and completing the check of the TRADOS word count report.
2. The method for quickly checking the word count report of TRADOS according to claim 1, characterized in that: The process of extracting the word count report in XML format specifically includes: Filter out files that are not word count reports in XML format and obtain the user's input path; Traversing each XML file under the input path of the user, and judging whether each XML file belongs to a word count report file according to the basic format structure characteristics of the XML word count report in Trados; If it is a word count report file, in the user-defined information item acquisition list, each information item is taken as a unit to traverse each word count report file in XML format input by the user to find the node where the current information item is located.
3. The method for quickly checking the word count report of TRADOS according to claim 2, characterized in that: The user-defined information items specifically include: The word count report contains the file name, total number of segments, total word count, number of files included, number of translation memory files used when analyzing the word count, name of the memory, number of locked segments and words, number of segments and words of repeated segments, and number of new segments and words.
4. The method for quickly checking the word count report of TRADOS according to claim 2, characterized in that: The method of finding the node where the current information item is located specifically includes: Traverse the nearest nodes of the current information item node; Compare the name of the current information item node with the name of the nearest node. If the names are consistent, it is considered to be found. If no consistent node is found at the end of the traversal, the default value is returned.
5. The present invention also discloses a system for quickly checking the word count report of Trados, the system is used to implement the method according to any one of claims 1 to 4, characterized in that: The system comprises: The initial data extraction module is used to extract and parse the word count report in XML format in Trados to obtain the word count report information in XML format; A table mapping module, used to generate a user interface table according to user needs based on the word count report information in XML format; The word count checking module is used to apply the exception checking rules, check and mark the user interface form, add the marking results to the user interface form, and complete the check of the TRADOS word count report.
6. The system for quickly checking the word count report of TRADOS according to claim 5, characterized in that: In the initial data extraction module, the process of extracting the word count report in XML format specifically includes: Filter out files that are not word count reports in XML format and obtain the user's input path; Traversing each XML file under the input path of the user, and judging whether each XML file belongs to a word count report file according to the basic format structure characteristics of the Trados XML word count report; If it is a word count report file, in the user-defined information item acquisition list, each information item is taken as a unit to traverse each word count report file in XML format input by the user to find the node where the current information item is located.
7. The system for quickly checking the word count report of TRADOS according to claim 6, characterized in that: The user-defined information items specifically include: The word count report contains the file name, total number of segments, total word count, number of files included, number of translation memory files used when analyzing the word count, name of the memory, number of locked segments and words, number of segments and words of repeated segments, and number of new segments and words.
8. The system for quickly checking the word count report of TRADOS according to claim 6, characterized in that: The method of finding the node where the current information item is located specifically includes: Traverse the nearest nodes of the current information item node; Compare the name of the current information item node with the name of the nearest node. If the names are consistent, it is considered to be found. If no consistent node is found at the end of the traversal, the default value is returned.