Cross-language specific content accuracy control method and system based on source-side explanation object and target-side understanding verification
By working together with the source-side interpretation module and the target-side understanding module, a structured source-side interpretation object and a colloquial mediation semantic representation are generated, which solves the problem that the target end does not understand the source-side specific content in machine translation, and achieves more accurate and traceable cross-language translation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEIYUAN EDUCATION TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2026-06-14
- Publication Date
- 2026-07-31
AI Technical Summary
Existing machine translation systems are prone to errors when dealing with cultural expressions, professional expressions, or other unique content. These errors can lead to the target end failing to understand the source-specific content and making speculative interpretations, over-domestication, mechanical translation, or semantic drift. Furthermore, they lack a traceable mechanism for controlling the accuracy of cross-language translation.
The source-side interpretation module identifies and interprets unique content, generating structured source-side interpretation objects and colloquial mediator semantic representations. The target-side understanding module verifies these representations. If they fail, a clarification protocol is used to request the source-side to supplement or correct them, ensuring that the accuracy threshold is met before generating the target language expression.
It improves the accuracy of cross-language translation, reduces semantic drift, enhances the traceability and process stability of target-side expression generation, adapts to various implementation forms, and avoids dependence on a single model.
Smart Images

Figure CN122491304A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of natural language processing, machine translation, semantic understanding, cross-language communication, and AI-assisted translation. More specifically, this invention relates to a method and system for controlling the accuracy of cross-language-specific content, which involves structured interpretation, target-side understanding verification, and protocol-based clarification of source-language-specific content such as cultural expressions, professional expressions, historical allusions, classical Chinese, internet slang, self-coined words, and names of new things in the source language before the natural expression of the target language is generated.
[0002] This invention does not necessarily presuppose a specific model architecture, a specific business model, or a specific neural network. The source-side interpretation module, the target-side understanding and expression module, the clarification protocol module, the external knowledge service, and the heterogeneous record module can be implemented by one or more natural language processing models, rule engines, retrieval services, knowledge bases, manual review interfaces, or combinations thereof. Background Technology
[0003] Existing machine translation technology can already output fluent translations from a large amount of everyday text, general text, and some professional text. However, in real-world cross-language communication scenarios, translation quality does not solely depend on the fluency of the target language. For idioms, allusions, classical Chinese, dialects, customs, internet slang, neologisms, names of new things, industry jargon, professional terms, and expressions with strong pragmatic functions, even if the target language output appears fluent, there may still be misunderstandings, omissions, or over-domestication of the source meaning.
[0004] For example, an expression in a source language may simultaneously possess literal meaning, implied meaning, emotional intensity, group identity conventions, historical allusions, and specific pragmatic functions. If the target translation system translates solely based on the target language's expression habits, it may convert the expression into seemingly natural idioms, colloquialisms, or internet slang that do not correspond semantically to the target language, thus causing semantic drift. Conversely, if the system only performs a literal translation, the target audience may not be able to understand the true function and context of the expression within the specific background of the source language.
[0005] Existing technologies commonly employ solutions such as terminology databases, translation memory, knowledge graph-enhanced translation, retrieval-enhanced translation, post-editing by humans, selection of multiple candidate translations, reverse translation, user feedback adaptation, and multi-agent collaborative translation. These solutions can improve terminology consistency, supplementation of professional knowledge, fluency of translations, or localization effects in specific scenarios. However, these solutions typically do not include "explicitly interpretable semantics of source-side content" as a mandatory control layer before target-side generation, nor do they require the target side to understand and verify the source-side interpretation object before generating natural expression candidates. Furthermore, they do not form a closed loop consisting of target-side low-confidence triggering, source-side supplementary interpretation response, interpretation object version update, gating state recording, and dual-end heterogeneous memory accumulation.
[0006] Therefore, a new cross-language translation accuracy control mechanism is still needed. This mechanism should prioritize "accurate understanding of source-specific content" over "target-side natural expression candidate generation or candidate expression strategy selection," so that target-side expression generation no longer depends on the language processing unit's guessing of source-specific meaning, but is completed based on structured source-side interpretation objects, colloquial mediation semantic representations, target-side understanding verification results, gating states, and necessary clarification loops. Summary of the Invention
[0007] The technical problem this invention aims to solve is that when the source language text contains cultural expressions, professionally specific expressions, or other unique content, existing machine translation systems tend to directly enter the target language generation stage without generating machine-readable source-side interpretable objects, completing target-side understanding verification, or establishing gating states and state register records. This leads to the target-side processing unit making speculative interpretations, over-dominantly interpreting, mechanically translating, or semantically drifting of the source-side unique content, resulting in invalid expression generation calls, untraceable interface states, and unstable exception handling. This problem is particularly prominent in scenarios such as cross-cultural communication, film and television subtitling, international education, translation of ancient books, cross-border customer service, professional document translation, and brand localization.
[0008] Furthermore, the technical problems to be solved by this invention also include: how to combine source-side explanation capabilities, target-side understanding and verification capabilities, clarification and completion capabilities, and expression generation capabilities into an executable technical process without being bound to a specific business model or specific neural network architecture; how to distinguish between source-side explanation libraries and target-side expression libraries to avoid degenerating this invention into a general bilingual glossary or translation memory; and how to form a structured, traceable, and version-updable cross-end clarification mechanism when the target side cannot understand it.
[0009] To address the aforementioned technical problems, this invention provides a method for controlling the accuracy of cross-language-specific content based on source-side interpretation objects and target-side understanding verification. The core idea of this method is as follows: First, the source-side interpretation module, executed by the processor, identifies and interprets the specific content in the source language text, transforming it into a structured source-side interpretation object and a colloquial intermediary semantic representation. Then, before generating a natural expression in the target language, the target-side understanding and expression module performs target-side understanding verification on the interpretation object and colloquial intermediary semantic representation based on interface data. If the verification fails, a structured clarification protocol is used to request the source side to supplement or correct the interpretation object and update the gating state. Only when the accuracy threshold is met and the gating state machine outputs a state allowing generation is the target side allowed to execute either natural expression generation or localization expression strategy selection.
[0010] In this invention, the "source end" is not limited to the source language user, but can also be a source language model, a source language rule engine, a source language knowledge base, a source language expert interface, or a combination thereof. Similarly, the "target end" is not limited to the target language user, but can also be a target language model, a target language rule engine, a target language localization module, a target language quality verification module, or a combination thereof. Furthermore, the "model" in this invention is not limited to a large language model, nor is it limited to a specific neural network architecture or a specific business model.
[0011] Compared with existing technologies, this invention has at least the following beneficial effects. First, by using a structured source-end interpretation object, this invention explicitly extracts source-end-specific content from the text to be translated and records information such as literal meaning, implied meaning, pragmatic function, risk of non-literal translation, and colloquial explanation in a field-based manner, so that the target-end processing unit no longer infers the source-end-specific meaning solely based on the probability distribution of the target language. Second, by using target-end understanding verification, accuracy thresholds, and a gating state machine, this invention places the generation of natural expressions in the target language after semantic accuracy verification, reducing the risk of semantic drift caused by the "first authentic, then corrective" approach from a process structure perspective. Third, through a structured clarification protocol, this invention enables the target end to request specific field completion from the source end when it has low confidence, does not understand, or fails to map, instead of performing general retranslation or open-ended questions and answers, thereby reducing invalid expression generation calls, invalid retranslations, and repeated interpretations. Fourth, this invention, through the heterogeneous association of source-side interpretation records, target-side expression records, and state registration records, can reuse both source-side unique content interpretations and target-side expression strategies, while avoiding confusion between the two as a general bilingual glossary and improving process state traceability. Fifth, this invention can be implemented in various forms, including dual-model, single-model multi-routing, rule engine, knowledge retrieval service, and manual supplementation interface, exhibiting strong deployment adaptability and the ability to avoid dependence on a single model.
[0012] Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings are briefly described below. The drawings are used to illustrate exemplary embodiments of the present invention.
[0014] Figure 1 This is a schematic diagram of the overall process of a cross-language-specific content accuracy control method according to the present invention. The diagram includes, in sequence, input acquisition, source-side specific content detection, generation of structured source-side interpretation objects, generation of colloquial mediator semantic representations, target-side understanding verification, clarification request and interpretation object update, accuracy threshold judgment, generation of natural expression in the target language, and heterogeneous record update.
[0015] Figure 2 This diagram illustrates the modular structure of a cross-language-specific content accuracy control system according to the present invention. The diagram includes an input acquisition module, a source detection module, a source interpretation object generation module, a target understanding verification module, a clarification protocol module, an expression generation module, a heterogeneous recording module, and a human-machine collaboration supplementation module.
[0016] Figure 3 This diagram illustrates the data relationships between the structured source-side explanation object, the colloquial mediation semantic representation, the target-side verification result, and the structured clarification request in this invention. The diagram demonstrates the association between the explanation object identifier, semantic signature, context identifier, and language pair identifier.
[0017] Figure 4 This diagram illustrates the heterogeneous association between the source-side explanation record and the target-side expression record in this invention. It demonstrates that the source-side explanation object library and the target-side expression mapping library are not the same bilingual terminology, but are associated through semantic signatures, context vectors, usage scenarios, and confidence levels. Implementation
[0018] The present invention will be further described below with reference to embodiments. The described embodiments are used to illustrate the implementation of the present invention. In different deployment environments, the module combination, data format or interface form can be adapted while maintaining the integrity of the technical chain of "source interpretation object - target understanding and verification - structured clarification - accuracy threshold - heterogeneous record".
[0019] For ease of understanding, several terms are defined in this specification as follows. "Source-specific content" refers to text content that possesses a special background, specific pragmatic function, specific group meaning, professionally defined meaning, or non-literal meaning in the source language or culture, and for which the target recipient or target processing unit has a default knowledge probability higher than a threshold, posing a risk of semantic deviation in literal translation, or posing a risk of pragmatic functional deviation. This content can be words, phrases, sentences, idioms, allusions, classical Chinese fragments, internet slang, industry terms, self-coined words, names of new things, or combinations thereof within the context.
[0020] A "structured source-end interpretation object" refers to a data object that stores source-specific content interpretation information in a field-based data structure. This object is neither a regular bilingual entry nor a target language translation, but rather an intermediate interpretation unit used by the target end to understand the meaning of the source content before translating or localizing it.
[0021] "Common intermediary semantic representation" refers to an intermediate semantic representation that, while retaining the non-specific content of the source language, replaces, annotates, or associates the source-specific content with structured source-specific interpretation object references. The purpose of this representation is not to generate elegant translations, but rather to ensure that the source-specific semantics enter into subsequent processing in a form verifiable by the target language.
[0022] "Target-side understanding verification" refers to the process by which the target-side understanding and expression module calculates the understanding confidence, semantic consistency, and target-side fit evaluation of the colloquial mediating semantic representation and the structured source-side explanation object before generating natural expression. The target-side fit evaluation can be calculated based on the target audience, target-side usage context, consistency of expression tone, cultural acceptability, and risk tolerance. "Target-side expression mapping candidates" refer to internal verification intermediate quantities, candidate strategy identifiers, or restricted slot mappings formed during target-side understanding verification. These are used to determine whether the source-side explanation object can be covered by the target-side semantics, and are not output as the final target language natural expression result under any gating state. "Accuracy threshold" refers to the control conditions that allow entry into the target language natural expression generation stage. The results can be expressed as gating states such as allowing generation, blocking generation, clarifying pending processing, or downgrading processing.
[0023] "Heterogeneous records" refer to bilingual tables where the source-side interpretation record and the target-side expression record are not of the same structure. The source-side interpretation record stores the interpretation objects containing source-specific content; the target-side expression record stores the target language expression results, expression strategies, applicable scenarios, and target-side adaptation evaluations. The two are linked through semantic signatures, context identifiers, language pair identifiers, context vectors, or usage scenarios.
[0024] In this invention, "culture" is defined broadly. Broadly defined, culture refers to the sum of all knowledge, expressions, and background information formed within a specific language community or group, transmitted through language, possessing specific meanings or functions for specific members of that language community or group, and for which the target recipient or processing unit may lack default background knowledge. Broadly defined culture can manifest as historical and traditional culture, contemporary popular culture, professional culture, organizational culture, and other linguistic expressions that possess specific meanings or functions within the source but lack default background knowledge at the target end. Historical and traditional culture can include idioms, allusions, classical Chinese, ancient script expressions, customs, or expressions related to historical events; contemporary popular culture can include internet slang, memes, self-coined terms, names of new things, or group-specific terms; professional culture can include industry terminology, domain jargon, professionally agreed-upon expressions, or occupational habits; and organizational culture can include internal company terminology, project codes, or internal organizational terms. This terminology definition covers the technical context referred to by the use of "culture" or compound terms containing "culture" in this application, which may include "cultural dependence", "cultural source", "cultural acceptance", "target culture", "source culture", "cultural expression", "cross-cultural", "cultural content" or "cultural transmission".
[0025] Overall Process Implementation Example. In one embodiment, the system acquires the source language text, target language identifier, and target usage context information input by the user. Target usage context information may include the target audience, dissemination purpose, formality level, domain tags, regional tags, platform type, text genre, whether to retain the unique unfamiliarity of the source language, whether to allow annotations, maximum output length, and risk tolerance, etc.
[0026] The system first segments the source language text into candidate segments. Candidate segments can be words, phrases, sentences, fixed expressions, quoted passages, or combinations across sentences. Then, the system calculates a specificity score for each candidate segment. The specificity score is determined by at least the literal translation risk and the target-side default knowledge probability, and can be further combined with rarity, cultural dependence, or contextual dependence. Rarity reflects the frequency of the segment's occurrence in a general corpus; cultural dependence reflects whether the segment depends on a specific region, history, custom, or group background; contextual dependence reflects whether the segment is easily misunderstood when taken out of context; literal translation risk reflects whether a literal translation might lead to misunderstanding by the target end; and the target-side default knowledge probability reflects whether the target language recipient or the target processing unit might lack relevant background knowledge.
[0027] When the specificity score meets the preset candidate criteria, the system generates a structured source-end explanation object for the candidate fragment. This object includes at least the following fields: object identifier, original text fragment, category label, context window, literal meaning, implied meaning, pragmatic function, popular explanation, risk of non-literal translation, and explanation confidence. It may further include enhanced fields such as cultural origin, usage scenario, emotional intensity, rhetorical type, suggestion for the degree of retention of source-end specificity, reason for prohibiting literal translation, explanation source, and explanation version number.
[0028] The system further generates a colloquial intermediary semantic representation based on the structured source-side interpretation objects. This colloquial intermediary semantic representation can retain the semantic representation of ordinary content in the original text and replace or annotate source-specific content with object references. For example, a popular internet slang term in the source language can be replaced with "[Interpretation Object ID: This expression represents self-deprecating humor in the current context, with moderate emotional intensity, and cannot be understood literally]". The target-side expression understanding module receives the colloquial intermediary semantic representation and interpretation objects, rather than just the original text or the target-side initial translation.
[0029] The target-side understanding and expression module then performs target-side understanding verification. This module calculates the understanding confidence, semantic consistency, and target-side adaptation of the explained object based on target language knowledge, target-specific expression habits, target audience background, target-side usage context, consistency of expression tone, cultural acceptability, and risk tolerance. It then generates gating results that allow generation, block generation, or downgrade the expression. If the target-side verification result meets a preset accuracy threshold and the gating result indicates that generation is allowed, the system enters the target language natural expression generation stage; otherwise, it generates a structured clarification request.
[0030] Structured clarification requests are not ordinary question-and-answer messages, but rather protocol-based messages with field constraints. They can include source-side explanation object identifiers, conflicting fields, missing context, target-side assumptions, failed expression mapping candidates, request supplementation fields, priority, and deadlines. Among these, failed expression mapping candidates are only used to indicate internal validation intermediates or candidate strategy identifiers that failed gating and do not constitute the final natural expression result in the target language. Upon receiving a clarification request, the source-side explanation module supplements, corrects, or updates the corresponding explanation object and returns the updated explanation object to the target-side understanding and expression module. The target then re-executes understanding and validation based on the updated object.
[0031] After the accuracy threshold is met, the target - side understanding and expression module selects an expression strategy according to the target - side usage context. The expression strategies can include literal translation, interpretive translation, functional equivalent expression, target - side analogical expression, retaining the original word with annotation, transliteration plus explanation, and replacement with target - side idiomatic expressions, etc. After the system generates the natural expression result in the target language, it records the structured source - side interpretation object or its updated version in the source - side interpretation record, and records the natural expression result in the target language, the target - side expression strategy, or the target - side expression mapping in the target - side expression record. The source - side interpretation record and the target - side expression record are associated through semantic signatures, context identifiers, or language - pair identifiers for subsequent reuse.
[0032] Data structure embodiments. In one embodiment, the structured source - side interpretation object can adopt the following field structure. Among them, object identifier, original text fragment, category label, context window, literal meaning, implicit meaning, pragmatic function, non - literal - translation risk, popular explanation, and interpretation confidence are the minimum field combination to support target - side understanding and verification; different implementation manners can add enhanced fields according to the scenario, but the minimum fields for explaining the source - side meaning and literal - translation risk should not be omitted.
[0033]
[0034] In one embodiment, the target - side verification result can adopt the following field structure.
[0035]
[0036] In one embodiment, the structured clarification request can adopt the following field structure.
[0037]
[0038] In one embodiment, the popular intermediary semantic representation can adopt the following exemplary JSON structure (for illustrative purposes only, not limiting the specific format):
[0039]
[0040] This example shows that in the text to be processed, "Hongmen banquet" and "full of twists and turns" are respectively replaced by references to the corresponding source - side interpretation objects, so that the target - side understanding and expression module must perform understanding and verification based on the interpretation objects before generating the target - language expression, rather than directly guessing their meanings.
[0041] Specificity scoring implementation example. In one embodiment, the system can calculate the specificity score as follows. Assuming a candidate text fragment is x, the system calculates the rarity score R(x), cultural dependence score C(x), context dependence score K(x), translation risk score L(x), and target-side default knowledge probability T(x). The system can calculate the specificity score S(x) using a weighted method:
[0042]
[0043] Among them, w1 to w5 are weight parameters, which ensure that at least the literal translation risk score L(x) and the target-side default knowledge probability T(x) are necessary scoring factors, i.e., w4>0 and w5>0 always hold true, ensuring that the literal translation risk and the target-side default knowledge probability always make a positive contribution to the specificity score; the weights of other factors (rarity, cultural dependence, context dependence) can be zero. The weight parameters can be adjusted according to the preset rule base, language pair configuration, domain configuration, user group, historical mistranslation records, manually annotated data, or manually calibrated records.
[0044] In another embodiment, the system may also introduce a context anomaly detection factor A(x). When the context window features of a candidate text fragment x (including at least one of contextual vocabulary distribution, semantic coherence, syntactic structure, and topic consistency) show statistical anomalies with the overall context distribution of the text to be processed—for example, an expression inconsistent with the overall style of the text appears in the current context—A(x) takes a positive value and participates in the specificity score calculation. The context anomaly detection factor can be weighted independently or added to the weighting formula as an adjustment coefficient of existing scoring factors. Specific implementation methods include context consistency quantification based on statistical language models, anomaly expression detection based on classifiers, or context matching degree evaluation based on semantic similarity calculation. When A(x) is higher than a preset anomaly detection threshold, the system marks the fragment as a context anomaly candidate and prioritizes its entry into the source-side specific content detection process.
[0045] When S(x) is higher than the first threshold, the system marks x as a source-specific content candidate; when S(x) is lower than the first threshold but higher than the second threshold, the system can mark x as an observation candidate and continue monitoring it during the target-side verification phase; when the target-side understanding verification fails, the observation candidate can be upgraded to a formal interpretation object. The above thresholds are derived from language pair configuration, domain configuration, risk level configuration, historical mistranslation records, or manual calibration records; the output of the context anomaly detection factor A(x) includes at least the anomaly score, anomaly type, and corresponding candidate state, and is written into the state register record as state transition input.
[0046] In another embodiment, the system need not use a linear weighting formula; it can also use a classifier, rule set, retrieval matching, manual verification, model output confidence, or a combination thereof to generate a specificity score. Any implementation of this invention is acceptable as long as its purpose is to identify source-specific content that the target end may not understand by default.
[0047] Accuracy Threshold and Expression Layer Activation Implementation Example. This invention emphasizes "accuracy first, expression later." In one embodiment, the target-side understanding and expression module only enters the target language natural expression generation stage after the accuracy threshold is met. The accuracy threshold includes: understanding confidence not lower than a first threshold, semantic consistency not lower than a second threshold, target-side adaptation evaluation not lower than a third threshold, the risk of non-literal translation being covered by the expression strategy, and the absence of unprocessed conflict fields; the above conditions together form the basic threshold combination for expression layer activation.
[0048] The execution flow of the structured clarification protocol includes the following stages: First, when the target-side understanding expression module detects conflicting fields or missing context during verification, it generates a structured clarification request containing conflicting field identifiers, target-side assumptions, and requested supplementary fields. Second, the source-side explanation module parses the structured fields in the clarification request and generates supplementary explanations or corrected versions targeting specific conflicting fields rather than general questions. Third, the source-side returns the updated explanation object to the target-side, which re-executes the understanding verification based on the updated explanation object. Fourth, if the verification passes, expression generation begins; if it fails and a preset number of rounds or deadlines are reached, degradation processing is triggered. This four-stage process ensures that clarification always revolves around specific fields, rather than degenerating into question-and-answer retries.
[0049] In one embodiment, the accuracy threshold and the expression layer activation are executed by a gating state machine, and the gating state transitions must at least satisfy the following table:
[0050]
[0051] Before the gating state machine executes, the system performs format verification on the received interface data. Format verification includes: checking that the object identifier conforms to a preset identifier format (such as UUID or structured numbering pattern), that the semantic signature length and character set conform to preset rules, that the context identifier is within a valid value range, and that the gating state field belongs to one of the following valid states: allowed generation, blocked generation, clarified pending processing, or downgraded processing. Interface data that fails format verification is marked as invalid interface data and triggers a status recording alarm, and does not enter the main process of the gating state machine until the interface data passes format verification or triggers downgrade processing.
[0052] After the gating state machine is executed, the system performs integrity checks on the generated gating states. The integrity checks include: (1) State transition path checks—checking whether the transition path between the current gating state and the previous state is a valid transition—for example, without a clarified response, conflict field closure, or re-verification record, it is not allowed to directly transition from the blocked generation state to the allowed generation state; (2) State record consistency checks—verifying that the gating state stored in the state register is consistent with the gating state actually output by the gating state machine, preventing the state record from being tampered with or overwritten. Gating states that fail the integrity checks are marked as state abnormalities and trigger alarms, and are not allowed to enter the expression generation stage.
[0053] The gated state machine can be executed as follows pseudocode:
[0054]
[0055] For example, if the common interpretation of a source idiom clearly states its true meaning as "doing the wrong thing at the wrong time," but the target candidate expression interprets it as "doing something romantically," then the semantic consistency is below the threshold. The system will not directly output the target candidate expression but will instead generate a clarification request or change the expression strategy. Similarly, if a source internet meme has self-deprecating and group identity expression functions, and the target candidate expression, while literally similar, changes its tone to attack others, or does not conform to the target audience, usage context, or risk tolerance, then the target fit evaluation is below the threshold, and the system should prevent that candidate from entering the final output.
[0056] In one embodiment of the downgrade process, the system enters the downgrade process when a preset round or a pre-set lower limit of reliability gain in the structured clarification protocol is triggered. The specific execution steps of the downgrade process include: First, the system marks the current gating state as a downgrade state and generates a downgrade processing identifier, which is recorded in the process log. Second, based on the original content type and dissemination purpose of the text to be processed, the system selects at least one downgrade strategy from preserving the source expression with additional explanation, requesting manual supplementary explanation, reducing the degree of expression naturalization, or marking high-risk translations. Third, the system adds a manual review mark and a failure to the gating state mark to the downgraded output, giving the downgraded output a traceable risk identifier in subsequent processes. Fourth, the downgraded output is sent to the target end or the manual review queue without triggering the standard output process of the expression generation module. When manual review confirms that the downgraded output is safe to use, the manual review result can update the fixed state of the corresponding record item, but this does not affect the independence of the core verification process in subsequent processing. The downgraded output can only be presented to the target recipient as a temporary substitute for the natural expression of the target language after manual review and confirmation. Before the manual review is completed, the downgraded output is limited to internal review process and cannot be directly used as the final natural expression of the target language.
[0057] An example of an expression strategy selector. In one embodiment, the target-side expression understanding module includes an expression strategy selector. This selector chooses an expression strategy based on content type, communication purpose, target audience, degree of retention of source-specificity, target-side usage context, and pragmatic function. The expression strategy is not limited to one type and can also be used in combination.
[0058]
[0059] The expression strategy selector does not simply replace source idioms with target idioms. For scenarios that require preserving the unique unfamiliarity of the source, the system can choose to retain the original words and add explanations; for advertising, film and television subtitles, or social media platforms, the system can choose functionally equivalent expressions for the target without compromising semantic accuracy; for educational or cultural dissemination scenarios, the system can prioritize explanatory translations to help the target audience understand the source culture.
[0060] Exemplary target-side naturalization conversion methods. In one embodiment, to convert source-side-specific content into a natural expression understandable to the target side, the system may employ one or more combinations of the following exemplary conversion methods after the target side has understood and verified the conversion:
[0061] (i) Substantive Content Method: Directly express the substantive meaning of the source-specific content in the target language. For example, the expression "Hongmen Banquet," which has a specific historical background in the source and is ostensibly a banquet but actually a dangerous situation, can be translated as "a meeting that is superficially friendly but contains hidden hostile intentions" or "a trap with pre-set risks," without being bound by the form of the banquet.
[0062] (ii) Analogy method: When the target end does not have a specific concept unique to the source end, find the closest analogy in terms of function or emotion on the target end. For example, translate "This is a Feast at Hongmen" as "This is a meeting that is friendly on the surface but contains hidden risks", using the concept of a risky meeting that is familiar to the target end readers to convey the sinister intentions behind the banquet from the source end.
[0063] (iii) Combining the real and the imaginary: Simultaneously conveying both the literal imagery and the actual meaning of the original content. For example, translating the original dish name "Ants Climbing a Tree" as "Minced Meat and Vermicelli (like ants climbing a tree)" retains the fun of the literal imagery while conveying the actual meaning.
[0064] (iv) Contextual Method: Integrate the meaning of the source-specific content into the actual context of the communication so that the target recipient understands it unconsciously. For example, if the source says, "Your perseverance in overcoming numerous difficulties is admirable," the meaning of the story of "The Foolish Old Man Who Moved Mountains" can be translated according to the current context as "Your spirit of perseverance in overcoming numerous difficulties and never giving up is admirable," rather than mechanically translating the literal meaning.
[0065] (5) Retention and Explanation Method: Retain the original form of the unique expressions at the source end, and attach short explanations in the form of footnotes, parentheses or embedding. For example, "Yu Gong Moves the Mountain" is directly translated as "Yu Gong Moves the Mountain" with the annotation "a Chinese fable about perseverance against impossible odds".
[0066] The described exemplary conversion method is used to illustrate how to select a candidate expression strategy after the understanding verification at the target end is passed. Other conversion methods can also be used as candidate strategies in the expression generation stage, but their activation is still restricted by the core process of "first source-end explanation, then target-end verification, and finally expression generation".
[0067] Dual-model Embodiment. In one embodiment, the source-end interpretation module is implemented by a native model of the source language or a dominant model of the source language, and the target-end understanding and expression module is implemented by a native model of the target language or a dominant model of the target language. The source-end model is responsible for identifying and interpreting the unique content at the source end, and the target-end model is responsible for verifying whether it is understandable at the target end and generating a natural expression result in the target language after the accuracy threshold is met.
[0068] The advantage of this dual-model embodiment is that the source-end model is more suitable for identifying the unique backgrounds, historical allusions, implicit meanings and usage scenarios in the source language, and the target-end model is more suitable for judging whether the target-language recipients can understand and which expression is more natural. The two are not simply connected in series for translation, but form a collaborative closed loop through structured source-end interpretation objects, popular intermediate semantic representations and clarification protocols.
[0069] Single-model Multiple-routing Embodiment. In another embodiment, the system can implement the source-end interpretation module and the target-end understanding and expression module through different functional routings in the same language processing unit. For example, the same language processing unit is configured as the source-end interpretation role in the first functional routing, outputting a structured source-end interpretation object; it is configured as the target-end verification role in the second functional routing, outputting the target-end verification result and the gating state; and it is configured as the expression generation role in the third functional routing, generating a natural expression result in the target language after the accuracy threshold is met and the gating state indicates permission to generate.
[0070] This embodiment avoids limiting the invention to the use of two independent models. Even in actual deployments using a single language processing unit, multiple prompt templates, multiple tool calls, multiple context windows, or multiple system roles, intermediate states such as structured source-side explanation objects, colloquial mediation semantic representations, target-side verification results, structured clarification requests, and gating states should be generated and transmitted in the process, and the processing boundaries of each functional route should be reflected through interface logs or state register records. The intermediate states can be represented by one or more equivalent data carriers such as object field tables, semantic graph nodes, slot sets, vector references, constrained prompt contexts, message queue payloads, or state register records, as long as they can make the source-side explanation, target-side verification, clarification request, and expression generation stages distinguishable and traceable in the processor execution flow.
[0071] External Knowledge Service Implementation Example. In another embodiment, the source-side interpretation module can call external knowledge services, terminology databases, encyclopedic knowledge bases, classical Chinese dictionaries, industry dictionaries, internal organizational knowledge bases, regionally specific knowledge databases, or online slang databases to supplement the structured source-side interpretation object fields. The role of external knowledge services is to provide background and explanation information within the source language scope for the source-side interpretation object, rather than directly replacing the target-side understanding and expression module in outputting the final natural expression result in the target language.
[0072] For example, when the text to be processed contains classical Chinese expressions, the system can first call upon ancient Chinese dictionaries or authoritative interpretive resources to obtain explanations of single characters, phrases, or allusions, and then the source-side interpretation module organizes them into structured source-side interpretation objects. The target-side understanding and expression module then performs verification and translation based on these interpretation objects and common mediating semantic representations, rather than directly letting the target-side model guess the meaning of the classical Chinese text.
[0073] A supplementary embodiment of human-computer collaboration. In another embodiment, when the target end's understanding verification fails and the source end's explanation module fails to identify a certain source-specific content, the system can trigger the human-computer collaboration supplementary module. This module sends a supplementary explanation request to the source end user, domain expert, content creator, or administrator. The supplementary explanation request may require the user to explain the true meaning, tone, usage scenario, whether it is a self-coined term, whether it involves internal puns, whether source-specific features should be retained, and whether there are any prohibited translations.
[0074] After receiving supplementary explanations from humans, the system converts them into structured source-end explanation objects or updates the fields of existing explanation objects. Subsequently, the target-end understanding and expression module re-executes understanding verification based on the updated explanation objects. If the verification passes, the system then enters the target language natural expression generation stage. When the received supplementary explanation has semantic conflicts with existing structured source-end explanation objects or system-built-in knowledge (including terminology bases, knowledge graphs, and existing verified explanation records), the human-machine collaborative supplementation module triggers a cross-validation process, requiring at least one additional source-end user or domain expert to independently confirm the supplementary explanation to ensure its accuracy and consistency. This embodiment is applicable to internal enterprise terminology, brand-created terms, social media memes, newly emerging online expressions, and names of new things not yet fully covered in model training data.
[0075] An embodiment of proactive input of source-side knowledge base. In another embodiment, the system further includes a source-side knowledge base management module, which receives source-side proprietary knowledge data proactively uploaded or input by source-side users. The source-side proprietary knowledge data includes source-side-specific content that needs to be proactively provided or confirmed by the source-side user when general knowledge services are not hit, domain access permissions are restricted, commercial access permissions are restricted, the probability of default knowledge on the target side is higher than a threshold, or there are conflicts in the confidence of existing explanations. Examples include specific operational procedures within an enterprise, organization-specific skill names, proprietary terminology systems, project codes, process abbreviations, and conventional terms circulating only within a specific group. Such knowledge may appear to be ordinary text, but it has a conventional meaning different from its literal meaning in the context of a specific organization, project, or group, and this conventional meaning is not fully based on existing word meanings in public knowledge bases or general training corpora.
[0076] The source-side knowledge base management module supports receiving source-side proprietary knowledge data through at least one of the following methods: file upload, form filling, API integration, or interactive Q&A. It organizes the received data into structured source-side explanation objects or supplementary fields for these objects and stores them in the source-side knowledge base. During the detection of source-side specific content, when a candidate text fragment in the text to be processed matches proprietary knowledge data in the source-side knowledge base, the system does not directly perform source-to-target word replacement. Instead, it generates or supplements the corresponding structured source-side explanation object based on the proprietary knowledge data and continues to perform target-side understanding verification.
[0077] This embodiment differs from the aforementioned human-machine collaborative supplementation mode. Human-machine collaborative supplementation occurs in a passive situation where the target end's understanding and verification fails, while the active input of the source-end knowledge base occurs before the translation process begins, with the source-end user proactively injecting their unique knowledge into the system beforehand. When the source-end user actively inputs the meaning, applicable context, or agreed-upon source of this unique content, the system can identify it as a candidate for source-end unique content in subsequent translations and generate or supplement a structured source-end explanation object. Furthermore, this process is still constrained by the structured source-end explanation object, the target end's understanding and verification, and accuracy thresholds, rather than directly generating the final translation based on preset translation replacement rules.
[0078] In another embodiment, the source-side knowledge base management module performs input validation, format validation, field validity validation, permission validation, and injection detection on the received source-side proprietary knowledge data to prevent malicious users from injecting false interpretation objects into the source-side knowledge base through privilege escalation or data forgery. Input validation includes verifying whether the content format of the knowledge data conforms to structured field requirements; format validation includes checking the validity of field length, field type, and field content; field validity validation includes checking whether the object identifier, context window, meaning explanation, and agreed source meet preset field constraints; injection detection includes cross-referencing the knowledge data to be entered with existing knowledge data to identify and block obviously conflicting false interpretations. Knowledge data that fails validation is rejected from being entered into the knowledge base and an alarm is triggered. Validated data is organized into structured source-side interpretation objects or supplementary fields for source-side interpretation objects and stored in the source-side knowledge base, generating call logs and conflict detection records when invoked.
[0079] Multi-target language concurrent implementation. In another embodiment, the system can provide the same structured source-side interpretation object and the same colloquial mediator semantic representation to multiple target-side understanding and expression modules respectively. For example, the same Chinese text can be translated into English, Japanese, French, and Arabic simultaneously. Each target-side understanding and expression module independently performs understanding verification, clarification request generation, and expression strategy selection based on the corresponding target language and the target-side specific context.
[0080] For content unique to the same source, different target languages may adopt different strategies. For example, a certain source allusion may be suitable for explanatory translation in English, may have a close analogy in Japanese, and may require retaining the original word with annotations in French. This invention decouples the source explanation object from the target expression record, allowing a single source explanation object to be associated with multiple target expression mappings, while retaining the verification results and adaptation evaluations of each target language.
[0081] In another embodiment, when performing concurrent translation of multiple target languages, the system performs cross-language consistency checks on the target-side validation results independently generated by each target-side understanding and expression module. When the gating result of a certain target language differs significantly from the gating results of other target languages—for example, the gating result of the target language blocks generation while other target languages allow generation, or the target language continuously triggers clarification requests while other target languages have passed validation—the system marks this difference as cross-language validation inconsistency. For target languages marked as having cross-language validation inconsistency, the system can lower the confidence level of their validation results, require the target to re-perform understanding and validation, or append the difference information to a supplementary field of the structured source-side interpretation object for subsequent audit reference. This mechanism helps identify abnormal validation behavior caused by differences in target-side model characteristics.
[0082] Heterogeneous Recording and Caching Implementation Example. In one embodiment, the system includes a source-side interpretation object library and a target-side expression mapping library. The source-side interpretation object library stores interpretation objects specific to the source, such as original text fragments, category tags, context windows, literal meanings, implied meanings, pragmatic functions, risks of non-literal translation, popular explanations, interpretation confidence levels, and version numbers. The target-side expression mapping library stores natural expression results in the target language, expression strategies, applicable scenarios, target-side adaptation evaluations, usage frequency, confirmation status, and records of manual corrections.
[0083] The source-end explanation object library and the target-end expression mapping library are not the same bilingual terminology. Ordinary bilingual terminologies typically store a "source word—target word" correspondence, while the source-end explanation object of this invention can associate multiple target languages, multiple target-end expression strategies, and multiple target-end scenarios. A single source-end explanation object can correspond to explanatory translations in cultural dissemination scenarios, short expressions in film and television subtitle scenarios, original word annotations in educational scenarios, and functionally equivalent expressions in marketing scenarios.
[0084] The system can update the confidence level or fixed status of record items based on the number of hits, user-initiated confirmation signals, target-side verification results, manual correction signals, and time decay factors. When a mapping between an explanation object and the target-side expression is verified multiple times and receives user-initiated confirmation, the system can increase its confidence level; when a record item is deemed inapplicable in a new context, the system can decrease its confidence level or limit it to a specific context. User-initiated confirmation signals refer to explicit confirmation signals issued by the source user—including clicking a confirmation button on the displayed translation result, submitting a confirmation form, or sending a confirmation command—rather than indirect signals implicitly inferred from user behavior feedback (such as page dwell time, click frequency, scrolling behavior), interaction history (such as historical translation selection preferences), or model fine-tuning results.
[0085] In this invention, "records" encompass both temporary and persistent storage. Session-level temporary memory also falls under the category of "records" in this invention—although it can be discarded after a session ends, it still exists during the session in a heterogeneous association where source-side interpretation records and target-side expression records are stored separately. Furthermore, the creation, updating, and querying of record items during a session all undergo the source-side content detection, target-side understanding verification, and accuracy threshold determination processes of this invention. Even when using entirely new temporary memory in each new session (without inheriting any records from previous sessions), the system still performs a complete source-side interpretation object generation, target-side understanding verification, accuracy threshold determination, and structured clarification process for the current text within each session—session-level temporary memory only affects the reuse efficiency of record items across sessions, not the integrity of the core verification process. Regardless of the form the record takes, the core feature of this invention—"heterogeneous association between source-side interpretation records and target-side expression records"—remains unchanged.
[0086] Safety and Risk Control Implementation Example. In one embodiment, the system performs risk control on the natural expression results in the target language. For translations that are low-confidence, controversial, may offend the target culture, may misuse group identity expressions, or may cause legal or compliance risks, the system may output a risk marker, provide multiple expression candidates, reduce the degree of naturalization, retain the source expression and add explanations, or request manual confirmation.
[0087] In another embodiment, the system performs field integrity verification and content specificity verification on structured clarification requests to identify and prevent spurious clarification protocols. A spurious clarification protocol refers to a false clarification where the target end simply replaces the structured clarification request with a generic question (such as "What does this mean?"), without including structured fields such as conflicting fields, missing context, or failed expression mapping candidates. Upon receiving a clarification request, the system verifies whether it contains at least two of the following structured fields: conflicting fields, missing context, and failed expression mapping candidates. If the request does not meet the field integrity requirements, it is identified as a spurious clarification and rejected from entering the source-end supplementation process or the target end is required to regenerate a structured clarification request that meets the field integrity requirements. Content specificity verification further requires that: when the clarification request contains conflicting fields, the conflicting fields must specify the specific literal meaning, implied meaning, or pragmatic functional conflict location, rather than merely declaring the existence of a conflict; when the clarification request contains missing context, the missing context must specify the specific domain, scope, or type of the required background information, rather than merely declaring insufficient context. This dual verification mechanism ensures that the clarification process is always based on field-level structured information and the specificity of field content, rather than general questions and answers or merely meeting the formal requirements for field existence.
[0088] In another embodiment, the system performs pre-trained knowledge mapping verification on the detection of source-specific content. When a candidate text segment appears less frequently in the general corpus than a preset frequency threshold and the language processing unit's confidence in its meaning is higher than a preset confidence threshold, or its output meaning is inconsistent with the current context window, the target-side default knowledge probability, the risk of non-translatability, or existing semantic signatures, the system marks the segment as "suspected pre-trained mapping". This marking is determined by at least one computable trigger condition among low-frequency high-confidence conflict, inconsistent context distribution, semantic signature conflict, abnormal target-side default knowledge probability, inconsistent non-translatability risk, or abnormal gating state, without relying on inferences about subjective reasons from the language processing unit. The system then submits the candidate segment to the target-side understanding verification for independent confirmation, or marks it as "pre-trained mapping warning" in the source-side interpretation object to prompt the target-side verification to strengthen the review. The preset frequency threshold and preset confidence threshold are functionally described: the preset frequency threshold is configured such that candidate text segments with a frequency below the threshold do not have a stable contextual distribution and semantic mapping in the general corpus; the preset confidence threshold is configured such that candidate text segments with a confidence level above the threshold have high output stability in a single inference. However, when the two are combined, a mapping suspicion flag is triggered—rather than a fixed value—to prevent competitors from designing bypasses of the parameters based on fixed values. This mechanism prevents the language processing unit from bypassing the source-specific content detection process based solely on superficial associations.
[0089] In another embodiment, the system can record clarification history, interpretation versions, and target verification results for auditing purposes. For professional domain texts, the system can retain records of terminology sources, interpretation sources, and human verification. For publicly disseminated texts, the system can record target adaptation evaluations and rejected candidates to avoid subsequently generating the same high-risk expressions repeatedly.
[0090] Example 1: Translation of Idioms and Allusions. Suppose the source text contains an idiom whose literal meaning differs significantly from its actual pragmatic function. The system first detects that the idiom has a high risk of being translated literally and a high degree of cultural dependence, therefore generating a structured source-side explanation object. This explanation object records the original idiom, its literal meaning, the source of the allusion, its true meaning, its pragmatic function, and the risk of not being able to translate it literally, and provides a popular explanation.
[0091] After receiving the colloquial, mediative semantic representation and the explanatory object, the target-side understanding module finds that a literal translation candidate would mislead the target recipient into believing that the expression describes a real event rather than the source's evaluative attitude; therefore, the semantic consistency does not meet the threshold. The target side generates a clarification request, asking the source to supplement the idiom's connotation in the current context and whether the source of the idiom should be retained. After the source returns an updated object, the target side selects either "explanatory translation" or "functionally equivalent expression" based on the dissemination purpose, and outputs the final translation after the accuracy threshold is met.
[0092] Example 2: Translation of Internet Slang. Assume the source text contains a popular internet slang term. This expression may not have a clear literal meaning, but it represents self-deprecation, banter, or irony within a specific community. The system detects that its rarity, context dependence, and probability of default knowledge in the target context are all high, thus generating a structured source-side explanation object. This object records the expression's online origin, user group, emotional intensity, pragmatic function, and the reason for prohibiting literal translation.
[0093] The target-side expression understanding module generates several target-side expression mapping candidates. These candidates serve only as internal verification intermediates, candidate strategy identifiers, or restricted slot mappings. If a candidate, while idiomatic, has an aggressive tone, the system classifies it as a failed candidate and prohibits its output. If another candidate retains self-deprecating and mildly satirical elements and has a high target-side adaptation rating, the system allows it to proceed to the expression generation stage. After output, the system records the target-side expression mapping in the target-side expression record and establishes an association with the source-side interpretation object.
[0094] Example 3: Translation of Technical Terms and Invented Terms. Suppose an inventive term appears in an internal company document, and there is no publicly available standard translation for this term. The source-side explanation module cannot directly obtain an explanation from a public terminology database, therefore generating a supplementary explanation request. The source-side user explains the naming intent, business scope, differences from existing terms, and prohibited translations for the inventive term. The system converts the supplementary explanation into a structured source-side explanation object and records it in the source-side explanation record.
[0095] The target-side understanding and expression module generates target-side expression candidates based on the interpretation object. If a target-side candidate conflicts with existing industry terminology, the target-side fit evaluation in the target-side verification result is reduced, and the system generates a clarification request or requests the source end to confirm whether the new translation is allowed. After confirmation, the system outputs the natural expression result in the target language and maps and records this expression to the organization-level target-side expression record for reuse in subsequent similar texts.
[0096] Example 4: Translation of Classical Chinese or Ancient Script Expressions. Assume the source language text contains Classical Chinese or ancient script expressions. The system first identifies this as source-specific content and then invokes authoritative interpretive resources or source-specific interpretation models for modern semantic interpretation. The structured source-specific interpretation object records the original text, word-by-word definitions, syntactic relationships, modern semantics, allusions, tone, and text style. Popular intermediary semantic representation uses modern semantic interpretation instead of direct word-for-word translation.
[0097] The target-side understanding and expression module does not directly translate the original classical Chinese sentences, but rather performs understanding and verification based on modern semantic interpretation and stylistic information. If the target-side candidate is overly modernized and loses its classical style, the system can choose to retain the original text with annotations or adopt a target language expression with a classical feel; if the target-side candidate retains the style but deviates from the semantics, the system prioritizes ensuring semantic accuracy and then adjusts the style.
[0098] Explanation of the differences from existing technologies: Compared to methods that simply translate subtle cultural differences, existing cultural translation methods typically focus on how to convert source language structure, cultural context, or information format into target language expression. The core of this invention lies in generating a structured source-end interpretation object, which is then understood, verified, and gated for accuracy by the target end before the expression is generated.
[0099] Compared to ordinary terminology databases and translation memories, which typically store standard terminology pairs and reuse translated historical sentence segments, the source-end interpretation record of this invention stores structured interpretation objects specific to the source, while the target-end expression record stores validated expression strategies and expression mappings within a specific target context. The two are heterogeneously linked through semantic signatures, contextual identifiers, and language pair identifiers, rather than a simple one-to-one bilingual comparison.
[0100] Compared to general retrieval-enhanced translation or knowledge graph-enhanced translation, retrieval-enhanced translation typically uses retrieval results as model context to directly participate in the translation. Even when using retrieval services, this invention requires that the retrieval results be organized as source-end interpretable objects and that the target end perform understanding verification and necessary clarification requests. Retrieval services are optional information sources, not the main inventive point of this invention.
[0101] Compared to typical multi-agent translation workflows, which can simulate translators, editors, localization experts, and reviewers in a translation company, the role division in this invention is not a typical workflow role. Instead, it is a computational closed loop formed between the source-end interpretation module and the target-end understanding and expression module, centered around structured interpretation objects, popular mediating semantic representations, clarification protocols, and accuracy thresholds.
[0102] Compared to ordinary ambiguity resolution or reverse translation selection, which focuses on issues such as word meaning, gender, candidate translations, or user selection, this invention focuses on how source-specific content is interpreted, verified, clarified, and recorded when the target lacks default background knowledge. It particularly emphasizes field-level supplementation of the source interpretation object when the target has low confidence, rather than simply selecting from several candidate translations.
[0103] Alternative implementation methods are available. This invention can be implemented on the server side, client side, in a local deployment environment, in a cloud service, in a private enterprise environment, or in a hybrid environment. The system can be deployed as a plugin, translation platform, content management system extension, subtitle generation tool, cross-border customer service system, knowledge base system, educational platform, or API service.
[0104] This invention can be used for text translation, and can also be extended to translation after speech-to-text transcription, subtitle translation, translation after image text recognition, document translation, webpage translation, chat message translation, and semantic translation of text in multimodal content. In multimodal scenarios, images, videos, audio, or contextual metadata can serve as contextual information used by the target end or auxiliary information generated by the source end's interpretive object. However, the core of this invention remains the closed loop of source end interpretive object and target end understanding and verification. The source end interpretive object, colloquial mediator semantic representation, target end verification result, structured clarification request, and gating state can be represented by field objects, semantic graphs, slot sets, vector references, constrained prompt context, interface messages, or state register records, respectively. The above data carriers are equivalent implementations and do not change the processing flow of this invention.
[0105] The modules of this invention can be implemented by software, hardware, or a combination of both. The computer program can be stored in a non-transitory computer-readable storage medium and executed by a processor to implement the method of this invention. The model, rule, knowledge base, and record modules in the system can be executed on the same device or deployed in a distributed manner.
[0106] In conclusion, this invention restructures the translation of cross-language-specific content from a process of "directly generating a target language translation" into a closed-loop process of "first explicitly explaining the special meaning of the source language, then verifying whether the target language understands it, clarifying across language boundaries when necessary, and finally generating a natural expression for the target language" by introducing structured source-end interpretation objects, colloquial mediation semantic representation, target-end understanding verification, structured clarification requests, accuracy thresholds, and heterogeneous records from both ends. This closed loop can cover various implementation methods such as dual-model, single-model multi-routing, external knowledge services, human-computer collaborative supplementation, and multi-target language concurrency, and provides an executable technical foundation for the accurate cross-language dissemination of professional content, cultural content, online content, self-created content, and classical Chinese content.
Claims
1. A method for controlling the accuracy of cross-language-specific content based on source-end interpretation object and target-end understanding verification, applied to one or more computing devices, characterized in that, include: Obtain the text to be processed expressed in the source language, the target language identifier, and the target context information; The text to be processed is subjected to source-specific content detection to obtain at least one candidate text fragment and a specificity score corresponding to the candidate text fragment. The specificity score is determined based at least on the direct translation risk and the target-side default knowledge probability of the candidate text fragment, and in combination with at least one of the rarity, cultural dependence or contextual dependence of the candidate text fragment in the source language. When the specificity score meets the preset candidate conditions, a structured source-end interpretation object is generated for the candidate text fragment. The structured source-end interpretation object includes at least the following fields: object identifier, original text fragment, category label, context window, literal meaning, implied meaning, pragmatic function, popular explanation, risk of non-literal translation, and explanation confidence. Based on the structured source-end interpretation object, the text to be processed is converted into a colloquial intermediary semantic representation, wherein the colloquial intermediary semantic representation contains a reference to the structured source-end interpretation object, and the source-end specific content is explicitly transformed into verifiable semantic information before entering the target language expression generation. The target-side understanding and expression module performs target-side understanding verification based on the common mediating semantic representation, the structured source-side interpretation object, and interface data including object identifier, semantic signature, context identifier, and gating state fields, generating target-side verification results. The target-side verification results include at least understanding confidence, semantic consistency, target-side adaptation evaluation, target-side expression mapping candidates, and gating results. The target-side expression mapping candidates, in any gating state, are only used as internal verification intermediates, candidate strategy identifiers, or restricted slot mappings, and are not output as the final natural expression result of the target language. The target-side adaptation evaluation includes scoring results with target audience, target-side usage context, expression tone consistency, and risk tolerance as inputs. The gating results are output by the gating state machine and include states of allowing generation, blocking generation, or downgrading processing. When the understanding confidence, semantic consistency, target adaptation evaluation, non-translatable risk coverage status, or conflict field handling status in the target verification results do not meet the preset accuracy threshold, a structured clarification request is generated and sent to the source interpretation module. The structured clarification request includes at least the source interpretation object identifier, conflict field, missing context, target hypothesis, failed expression mapping candidate, request supplement field, priority, and deadline. In response to the structured clarification request, the source-end interpretation module performs supplementation, correction, or version update on the corresponding structured source-end interpretation object, and provides the updated structured source-end interpretation object to the target-end understanding and expression module to re-execute the target-end understanding verification; After the target-side verification result meets the preset accuracy threshold and the gating state machine outputs a state allowing generation, the target-side understanding and expression module generates a natural expression result in the target language based on the colloquial mediating semantic representation, the structured source-side interpretation object, and the target-side contextual information; and The structured source-end interpretation object or its updated version is recorded in the source-end interpretation record, and the target language natural expression result, target-end expression strategy or target-end expression mapping is recorded in the target-end expression record. At the same time, the target-end understanding verification, accuracy threshold judgment, structured clarification request and gated state transition are written into the state register record, so that the source-end interpretation record and the target-end expression record are associated through at least one of semantic signature, context identifier or language pair identifier. The target-side understanding verification and the preset accuracy threshold determination are prerequisites for starting the target language natural expression result generation process. The output of the target-side understanding verification module serves as the pre-input and gating state of the expression generation module. When the gating result does not indicate that generation is allowed, the expression generation module remains in a blocked state, or only triggers a downgrade process or manual review process, without generating the final target language natural expression result that has not passed verification, thereby reducing invalid expression generation calls and maintaining traceability of the processing status.
2. The method according to claim 1, characterized in that, The source-end specific content detection includes: calculating the rarity, cultural dependence, context dependence, literal translation risk, and target-end default knowledge probability for candidate text segments in the text to be processed; using the literal translation risk and target-end default knowledge probability as necessary scoring factors; and combining them with at least one of the rarity, cultural dependence, or context dependence for weighted fusion to obtain the specificity score; wherein the weight parameters of the weighted fusion are derived from at least one of a preset rule base, language pair configuration, domain configuration, manually labeled samples, historical mistranslation records, or manually calibrated records; the specificity score also includes context anomaly detection results. When there is a statistical anomaly between the context window features of the candidate text segment and the overall context distribution of the text to be processed, the context anomaly detection results are used as a weighting factor or adjustment coefficient for the specificity score.
3. The method according to claim 1, characterized in that, The source-specific content includes at least one of the following: idioms, allusions, classical Chinese or ancient script expressions, local customs expressions, internet slang, group-conventional terms, self-coined words, names of new things, internal organizational terms, professional terms, fixed expressions with pragmatic implied meanings, and other source-specific expressions that satisfy the target-side default knowledge probability being higher than a threshold and have the risk of literal translation or pragmatic deviation.
4. The method according to claim 1, characterized in that, In addition to including fields such as object identifier, original text fragment, category label, context window, literal meaning, implied meaning, pragmatic function, popular explanation, risk of non-literal translation, and explanation confidence in the structured source-end explanation object, the structured source-end explanation object also includes at least one of the following enhanced fields: specific source, usage scenario, emotional intensity, rhetorical type, suggestion for retaining source-end specificity, degree of substitutability, reason for prohibiting literal translation, explanation source, explanation version number, or update time. When the enhanced field is included, its field content meets the content sufficiency requirements that match the corresponding field name and purpose, including at least one of the following: field value is not empty, field format conforms to preset constraints, and field content has a semantic relationship with the original text fragment.
5. The method according to claim 1, characterized in that, The colloquial mediation semantic representation includes the semantic representation of non-source-specific content in the text to be processed and object references to the structured source-end interpretation object; wherein, the colloquial mediation semantic representation does not use the final target language natural expression candidate as an input condition, and the target-end expression mapping candidate only participates in the gating calculation as an internal verification intermediate quantity, candidate strategy identifier or restricted slot mapping, so that the generation of the target language natural expression result is placed after the target-end understanding verification.
6. The method according to claim 1, characterized in that, The target-side understanding verification includes: converting the colloquial mediation semantic representation and the structured source-side interpretation object into target-side evaluable semantic units; calculating the understanding confidence, semantic consistency, and target-side fit evaluation for each target-side evaluable semantic unit, wherein the target-side fit evaluation is calculated based at least on the target audience, the target-side usage context, the consistency of expression tone, and the risk tolerance to obtain a scoring result; limiting the target-side expression mapping candidates to internal verification intermediate quantities and using them to identify whether the candidate strategies cover the risk of non-translatability; and generating gating results that allow generation, block generation, or downgrade processing based on the understanding confidence, semantic consistency, target-side fit evaluation, non-translatability risk coverage status, and conflict field handling status.
7. The method according to claim 1, characterized in that, The preset accuracy thresholds include: comprehension confidence not lower than a first threshold, semantic consistency not lower than a second threshold, target-side adaptation evaluation not lower than a third threshold, non-translatable risk being covered by the target-side expression strategy, and no unprocessed conflict fields; wherein the first, second, and third thresholds are determined by language pair configuration, domain configuration, historical mistranslation records, manual calibration records, or risk level configuration, respectively; when any of the following conditions are met: comprehension confidence lower than the first threshold, semantic consistency lower than the second threshold, target-side adaptation evaluation lower than the third threshold, non-translatable risk not being covered, semantic signature conflict, or unprocessed conflict fields, the preset accuracy thresholds are determined to be unmet and the gating state is set to block generation or clarify pending processing.
8. The method according to claim 1, characterized in that, The structured clarification request is generated according to a preset clarification protocol, which includes at least two types of messages: request message, supplementary explanation response message, update object message, confirmation message, deadline condition message, and manual review status message. The supplementary explanation response message includes at least one of the following: supplementary explanation for the conflicting field or requested supplementary field, corrected pragmatic function, supplementary context, prohibited or recommended expression strategies, and updated explanation version number. The failed expression mapping candidates in the structured clarification request are internal verification intermediate quantities that have not passed gating, used to explain the cause of the conflict but not as external output content.
9. The method according to claim 1, characterized in that, During the process of re-executing the target end understanding verification, if the target end verification result fails to meet the preset accuracy threshold continuously and reaches the preset number of rounds, preset time consumption, or preset reliability benefit lower limit, the cutoff condition is triggered. When the cutoff condition is triggered, the output is at least one of the following degradation processing results: retaining the source expression with additional explanation, requesting manual supplementary explanation, reducing the naturalization degree of the expression, or marking high-risk translations; wherein, the degradation processing result carries a failed gating status and a manual review mark, the target expression mapping candidate still retains the internal verification intermediate quantity attribute, and does not trigger the expression generation module to output the final target language natural expression result that has not passed verification.
10. The method according to claim 1, characterized in that, The source-end interpretation record is used to store interpretation objects with source-end unique content, and the target-end expression record is used to store the target language natural expression results, target-end expression strategies, or target-end expression mappings. The source-end interpretation record and the target-end expression record are not the same bilingual terminology, but are respectively indexed by the source-end interpretation object identifier and the target-end expression mapping identifier, and heterogeneous associations are established through at least one of semantic signature, context vector, language pair identifier, usage scenario, or confidence level.
11. The method according to claim 1, characterized in that, Generating the natural expression result of the target language includes: selecting at least one expression strategy from literal translation, explanatory translation, functional equivalence expression, target analogy expression, retaining the original word with annotation, transliteration with explanation, and replacement of the target language's common expression, based on the content type, dissemination purpose, target audience, degree of retention of source-end specificity, target-end usage context information, and pragmatic function.
12. The method according to claim 1, characterized in that, The source-end interpretation module and the target-end understanding and expression module are each independently implemented by at least one language processing unit. The language processing unit is functionally configured to perform at least one of the following tasks: generating a structured source-end interpretation object, verifying the target-end understanding, or generating a natural expression of the target language. It exchanges the colloquial mediating semantic representation, the target-end verification result, or the structured clarification request between the source-end interpretation module and the target-end understanding and expression module through interface data containing object identifiers, semantic signatures, context identifiers, and gating states. The source-end specific content detection, structured source-end interpretation object generation, target-end understanding verification, preset accuracy threshold determination, and cross-end clarification loop each generate a recordable intermediate state. When the process is not completed or the verification fails, the gating state is set to block generation or downgrade processing.
13. The method according to claim 12, characterized in that, When the source-end interpretation module and the target-end understanding and expression module are implemented by different functional routes in the same language processing unit, the different functional routes execute distinct and independent processing stages in computational logic and respectively execute the structured source-end interpretation object generation task and the target-end understanding verification and expression generation task. The intermediate output results of each stage—including the structured source-end interpretation object, the colloquial mediation semantic representation, the target-end verification result, the structured clarification request, and the gating state—are generated, transmitted, and recorded in the interface log or state register during the process execution, rather than simply adding metadata tags to the final output.
14. The method according to claim 12, characterized in that, When all or part of the functions of the source-end interpretation module are implemented by an external knowledge service, the external knowledge service is limited to providing interpretation information, background information or structured interpretation object field completion information within the scope of the source language, and does not use the natural expression result of the target language as the final output of the target-end understanding and expression module.
15. The method according to claim 1, characterized in that, At least one of the source-end interpretation record and the target-end expression record is a session-level temporary memory, a user-level persistent memory, an organization-level shared memory, or a multi-user aggregated memory. Regardless of whether the source-end interpretation record and the target-end expression record are in transient or persistent form, they maintain the heterogeneous association characteristic that the source-end interpretation record stores structured source-end interpretation objects, and the target-end expression record stores natural expression results of the target language or target-end expression strategies. Furthermore, the record items are updated with confidence or fixed status based on at least one of the following: hit count, user active confirmation signal, target-end verification result, manual correction signal, or time decay factor. The user active confirmation signal refers to an active confirmation signal issued by the source-end user through an explicit confirmation operation, rather than an indirect signal implicitly inferred from user behavior feedback or interaction history.
16. The method according to claim 1, characterized in that, The method further includes: when the source-end interpretation module fails to identify a certain text fragment as source-end specific content, and the target-end understanding and expression module's understanding confidence or target-end adaptation evaluation for the text fragment is lower than a preset threshold, generating a source-end supplementary interpretation request, receiving supplementary interpretation input by the source-end user or domain expert, and generating or updating the corresponding structured source-end interpretation object based on the supplementary interpretation.
17. The method according to claim 1, characterized in that, When the target language identifier corresponds to multiple target languages, the method provides the same structured source-end interpretation object and the same common mediating semantic representation to multiple target-end understanding and expression modules respectively, so that each target-end understanding and expression module performs target-end understanding verification, clarification request generation and target language natural expression result generation respectively.
18. The method according to claim 1, characterized in that, At least one of the source-specific content detection, target-side understanding verification, structured clarification request generation, or record update is further adjusted based on an evaluation metric associated with the gating result, which is at least one of the following: mistranslation rate, semantic consistency, clarification benefit, target-side comprehensibility, manual correction rate, or user confirmation rate. The adjusted parameters are then used to update candidate conditions, accuracy thresholds, clarification priorities, or record confidence.
19. The method according to claim 1, characterized in that, The method further includes: receiving source-side proprietary knowledge data actively input by the source-side user; organizing the source-side proprietary knowledge data into at least one structured source-side explanation object or supplementary field of the source-side explanation object; and storing it in the source-side knowledge base; wherein the source-side proprietary knowledge data includes source-side-specific content that needs to be actively provided or confirmed by the source-side user when there is a general knowledge service miss, domain access restrictions, commercial access restrictions, the probability of the target-side default knowledge is higher than a threshold, or there is a conflict in the confidence of existing explanations; and the source-side proprietary knowledge data is not merely a dual-source knowledge base that stores the correspondence between source words and target words. The terminology is not a glossary, but rather includes at least one of the following: meaning explanation, applicable context, or conventional source of source-specific content. Furthermore, the source-specific knowledge data is acquired through active uploading, input, or confirmation by the source user, rather than implicitly inferred from user behavior feedback, interaction history, or model fine-tuning results. During the source-specific content detection process, when a candidate text fragment in the text to be processed matches the source-specific knowledge data, a corresponding structured source-specific explanation object is generated or supplemented based on the source-specific knowledge data, and the target-side understanding verification module continues to perform target-side understanding verification.
20. The method according to claim 1, characterized in that, The source-side specific content detection also includes pre-trained knowledge mapping verification: when the frequency of a candidate text fragment in the general corpus is lower than a preset frequency threshold and the confidence level of the language processing unit's output of its meaning is higher than a preset confidence threshold, or when the output meaning of the candidate text fragment is inconsistent with the current context window, the target-side default knowledge probability, the risk of non-translatability, or the existing semantic signature, the candidate text fragment is marked as having a suspicious pre-trained mapping; wherein, the suspicious pre-trained mapping is determined by at least one calculable trigger condition among low-frequency high-confidence conflict, inconsistent context distribution, semantic signature conflict, abnormal target-side default knowledge probability, or inconsistent risk of non-translatability; the system submits the candidate text fragment to the target-side understanding verification for independent confirmation, or records a pre-trained mapping warning field in the corresponding structured source-side interpretation object.
21. A cross-language-specific content accuracy control system based on source-end interpretation object and target-end understanding verification, characterized in that, include: The input acquisition module is configured to acquire the text to be processed in the source language, the target language identifier, and the target context information. The source detection module is configured to perform source-specific content detection on the text to be processed, and obtain at least one candidate text segment and a specificity score corresponding to the candidate text segment. The specificity score is determined based at least on the direct translation risk and the target default knowledge probability, and combined with at least one of rarity, cultural dependence or context dependence. The source-end interpretation object generation module is configured to generate a structured source-end interpretation object for the candidate text fragment when the specificity score meets the preset candidate conditions. The structured source-end interpretation object includes at least the fields of object identifier, original text fragment, category label, context window, literal meaning, implied meaning, pragmatic function, colloquial interpretation, risk of non-literal translation, and interpretation confidence. The module then converts the text to be processed into a colloquial mediation semantic representation based on the structured source-end interpretation object. The target-side understanding verification module is configured to perform target-side understanding verification based on the common mediator semantic representation and the structured source-side interpretation object, and generate target-side verification results including understanding confidence, semantic consistency, target-side adaptation evaluation, target-side expression mapping candidates and gating results, wherein the gating results include states of allowing generation, blocking generation or downgrading processing; The clarification protocol module is configured to generate a structured clarification request that includes at least the source explanation object identifier, conflict field, missing context, target assumption, failed expression mapping candidate, request supplement field, priority and deadline when the target verification result does not meet the preset accuracy threshold. The module provides the structured clarification request to the source explanation object generation module and receives the updated structured source explanation object. The expression generation module is configured to, after the target-side verification result meets the preset accuracy threshold and the gating result indicates that generation is allowed, generate a natural expression result in the target language based on the colloquial mediating semantic representation, the structured source-side interpretation object, and the target-side contextual information; and The heterogeneous recording module is configured to record the structured source-end interpretation object or its updated version, the target-end verification result, the gating result, and the target language natural expression result, the target-end expression strategy, or the target-end expression mapping, so that the source-end interpretation record and the target-end expression record establish a heterogeneous association through at least one of semantic signature, context identifier, or language pair identifier.
22. The system according to claim 21, characterized in that, The source detection module, source interpretation object generation module, target understanding verification module, clarification protocol module, and expression generation module are each independently implemented by at least one language processing unit. Each language processing unit includes at least one of a natural language processing model, a rule engine, a retrieval service, or an external knowledge service. The system transmits interface data containing object identifiers, semantic signatures, context identifiers, gating states, and state transition identifiers between modules, and records the intermediate states of source-specific content detection, structured source interpretation object generation, target understanding verification, preset accuracy threshold determination, and cross-end clarification loops. The system also writes the interface data, intermediate states, and gating state transitions into interface logs or state register records.
23. The system according to claim 21, characterized in that, The clarification protocol module is configured to generate a structured clarification request that includes a request identifier, a source-side explanation object identifier, a conflict field, a missing context, a target-side hypothesis, a failed expression mapping candidate, a request supplement field, a priority, and a deadline.
24. The system according to claim 21, characterized in that, The heterogeneous recording module includes a source-side interpretation object library, a target-side expression mapping library, and a state register record. The source-side interpretation object library stores object identifiers, original text fragments, category labels, context windows, literal meanings, implied meanings, pragmatic functions, popular explanations, risks of non-literal translation, and interpretation confidence fields. The target-side expression mapping library stores the target language natural expression results, expression strategies, applicable scenarios, target-side adaptation evaluation, usage frequency, and confirmation status. The state register record stores the target-side understanding verification results, gated state transitions, cross-language verification inconsistency markers, and rejected internal expression mapping candidates.
25. The system according to claim 21, characterized in that, It also includes a human-machine collaboration supplement module, which is configured to send a supplementary explanation request to the source user, domain expert, or administrator when the target end fails to understand and verify and the source end detection module fails to identify source-specific content, and convert the received supplementary explanation into a structured source end explanation object or its field update; when the received supplementary explanation has a semantic conflict with the existing structured source end explanation object or system built-in knowledge, a cross-validation process is triggered, requiring at least one additional source end user or domain expert to independently confirm the supplementary explanation.
26. The system according to claim 21, characterized in that, It also includes a source-side knowledge base management module, which is configured to receive source-side proprietary knowledge data actively input by source-side users, organize the source-side proprietary knowledge data into at least one structured source-side interpretation object or source-side interpretation object supplementary fields, and store it in the source-side knowledge base. The source-side proprietary knowledge data includes source-side-specific content that needs to be actively provided or confirmed by the source-side user when general knowledge services are not hit, domain access permissions are restricted, commercial access permissions are restricted, the probability of the target-side default knowledge is higher than the threshold, or there is a conflict in the confidence of existing explanations. It is not a bilingual glossary that only stores the correspondence between source words and target words. The source-side proprietary knowledge data is obtained by the source-side user actively uploading, inputting, or confirming it, rather than being implicitly inferred from user behavior feedback, interaction history, or model fine-tuning results. The source-side knowledge base management module performs input validation, format validation, field validity validation, permission validation, and injection detection on the received source-side proprietary knowledge data. Knowledge data that fails validation is rejected and an alarm is triggered. When the entered knowledge data is invoked, a call log and conflict detection record are generated. During the source-side specific content detection process, when a candidate text fragment in the text to be processed matches the proprietary knowledge data in the source-side knowledge base, the source-side detection module generates or supplements a corresponding structured source-side interpretation object based on the proprietary knowledge data, and causes the target-side understanding verification module to continue to perform target-side understanding verification.
27. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it causes one or more computing devices to implement the method of any one of claims 1 to 20, or to implement the functions of the input acquisition module, source detection module, source interpretation object generation module, target understanding verification module, clarification protocol module, expression generation module, heterogeneous record module, human-computer collaboration supplement module, or source knowledge base management module in the system of any one of claims 21 to 26.