An artificial intelligence translation optimization device for professional terms
By employing cross-modal fusion and domain adaptation technologies, the problems of ambiguous polysemous terms and confusion of homographs across domains in existing translation devices have been solved, achieving accuracy and coherence in the translation of professional terms and reducing the cost of manual correction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANWEI DINGLI TECHNOLOGY CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-21
Smart Images

Figure CN122433754A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence translation technology, specifically to an artificial intelligence translation optimization device for specialized terminology. Background Technology
[0002] As a core support for cross-language communication, artificial intelligence translation technology has been widely applied in on-site communication and text conversion scenarios in professional fields such as engineering construction, medical surgery, financial transactions, legal consultation, and circuit research and development. Portable translation devices designed for professional scenarios have become important equipment for solving cross-language collaboration barriers in professional fields. The accuracy of their translation of professional terms directly affects the efficiency of professional communication and the implementation of results.
[0003] Most existing professional translation devices rely on general translation models and basic terminology databases to achieve translation functions. They mainly depend on the single textual semantic features of the source language text for terminology recognition and translation matching, without combining on-site visual scene information to construct scene semantic constraints. At the same time, the general models have not been specifically optimized for vertical fields such as circuit design, orthopedic surgery, and financial derivatives. They are not capable of distinguishing the meanings of polysemous professional terms and are difficult to match the standard terminology translations corresponding to the scene. This leads to problems such as ambiguity and non-standard translations in the translation of professional terms, and fails to meet the accuracy requirements of terminology translation in professional scenarios. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides an AI-powered translation optimization device for technical terms. This device solves the problems of existing technical translation devices, which rely solely on text semantics to resolve ambiguities of polysemous terms, easily lead to translation confusion of homographs across disciplines, and lack of accuracy and scenario adaptability in technical terminology translation, as well as the inability to continuously iterate and optimize.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an AI translation optimization device for technical terminology, comprising a main body of the translation device, wherein the top of the main body of the translation device is provided with a touch screen for displaying an interactive interface and buttons for inputting operation commands, arranged sequentially from back to front; both the front and rear sides of the top of the main body of the translation device are fixedly connected to microphones, the rear microphone is located at the rear of the display screen to record the other party's speech, and the front microphone is located at the front of the buttons to record our own speech; the rear left side and the front right side of the main body of the translation device are provided with amplifiers; a switch is provided at the front left side of the main body of the translation device; and a camera for scanning and acquiring visual text and scene images is fixedly connected to the rear of the main body of the translation device. The translation device has a processor and a memory inside. The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, it implements the following functional modules: The voice acquisition and recognition module is used to acquire voice signals through the microphone and convert them into source language text. A visual scene acquisition module is used to acquire real-time visual scene information through the camera; The scene semantic extraction module is used to extract scene semantic constraint vectors and scene semantic topic tags based on the visual scene information; The terminology candidate recognition module is used to identify candidate terms in the source language text and generate candidate translations and their domain knowledge graph vectors for each candidate term. The cross-modal disambiguation matching module is used to perform cross-modal similarity calculation on the candidate translation terms based on the scene semantic constraint vector, the scene semantic topic tag and the domain knowledge graph vector, so as to determine the target term translation; The translation synthesis module is used to generate the final translation result based on the target term translation.
[0006] Preferably, the formula for cross-modal similarity calculation performed by the cross-modal disambiguation matching module is as follows: , in, For the first The overall matching score of each candidate translation term; This is a scene semantic constraint vector; For the first Domain knowledge graph vectors of candidate translations; Indicates cosine similarity; A collection of scene semantic topic tags; This is the set of domain labels for the candidate translations. express Similarity coefficient; A vector of scene entity categories; This is the entity category vector corresponding to the candidate translated meaning; For a category matching indicator function, if and only if and The value is 1 if the match is consistent, otherwise it is 0. All are preset weighting coefficients, satisfying ,and The value ranges from 0.4 to 0.6. The value ranges from 0.2 to 0.3. The value ranges from 0.2 to 0.3.
[0007] Preferably, the scene semantic extraction module further includes a scene topic tag generation unit, used to output the scene semantic topic tags, wherein the scene semantic topic tags are selected from at least one of circuit design, network communication, orthopedic surgery, financial derivatives, legal contracts, and machining; the scene semantic constraint vector is updated synchronously when the camera detects a change in the translation scene.
[0008] Preferably, the processor further includes the following when executing computer program instructions: The domain model switching module is used to dynamically activate the corresponding domain-adaptive translation sub-model based on the scene semantic topic tags output by the scene semantic extraction module. Multiple domain-adaptive translation sub-models are provided, each of which is obtained by fine-tuning and training a general translation base model with bilingual parallel corpora of a specific domain and stored in the memory. A domain terminology database, corresponding one-to-one with each domain-adapted translation sub-model, is used to store the standard translations of professional terms in the corresponding domain.
[0009] Preferably, the domain model switching module further includes a scene switching transition processing unit for maintaining a cross-domain terminology homomorphic conflict resolution table; when the scene semantic topic label changes, the scene switching transition processing unit explicitly marks the homomorphic conflict terms appearing in the current translation result and generates transition prompt information through the touch screen.
[0010] Preferably, the term candidate recognition module includes: The term polysemy determination unit is used to determine whether there are multiple unrelated translation meanings of words in the source language text to be translated by querying the polysemous term knowledge base preset in the memory. The domain knowledge graph retrieval unit is used to retrieve the associated domain category, superordinate concept and synonym information of candidate terms that have been determined to be ambiguous in the domain knowledge graph, and generate the domain knowledge graph vector of the candidate translation meaning.
[0011] Preferably, the visual scene acquisition module includes: An image frame extraction unit is used to extract key image frames from the real-time video stream captured by the camera at a preset sampling frequency. The text recognition unit is used to perform optical character recognition on the text content appearing in the key image frame and generate visual text information. A visual text vectorization unit is used to encode the visual text information into a text feature vector; The visual scene feature extraction unit is used to extract scene-level visual feature vectors from the key image frames.
[0012] Preferably, when the processor executes the computer program instructions, it further includes a manual confirmation module, which is used to push the candidate terms and their top-ranked candidate translations to the touch screen when the highest comprehensive matching score output by the cross-modal disambiguation matching module is lower than a preset confidence threshold, and receive confirmation from the user through manual correction on the screen or selection by button, and update the corresponding term association weights in the domain knowledge graph in the memory according to the user's operation.
[0013] Preferably, the preset information threshold is set to When the highest similarity score When, the manual confirmation module is triggered; when When the highest similarity score is obtained, the candidate translation term is directly used as the translation of the target term.
[0014] Preferably, the scene semantic extraction module includes a pre-trained multimodal semantic understanding model. The pre-trained multimodal semantic understanding model adopts a dual-tower encoder architecture, including a visual encoder and a text encoder. The visual encoder is used to extract visual feature vectors of visual information, and the text encoder is used to extract text feature vectors of the source language text to be translated. The multimodal semantic understanding model is trained through a cross-modal contrastive learning loss function to maximize the semantic consistency between the visual modality and the text modality in the same scene.
[0015] Working Principle: After the device is activated by the switch, the front microphone captures the user's voice, and the rear microphone captures the other party's voice. These two voice signals are converted into source language text by the voice acquisition and recognition module and sent to the processor. Simultaneously, the rear camera captures a real-time video stream of the surrounding scene. The visual scene acquisition module extracts key image frames from the video stream at a preset frequency and processes them in two paths. One path uses a text recognition unit to recognize text such as device nameplates, screen menus, and document titles in the image, generating visual text information and vectorizing it. The other path uses a visual scene feature extraction unit to extract scene-level visual feature vectors. The processor relies on a multimodal semantic understanding module based on a dual-tower encoder architecture. The system integrates visual and textual features to extract scene semantic constraint vectors and subject tags for professional fields such as circuit design, network communication, orthopedic surgery, financial derivatives, legal contracts, and mechanical processing. The scene semantic constraint vectors are updated synchronously as the scene changes. Next, the terminology candidate recognition module uses a polysemous terminology knowledge base in memory to identify polysemous professional terms in the source language text, and then retrieves relevant information from the domain knowledge graph to generate a domain knowledge graph vector for each candidate translation. Finally, the cross-modal disambiguation matching module combines the scene semantic constraint vectors, scene subject tags, and scene entity categories, and calculates the comprehensive matching score for each candidate translation using a preset weighted formula. A score higher than 0.6 is considered a match. The system directly assigns the highest-scoring translation as the target term translation once the confidence threshold is reached. If the score is below the threshold, a manual confirmation module is triggered, pushing candidate terms and top-ranked alternative translations to the touchscreen. Users can manually correct the translations via the touchscreen or confirm them using buttons. The system updates the term association weights in the domain knowledge graph based on user actions. Simultaneously, the domain model switching module automatically activates the corresponding domain-adaptive translation sub-model and calls the dedicated domain terminology library based on the scene theme tags. If cross-domain homograph conflicts occur during scene switching, the conflicting terms are automatically marked and a transition prompt is generated on the display screen. Finally, the translation synthesis module integrates the accurate target term translation with the domain model translation content to generate a fluent and professional final translation result, which is output through a speaker and text display on the touchscreen. User manual correction data continuously optimizes the device's domain knowledge graph, constantly improving the accuracy of professional terminology translation.
[0016] This invention provides an AI-powered translation optimization device for specialized terminology. It offers the following advantages: 1. This invention achieves cross-modal disambiguation by fusing visual scenes and speech text, and combines scene semantic constraint vectors and domain knowledge graph vectors for weighted matching calculation. This can accurately resolve the problem of polysemy in professional terms and significantly improve the scene adaptability and translation accuracy of professional terminology translation.
[0017] 2. This invention can dynamically activate the corresponding domain-adaptive translation sub-model based on the scene theme tag and call the exclusive domain terminology library. At the same time, it completes the scene switching transition processing through the cross-domain terminology homograph conflict resolution table, effectively avoiding confusion in the translation of cross-domain homograph terms and ensuring the consistency and standardization of translation in multiple professional fields.
[0018] 3. This invention triggers a manual confirmation mechanism by setting a pre-set threshold, and combines this with real-time updates of the domain knowledge graph term association weights to form a human-machine collaborative optimization closed loop. This continuously iterates and improves the device's ability to recognize and translate professional terms, reducing the cost of manual correction for professional translations. Attached Figure Description
[0019] Figure 1 This is a perspective view of the present invention; Figure 2 This is a schematic diagram of the sound amplification hole structure of the present invention; Figure 3 This is a system diagram of the present invention.
[0020] The components include: 1. the main body of the translation device; 2. the display screen; 3. the buttons; 4. the switch; 5. the speaker grille; 6. the microphone grille; and 7. the camera. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example: Please see the appendix Figure 1 - Appendix Figure 3 This invention provides an AI translation optimization device for technical terms, including a translation device body 1. The top of the translation device body 1 is provided with a touch screen 2 for displaying an interactive interface and a button 3 for inputting operation commands, arranged sequentially from back to front. The front and rear sides of the top of the translation device body 1 are fixedly connected to a microphone 6. The rear microphone 6 is located at the rear of the display screen 2 to record the other party's voice, and the front microphone 6 is located at the front of the button 3 to record the user's voice. The rear left side and the front right side of the translation device body 1 are provided with a speaker 5. The front left side of the translation device body 1 is provided with a switch 4. The rear end of the translation device body 1 is fixedly connected to a camera 7 for scanning and collecting visual text and scene images. The main body 1 of the translation device contains a processor and a memory. The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, it implements the following functional modules: The voice acquisition and recognition module is used to acquire voice signals through the microphone hole 6 and convert them into source language text. The visual scene acquisition module is used to acquire real-time visual scene information through camera 7; The scene semantic extraction module is used to extract scene semantic constraint vectors and scene semantic topic tags based on visual scene information; The terminology candidate recognition module is used to identify candidate terms in the source language text and generate candidate translations and their domain knowledge graph vectors for each candidate term. The cross-modal disambiguation matching module is used to calculate the cross-modal similarity of candidate translation terms based on scene semantic constraint vectors, scene semantic topic tags, and domain knowledge graph vectors in order to determine the target term translation. The translation synthesis module is used to generate the final translation result based on the target term translation. Specifically, during runtime, the cross-modal disambiguation matching module first substitutes the scene semantic constraint vector, scene semantic topic tag set, and scene entity category vector output by the scene semantic extraction module, along with the domain knowledge graph vector, domain tag set, and corresponding entity category vector of each candidate translation meaning generated by the terminology candidate recognition module, into the weighted calculation formula. The weight coefficients can be dynamically adjusted according to the scenario requirements. In typical scenarios, α is 0.5, β is 0.25, and γ is 0.25, which strengthens the core leading role of scene semantic constraints while also taking into account the auxiliary verification effect of domain tag matching and entity category consistency. After the calculation is completed, the system automatically sorts the comprehensive matching scores, selects the highest-scoring item and compares it with a preset confidence threshold of 0.6: when the score is ≥0.6, the highest-scoring translation is directly set as the target term translation, without any manual intervention; when the score is <0.6, the manual confirmation module is immediately triggered, and the original text of the candidate term and the top three candidate translations are pushed to the touch screen. Users can manually correct the translation through the touch screen or select confirmation by button. The system will update the term association weights of the domain knowledge graph in the memory in real time according to the user's operation, and continuously optimize the accuracy of subsequent term matching.
[0023] The formula for cross-modal similarity calculation performed by the cross-modal disambiguation matching module is as follows: , in, For the first The overall matching score of each candidate translation term; This is a scene semantic constraint vector; For the first Domain knowledge graph vectors of candidate translations; Indicates cosine similarity; A collection of scene semantic topic tags; This is the set of domain labels for the candidate translations. express Similarity coefficient; A vector of scene entity categories; This is the entity category vector corresponding to the candidate translated meaning; For a category matching indicator function, if and only if and The value is 1 if the match is consistent, otherwise it is 0. All are preset weighting coefficients, satisfying ,and The value ranges from 0.4 to 0.6. The value ranges from 0.2 to 0.3. The value ranges from 0.2 to 0.3; When the processor executes computer program instructions, it also includes a manual confirmation module, which is used to push candidate terms and their top-ranked candidate translations to the touch screen 2 when the highest comprehensive matching score output by the cross-modal disambiguation matching module is lower than the preset confidence threshold. The user can manually correct the term or select confirmation via the button 3, and update the corresponding term association weights in the domain knowledge graph in the memory according to the user's operation. The preset threshold is set to When the highest similarity score When, the manual confirmation module is triggered; when When the highest similarity score is used, the candidate translation term is directly adopted as the translation of the target term. The scene semantic extraction module also includes a scene topic tag generation unit, which outputs scene semantic topic tags. The scene semantic topic tags are selected from at least one of circuit design, network communication, orthopedic surgery, financial derivatives, legal contracts, and machining. The scene semantic constraint vector is updated synchronously when the camera 7 detects a change in the translation scene. When a processor executes computer program instructions, it also includes: The domain model switching module is used to dynamically activate the corresponding domain-adaptive translation sub-model based on the scene semantic topic tags output by the scene semantic extraction module. Multiple domain-adapted translation sub-models are provided. Each domain-adapted translation sub-model is obtained by fine-tuning and training a general translation base model with bilingual parallel corpora of a specific domain and stored in memory. A domain terminology database, which corresponds one-to-one with the translation sub-models adapted to each domain, is used to store the standard translations of professional terms in the corresponding domain. The domain model switching module also includes a scene switching transition processing unit, which is used to maintain a cross-domain terminology homomorphic conflict resolution table. When the scene semantic topic label changes, the scene switching transition processing unit explicitly marks the homomorphic conflict terms that appear in the current translation result and generates transition prompt information through the touch screen 2. Specifically, after receiving the scene semantic topic tag, the domain model switching module quickly retrieves the memory, accurately matches and activates the corresponding professional domain-adaptive translation sub-model. This sub-model is based on a general translation foundation model and fine-tuned through massive bilingual professional parallel corpora in vertical fields such as circuit design, orthopedic surgery, and financial derivatives, possessing the ability to understand and translate terminology specific to its domain. At the same time, the module automatically calls the domain-specific terminology library bound to the sub-model. The terminology library contains bilingual translations of all standard professional terms in the domain, directly replacing non-standard and colloquial expressions in general translation. When the scene topic tag undergoes a cross-domain switch, the scene switching transition processing unit immediately retrieves the cross-domain terminology homograph conflict resolution table, highlights and explicitly marks homophones and homographs in the translation results that have different meanings across domains, and pops up a text transition prompt at the top of the touch screen, clearly informing the user that the current translation domain has been switched and conflicting terms have been marked, completely avoiding confusion in cross-domain terminology translation.
[0024] The terminology candidate identification module includes: The term polysemy determination unit is used to determine whether there are multiple unrelated translation meanings of words in the source language text to be translated by querying a polysemous term knowledge base pre-stored in the memory. The domain knowledge graph retrieval unit is used to retrieve the associated domain categories, superordinate concepts, and synonym information of candidate terms that have been determined to be ambiguous in the domain knowledge graph, and generate domain knowledge graph vectors of candidate translations. Specifically, after receiving the source language text, the terminology polysemy determination unit traverses the text word by word and performs precise string matching with the pre-stored polysemous terminology knowledge base in memory. It quickly determines whether a word has multiple unrelated translated meanings and only sends polysemous terms to subsequent processing, reducing unnecessary computation. The domain knowledge graph retrieval unit, for polysemous candidate terms, deeply retrieves information from the domain knowledge graph, including its professional domain category, higher-level concept hierarchy, synonyms, near-synonyms, and fixed collocations. This multi-dimensional information is fused and encoded into a high-dimensional domain knowledge graph vector. This vector fully carries the semantic features, attribute information, and relationships of the term in the corresponding domain, providing accurate and comprehensive data support for subsequent cross-modal disambiguation matching.
[0025] The visual scene acquisition module includes: The image frame extraction unit is used to extract key image frames from the real-time video stream captured by the camera 7 at a preset sampling frequency. The text recognition unit is used to perform optical character recognition on the text content appearing in key image frames and generate visual text information. Visual text vectorization unit is used to encode visual text information into text feature vectors; The visual scene feature extraction unit is used to extract scene-level visual feature vectors from key image frames. Specifically, the image frame extraction unit selects key image frames with high clarity, complete information, and strong representativeness from the real-time video stream captured by the camera at a preset sampling frequency of 15 frames per second, eliminating blurry, jittery, and redundant invalid frames to reduce the processing load of subsequent modules. The text recognition unit performs high-precision optical character recognition on visual text such as equipment nameplates, operation panel text, document titles, professional logos, and drawing annotations in the key image frames, outputting standardized visual text information without garbled characters or missing parts. The visual text vectorization unit uses a pre-trained language model to encode visual text information into fixed-dimensional text feature vectors, realizing the digital and vectorized representation of visual text information. The visual scene feature extraction unit extracts scene-level visual features from key image frames through a deep convolutional neural network, covering overall visual information such as scene environment, equipment type, operating objects, and spatial layout, generating high-dimensional scene-level visual feature vectors to provide multi-dimensional and comprehensive visual input for scene semantic extraction.
[0026] The scene semantic extraction module includes a pre-trained multimodal semantic understanding model. The pre-trained multimodal semantic understanding model adopts a dual-tower encoder architecture, including a visual encoder and a text encoder. The visual encoder is used to extract visual feature vectors of visual information, and the text encoder is used to extract text feature vectors of the source language text to be translated. The multimodal semantic understanding model is trained by a cross-modal contrastive learning loss function to maximize the semantic consistency between the visual modality and the text modality in the same scene. Specifically, the pre-trained multimodal semantic understanding model employs a dual-tower encoder architecture for both visual and textual modes. The visual encoder, based on a deep residual network, performs deep encoding on the scene-level visual feature vectors output by the visual scene acquisition module, extracting high-level semantic features of the visual modality. The text encoder, based on a Transformer architecture, simultaneously encodes both the source language text and visual text information, extracting core semantic features of the text modality. During model training, a cross-modal contrastive learning loss function is used to force the semantic feature distance between the visual and textual modalities within the same scene to be closer, and to widen the feature distance between different scenes, maximizing the semantic consistency between the visual and textual modalities. During inference, the model integrates dual-modal features to accurately generate scene semantic constraint vectors and professional domain topic labels. When the camera detects a change in the translation scene, the scene semantic constraint vectors are immediately updated synchronously, ensuring the real-time, accuracy, and adaptability of the scene semantic information.
[0027] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An AI-powered translation optimization device for specialized terminology, characterized in that, The device includes a translation device body (1). The top of the translation device body (1) is provided with a touch screen (2) for displaying an interactive interface and a button (3) for inputting operation commands. The front and rear sides of the top of the translation device body (1) are fixedly connected with sound holes (6). The sound hole (6) on the rear side is located at the rear end of the screen (2) to record the other party's voice, and the sound hole (6) on the front side is located at the front end of the button (3) to record our own voice. The rear left side and the front right side of the translation device body (1) are provided with amplification holes (5). The front left side of the translation device body (1) is provided with a switch (4). The rear end of the translation device body (1) is fixedly connected with a camera (7) for scanning and collecting visual text and scene images. The main body (1) of the translation device is equipped with a processor and a memory. The memory stores computer program instructions that can be run on the processor. When the processor executes the computer program instructions, it implements the following functional modules: The voice acquisition and recognition module is used to acquire voice signals through the microphone (6) and convert them into source language text; A visual scene acquisition module is used to acquire real-time visual scene information through the camera (7); The scene semantic extraction module is used to extract scene semantic constraint vectors and scene semantic topic tags based on the visual scene information; The terminology candidate recognition module is used to identify candidate terms in the source language text and generate candidate translations and their domain knowledge graph vectors for each candidate term. The cross-modal disambiguation matching module is used to perform cross-modal similarity calculation on the candidate translation terms based on the scene semantic constraint vector, the scene semantic topic tag and the domain knowledge graph vector, so as to determine the target term translation; The translation synthesis module is used to generate the final translation result based on the target term translation.
2. The artificial intelligence translation optimization device for technical terminology according to claim 1, characterized in that, The formula for cross-modal similarity calculation performed by the cross-modal disambiguation matching module is as follows: , in, For the first The overall matching score of each candidate translation term; This is a scene semantic constraint vector; For the first Domain knowledge graph vectors of candidate translations; Indicates cosine similarity; A collection of scene semantic topic tags; This is the set of domain labels for the candidate translations. express Similarity coefficient; A vector of scene entity categories; This is the entity category vector corresponding to the candidate translated meaning; For a category matching indicator function, if and only if and The value is 1 if the match is consistent, otherwise it is 0. All are preset weighting coefficients, satisfying ,and The value ranges from 0.4 to 0.
6. The value ranges from 0.2 to 0.
3. The value ranges from 0.2 to 0.
3.
3. The artificial intelligence translation optimization device for technical terminology according to claim 2, characterized in that, The scene semantic extraction module also includes a scene theme tag generation unit, which is used to output the scene semantic theme tag. The scene semantic theme tag is selected from at least one of circuit design, network communication, orthopedic surgery, financial derivatives, legal contracts, and mechanical processing. The scene semantic constraint vector is updated synchronously when the camera (7) detects a change in the translation scene.
4. The artificial intelligence translation optimization device for technical terminology according to claim 3, characterized in that, The processor, when executing computer program instructions, also includes: The domain model switching module is used to dynamically activate the corresponding domain-adaptive translation sub-model based on the scene semantic topic tags output by the scene semantic extraction module. Multiple domain-adaptive translation sub-models are provided, each of which is obtained by fine-tuning and training a general translation base model with bilingual parallel corpora of a specific domain and stored in the memory. A domain terminology database, corresponding one-to-one with each domain-adapted translation sub-model, is used to store the standard translations of professional terms in the corresponding domain.
5. The artificial intelligence translation optimization device for technical terminology according to claim 4, characterized in that, The domain model switching module also includes a scene switching transition processing unit, which is used to maintain a cross-domain terminology homomorphic conflict resolution table. When the scene semantic topic label changes, the scene switching transition processing unit explicitly marks the homomorphic conflict terms that appear in the current translation result and generates transition prompt information through the touch screen (2).
6. The artificial intelligence translation optimization device for technical terminology according to claim 1, characterized in that, The terminology candidate identification module includes: The term polysemy determination unit is used to determine whether there are multiple unrelated translation meanings of words in the source language text to be translated by querying the polysemous term knowledge base preset in the memory. The domain knowledge graph retrieval unit is used to retrieve the associated domain category, superordinate concept and synonym information of candidate terms that have been determined to be ambiguous in the domain knowledge graph, and generate the domain knowledge graph vector of the candidate translation meaning.
7. The artificial intelligence translation optimization device for technical terminology according to claim 1, characterized in that, The visual scene acquisition module includes: The image frame extraction unit is used to extract key image frames from the real-time video stream acquired by the camera (7) at a preset sampling frequency; The text recognition unit is used to perform optical character recognition on the text content appearing in the key image frame and generate visual text information. A visual text vectorization unit is used to encode the visual text information into a text feature vector; The visual scene feature extraction unit is used to extract scene-level visual feature vectors from the key image frames.
8. The artificial intelligence translation optimization device for technical terminology according to claim 2, characterized in that, When the processor executes the computer program instructions, it also includes a manual confirmation module, which is used to push the candidate terms and their top-ranked candidate translations to the touch screen (2) when the highest comprehensive matching score output by the cross-modal disambiguation matching module is lower than a preset confidence threshold. The user can manually correct the term or select confirmation by pressing a button (3) on the screen (2), and update the corresponding term association weights in the domain knowledge graph in the memory according to the user's operation.
9. The artificial intelligence translation optimization device for technical terminology according to claim 8, characterized in that, The preset threshold is set to ; When the highest similarity score When, the manual confirmation module is triggered; when When the highest similarity score is obtained, the candidate translation term is directly used as the translation of the target term.
10. The artificial intelligence translation optimization device for technical terminology according to claim 1, characterized in that, The scene semantic extraction module includes a pre-trained multimodal semantic understanding model. The pre-trained multimodal semantic understanding model adopts a dual-tower encoder architecture, including a visual encoder and a text encoder. The visual encoder is used to extract visual feature vectors of visual information, and the text encoder is used to extract text feature vectors of the source language text to be translated. The multimodal semantic understanding model is trained through a cross-modal contrastive learning loss function to maximize the semantic consistency between the visual modality and the text modality in the same scene.