Online interactive translation method based on factory intelligent interaction system and related equipment

By utilizing the online interactive translation method of the factory intelligent interaction system, and employing computing node clusters and voice translation models, the problems of communication delays and insufficient adaptation of professional terminology in cross-language industrial maintenance have been solved. This has enabled low-latency two-way voice interaction and data linkage, thereby improving the efficiency of remote maintenance.

CN122113939APending Publication Date: 2026-05-29SOYO TECH DEV CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOYO TECH DEV CO LTD
Filing Date
2026-04-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing cross-language industrial maintenance solutions, language conversion relies on manual or general translation tools, which leads to communication delays and insufficient adaptation of professional terminology.

Method used

By using an online interactive translation method based on a factory intelligent interaction system, and leveraging intelligent interactive servers, computing node clusters, and speech translation models, real-time speech interaction and translation between on-site interactive terminals and expert interactive terminals are achieved. This includes speech separation, preprocessing, language mapping, and translation, while linking fault data with historical data to optimize the translation process.

Benefits of technology

It enables low-latency two-way voice interaction in cross-language industrial maintenance scenarios, improves the interaction efficiency and diagnostic decision-making efficiency in remote maintenance scenarios, and reduces the load on the intelligent interaction server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113939A_ABST
    Figure CN122113939A_ABST
Patent Text Reader

Abstract

The application provides an online interactive translation method based on a factory intelligent interaction system and related equipment, which is applied to an intelligent interaction server in the factory intelligent interaction system, the factory intelligent interaction system comprising the intelligent interaction server, a field interaction terminal and an expert interaction terminal; the method comprises the following steps: obtaining an interaction request of the field interaction terminal for the expert interaction terminal; determining a first computing node corresponding to the field interaction terminal and a second computing node corresponding to the expert interaction terminal according to the interaction request; performing voice interaction processing between the expert interaction terminal and the field interaction terminal through the intelligent interaction server, the first computing node and the second computing node, wherein the voice interaction processing comprises translation; and when a maintenance result instruction is obtained, ending a maintenance interaction process between the field interaction terminal and the expert interaction terminal according to the maintenance result instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of speech processing technology, specifically relating to an online interactive translation method and related equipment based on a factory intelligent interactive system. Background Technology

[0002] In globalized industrial production scenarios, equipment failure repair in multinational factories often involves cross-language collaboration between on-site operators and remote experts. Such repairs rely on historical data from production management systems and real-time equipment alarm information to support decision-making, and have clear requirements for real-time communication and seamless data interaction. Current industrial interaction systems need to accommodate multilingual real-time communication, multi-source equipment data linkage, and closed-loop management of repair processes to meet the professional and efficient requirements of equipment repair in industrial environments.

[0003] In existing cross-language industrial maintenance solutions, language conversion relies on manual or general translation tools, which leads to communication delays and insufficient adaptation of professional terminology. Summary of the Invention

[0004] This application provides an online interactive translation method and related equipment based on a factory intelligent interactive system, aiming to improve the efficiency of cross-language interaction in remote maintenance scenarios.

[0005] Firstly, this application provides an online interactive translation method based on a factory intelligent interaction system, applied to an intelligent interaction server within the factory intelligent interaction system, wherein the factory intelligent interaction system includes the intelligent interaction server, on-site interaction terminals, and expert interaction terminals; the method includes: Obtain the interaction request from the on-site interactive terminal to the expert interactive terminal; The first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal are determined based on the interactive request. Voice interaction processing between the expert interaction terminal and the on-site interaction terminal is performed through the intelligent interaction server, the first computing node, and the second computing node, and the voice interaction processing includes translation. When a repair result instruction is received, the repair interaction process between the on-site interactive terminal and the expert interactive terminal is terminated according to the repair result instruction.

[0006] In conjunction with the first aspect, in one possible embodiment, the voice interaction processing between the expert interaction terminal and the on-site interaction terminal via the intelligent interaction server, the first computing node, and the second computing node includes: determining a language mapping relationship between the on-site interaction terminal and the expert interaction terminal, the language mapping relationship including a mapping relationship between a first language and a second language, wherein the first language is a preset language of the on-site interaction terminal, and the second language is a preset language of the expert interaction terminal; calling the intelligent interaction server and the first computing node to translate the first voice data in the first language uploaded by the on-site interaction terminal into second voice data in the second language, and forwarding the second voice data to the expert interaction terminal; calling the intelligent interaction server and the second computing node to translate the third voice data in the second language uploaded by the expert interaction terminal into fourth voice data corresponding to the first language, and forwarding the fourth voice data to the on-site interaction terminal.

[0007] In conjunction with the first aspect, in one possible embodiment, the step of calling the intelligent interaction server and the first computing node to translate the first speech data in the first language uploaded by the on-site interaction terminal into the second speech data in the second language includes: obtaining at least one first target clean speech stream fed back by the first computing node; the at least one first target clean speech stream is obtained based on the first speech data through speech separation and preprocessing, wherein the speech separation refers to splitting at least one corresponding first clean speech stream from the first speech data according to the first voiceprint features, the number of first clean speech streams being consistent with the number of first voiceprint features, and the preprocessing refers to standardizing the at least one first clean speech stream to obtain the at least one first target clean speech stream; associating the at least one first target clean speech stream with the first voiceprint features and inputting them into at least one corresponding first speech translation model, wherein the at least one first speech translation model translates the at least one first target clean speech stream into the second speech data.

[0008] In conjunction with the first aspect, in one possible embodiment, the step of calling the intelligent interaction server and the second computing node to translate the third speech data in the second language uploaded by the expert interaction terminal into the fourth speech data corresponding to the first language, and forwarding the fourth speech data to the on-site interaction terminal, includes: obtaining a second target clean speech stream fed back by the second computing node; the second target clean speech stream is obtained by preprocessing the third speech data; inputting the second target clean speech stream into a second speech translation model, and having the second speech translation model translate the second target clean speech stream into the fourth speech data.

[0009] In conjunction with the first aspect, in one possible embodiment, before obtaining the maintenance result instruction, the method further includes: obtaining first fault data of the corresponding faulty device according to the first device number in the interaction request, the first fault data including an alarm code and real-time operating parameters of the faulty device; retrieving historical fault data of the faulty device from the database according to the first device number in the interaction request; translating the first fault data and the historical fault data into multi-version text data respectively, the multi-version text data including first text data corresponding to the first language and second text data corresponding to the second language, each version of the multi-version text data including unified language text corresponding to the first fault data and the historical fault data; sending the first text to the on-site interaction terminal; and sending the second text to the expert interaction terminal.

[0010] In conjunction with the first aspect, in one possible embodiment, the method further includes: when a part model of a target faulty part determined by an expert is detected from the fourth voice data, retrieving the part number and assembly drawing of the target faulty part from a database based on the part model; converting the part number and the assembly drawing into a first version corresponding to a first language; sending the first version of the part model and the assembly drawing to the on-site interactive terminal; generating a corresponding repair plan based on the repair description information about the target faulty part in the second voice data and the fourth voice data; converting the repair plan into a first repair plan corresponding to a first language and a second repair plan corresponding to a second language; and sending the first repair plan and the second repair plan to the on-site interactive terminal and the expert interactive terminal, respectively.

[0011] In conjunction with the first aspect, in one possible embodiment, determining the first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal based on the interaction request includes: determining the first location and the second location of the on-site interactive terminal and the expert interactive terminal based on the physical address in the interaction request; querying the first computing node cluster and the second computing node cluster closest to the first location and the second location; and determining the first computing node of the on-site interactive terminal and the second computing node of the expert interactive terminal from the first computing node cluster and the second computing node cluster, respectively.

[0012] In conjunction with the first aspect, in one possible embodiment, the method further includes: after determining that the maintenance is completed, obtaining the industrial terminology used in the current interaction; and adding the industrial terminology to the terminology database between the first language and the second language to optimize the translation accuracy in similar scenarios.

[0013] Secondly, this application provides an online interactive translation device based on a factory intelligent interaction system, applied to an intelligent interaction server within the factory intelligent interaction system, wherein the factory intelligent interaction system includes the intelligent interaction server, a field interaction terminal, and an expert interaction terminal; the device includes: The acquisition unit is used to acquire the interaction requests from the on-site interactive terminal to the expert interactive terminal. The determining unit is configured to determine, based on the interaction request, the first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal; The processing unit is configured to perform voice interaction processing between the expert interaction terminal and the field interaction terminal through the intelligent interaction server, the first computing node, and the second computing node, wherein the voice interaction processing includes translation; and when a maintenance result instruction is obtained, to terminate the maintenance interaction process between the field interaction terminal and the expert interaction terminal according to the maintenance result instruction.

[0014] Thirdly, this application provides an electronic device including a processor, a memory, a communication interface, and one or more programs, said one or more programs being stored in the memory and configured to be executed by the processor, said programs including instructions for performing the steps of either the first or second aspect of this application.

[0015] Fourthly, this application provides a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in either the first or second aspect of this application.

[0016] Fifthly, this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in either the first or second aspect of this application. The computer program product may be a software installation package.

[0017] As can be seen, in this application, the interaction request from the field interaction terminal to the expert interaction terminal is first obtained; based on the interaction request, the first computing node corresponding to the field interaction terminal and the second computing node corresponding to the expert interaction terminal are determined; voice interaction processing between the expert interaction terminal and the field interaction terminal is performed through the intelligent interaction server, the first computing node, and the second computing node, and the voice interaction processing includes translation; when a maintenance result instruction is obtained, the maintenance interaction process between the field interaction terminal and the expert interaction terminal is terminated according to the maintenance result instruction. In this way, by inserting a cluster of nearby computing nodes between the field end and the expert end, the multi-speaker speech is separated and translated in parallel, reducing the load on the intelligent interaction server, realizing low-latency two-way voice interaction in cross-language industrial maintenance scenarios, and linking fault data, historical data, and the translation process, the efficiency of diagnostic decision-making is improved, and the efficiency of cross-language interaction in remote maintenance scenarios is enhanced. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the system architecture of the first intelligent factory interaction system provided in the embodiments of this application; Figure 2 This is a schematic diagram of the system architecture of the second type of intelligent factory interaction system provided in the embodiments of this application; Figure 3 This is a schematic block diagram of the structure of the first computing node cluster provided in the embodiments of this application; Figure 4 This is a schematic block diagram of the structure of the second computing node cluster provided in the embodiments of this application; Figure 5 This is a schematic diagram illustrating an application scenario of an online interactive translation method based on a factory intelligent interactive system provided in an embodiment of this application; Figure 6 This is a flowchart illustrating an online interactive translation method based on a factory intelligent interactive system, as improved in an embodiment of this application. Figure 7 This is a schematic diagram of the structure of an online interactive translation device based on a factory intelligent interaction system provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0021] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, systems, products, or apparatuses.

[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0023] Currently, in existing cross-language industrial maintenance solutions, language conversion relies on manual or general translation tools, which results in communication delays and insufficient adaptation of professional terminology.

[0024] To address the aforementioned issues, this application provides an online interactive translation method based on a factory intelligent interaction system. This method can be applied to cross-language interaction scenarios in remote maintenance. It involves acquiring interaction requests from a field interaction terminal to an expert interaction terminal; determining a first computing node corresponding to the field interaction terminal and a second computing node corresponding to the expert interaction terminal based on the interaction request; performing voice interaction processing between the expert interaction terminal and the field interaction terminal through the intelligent interaction server, the first computing node, and the second computing node, including translation; and terminating the maintenance interaction process between the field interaction terminal and the expert interaction terminal based on the maintenance result instruction when a maintenance result instruction is received. This improves the efficiency of cross-language interaction in remote maintenance scenarios. This solution is applicable to various scenarios, including but not limited to the application scenarios mentioned above.

[0025] The following is a description of the technical terms used in the embodiments of this application.

[0026] MES (Manufacturing Execution System) is a real-time information management system for the workshop level in manufacturing. It is located between the upper-level ERP (Enterprise Resource Planning) and the lower-level industrial control (such as PLC, SCADA) of an enterprise. It is the core bridge connecting "planning" and "execution" and is known as the "digital central nervous system" of the factory.

[0027] The system architecture involved in the embodiments of this application is described below.

[0028] Please see Figure 1 and Figure 2 This application provides a factory intelligent interaction system, which includes an intelligent interaction server, field interaction terminals, and expert interaction terminals. The field interaction terminals are deployed in the factory, while the expert interaction terminals are deployed at the maintenance company or after-sales service provider. When a target fault occurs in the factory equipment, factory employees send an interaction request to the expert interaction terminal via the field interaction terminal to the intelligent interaction server. The intelligent interaction server then establishes a communication link between the field interaction terminal and the expert interaction terminal. The intelligent interaction server then performs real-time translation of the data exchange between the field interaction terminal and the expert interaction terminal to facilitate communication and interaction between the factory and after-sales service provider regarding the repair of the target faulty equipment.

[0029] Furthermore, the factory intelligent interaction system also includes multiple computing node clusters distributed around the world. When a field interaction terminal initiates an interaction request to an expert interaction terminal, the intelligent interaction server selects the computing node closest to the field interaction terminal as the dedicated computing node for processing the data from the field interaction terminal.

[0030] The specific methods will be described in detail below.

[0031] Please see Figure 6 This application also provides an online interactive translation method based on a factory intelligent interaction system, applied to an intelligent interaction server in the factory intelligent interaction system, wherein the factory intelligent interaction system includes the intelligent interaction server, on-site interaction terminals, and expert interaction terminals; the method includes: Step S601: Obtain the interaction request from the on-site interactive terminal to the expert interactive terminal.

[0032] In practice, when a target fault occurs in equipment within the factory, factory employees send an interaction request to the intelligent interaction server, targeting an expert interaction terminal, via a field interaction terminal. The intelligent interaction server then establishes a communication link between the field interaction terminal and the expert interaction terminal, enabling them to interact.

[0033] For example, in industrial settings, the value of a factory's intelligent interactive system lies in enabling two-way audio interaction. For instance, in a Brazilian factory, on-site personnel speak Portuguese, while remote audio-accessed instructors speak Chinese. Translation models deployed on intelligent interactive servers or local terminals (such as on-site interactive terminals or expert interactive terminals) can translate in real time.

[0034] Step S602: Determine the first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal based on the interaction request.

[0035] In one possible embodiment, determining the first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal based on the interaction request includes: determining the first location and the second location of the on-site interactive terminal and the expert interactive terminal based on the physical addresses in the interaction request; querying the cluster of first computing nodes (e.g., the cluster of first computing nodes closest to the first location and the second location) that are closest to the first location and the second location. Figure 3 The first computing node cluster shown includes the first computing node, computing node 11, computing node 12 to computing node 1n, etc.) and the second computing node cluster (such as... Figure 4 The second computing node cluster shown includes a second computing node, computing node 21, computing node 22 to computing node 2n, etc.); the first computing node of the field interactive terminal and the second computing node of the expert interactive terminal are determined from the first computing node cluster and the second computing node cluster, respectively.

[0036] For details, please refer to Figure 3 , Figure 4 and Figure 5In each country or region where the field interactive terminal is located, one or more corresponding computing node clusters are configured. When the intelligent interactive server receives an interaction request from the field interactive terminal, it determines the first location of the field interactive terminal and the second location of the expert interactive terminal based on the first and second addresses carried in the interaction request. Then, it queries the address database for the third address closest to the first location (i.e., the first address) and sends a communication link request to the first computing node cluster corresponding to the third address. If the first computing node cluster responds to the communication link request, it determines an idle computing node as the first computing node, and the intelligent interactive server adds the first computing node to the communication group. Similarly, it queries the address database for the fourth address closest to the second location (i.e., the second address) and sends a communication link request to the second computing node cluster corresponding to the fourth address. If the second computing node cluster responds to the communication link request, it determines an idle computing node as the second computing node, and the intelligent interactive server adds the second computing node to the communication group.

[0037] Optionally, the communication group may include a field interactive terminal, a first computing node, a second computing node, an intelligent interactive server, and an expert interactive terminal, enabling information exchange between the field interactive terminal, the first computing node, the second computing node, the intelligent interactive server, and the expert interactive terminal according to preset communication rules. In this way, the voice data uploaded by the field interactive terminal and the expert interactive terminal is processed (such as multi-speaker mixed speech separation and preprocessing) by the first and second computing nodes closest to the field interactive terminal before being transmitted to the intelligent interactive server for real-time translation, improving data processing efficiency and reducing the data processing load on the intelligent interactive server.

[0038] In practice, when a factory operator detects equipment malfunction information, such as an injection molding machine stopping or an alarm indicator light illuminating, the operator enters the unique equipment number (i.e., the first equipment number) of the target malfunctioning device into the on-site interactive terminal, such as "IM-2024-035". The intelligent interactive server assigns "fault repair" mode trigger permission to the on-site interactive terminal and simultaneously assigns call receiving permission to the remote expert interactive terminal (expert end). The operator triggers the "fault repair" mode by clicking "Initiate Repair Call" on the first interface of the on-site interactive terminal. The intelligent interactive server prioritizes calling the factory's local edge computing node (i.e., the first computing node) to perform initial processing on the first language voice data uploaded by the on-site interactive terminal (such as multi-speaker mixed speech separation, preprocessing, etc.), while activating cloud backup nodes for redundancy processing (ensuring redundancy), completing the activation and resource allocation of distributed audio processing nodes. The expert interactive terminal (expert end) pops up a call reminder, and the expert clicks "Accept" to establish a two-way communication link. After the first computing node completes its initial processing, it transmits the processed voice data to the intelligent interactive terminal for real-time translation, and finally sends the translated voice data back to the expert interactive terminal. Understandably, the redundancy of the cloud backup node refers to backing up the voice data in the first language. It also allows the cloud backup node to perform the initial processing when the first computing node is unable to complete the initial processing. After the intelligent interaction server sends the second voice data to the expert interaction terminal, the expert responds to the second voice data through the expert interaction terminal, sending the third voice data to the second computing node. The second computing node preprocesses the third voice data and then feeds it back to the intelligent interaction server, which translates the third voice data into fourth voice data and sends it to the on-site interaction terminal.

[0039] As can be seen, in this embodiment, by introducing a first computing node and a second computing node deployed nearby between the on-site interactive terminal and the expert interactive terminal, voiceprint separation and preprocessing of mixed speech from multiple speakers are performed. Combined with the end-to-end speech translation model of the intelligent interactive server, low-latency bidirectional multi-voice path real-time translation is achieved in cross-language industrial maintenance scenarios.

[0040] Step S603: Perform voice interaction processing between the expert interaction terminal and the field interaction terminal through the intelligent interaction server, the first computing node, and the second computing node.

[0041] The voice interaction processing includes translation.

[0042] In one possible embodiment, the voice interaction processing between the expert interaction terminal and the on-site interaction terminal via the intelligent interaction server, the first computing node, and the second computing node includes: determining a language mapping relationship between the on-site interaction terminal and the expert interaction terminal, the language mapping relationship including a mapping relationship between a first language and a second language, wherein the first language is a preset language of the on-site interaction terminal, and the second language is a preset language of the expert interaction terminal; calling the intelligent interaction server and the first computing node to translate the first voice data in the first language uploaded by the on-site interaction terminal into second voice data in the second language, and forwarding the second voice data to the expert interaction terminal; calling the intelligent interaction server and the second computing node to translate the third voice data in the second language uploaded by the expert interaction terminal into fourth voice data corresponding to the first language, and forwarding the fourth voice data to the on-site interaction terminal.

[0043] Optionally, the on-site interactive terminal can have one or more preset languages, meaning the first language includes one or more first sub-languages; similarly, the expert interactive terminal can also have one or more preset languages, meaning the second language includes one or more second sub-languages. In other words, both the on-site and expert sides can be configured with only one language or multiple languages. When multiple languages ​​are configured, there can be speakers speaking different languages ​​on-site, and the on-site interactive terminal uniformly records the voices of all speakers and uploads them to the first computing node for initial processing; while the expert interactive terminal can also have speakers speaking different languages, which are then processed by the intelligent interactive server.

[0044] In the specific implementation, after the intelligent interaction server receives the interaction request, it automatically matches the languages ​​of both parties (operator's end: Portuguese, expert's end: Chinese) through the terminal language recognition module, generating a two-way mapping relationship of "Portuguese → Chinese, Chinese → Portuguese". The on-site interaction terminal activates its built-in microphone array to collect the operator's Portuguese speech in real time (i.e., the first language), while simultaneously running an industrial-grade environmental noise reduction algorithm to filter out noise from the workshop machines, assembly line noise, and other background disturbances. In addition, the expert interaction terminal collects the Chinese speech of the Chinese expert and uploads it synchronously to the activated distributed audio processing node (i.e., the second computing node) for processing via an encrypted channel. The first or second computing node calls a multi-speaker speech separation algorithm. If a local maintenance worker interrupts (≤3 overlapping speech streams), it separates the first voiceprint feature of each speech stream in real time, removes background noise, and generates a clean speech stream with a single path. The first or second computing node standardizes the format of the preprocessed audio data (sampling rate 16kHz, bit depth 16bit) to ensure compatibility with subsequent translation models. Understandably, if there are no multiple speakers, the language data can be directly standardized.

[0045] Specifically, the step of calling the intelligent interaction server and the first computing node to translate the first speech data in the first language uploaded by the on-site interaction terminal into the second speech data in the second language includes: obtaining at least one first target clean speech stream fed back by the first computing node; the at least one first target clean speech stream is obtained by speech separation and preprocessing based on the first speech data, wherein the speech separation refers to splitting at least one corresponding first clean speech stream from the first speech data according to the first voiceprint features, the number of first clean speech streams being consistent with the number of first voiceprint features, and the preprocessing refers to standardizing the at least one first clean speech stream to obtain the at least one first target clean speech stream; associating the at least one first target clean speech stream with the first voiceprint features and inputting them into at least one corresponding first speech translation model, wherein the at least one first speech translation model translates the at least one first target clean speech stream into the second speech data.

[0046] The first clean speech stream refers to the single-speaker audio stream after speech separation and standardization.

[0047] In practice, after the communication group is established, the on-site interactive terminal can begin voice interaction with the expert interactive terminal. First, the on-site interactive terminal collects the first voice data from the on-site personnel and sends it to the first computing node. Then, the first computing node initiates a voice separation algorithm to separate the multi-person voices in the first voice data. Understandably, before this, each on-site participant needs to record corresponding verification voice in the on-site interactive terminal. The on-site interactive terminal or the first computing node extracts the corresponding first voiceprint feature from the verification voice, obtaining at least one first voiceprint feature, and stores this first at least one first voiceprint feature in the first computing node. When the first computing node initiates the voice separation algorithm, it extracts all target voiceprint features from the first voice data and compares each target voiceprint feature with the at least one first voiceprint feature stored in the first computing node. If a match is found between the target voiceprint feature and any one of the at least one first voiceprint features, the voice stream corresponding to that target voiceprint feature is separated from the first voice data, resulting in a first clean voice stream. The number of first clean voice streams is the same as the number of participants speaking at the on-site interactive terminal. It is understandable that there may be cases where the number of first clean speech streams matches the number of first voiceprint features (i.e., only some of the people who registered their first voiceprint features participated in the speech), and there may also be cases where they do not match (i.e., all of the people who registered their first voiceprint features participated in the speech). After obtaining at least one first clean speech stream, each of the at least one first clean speech stream is standardized, for example, by unifying the sampling rate and bit depth, ultimately obtaining at least one first target clean speech stream. Finally, the first computing node sends at least one first target clean speech stream to the intelligent interaction server.

[0048] In one example, after the first computing node has stored at least one first voiceprint feature, it sends the first first voiceprint feature to the intelligent interaction server. The intelligent interaction server then generates at least one first speech translation model, and associates the at least one first voiceprint feature with the at least one first speech translation model in a one-to-one correspondence. When the intelligent interaction server receives at least one first target clean speech stream, it inputs the at least one first target clean speech stream into the at least one first speech translation model for translation based on the first voiceprint feature, achieving one-to-one parallel translation, and finally obtaining at least one second speech data, thus improving translation efficiency.

[0049] Specifically, the step of calling the intelligent interaction server and the second computing node to translate the third speech data in the second language uploaded by the expert interaction terminal into the fourth speech data corresponding to the first language includes: obtaining the second target clean speech stream fed back by the second computing node; the second target clean speech stream is obtained by preprocessing the third speech data; the second target clean speech stream is input into the second speech translation model, and the second speech translation model translates the second target clean speech stream into the fourth speech data.

[0050] It is understandable that the second computing node performs speech separation and standardization processing on the third speech data in the same way as the first computing node, which will not be elaborated here.

[0051] In practice, the intelligent interaction server assigns the clean speech stream to the translation model according to its "voiceprint identifier," calls the "Portuguese-Chinese" transfer learning pre-training parameters (to optimize the translation accuracy of low-resource languages), and directly completes the end-to-end "speech → speech" translation, skipping the intermediate "speech-text-speech" step. The speech translation model simultaneously generates target language text (Portuguese / Chinese) and automatically associates it with the first device number "IM-2024-035" input by the operator from the on-site interaction terminal. The on-site interaction terminal plays the Portuguese translation of the expert's Chinese speech through a high-fidelity speaker, and the screen simultaneously displays the Portuguese text + device number (to prevent missed speech due to workshop noise). The expert interaction terminal plays the Chinese translation of the operator's Portuguese speech, and the screen simultaneously displays the Chinese text + preliminary fault description (e.g., "injection molding machine shutdown alarm"); the server sends a "translation complete" confirmation signal to both terminals.

[0052] In one possible embodiment, before obtaining the maintenance result instruction, the method further includes: obtaining first fault data of the corresponding faulty device according to the first device number in the interaction request, the first fault data including an alarm code and real-time operating parameters of the faulty device; retrieving historical fault data of the faulty device from the database according to the first device number in the interaction request; translating the first fault data and the historical fault data into multi-version text data, the multi-version text data including first text data corresponding to the first language and second text data corresponding to the second language, each version of the multi-version text data including unified language text corresponding to the first fault data and the historical fault data; sending the first text to the on-site interaction terminal; and sending the second text to the expert interaction terminal.

[0053] Among them, multi-version text data refers to generating multiple target language versions of the same unified language text.

[0054] In practice, the intelligent interactive server retrieves historical fault data (e.g., "Repaired part B-0892 in August 2024 due to bearing wear") from the MES system via an API interface, using the device number "IM-2024-035" as an index. Simultaneously, the intelligent interactive server reads real-time PLC data (i.e., the first fault data) via a Modbus TCP interface, extracting alarm codes (e.g., "E023") and real-time operating parameters (e.g., corresponding descriptions ("motor overload"), key operating parameters (e.g., motor temperature 85℃, speed 1200r / min)). The intelligent interactive server converts the retrieved data into multilingual versions according to the language mapping rules between the field interactive terminal and the expert interactive terminal: historical fault data is converted to Portuguese, and PLC real-time data is converted to Chinese. The intelligent interactive server then pushes the multilingual data to both terminals. The field interactive terminal displays the Portuguese version of the historical fault records, while the expert interactive terminal displays the Chinese version of the alarm codes and real-time operating parameters. The intelligent interactive server generates a "Device data association complete" signal, confirming that the data has been synchronized.

[0055] Step S604: When a maintenance result instruction is obtained, the maintenance interaction process between the field interactive terminal and the expert interactive terminal is terminated according to the maintenance result instruction.

[0056] Specifically, the method further includes: when the part model of the target faulty part determined by the expert is detected from the fourth voice data, retrieving the part number and assembly drawing of the target faulty part from the database according to the part model; converting the part number and the assembly drawing into a first version corresponding to a first language; sending the first version of the part model and the assembly drawing to the on-site interactive terminal; generating a corresponding repair plan based on the repair description information of the target faulty part in the second voice data and the fourth voice data; converting the repair plan into a first repair plan corresponding to a first language and a second repair plan corresponding to a second language; and sending the first repair plan and the second repair plan to the on-site interactive terminal and the expert interactive terminal, respectively.

[0057] In practice, the repair details for the target faulty part in the second voice data and the fourth voice data are monitored in real time to generate a corresponding repair plan.

[0058] Based on the alarm parameters (E023: motor overload) and historical faults (bearing wear) displayed on the expert interactive terminal, the expert determines the cause of the fault ("It is suspected that the motor overload caused secondary bearing wear, and the motor end cover needs to be disassembled and inspected"). The expert then issues maintenance guidance voice commands through the expert interactive terminal. The intelligent interactive server translates the expert's Chinese guidance voice commands into Portuguese in real time and pushes them to the industrial site intelligent interactive terminal (simultaneous voice and text output). After the operator executes the guidance steps corresponding to the voice commands (e.g., "Disconnect the main power supply to the equipment and remove the motor end cover"), they report the current status of the faulty equipment through the on-site interactive terminal (Portuguese: "End cover removed, bearing has obvious scratches"). The intelligent interactive server translates the feedback into Chinese and pushes it to the expert interactive terminal. When the expert needs to confirm the part model, the intelligent interactive server retrieves the corresponding part number ("EC-0321") and 3D assembly drawing for the "motor end cover" through the MES BOM table interface, converts it to Portuguese annotations, and pushes it to the on-site interactive terminal. Both parties confirm the details of the maintenance plan through the translation channel (e.g., "Replace part EC-0321, assemble according to step 3 of the drawing"), and the server records the key interaction content.

[0059] As can be seen, in this embodiment, by uniformly converting the real-time operating parameters, alarm codes, and historical fault records of the faulty equipment into multilingual text, and linking the BOM library to automatically push the target part number and assembly drawing, the translation process is synchronized with the production management data, reducing the time spent manually searching for data.

[0060] In one possible embodiment, the method further includes: after determining that the maintenance is completed, obtaining the industrial terminology used in the interaction; and adding the industrial terminology to the terminology database between the first language and the second language to optimize the translation accuracy in similar scenarios.

[0061] In practice, the on-site interactive terminal translates the operator's Portuguese "Repair Complete" feedback into Chinese and pushes it to the expert interactive terminal. After expert confirmation, a "closed-loop" command is issued. The intelligent interactive server writes the repair results into the MES system via API interface: fault cause ("motor overload caused secondary bearing wear," including Portuguese translation), handling steps ("power off, remove end cover, replace bearing, restart test"), time taken (e.g., "25 minutes"), and participating personnel (operator + expert), updating the equipment maintenance file. The intelligent interactive server extracts industrial terminology from this interaction (such as "motor end cover," "bearing scratches," "alarm code, such as E023") and adds it to the "Industrial Portuguese-Chinese Terminology Database" to optimize the translation accuracy of subsequent similar scenarios. The intelligent interactive server reads the current equipment status (such as "operating status code 01") through the PLC, converts it into multiple languages, and pushes it to both terminals to confirm normal equipment operation. The intelligent interactive server shuts down the distributed audio processing node (keeping cloud backup), releases temporary resources, and generates a "Transaction Closed-Loop Complete" report. By automatically extracting industry-specific terms from the interaction process after the maintenance loop is closed, and dynamically updating the bilingual terminology database, the translation model in similar scenarios can gradually adapt to industry terminology, thereby improving the accuracy of subsequent translations.

[0062] As can be seen, in this embodiment, the interaction request from the field interaction terminal to the expert interaction terminal is obtained; the first computing node corresponding to the field interaction terminal and the second computing node corresponding to the expert interaction terminal are determined according to the interaction request; voice interaction processing between the expert interaction terminal and the field interaction terminal is performed through the intelligent interaction server, the first computing node, and the second computing node, and the voice interaction processing includes translation; when a maintenance result instruction is obtained, the maintenance interaction process between the field interaction terminal and the expert interaction terminal ends according to the maintenance result instruction. This improves the efficiency of cross-language interaction in remote maintenance scenarios. The intelligent interaction server links MES / PLC / BOM data and translation channels, automatically pushes multilingual fault data and part information, and extracts terms from the dialogue to fill the terminology database, forming a closed loop of "translation, data, and terminology".

[0063] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, mobile electronic devices include corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0064] This application embodiment can divide the electronic device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0065] Please see Figure 7 This application also provides an online interactive translation device based on a factory intelligent interaction system, applied to an intelligent interaction server within the factory intelligent interaction system. The factory intelligent interaction system includes the intelligent interaction server, on-site interaction terminals, and expert interaction terminals. The device includes: The acquisition unit is used to acquire the interaction requests from the on-site interactive terminal to the expert interactive terminal. The determining unit is configured to determine, based on the interaction request, the first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal; The processing unit is configured to perform voice interaction processing between the expert interaction terminal and the field interaction terminal through the intelligent interaction server, the first computing node, and the second computing node, wherein the voice interaction processing includes translation; and, when a maintenance result instruction is obtained, terminate the maintenance interaction process between the field interaction terminal and the expert interaction terminal according to the maintenance result instruction.

[0066] As can be seen, in this application, the interaction request from the on-site interactive terminal to the expert interactive terminal is first obtained; based on the interaction request, a first computing node corresponding to the on-site interactive terminal and a second computing node corresponding to the expert interactive terminal are determined; voice interaction processing between the expert interactive terminal and the on-site interactive terminal is performed through the intelligent interactive server, the first computing node, and the second computing node, the voice interaction processing including translation; when a maintenance result instruction is obtained, the maintenance interaction process between the on-site interactive terminal and the expert interactive terminal is terminated according to the maintenance result instruction. This improves the efficiency of cross-language interaction in remote maintenance scenarios.

[0067] In one possible embodiment, the aspect of processing voice interaction between the expert interaction terminal and the on-site interaction terminal through the intelligent interaction server, the first computing node, and the second computing node, wherein the processing unit is specifically configured to: determine the language mapping relationship between the on-site interaction terminal and the expert interaction terminal, the language mapping relationship including the mapping relationship between the first language and the second language, wherein the first language is a preset language of the on-site interaction terminal, and the second language is a preset language of the expert interaction terminal; invoke the intelligent interaction server and the first computing node to translate the first voice data in the first language uploaded by the on-site interaction terminal into the second voice data in the second language, and forward the second voice data to the expert interaction terminal; invoke the intelligent interaction server and the second computing node to translate the third voice data in the second language uploaded by the expert interaction terminal into the fourth voice data corresponding to the first language, and forward the fourth voice data to the on-site interaction terminal.

[0068] In one possible embodiment, the aspect of calling the intelligent interaction server and the first computing node to translate the first speech data in the first language uploaded by the on-site interactive terminal into the second speech data in the second language, the processing unit is specifically used to: obtain at least one first target clean speech stream fed back by the first computing node; the at least one first target clean speech stream is obtained based on the first speech data through speech separation and preprocessing, wherein the speech separation refers to splitting at least one corresponding first clean speech stream from the first speech data according to the first voiceprint features, the number of first clean speech streams being consistent with the number of first voiceprint features, and the preprocessing refers to standardizing the at least one first clean speech stream to obtain the at least one first target clean speech stream; and associate the at least one first target clean speech stream with the first voiceprint features and input them into at least one corresponding first speech translation model, wherein the at least one first speech translation model translates the at least one first target clean speech stream into the second speech data.

[0069] In one possible embodiment, the aspect of calling the intelligent interaction server and the second computing node to translate the third speech data in the second language uploaded by the expert interaction terminal into the fourth speech data corresponding to the first language, the processing unit is specifically used to: obtain a second target clean speech stream fed back by the second computing node; the second target clean speech stream is obtained by preprocessing the third speech data; input the second target clean speech stream into a second speech translation model, and have the second speech translation model translate the second target clean speech stream into the fourth speech data.

[0070] In one possible embodiment, before obtaining the maintenance result instruction, the online interactive translation device based on the factory intelligent interaction system further includes: an acquisition unit, configured to acquire first fault data of the corresponding faulty device according to the first device number in the interaction request, the first fault data including an alarm code and real-time operating parameters of the faulty device; a processing unit, configured to retrieve historical fault data of the faulty device from the database according to the first device number in the interaction request; and to translate the first fault data and the historical fault data into multi-version text data, the multi-version text data including first text data corresponding to the first language and second text data corresponding to the second language, each version of the multi-version text data including unified language text corresponding to the first fault data and the historical fault data; and to send the first text to the on-site interactive terminal; and to send the second text to the expert interactive terminal.

[0071] In one possible embodiment, the online interactive translation device based on the factory intelligent interaction system further includes: an acquisition unit, configured to, when a part number of a target faulty part determined by an expert is detected from the fourth voice data, acquire the part number and assembly drawing of the target faulty part from the database according to the part number; a processing unit, configured to convert the part number and the assembly drawing into a first version corresponding to a first language; send the first version of the part number and the assembly drawing to the on-site interactive terminal; generate a corresponding repair plan based on the repair description information about the target faulty part in the second voice data and the fourth voice data; convert the repair plan into a first repair plan corresponding to a first language and a second repair plan corresponding to a second language; and send the first repair plan and the second repair plan to the on-site interactive terminal and the expert interactive terminal, respectively.

[0072] In one possible embodiment, the aspect of determining the first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal based on the interaction request is specifically configured to: determine the first location and the second location of the on-site interactive terminal and the expert interactive terminal based on the physical address in the interaction request; query the first computing node cluster and the second computing node cluster closest to the first location and the second location; and determine the first computing node of the on-site interactive terminal and the second computing node of the expert interactive terminal from the first computing node cluster and the second computing node cluster, respectively.

[0073] In one possible embodiment, the online interactive translation device based on the factory intelligent interaction system further includes: the acquisition unit, used to acquire the industrial professional terms in this interaction after determining that the maintenance is completed; and the processing unit, used to supplement the industrial professional terms into the terminology database between the first language and the second language to optimize the translation accuracy in similar scenarios.

[0074] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0075] This application also provides an electronic device 80, such as... Figure 8As shown, it includes at least one processor 81; a display screen 82; and a memory 83, and may also include a communications interface 85 and a bus 84. The processor 81, display screen 82, memory 83, and communications interface 85 can communicate with each other via the bus 84. The display screen 82 is configured to display a preset user guide interface in the initial setup mode. The communications interface 85 can transmit information. The processor 81 can call logical instructions in the memory 83 to execute the methods described in the above embodiments.

[0076] Optionally, the electronic device 80 may be a mobile electronic device, an electronic device, or other devices, and is not limited to any particular type.

[0077] Furthermore, the logic instructions in the aforementioned memory 83 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0078] The memory 83, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of this disclosure. The processor 81 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 83, thereby implementing the methods in the above embodiments.

[0079] The memory 83 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the electronic device 80. Furthermore, the memory 83 may include high-speed random access memory (RAM) and may also include non-volatile memory. For example, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, may be used, or they may be transient storage media.

[0080] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.

[0081] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.

[0082] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0083] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and other division methods may exist in actual implementation; for example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0084] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0085] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can be physically comprised separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or in the form of hardware plus software functional units.

[0086] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, volatile memory, or non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM), etc., which are various media capable of storing program code.

[0087] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can easily conceive of variations or substitutions without departing from the spirit and scope of the present invention, and various modifications and alterations can be made, including combinations of the different functions and implementation steps described above, as well as software and hardware implementation methods, all of which are within the protection scope of the present invention.

Claims

1. An online interactive translation method based on a factory intelligent interactive system, characterized in that, An intelligent interaction server is applied to a factory intelligent interaction system, the factory intelligent interaction system including the intelligent interaction server, field interaction terminals, and expert interaction terminals; the method includes: Obtain the interaction request from the on-site interactive terminal to the expert interactive terminal; The first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal are determined based on the interactive request. Voice interaction processing between the expert interaction terminal and the on-site interaction terminal is performed through the intelligent interaction server, the first computing node, and the second computing node, and the voice interaction processing includes translation. When a repair result instruction is received, the repair interaction process between the on-site interactive terminal and the expert interactive terminal is terminated according to the repair result instruction.

2. The method according to claim 1, characterized in that, The voice interaction processing between the expert interaction terminal and the on-site interaction terminal through the intelligent interaction server, the first computing node, and the second computing node includes: Determine the language mapping relationship between the on-site interactive terminal and the expert interactive terminal. The language mapping relationship includes the mapping relationship between a first language and a second language. The first language is the language preset by the on-site interactive terminal, and the second language is the language preset by the expert interactive terminal. The intelligent interactive server and the first computing node are invoked to translate the first voice data in the first language uploaded by the on-site interactive terminal into the second voice data in the second language, and the second voice data is forwarded to the expert interactive terminal. The intelligent interaction server and the second computing node are invoked to translate the third voice data in the second language uploaded by the expert interaction terminal into the fourth voice data in the first language, and then forward the fourth voice data to the on-site interaction terminal.

3. The method according to claim 2, characterized in that, The step of calling the intelligent interaction server and the first computing node to translate the first speech data in the first language uploaded by the on-site interaction terminal into the second speech data in the second language includes: At least one first target clean speech stream is obtained from the feedback of the first computing node; the at least one first target clean speech stream is obtained by speech separation and preprocessing based on the first speech data. The speech separation refers to splitting at least one corresponding first clean speech stream from the first speech data according to the first voiceprint feature. The number of first clean speech streams is consistent with the number of first voiceprint features. The preprocessing refers to standardizing the at least one first clean speech stream to obtain the at least one first target clean speech stream. The at least one first target clean speech stream is associated with the first voiceprint feature and input into the corresponding at least one first speech translation model, and the at least one first speech translation model translates the at least one first target clean speech stream into at least one second speech data.

4. The method according to claim 2, characterized in that, The step of calling the intelligent interaction server and the second computing node to translate the third speech data in the second language uploaded by the expert interaction terminal into the fourth speech data corresponding to the first language includes: Obtain the second target clean speech stream fed back by the second computing node; the second target clean speech stream is obtained by preprocessing the third speech data; The second target clean speech stream is input into the second speech translation model, which then translates the second target clean speech stream into the fourth speech data.

5. The method according to claim 1, characterized in that, Before obtaining the repair result instruction, the method further includes: Based on the first device number in the interaction request, obtain the first fault data of the corresponding faulty device. The first fault data includes the alarm code and the real-time operating parameters of the faulty device. Retrieve historical fault data of the faulty device from the database based on the first device number in the interaction request; The first fault data and the historical fault data are translated into multi-version text data respectively. The multi-version text data includes first text data corresponding to a first language and second text data corresponding to a second language. Each version of the multi-version text data includes unified language text corresponding to the first fault data and the historical fault data. Send the first text to the on-site interactive terminal; The second text is sent to the expert interaction terminal.

6. The method according to claim 2, characterized in that, The method further includes: When the part model of the target faulty part determined by the expert is detected from the fourth voice data, the part number and assembly drawing of the target faulty part are obtained from the database according to the part model. Convert the part number and the assembly drawing into a first version corresponding to the first language; Send the first version of the part model and the assembly drawing to the on-site interactive terminal; Based on the repair description information about the target faulty part in the second voice data and the fourth voice data, a corresponding repair plan is generated; The repair plan is then converted into a first repair plan in a first language and a second repair plan in a second language. The first repair plan and the second repair plan are sent to the on-site interactive terminal and the expert interactive terminal, respectively.

7. The method according to any one of claims 1-6, characterized in that, The step of determining the first computing node corresponding to the on-site interactive terminal and the second computing node corresponding to the expert interactive terminal based on the interaction request includes: The first and second locations of the on-site interactive terminal and the expert interactive terminal are determined based on the physical addresses in the interactive request. Query the first and second compute node clusters that are closest to the first and second positions, respectively. The first computing node of the field interactive terminal and the second computing node of the expert interactive terminal are determined from the first computing node cluster and the second computing node cluster, respectively.

8. The method according to any one of claims 1-6, characterized in that, The method further includes: After confirming the completion of the repair, obtain the industry-specific terminology used in this interaction; The industrial terminology is added to the terminology database between the first and second languages ​​to optimize translation accuracy in similar scenarios.

9. An electronic device, characterized in that, The method includes a processor, a memory, a communication interface, and one or more programs, said one or more programs being stored in the memory and configured to be executed by the processor, said programs including instructions for performing the steps of the method as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, A computer program for electronic data interchange is stored, wherein the computer program causes a computer to execute instructions for the steps of the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Complex equipment expert remote assistance system based on augmented reality

    CN117369625A

  • Intelligent glasses independent communication method and system based on RedCap and eSIM, and medium

    CN121568207A

  • Intelligent speech translation mobile phone and system capable of realizing multi-language inter-translation

    CN121711426A

  • Multi-device multi-language synchronous translation device and method based on block chain and double AIs

    CN121882063A