Device-side artificial intelligence (AI) apparatus and method for providing multi-party translation service
The device-side AI device preprocesses and classifies speech speech, recognizes and translates the speech of specific speakers, and uses external AI translation processing devices to perform distributed processing when processing capabilities are insufficient, solving the problem that the device-side AI cannot translate dialogues of multiple speakers accurately in real time and improves translation efficiency and accuracy.
Patent Information
- Application Number
- CN202510123672.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-04
- Filing Date
- 2025-01-26
- Publication Date
- 2025-08-05
AI Technical Summary
Existing device-side AI cannot translate conversations between multiple speakers accurately in real time, and cannot recognize and selectively translate voices of specific speakers.
The device-side AI device receives speech voice through the input module, and the processor preprocesses and classifies speech voice, recognizes a specific speaker, and translates it into a target language. At the same time, when processing power is insufficient, use an external AI translation processing device to perform distributed processing.
Real-time accurate translation of conversations with multiple speakers is achieved, improving the speed and accuracy of translation processing, reducing power consumption and extending the service life of the device.
Smart Images

Figure CN120430318A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an on-device AI (AI) apparatus capable of performing real-time translation (or interpretation) of conversations between multiple speakers and a method for providing a multi-party translation service. Background Art
[0002] Generally, artificial intelligence is a field of computer engineering and information technology that studies methods for enabling computers to perform thinking, learning, self-development, etc. based on human intelligence, and is intended to enable computers to imitate human intelligent behavior.
[0003] Furthermore, artificial intelligence does not exist in itself, but is directly or indirectly related to other fields of computer science. In particular, in modern times, attempts are being made to introduce elements of artificial intelligence into various fields of information technology and to use them to solve problems in those fields.
[0004] At the same time, research is being actively conducted on technologies that use artificial intelligence to recognize and learn surrounding situations and provide the user's desired information in the desired form or perform the user's desired actions or functions.
[0005] In addition, electronic devices that provide such various actions and functions can be called artificial intelligence devices.
[0006] Recently, on-device AI, which can process information on terminal devices themselves without connecting to a server or the cloud, has attracted attention.
[0007] Since device-side AI does not send data to a server or the cloud, but instead performs AI calculations on the user's terminal device, it has the advantages of high speed and advantages in terms of privacy protection and cost.
[0008] However, there are still limitations on the completeness of real-time translation services currently available through on-device AI.
[0009] That is to say, current device-side AI is not only unable to accurately translate conversations between multiple speakers in real time, but also unable to recognize the voices of multiple speakers to provide selective translation services.
[0010] Therefore, in the future, there is a need to develop an on-device AI device that can accurately translate conversations between multiple speakers in real time without the need for separate equipment. Summary of the Invention
[0011] Technical issues
[0012] The present disclosure is directed to solving the above-referenced problems and other problems.
[0013] The present disclosure aims to provide a device-side AI apparatus and a method for providing a multi-party translation service. When a specific speaker is selected from multiple speakers, the device-side AI apparatus and the method for providing the multi-party translation service can accurately translate the conversation between multiple speakers in real time by extracting the speech speech of the specific speaker from the speech speech classified by speaker and translating it into a target language.
[0014] Technical Solution
[0015] According to one embodiment of the present disclosure, a device-side AI apparatus may include: an input module, into which a speaker's speech voice is input; and a processor, which is configured to perform AI processing to translate the speaker's speech voice into a target language, wherein the processor is configured to: when speech voice is input from multiple speakers, preprocess the speech voice, classify the preprocessed speech voice by speaker, and when a specific speaker is selected from multiple speakers, extract the speech voice of the selected specific speaker from the speech voice classified by speaker, and translate the speech voice of the specific speaker into the target language and output it.
[0016] An AI translation processing device according to one embodiment of the present disclosure is an AI translation processing device that is communicatively connected to a device-side AI device, and may include: a communication module that is connected to the device-side AI device; a memory that stores an AI model for AI translation processing; and a processor that is configured to: perform AI translation processing in response to a translation proxy service request from the device-side AI device; when receiving a translation proxy service request from the device-side AI device, determine whether self-translation processing can be performed by measuring the AI translation processing amount corresponding to the translation proxy service request; when it is determined that self-translation processing can be performed, transmit approval of the translation proxy service request to the device-side AI device; and when receiving speaker's speech voice data and target language information to be translated from the device-side AI device, translate the speaker's speech voice into the target language and output it.
[0017] According to one embodiment of the present disclosure, a device-side AI system includes at least one AI translation processing device communicatively connected to a device-side AI apparatus, and may include: a device-side AI apparatus configured to perform AI translation processing of translating a speaker's speech voice into a target language; and at least one AI translation processing device configured to perform AI translation processing in response to a translation proxy service request from the device-side AI apparatus, wherein, when speech voice is input from a plurality of speakers, the device-side AI apparatus preprocesses the speech voice, classifies the preprocessed speech voice by speaker, and when a specific speaker is selected from the plurality of speakers, extracts the speech voice of the selected specific speaker from the speech voice classified by speaker, translates the speech voice of the specific speaker into a target language, and outputs it; and when a user command requesting a translation proxy service is input, the device-side AI apparatus requests the AI translation processing apparatus to perform a translation proxy service on the speech voice of the specific speaker, and when approval of the translation proxy service request is received from the AI translation processing apparatus, sends speech data of the specific speaker and target language information to be translated to the AI translation processing apparatus to be translated and output, so that the AI translation processing apparatus translates the speech voice of the specific speaker into the target language.
[0018] According to one embodiment of the present disclosure, a method for providing a multi-party translation service for a device-side AI device may include: receiving speech voices from multiple speakers; preprocessing the speech voices; classifying the preprocessed speech voices by speaker; when a specific speaker is selected from the multiple speakers, extracting the speech voice of the selected specific speaker from the speech voices classified by speaker; and translating the speech voice of the specific speaker into a target language and outputting it.
[0019] Technical Effects
[0020] According to one embodiment of the present disclosure, when a specific speaker is selected from multiple speakers, the device-side AI device extracts the speech voice of the specific speaker from the speech differentiated by speaker and translates it into the target language, thereby identifying the conversation between multiple speakers by speaker and accurately interpreting the conversation in real time.
[0021] In addition, the present disclosure can improve the processing speed of AI translation processing and the accuracy and service quality of translation processing result values by selecting an external AI translation processing device and requesting distributed processing of AI translation processing when the amount of AI processing used to translate the speaker's spoken speech exceeds the amount that the device itself can process.
[0022] Furthermore, the present disclosure can minimize power consumption and reduce heat generation by distributing AI translation processing with an externally located AI translation processing device, thereby improving performance and lifespan. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 An artificial intelligence device according to an embodiment of the present disclosure is shown.
[0024] Figure 2 An artificial intelligence server according to an embodiment of the present disclosure is shown.
[0025] Figure 3 An artificial intelligence system according to an embodiment of the present disclosure is shown.
[0026] Figures 4 to 6 2 is a diagram for illustrating a device-side AI system according to an embodiment of the present disclosure.
[0027] Figure 7 and Figure 8 2 is a diagram for illustrating a device-side AI apparatus according to an embodiment of the present disclosure.
[0028] Figures 9 to 11 3 is a diagram for explaining a speaker recognition process of a device-side AI apparatus according to an embodiment of the present disclosure.
[0029] Figures 12 to 15 3 is a diagram for explaining a speaker selection process of a device-side AI apparatus according to an embodiment of the present disclosure.
[0030] Figure 16 2 is a diagram for explaining a process of setting a target language of a device-side AI apparatus according to an embodiment of the present disclosure.
[0031] Figure 17 and Figure 18 1 is a diagram for explaining a process of measuring the AI translation processing capacity of a device-side AI apparatus according to an embodiment of the present disclosure.
[0032] Figures 19 to 23 3 is a diagram for explaining a method for providing a multi-party translation service of an AI device on a device side according to an embodiment of the present disclosure.
[0033] Figure 24 and Figure 25 3 is a diagram for explaining a method for providing a multi-party translation service by an AI translation processing device connected to an AI device on a device side according to an embodiment of the present disclosure.
[0034] Figure 26 3 is a diagram for explaining a method for providing a multi-party translation service of a device-side AI system according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] Hereinafter, the embodiments disclosed in this specification will be described in detail with reference to the accompanying drawings. Regardless of the figure numbers, the same or similar parts will be given the same figure numbers, and their redundant descriptions will be omitted. The suffixes "module" and "part" used for components in the following description are given or used interchangeably only for the convenience of writing the description, and they themselves do not have different meanings or functions. In addition, when describing the embodiments disclosed in this specification, if it is determined that the specific description of the relevant known technology may obscure the main purpose of the embodiments disclosed in this specification, its detailed description will be omitted. In addition, the drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical concepts disclosed in this specification are not limited by the drawings, and should be understood to include all modifications, equivalents and substitutes included in the concepts and technical scope of the present disclosure.
[0036] Terms including ordinal numbers such as first, second, etc. may be used to describe various components, but the components are not limited by the terms. These terms are only used to distinguish one component from another.
[0037] When a component is referred to as being “connected” or “connected to” another component, it should be understood that it can be directly connected or connected to the other component, but other components may exist between them. On the other hand, when a component is referred to as being “directly connected” or “directly connected to” another component, it should be understood that no other components exist between them.
[0038] In addition, throughout this specification, the terms neural network, neural network, and network function are used interchangeably. A neural network can be composed of a set of interconnected computing units, which can generally be referred to as "nodes." These "nodes" can also be referred to as "neurons." A neural network is composed of at least two or more nodes. The nodes (or neurons) that make up a neural network can be interconnected through one or more "links."
[0039] Figure 1 The AI device 100 according to an embodiment of the present disclosure is illustrated.
[0040] The AI device 100 can be implemented as a fixed device or a mobile device, such as a TV, a projector, a mobile phone, a smart phone, a desktop computer, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a tablet PC, a wearable device, a set-top box (STB), a DMB receiver, a radio, a washing machine, a refrigerator, a desktop computer, a digital signage, a robot, a vehicle, etc.
[0041] Reference Figure 1, the AI device 100 may include a communication unit 110 , an input unit 120 , a learning processor 130 , a sensing unit 140 , an output unit 150 , a memory 170 , and a processor 180 .
[0042] The communication unit 110 may use wired or wireless communication technology to transmit and receive data with an external device such as other AI devices 100a to 100e or the AI server 200. For example, the communication unit 110 may transmit and receive sensor information, user input, learning models, control signals, etc. with the external device.
[0043] At this time, the communication technology used by the communication unit 110 includes Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Long Term Evolution (LTE), 5G, Wireless LAN (WLAN), Wi-Fi (Wireless Fidelity), BluetoothTM, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), ZigBee, Near Field Communication (NFC), etc.
[0044] The input unit 120 may obtain various types of data.
[0045] At this time, the input unit 120 can include a camera for inputting video signals, a microphone for receiving audio signals, a user input module for receiving information from the user, etc. Here, the camera or microphone can be regarded as a sensor, and the signal obtained from the camera or microphone can be referred to as sensing data or sensor information.
[0046] The input unit 120 may obtain input data to be used when obtaining output using the learning data and the learning model for model learning. The input unit 120 may obtain unprocessed input data, and in this case, the processor 180 or the learning processor 130 may extract input features as preprocessing of the input data.
[0047] The learning processor 130 can use the learning data to learn a model composed of an artificial neural network. The learned artificial neural network can be referred to as a learning model. The learning model can be used to infer the result value for new input data that is not the learning data, and the inferred value can be used as the basis for determining a certain action.
[0048] At this time, the learning processor 130 can Figure 2 The AI server 200 and the learning processor 240 together perform AI processing.
[0049] At this time, the learning processor 130 may include a memory integrated or implemented in the AI device 100. Alternatively, the learning processor 130 may be implemented using the memory 170, an external memory directly connected to the AI device 100, or a memory retained in an external device.
[0050] The sensing unit 140 may obtain at least one of internal information of the AI device 100 , surrounding environment information of the AI device 100 , and user information using various sensors.
[0051] At this time, the sensors included in the sensing unit 140 include a proximity sensor, an illumination sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a lidar, a radar, and the like.
[0052] The output unit 150 may generate output relevant to the senses of sight, hearing, or touch.
[0053] At this time, the output unit 150 may include a display module that outputs visual information, a speaker that outputs auditory information, a haptic module that outputs tactile information, and the like.
[0054] The memory 170 may store data supporting various functions of the AI device 100. For example, the memory 170 may store input data obtained from the input unit 120, learning data, a learning model, a learning history, and the like.
[0055] The processor 180 may determine at least one executable operation of the AI device 100 based on information determined or generated using a data analysis algorithm or a machine learning algorithm. In addition, the processor 180 may control components of the AI device 100 to perform the determined operation.
[0056] To this end, the processor 180 may request, search, receive, or utilize data from the learning processor 130 or the memory 170 and control the components of the AI device 100 to perform a predicted operation or an operation determined to be desirable among at least one executable operation.
[0057] At this time, if the link of the external device is required to perform the determined operation, the processor 180 may generate a control signal for controlling the external device and transmit the generated control signal to the external device.
[0058] The processor 180 may obtain intention information for the user input and determine the user's requirement based on the obtained intention information.
[0059] At this time, the processor 180 can obtain intent information corresponding to the user input by using at least one of a speech-to-text (STT) engine for converting voice input into a character string or a natural language processing (NLP) engine for obtaining intent information of a natural language.
[0060] At this time, at least one of the STT engine or the NLP engine can be configured with an artificial neural network that is learned at least in part based on a machine learning algorithm. In addition, at least one of the STT engine or the NLP engine can be learned by the learning processor 130, by the learning processor 240 of the AI server 200, or through distributed processing of these processors.
[0061] The processor 180 may collect historical information including the operation content of the AI device 100 or user feedback on the operation, and store it in the memory 170 or the learning processor 130, or transmit it to an external device such as the AI server 200. The collected historical information may be used to update the learning model.
[0062] The processor 180 may control at least some components of the AI device 100 to drive the application stored in the memory 170. In addition, the processor 180 may operate two or more components included in the AI device 100 in combination to drive the application.
[0063] Figure 2 An AI server 200 according to an embodiment of the present disclosure is shown.
[0064] refer to Figure 2 , the AI server 200 may refer to a device that uses a machine learning algorithm to train an artificial neural network or uses a trained artificial neural network. Here, the AI server 200 may be composed of multiple servers to perform distributed processing and may be defined as a 5G network. In this case, the AI server 200 may be included as part of the AI device 100 and may perform at least part of the AI processing together.
[0065] The AI server 200 may include a communication unit 210 , a memory 230 , a learning processor 240 , a processor 260 , and the like.
[0066] The communication unit 210 may transmit and receive data with an external device such as the AI device 100 .
[0067] The memory 230 may include a model storage unit 231. The model storage unit 231 may store a model (or an artificial neural network, 231a) being learned or learned by the learning processor 240.
[0068] The learning processor 240 may learn the artificial neural network 231a using the learning data. The learning model may be used while being installed in the AI server 200 of the artificial neural network, or may be used while being installed in an external device such as the AI device 100.
[0069] The learning model may be implemented by hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented by software, one or more instructions constituting the learning model may be stored in the memory 230 .
[0070] Processor 260 may use the learned model to infer result values for new input data and generate responses or control commands based on the inferred result values.
[0071] Figure 3 An AI system 1 according to an embodiment of the present invention is shown.
[0072] Reference Figure 3 , the AI system 1 is connected to at least one of the AI server 200, the robot 100a, the self-driving vehicle 100b, the XR device 100c, the smartphone 100d, or the home appliance 100e through the cloud network 10. Here, the robot 100a, the self-driving vehicle 100b, the XR device 100c, the smartphone 100d, or the home appliance 100e that applies AI technology may be referred to as AI devices 100a to 100e.
[0073] The cloud network 10 may mean a network that constitutes a part of a cloud computing infrastructure or exists within a cloud computing infrastructure. Here, the cloud network 10 may be configured using a 3G network, a 4G or LTE network, a 5G network, or the like.
[0074] That is, the devices 100a to 100e and 200 constituting the AI system 1 can be connected to each other via the cloud network 10. Specifically, the devices 100a to 100e and 200 can communicate with each other via a base station, but can also communicate with each other directly without going through the base station.
[0075] The AI server 200 may include a server that performs AI processing and a server that performs calculations on big data.
[0076] The AI server 200 is connected to at least one or more of the AI devices constituting the AI system 1 (such as a robot 100a, a self-driving vehicle 100b, an XR device 100c, a smart phone 100d, or a home appliance 100e) through the cloud network 10, and can assist at least a portion of the AI processing of the connected AI devices 100a to 100e.
[0077] At this time, the AI server 200 can train the artificial neural network according to the machine learning algorithm on behalf of the AI devices 100a to 100e, and can directly store the learning model or send it to the AI devices 100a to 100e.
[0078] At this time, the AI server 200 may receive input data from the AI devices 100a to 100e, infer result values of the received input data using a learning model, and generate a response or control command based on the inferred result value and send it to the AI devices 100a to 100e.
[0079] Alternatively, the AI devices 100 a to 100 e may directly infer result values of the input data using a learning model, and generate a response or control command based on the inferred result values.
[0080] Figures 4 to 6 2 is a diagram for illustrating a device-side AI system according to an embodiment of the present disclosure.
[0081] like Figure 4 As shown, the device-side AI system of the present disclosure may include a device-side AI apparatus 500 that performs AI translation processing to translate speech of multiple speakers 600 into a target language.
[0082] Here, the device-side AI device 500 is an artificial intelligence device capable of performing device-side AI processing, and may include permanent devices such as personal computers (PCs), network TVs, hybrid broadcast broadband TVs (HBBTVs), smart TVs, Internet Protocol TVs (IPTVs), etc., as well as mobile devices (or handheld devices) such as smart phones, tablet PCs, notebooks, PDAs, smart watches, smart glasses, robots, etc.
[0083] When speech voices are input from multiple speakers 600, the device-side AI device 500 preprocesses the speech voices, classifies the preprocessed speech voices by speaker, and when a specific speaker is selected from the multiple speakers 600, extracts the speech voice of the selected specific speaker from the speech voices classified by speaker, translates the speech voice of the specific speaker into the target language, and outputs it.
[0084] In some cases, such as Figure 5 and Figure 6 As shown, the device-side AI system of the present disclosure may include a device-side AI device 500 that performs AI translation processing to translate the speech of multiple speakers 600 into a target language, and at least one AI translation processing device 700 that performs AI translation processing in response to a translation proxy service request or a distributed processing request of the device-side AI device 500.
[0085] Here, the AI translation processing device 700 can be the same device as the device-side AI device 500 or a different device.
[0086] For example, Figure 5As shown, the AI translation processing device 700 may include at least one of the following: a permanent device including a PC, a network TV, an HBBTV, a smart TV, an IPTV, etc., which performs translation processing based on a pre-learned AI model and outputs the translation processing execution result; and a mobile device (handheld device) including a smart phone, a tablet PC, a laptop computer, a PDA, a robot, etc.
[0087] As another example, Figure 6 As shown, the AI translation processing device 700 may include a battery pack 710, headphones 720, smart glasses 730, smart watches 740, etc. that store at least one AI model and perform translation processing only based on the pre-learned AI model and do not output the translation processing execution result, and may further include various peripheral devices such as external memory, smart health band, adapter, global positioning system (GPS), PDA, barcode reader, character recognition device, voice recognition device, etc.
[0088] Here, the AI translation processing device 700 can be connected to the device-side AI device 500 in a wired and wireless manner.
[0089] As an example, when a user command requesting a translation proxy service is input, the device-side AI device 500 requests a translation proxy service for the speech voice of a specific speaker from the AI translation processing device 700, and when the device-side AI device 500 receives approval for the translation proxy service request, the device-side AI device 500 is able to send the speech voice data of the specific speaker and the target language information to be translated to the AI translation processing device 700, so that the AI translation processing device 700 translates the speech voice of the specific speaker into the target language and outputs it.
[0090] Here, when a user command requesting a translation agent service is input, the device-side AI device 500 can check whether the translation of the speech voice of a specific speaker is currently being performed, and if the translation of the speech voice of the specific speaker is currently being performed, the translation of the speech voice of the specific speaker can be stopped, and the speech voice data of the specific speaker input after the time point when the translation is stopped can be obtained.
[0091] Then, when a translation proxy service request is received from the device-side AI device 500, the AI translation processing device 700 measures the AI translation processing amount corresponding to the translation proxy service request to determine whether self-translation processing can be performed, and if it is determined that self-translation processing can be performed, the AI translation processing device 700 sends approval of the translation proxy service request to the device-side AI device 500, and when the speech voice data of the speaker 600 and the target language information to be translated are received from the device-side AI device 500, the AI translation processing device 700 is able to translate the speech voice of the speaker 600 into the target language and output it.
[0092] As another embodiment, when a user command requesting to suspend the translation of some speakers corresponding to the translation proxy service is input, the device-side AI device 500 can request the AI translation processing device 700 to suspend the translation proxy service of some speakers, and receive information about the completion of the suspension of the translation proxy service of some speakers and information about the continuation of the translation proxy service of the remaining speakers from the AI translation processing device 700.
[0093] Here, when the AI translation processing device 700 receives a request to suspend the translation proxy service for some speakers from the device-side AI device 500, the AI translation processing device 700 can suspend the translation proxy service for some speakers and transmit information on the completion of the suspension of the translation proxy service for some speakers and the continuation of the translation proxy service for the remaining speakers to the device-side AI device 500.
[0094] As another embodiment, the device-side AI device 500 can measure the AI translation processing amount used to convert the speaker's speech voice into text to translate the current language into the target language, and if the measured AI translation processing amount exceeds the amount that can be self-processed, request distributed processing of the AI translation processing from the AI translation processing device 700, and if a first AI translation processing result value is received from the AI translation processing device 700, a final translation result value can be provided based on the first AI translation processing result value and the self-processed second AI translation processing result value.
[0095] Here, when the AI translation processing device 700 receives a distributed processing request from the device-side AI device 500, it extracts distributed processing information from the distributed processing request, performs AI translation processing based on the distributed processing information to generate an AI translation processing result value, and provides the generated AI translation processing result value to the device-side AI device 500.
[0096] At this time, the AI translation processing device 700 can extract distributed processing information corresponding to the distributed processing part, including AI model information, location information, distributed processing amount information, and input data, from the distributed processing request.
[0097] In this way, if the device-side AI device 500 is a device that performs translation processing based on the AI model that has been pre-learned by the AI translation processing device 700 and outputs the translation processing execution result, then Figure 5 As shown, the device-side AI device 500 can request the AI translation processing device 700 for a translation proxy service or can request the AI translation processing device 700 for distributed processing of the AI translation process.
[0098] In addition, if the device-side AI device 500 is a device that only performs translation processing based on the AI model that has been pre-learned by the AI translation processing device 700 without outputting the translation processing execution result, then Figure 6 As shown, the device-side AI device 500 can request the AI translation processing device 700 for distributed processing of AI translation processing.
[0099] Meanwhile, the device-side AI apparatus 500 stores at least one AI model and may provide various service results through the pre-learned AI model corresponding to the command of the speaker 600 .
[0100] Here, the AI model may be a deep neural network (DNN) including a plurality of hidden layers in addition to an input layer and an output layer.
[0101] AI models can identify the underlying structure of data such as photos, text, video, speech, and music.
[0102] Figure 7 and Figure 8 2 is a diagram for illustrating a device-side AI apparatus according to an embodiment of the present disclosure.
[0103] like Figure 7 As shown, the device-side AI device 500 of the present disclosure may include: an input module 510, into which the speaker's speech voice is input; a processor 520, which performs AI processing to translate the speaker's speech voice into a target language; and a memory 530, which stores at least one AI model 540.
[0104] The processor 520 of the present disclosure may perform translation processing based on a pre-learned AI model.
[0105] When speech sounds are input from multiple speakers, the processor 520 may pre-process the speech sounds, classify the pre-processed speech sounds by speaker, and when a specific speaker is selected from the multiple speakers, extract the speech sounds of the selected specific speaker from the speech sounds classified by speaker, and translate the speech sounds of the specific speaker into a target language and output it.
[0106] When pre-processing the utterance voice, the processor 520 may perform pre-processing by analyzing a frequency corresponding to the utterance voice when the utterance voice of the speaker is input and removing a noise frequency.
[0107] Here, when analyzing the frequency corresponding to the utterance voice, the processor 520 may check whether a specific frequency outside the human voice frequency range exists within the frequency corresponding to the utterance voice, and if so, identify the specific frequency as a noise frequency.
[0108] In some cases, the processor 520 may input the speaker's utterance speech into a pre-learned noise classification model to classify and remove noise frequencies.
[0109] In addition, when the processor 520 classifies the preprocessed speech speech according to the speaker, it can extract speech features of the preprocessed speech speech, identify the speaker of the preprocessed speech speech based on the extracted speech features, and classify the preprocessed speech speech according to the identified speaker.
[0110] Here, the processor 520 may input the preprocessed utterance speech into a pre-learned feature extraction model to extract speech features of the preprocessed utterance speech, and input the extracted speech features into a pre-learned speaker recognition model to recognize the speaker of the preprocessed utterance speech.
[0111] In some cases, when the voice features of the preprocessed speech voice are extracted, the processor 520 may select a speaker whose voice is most similar to the voices of a pre-registered speaker voice list based on the extracted voice features, and may match the preprocessed speech voice to the selected speaker to distinguish the speakers.
[0112] In addition, when extracting the voice features of the preprocessed speech voice, the processor 520 may check whether the preprocessed speech voice is a mixed speech in which the speech voices of multiple speakers are mixed, and if the speech voice is a mixed speech, perform speaker separation on the mixed speech to separate the mixed speech into individual speech, and extract voice features for each separated individual speech.
[0113] In some cases, when extracting speech features of preprocessed speech speech, the processor 520 may check whether the preprocessed speech speech is continuous speech of speech speech of multiple speakers, and if the speech speech is continuous speech, perform speaker classification on the continuous speech to separate the continuous speech into speaker unit speech, group the separated speaker unit speech by speaker, and extract speech features for each speaker unit speech grouped by speaker.
[0114] Next, when a specific speaker is selected, if the utterance voice is separated by speaker, the processor 520 may generate and provide a speaker list corresponding to the utterance voice, and when a user input for selecting at least one speaker included in the speaker list is received, the speaker selected by the user input may be selected as the specific speaker.
[0115] Then, when selecting a specific speaker, if the utterance voice is divided by speaker, the processor 520 analyzes the utterance voice data volume of each speaker within a predetermined period of time and selects the specific speaker from a plurality of speakers based on the utterance voice data volume of each speaker.
[0116] Here, the processor 520 compares the utterance voice data volume of each speaker with a preset standard data volume, and selects a speaker whose utterance voice data volume is greater than or equal to the standard data volume as a specific speaker.
[0117] In addition, if there are a plurality of speakers whose utterance voice data amount is greater than or equal to the standard data amount, the processor 520 may select a speaker whose utterance voice data amount is the largest as a specific speaker.
[0118] Here, if there are multiple speakers whose speech voice data is greater than or equal to the standard data amount, the processor 520 checks whether the standard number of speakers for the specific speaker is preset, and if the standard number of speakers for the specific speaker is preset, if the number of speakers whose speech voice data is greater than or equal to the standard data amount is greater than or equal to the standard number of speakers, the processor selects the specific speaker in the order of the largest speech voice data amount, and if the number of speakers whose speech voice data is greater than or equal to the standard data amount is less than the standard number of speakers, the processor may select the specific speaker in the order of the number of speakers whose speech voice data is greater than or equal to the standard data amount.
[0119] At this time, when the standard number of speakers is preset, the processor 520 may preset the standard number of speakers based on the translation processing amount for the utterance voice of the speaker.
[0120] In some cases, when selecting a specific speaker, if the speaker's speech voice is classified by speaker, the processor 520 can convert the speaker's speech voice for a predetermined time period into text, analyze the converted text to extract common associated keywords by speaker, and select the specific speaker among multiple speakers based on the common associated keywords by speaker.
[0121] Here, when extracting common related keywords, the processor 520 may extract common related keywords including conference topic keywords and related keywords from the text corresponding to the speaker's speech, and group the extracted common related keywords by speaker.
[0122] In addition, the processor 520 may compare the number of commonly associated keywords per speaker with a preset standard keyword number, and select a speaker whose number of commonly associated keywords is greater than the standard keyword number as a specific speaker.
[0123] Here, if there are multiple speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords, the processor 520 may select the speaker with the largest number of commonly associated keywords as the specific speaker.
[0124] In addition, if there are multiple speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords, the processor 520 can check whether the standard number of speakers for the specific speaker has been preset, and if the standard number of speakers for the specific speaker has been preset, if the number of speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords is greater than or equal to the standard number of speakers, the specific speaker is selected in the order of the largest number of commonly associated keywords, and if the number of speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords is less than or equal to the standard number of speakers, the specific speaker is selected in the order of the number of speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords.
[0125] Here, when presetting the number of standard speakers, the processor 520 may pre-set the number of standard speakers based on the translation processing capability of the speaker's speech.
[0126] Next, when extracting the speech voice of a specific speaker, the processor 520 may, when a specific speaker is selected from multiple speakers, leave only the speech voice of the selected specific speaker among the speech voices classified by speaker, and remove the speech voices of the remaining speakers except the specific speaker.
[0127] Here, when removing the utterance voices of the remaining speakers, the processor 520 may leave only the utterance voices of other speakers that are continuous before or after the utterance voice time of the specific speaker and remove the utterance voices of the remaining speakers.
[0128] In addition, in order to extract reference inference keywords when translating the utterance of a specific speaker, the processor 520 may store only the utterances of other speakers that are continuous before or after the utterance of the specific speaker in the database.
[0129] Then, when translating the speech voice of the specific speaker, the processor 520 converts the speech voice of the specific speaker into text to check the current language, and when the target language is set, measures the AI translation processing amount used to translate the current language into the target language to determine whether self-translation processing can be performed, and when it is determined that self-translation processing can be performed, the processor 520 can translate the speech voice of the specific speaker into the target language by self-processing the AI translation processing.
[0130] Here, when setting the target language, if the current language is confirmed, the processor 520 checks whether the target language to be translated is set. If the target language is not set, the processor 520 generates and provides a target language list window. If a user input for selecting a specific language is received through the target language list window, the selected specific language can be set as the target language.
[0131] In some cases, if the target language is not set, the processor 520 may generate and provide a user notification requesting that the target language be set.
[0132] For example, the processor 520 may generate and output a user notification in at least one of a text form and a sound form.
[0133] In addition, when determining whether translation processing can be performed, the processor 520 can measure the AI translation processing volume used to translate the current language into the target language. If the measured AI translation processing volume is less than or equal to the self-processing capability (or self-processing volume), the processor 520 can determine that self-translation processing can be performed.
[0134] In some cases, if the measured AI translation processing amount exceeds the self-processing capability, the processor 520 may determine that self-translation processing is not possible, select an AI translation processing device for distributed processing of the AI translation processing, request distributed processing of the AI translation processing from the selected AI translation processing device, and upon receiving a first AI translation processing result value from the AI translation processing device, the processor may provide a final translation result value based on the first AI translation processing result value and a second AI translation processing result value of the self-processing.
[0135] Here, when selecting an AI translation processing device for distributed processing of AI translation processing, the processor 520 checks whether there is a communication connection with an external device, and if the external device is connected, obtains identification information of the external device from the external device, and checks whether the external device is an AI translation processing device based on the identification information, and if the external device is an AI translation processing device, selects the external device as the AI translation processing device for distributed processing of AI translation processing.
[0136] like Figure 8As shown, the present disclosure may further include a communication module 550 connected to an external device by wire or wirelessly, and the processor 520 may check whether there is a communication connection with the external device through the communication module 550 .
[0137] When checking whether the external device is an AI translation processing device, if the external device stores an AI model for translation processing, the processor 520 may identify the external device as an AI translation processing device.
[0138] For example, the AI translation processing device may include at least one of permanent devices including PC, network TV, HBBTV, smart TV and IPTV, and mobile devices or handheld devices including smart phones, tablet PCs, notebooks, PDAs and robots that perform translation processing based on the AI model and output translation processing results.
[0139] For another example, the AI translation processing device may include at least one of a battery pack, headphones, external memory, smart watch, smart glasses, smart health band, adapter, GPS, PDA, barcode reader, character recognition device and speech recognition device, which only performs translation processing based on the AI model and does not output the translation processing result.
[0140] When selecting an AI translation processing device, if there are multiple external devices, the processor 520 may select all external devices identified as AI translation processing devices as AI translation processing devices, and may assign priorities to the multiple external devices based on processing performance indicators of the multiple external devices selected as AI translation processing devices.
[0141] Here, the processor 520 may assign the highest priority to an external device having the highest processing performance index among the plurality of external devices, and may assign the lowest priority to an external device having the lowest processing performance index.
[0142] Therefore, if Figure 8 As shown, if the measured AI translation processing amount exceeds its own processing capacity, the processor 520 can check the processing capacity of the external device for the excess amount, and if multiple external devices are required for the excess amount, it can request distributed processing of the AI translation processing for the excess amount from the external devices according to the priority assigned to them.
[0143] Here, when requesting AI translation processing for an excess amount in a distributed manner, the processor 520 may request a plurality of external devices to process different excess amounts.
[0144] That is, when requesting processing of different excess amounts, the processor 520 may request processing of a first excess amount from an external device with a high priority, and may request processing of a second excess amount from an external device with a low priority, the second excess amount being the remaining portion of the total excess amount other than the first excess amount.
[0145] In some cases, when AI translation processing for an excess amount is requested in a distributed manner, the processor 520 may request processing of the same excess amount from a plurality of external devices.
[0146] Here, the processor 520 may equally distribute the total excess amount to a plurality of external devices, request processing of a first excess amount from external devices having a high priority, and request processing of a second excess amount the same as the first excess amount from external devices having a low priority.
[0147] In addition, when requesting distributed processing of AI translation processing, the processor 520 calculates the excess amount of the measured AI translation processing amount excluding its own processing capacity, checks the processing capacity of the selected AI translation processing device, and if the processing capacity of the AI translation processing device is greater than the excess amount, the processor 520 can request the AI translation processing device to perform distributed processing of the AI translation processing for the excess amount.
[0148] Here, if the processing capacity of the AI translation processing device is less than the excess amount, the processor 520 may additionally select another AI translation processing device and request distributed processing of the AI translation processing for the excess amount from the plurality of AI translation processing devices.
[0149] In addition, when requesting distributed processing of AI translation processing, the processor 520 can calculate the excess amount of the measured AI translation processing amount excluding the self-processing capacity, extract the distributed processing part corresponding to the excess amount in the AI translation processing, and request the selected AI translation processing device to perform distributed processing of the AI translation processing for the extracted distributed processing part.
[0150] Here, when extracting the distributed processing part corresponding to the excess amount, the processor 520 can analyze the AI model used for translation processing to check whether there are branch points connecting one upper-level operator to multiple lower-level operators and junction points joining multiple upper-level operators to one lower-level operator, and if there are branch points and junction points, it checks whether there is at least one parallel processing part based on the branch points and junction points, and can extract the parallel processing part as the distributed processing part.
[0151] In some cases, when extracting the distributed processing part corresponding to the excess amount, if there are multiple AI models for translation processing, the processor 520 can check whether there is an AI model that can be processed in parallel among the multiple AI models, and if there is an AI model that can be processed in parallel, it can extract the processing part performed by the parallel processing AI model as the distributed processing part.
[0152] Here, when requesting distributed processing of the AI translation process, the processor 520 may provide the AI translation processing device with a distributed processing request including AI model information, location information, distributed processing amount information, and input data corresponding to the distributed processing part.
[0153] In addition, when providing a final translation result value, the processor 520 can map the first AI translation processing result value received from the AI translation processing device and the second AI translation processing result value processed by itself to provide a final translation result value that translates the speech voice of a specific speaker into the target language.
[0154] In addition, when translating and outputting the utterance voice of a specific speaker, the processor 520 may output the translation result in at least one of a first method of outputting the translation result as voice through a speaker and a second method of outputting the translation result as text through a display.
[0155] At the same time, if Figure 8 As shown, the present disclosure may further include a communication module 550, which is connected to an AI translation processing device by wire or wirelessly, and the AI translation processing device performs translation processing based on an AI model, wherein when a user command requesting a translation proxy service is input, the processor 520 checks whether a communication connection is established with the AI translation processing device, and when a communication connection is established with the AI translation processing device, requests the AI translation processing device to provide a translation proxy service for the speech voice of a specific speaker, and when approval of the translation proxy service request is received from the AI translation processing device, the processor 520 may send the speech voice data of the specific speaker and information about the target language to be translated to the AI translation processing device, so that the AI translation processing device translates the speech voice data of the specific speaker into the target language and outputs it.
[0156] Here, when a user command requesting a translation agent service is input, the processor 520 checks whether translation of the speech voice of a specific speaker is currently being performed, and if translation of the speech voice of the specific speaker is currently being performed, the processor stops the translation of the speech voice of the specific speaker, and can obtain the speech voice data of the specific speaker input after the translation is stopped.
[0157] As an example, when a user command requesting interruption of translation for some speakers corresponding to the translation proxy service is input, the processor 520 may request the AI translation processing device to interrupt the translation proxy service for some speakers, and receive information about completion of the interruption of the translation proxy service for some speakers and information about continuation of the translation proxy service for the remaining speakers from the AI translation processing device.
[0158] In another case, when the processor 520 receives a distributed processing request from the AI translation processing device, it can extract distributed processing information from the distributed processing request, perform AI translation processing based on the distributed processing information to generate an AI translation processing result value, and provide the generated AI translation processing result value to the AI translation processing device.
[0159] Here, when extracting distributed processing information, the processor 520 may extract distributed processing information corresponding to the distributed processing part including AI model information, location information, distributed processing amount information, and input data from the distributed processing request.
[0160] In another case, when the processor 520 receives a translation proxy service request from the AI translation processing device, it measures the AI translation processing amount corresponding to the translation proxy service request to determine whether self-translation processing can be performed, and if it determines that self-translation processing can be performed, it sends approval of the translation proxy service request to the AI translation processing device, and when it receives the speaker's speech voice data and the target language information to be translated from the AI translation processing device, it is able to translate the speaker's speech voice into the target language and output it.
[0161] Here, the processor 520 may suspend the translation proxy service for some speakers upon receiving a request to suspend the translation proxy service for some speakers from the AI translation processing device, and send information on completion of suspending the translation proxy service for some speakers and information on continuing the translation proxy service for the remaining speakers to the AI translation processing device.
[0162] In this way, when the device-side AI device of the present invention selects a specific speaker from multiple speakers, it extracts the speech voice of the specific speaker from the speech voice classified by speaker and translates it into the target language, thereby being able to identify the conversation between multiple speakers by speaker and translate it accurately in real time.
[0163] In addition, when the amount of AI processing used to translate the speaker's speech exceeds the self-processing capacity, the present invention selects an external AI translation processing device and requests distributed processing of the AI translation processing, thereby improving the speed of the AI translation processing and the accuracy of the translation processing results and the service quality.
[0164] Furthermore, the present disclosure can minimize power consumption and improve performance and lifespan by distributing AI translation processing with externally located autonomous AI translation processing.
[0165] Figures 9 to 11 3 is a diagram for explaining a speaker recognition process of a device-side AI apparatus according to an embodiment of the present disclosure.
[0166] like Figure 9 As shown, the present disclosure can pre-process utterance speech when utterance speech is input from a speaker, and recognize the pre-processed utterance speech by speaker.
[0167] The present disclosure may include a preprocessing module 810 , a speech feature extraction module 820 , and a speaker recognition module 830 , wherein the preprocessing module 810 preprocesses the speech, the speech feature extraction module 820 extracts features of the preprocessed speech, and the speaker recognition module 830 recognizes the speaker of the speech based on the speech features.
[0168] Here, the preprocessing module 810 may perform preprocessing by analyzing a frequency corresponding to an utterance voice of a speaker when the utterance voice is input and removing a noise frequency.
[0169] For example, when analyzing the frequency corresponding to the speech voice, the preprocessing module 810 can check whether there is a specific frequency outside the human speech frequency range within the frequency corresponding to the speech voice, and if there is a specific frequency, it can identify the specific frequency as a noise frequency.
[0170] In some cases, the pre-processing module 810 may input the speaker's utterance speech into a pre-learned noise classification model to classify and remove noise frequencies.
[0171] In addition, the speech feature extraction module 820 extracts speech features of the preprocessed speech speech, and the speaker recognition module 830 may recognize the speaker of the preprocessed speech speech based on the extracted speech features and classify the preprocessed speech speech by the recognized speaker.
[0172] Here, the speech feature extraction module 820 inputs the preprocessed speech speech into a pre-learned feature extraction model to extract speech features of the preprocessed speech speech, and the speaker recognition module 830 inputs the extracted speech features into a pre-learned speaker recognition model to identify the speaker of the preprocessed speech speech.
[0173] In some cases, when the speech features of the preprocessed speech voice are extracted, the speaker recognition module 830 may select a speaker whose speech has the highest similarity in a pre-registered speaker speech list based on the extracted speech features, and match the preprocessed speech voice with the selected speaker to distinguish each speaker.
[0174] In addition, if Figure 10 As shown, the present disclosure can check whether the preprocessed speech voice is a mixed speech in which speech voices of multiple speakers are mixed, and if the speech voice is a mixed speech, perform speaker separation on the mixed speech voice to separate the mixed speech voice into individual speech, and can extract speech features for each separated individual speech.
[0175] In some cases, such as Figure 11 As shown, the present disclosure can check whether the preprocessed speech speech is a continuous speech of speech speech of multiple speakers, and if the speech speech is a continuous speech, perform speaker classification on the continuous speech to separate the continuous speech into speaker unit speech, and the separated speaker unit speech can be grouped by speaker, and speech features can be extracted for each speaker unit speech grouped by speaker.
[0176] Figures 12 to 15 3 is a diagram for explaining a speaker selection process of a device-side AI apparatus according to an embodiment of the present disclosure.
[0177] like Figure 12 As shown in , the present disclosure may select a specific speaker by classifying 842 the utterance voice by speaker, generating 844 a speaker list corresponding to the utterance voice, and when receiving a user input for selecting at least one speaker included in the speaker list, selecting 846 the speaker selected by the user input as the specific speaker.
[0178] like Figure 13 As shown, the present disclosure may generate a speaker list window 910 and output it on the screen of the device-side AI apparatus 500 .
[0179] Here, the speaker list window 910 may include various items such as a speaker selection item, an utterance voice providing item for each speaker, and an utterance voice data amount for each speaker.
[0180] Furthermore, when the speaker list window 910 is provided, the present disclosure may provide a speaker selection request notification message 920 together with the speaker list window 910 .
[0181] Here, the present disclosure may generate and output the speaker selection request notification message 920 in at least one of a text form and a sound form.
[0182] like Figure 14 As shown, the present disclosure can select a specific speaker by dividing speech voice by speaker 852, analyzing speech voice data volume of the speaker within a predetermined time period 854, and selecting a specific speaker 856, 858 from multiple speakers based on the speech voice data volume of the speaker.
[0183] Here, the present disclosure may compare the utterance voice data amount of the speaker with a preset reference data amount, and select a speaker whose utterance voice data amount is greater than or equal to the reference data amount as a specific speaker 856 .
[0184] Furthermore, the present disclosure may select a speaker having the largest amount of utterance voice data as a specific speaker in the case where there are a plurality of speakers whose utterance voice data amounts are greater than or equal to a reference data amount.
[0185] Here, the present disclosure checks whether a standard number of speakers for a specific speaker is preset when there are multiple speakers whose speech voice data is greater than or equal to the standard data amount, and if the standard number of speakers for the specific speaker is preset, if the number of speakers whose speech voice data is greater than or equal to the standard data amount is greater than or equal to the standard number of speakers, the specific speaker is selected in the order of the largest speech voice data amount, and if the number of speakers whose speech voice data is greater than or equal to the standard data amount is less than the standard number of speakers, the specific speaker 858 can be selected in the order of the number of speakers whose speech voice data is greater than or equal to the standard data amount.
[0186] At this time, the present disclosure may preset the standard number of speakers based on the translation processing capability of the speaker's utterance voice when presetting the standard number of speakers.
[0187] like Figure 15 As shown, when selecting a specific speaker, the present disclosure classifies the speech voices by speaker 862, converts the speech voices of the speakers in a predetermined time period into text 864, analyzes the converted text to extract common associated keywords of the speakers 866, and selects a specific speaker from multiple speakers based on the common associated keywords of the speakers 868.
[0188] Here, when extracting commonly related keywords, the present disclosure extracts commonly related keywords including conference topic keywords and related keywords from the text corresponding to the speaker's speech, and groups the extracted commonly related keywords by speaker.
[0189] In addition, the present disclosure compares the number of associated keywords in common among speakers with a preset number of reference keywords, and selects a speaker whose number of associated keywords in common is greater than the number of reference keywords as a specific speaker 868 .
[0190] Here, the present disclosure may select a speaker with the largest number of commonly associated keywords as a specific speaker when there are multiple speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords.
[0191] Furthermore, if there are multiple speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords, the present disclosure may select the speaker with the largest number of commonly associated keywords as the specific speaker. If the standard number of speakers for a specific speaker is preset, then if the number of speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords is greater than or equal to the standard number of speakers, the present disclosure may select the specific speaker in the order of the largest number of commonly associated keywords. If the number of speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords is less than or equal to the standard number of speakers, the present disclosure may select the specific speaker in the order of the number of speakers whose number of commonly associated keywords is greater than or equal to the standard number of keywords.
[0192] Here, the present disclosure may pre-set the standard number of speakers based on the translation processing capability of the speaker's utterance voice when pre-setting the standard number of speakers.
[0193] Figure 16 2 is a diagram for explaining a process of setting a target language of a device-side AI apparatus according to an embodiment of the present disclosure.
[0194] In the present disclosure, when translating a specific speaker's spoken speech, the specific speaker's spoken speech is converted into text, the current language is confirmed through the converted text, and if the current language is confirmed, it can be confirmed whether a target language to be translated is set.
[0195] Here, as Figure 16 As shown, in the present disclosure, if the target language is not set, a translation language list window 1010 may be generated and output on the screen of the device-side AI apparatus 500 .
[0196] At this time, the translation language list window 1010 may include various items such as a current language item and a target language selection item.
[0197] Furthermore, when the translation language list window 1010 is provided, the present disclosure may generate and provide a user notification 1020 requesting to set a target language together with the translation language list window 1010 .
[0198] As an example, the present disclosure may generate and output the user notification 1020 in at least one of a text form and a sound form.
[0199] Figure 17 and Figure 181 is a diagram for explaining a process of measuring the AI translation processing capacity of a device-side AI apparatus according to an embodiment of the present disclosure.
[0200] The present disclosure measures an AI translation processing amount for translating a current language into a target language by converting a speaker's utterance voice into text, and if the measured AI translation processing amount exceeds a self-processing capacity, requests distributed processing of the AI translation processing to an AI translation processing device, and if a first AI translation processing result value is received from the AI translation processing device, a final translation result value can be provided based on the first AI translation processing result value and a second AI translation processing result value that has been processed by itself.
[0201] Here, when measuring the AI translation processing amount, the present disclosure can measure the AI translation processing amount used for translation based on the calculation amount used by the AI model 540 for translation.
[0202] like Figure 17 As shown, if there is one AI model 540 for translation, the present disclosure may measure the AI translation processing amount based on the amount of computation processed by one AI model 640 .
[0203] That is, the present disclosure may analyze the AI model 540 to determine whether there are branch points 542 connecting one upper-level operator to multiple lower-level operators and joining points 544 joining multiple upper-level operators to one lower-level operator.
[0204] In addition, if a branch point 542 and a junction point 544 exist, the present disclosure can determine whether there is at least one parallel processing part 546 based on the branch point 542 and the junction point 544, and if a parallel processing part 546 exists, the AI translation processing amount can be measured based on the first computing amount of the parallel processing part 546 and the second computing amount of the remaining parts except the parallel processing part.
[0205] In addition, the present disclosure may determine whether to perform distributed processing of the parallel processing section 546 in the entire AI translation processing volume for one AI model when the AI translation processing volume for translation exceeds the self-processing capability.
[0206] As another example, Figure 18 As shown, if there are multiple AI models 540 , the present disclosure may measure the AI translation processing amount based on the total computational amount processed by the multiple AI models 540 .
[0207] Here, the present disclosure analyzes the AI model group to determine whether there is a branch point 548 connecting one upper-level AI model to multiple lower-level AI models and a junction point 549 joining multiple upper-level AI models to one lower-level AI model, and if there are branch points 548 and junction points 549, it is determined based on the branch points 548 and junction points 549 whether there is at least one parallel processing part, and if there are parallel processing parts, the AI translation processing amount can be measured based on the first computing amount of the parallel processing part and the second computing amount of the remaining parts except the parallel processing part.
[0208] In addition, the present disclosure can determine whether to perform distributed processing on the parallel processing portion of the entire AI translation processing when the AI translation processing volume for performing translation on an AI model group including multiple AI models exceeds the self-processing capacity.
[0209] In addition, the present disclosure can determine, for each AI model in the AI model group, whether there are branch points connecting one upper-level operator to multiple lower-level operators and joining points joining multiple upper-level operators to one lower-level operator, and if there are branch points and joining points, it can be determined based on the branch points and joining points whether there is at least one parallel processing part, and whether distributed processing is performed on the parallel processing parts.
[0210] In addition, when requesting distributed processing of AI translation processing, the present disclosure can request distributed processing of the parallel processing parts between the branch points and the junction points within each AI model at a later time, and can also request distributed processing of the overall processing part performed by the parallel processing parts between the branch points and the junction points within the AI model group from the AI translation processing device.
[0211] Figures 19 to 23 3 is a diagram for explaining a method for providing a multi-party translation service of an AI device on a device side according to an embodiment of the present disclosure.
[0212] like Figure 19 As shown, when speech voices of multiple speakers are input ( S10 ), the device-side AI apparatus of the present disclosure can pre-process the speech voices ( S20 ).
[0213] Here, the present disclosure may perform pre-processing by analyzing a frequency corresponding to an utterance voice when an utterance voice of a speaker is input and removing a noise frequency.
[0214] Furthermore, the present disclosure may distinguish the pre-processed utterance voice by speaker ( S30 ).
[0215] Here, the present disclosure may extract voice features of the preprocessed utterance voice, identify a speaker of the preprocessed utterance voice based on the extracted voice features, and distinguish the preprocessed utterance voice by the identified speaker.
[0216] Next, the present disclosure may check whether a specific speaker is selected among a plurality of speakers ( S40 ).
[0217] Here, when distinguishing utterance voices by speakers, the present disclosure may generate and provide a speaker list corresponding to the utterance voices, and when receiving a user input for selecting at least one speaker included in the speaker list, select the speaker selected by the user input as a specific speaker.
[0218] In some cases, the present disclosure may convert the speech speech of each speaker within a predetermined time period into text when distinguishing speech speech by speaker, analyze the converted text to extract common associated keywords of each speaker, and select a specific speaker among multiple speakers based on the common associated keywords of each speaker.
[0219] In other cases, the present disclosure may convert the speech speech of each speaker within a predetermined time period into text when distinguishing speech speech by speaker, analyze the converted text to extract common associated keywords of each speaker, and select a specific speaker from multiple speakers based on the common associated keywords of each speaker.
[0220] Next, when a specific speaker is selected from a plurality of speakers, the present disclosure may extract the utterance voice of the selected specific speaker from the utterance voices classified by speakers ( S50 ).
[0221] Here, when a specific speaker is selected from a plurality of speakers, the present disclosure may remove utterance voices of the remaining speakers except the specific speaker by leaving only the utterance voice of the selected specific speaker from the utterance voices classified by speakers.
[0222] Furthermore, the present disclosure can extract the speech of a specific speaker and then translate the speech of the specific speaker into a target language and output the translated speech ( S60 ).
[0223] Here, the present disclosure converts the speech voice of a specific speaker into text to check the current language, and when a target language is set, measures the amount of AI translation processing used to translate the current language into the target language to determine whether self-translation processing can be performed, and if it is determined that self-translation processing can be performed, the AI translation processing can be self-processed to translate the speech voice of the specific speaker into the target language.
[0224] Furthermore, if a specific speaker is not selected among a plurality of speakers, the present disclosure may output the voices of all speakers by translating the voices of all speakers into a target language ( S70 ).
[0225] like Figure 20 As shown, the device-side AI apparatus of the present disclosure can also execute distributed processing requests based on the AI translation processing volume for translating the current language into the target language.
[0226] like Figure 20 As shown, the present disclosure may convert a specific speaker's utterance voice into text to confirm a current language, and if a target language is set, may measure an AI translation processing amount for translating the current language into the target language ( S61 ).
[0227] Furthermore, the present disclosure may confirm whether the measured AI translation processing volume exceeds the self-processing capacity ( S62 ).
[0228] Next, the present disclosure can select an AI translation processing device for distributed processing of the AI translation process by determining that self-translation is impossible when the measured AI translation processing amount exceeds the self-processing capacity ( S63 ).
[0229] Here, the present disclosure checks whether there is a communication connection with an external device, and if the external device is connected to the communication, obtains identification information of the external device from the external device, and checks whether the external device is an AI translation processing device based on the identification information, and if the external device is an AI translation processing device, the external device can be selected as the AI translation processing device for distributed processing of AI translation processing.
[0230] Next, the present disclosure may request the selected AI translation processing device to perform distributed processing of the AI translation process ( S64 ).
[0231] Here, the present disclosure calculates an excess amount of the measured AI translation processing amount other than its own processing capacity, checks the processable amount of the selected AI translation processing device, and if the processable amount of the AI translation processing device is greater than the excess amount, a distributed processing of the AI translation processing for the excess amount can be requested from the AI translation processing device.
[0232] At this time, the present disclosure may select another AI translation processing device if the processing capacity of the AI translation processing device is less than the excess amount, and request that the AI translation processing for the excess amount be distributed to multiple AI translation processing devices.
[0233] In addition, the present disclosure can calculate the excess amount of the measured AI translation processing amount excluding the self-processing capacity, extract the distributed processing part corresponding to the excess amount in the AI translation processing, and request the selected AI translation processing device to perform AI translation processing for the extracted distributed processing part.
[0234] In addition, the present disclosure may receive a first AI translation processing result value from the AI translation processing device ( S65 ).
[0235] Next, the present disclosure may provide a final translation result value based on the first AI translation processing result value and the self-processed second AI translation processing result value ( S66 ).
[0236] Here, the present disclosure may provide a final translation result value of translating an utterance voice of a specific speaker into a target language by mapping a first AI translation processing result value received from an AI translation processing device and a self-processed second AI translation processing result value.
[0237] In addition, when translating and outputting the speech voice of a specific speaker, the present disclosure may output the translation result in at least one of a first method of outputting the translation result as speech through a speaker and a second method of outputting the translation result as text through a display.
[0238] At the same time, if the measured AI translation processing amount is less than the self-processing capability, the present disclosure can determine whether self-translation processing can be performed (S67), and if it is determined that self-translation processing can be performed, it can provide an AI translation processing result value by processing the AI translation processing by itself (S68).
[0239] like Figure 21 As shown, the device-side AI apparatus of the present disclosure may receive a user command requesting a translation proxy service ( S111 ).
[0240] Next, when a user command requesting the translation proxy service is input, the present disclosure may check whether a communication connection is established with the AI translation processing device ( S113 ).
[0241] Here, when a user command requesting a translation agent service is input, the present disclosure may check whether translation of the speech voice of a specific speaker is currently being performed, and if translation of the speech voice of the specific speaker is currently being performed, stop the translation of the speech voice of the specific speaker, and obtain speech voice data of the specific speaker input after the translation is stopped.
[0242] Next, the present disclosure can request a translation proxy service of an utterance of a specific speaker from the AI translation processing device when a communication connection is established with the AI translation processing device ( S115 ).
[0243] Furthermore, the present disclosure may receive approval of the translation proxy service request from the AI translation processing device ( S117 ).
[0244] Then, upon receiving approval for the translation proxy service request from the AI translation processing device, the present disclosure may send the speech voice data of the specific speaker and the target language information to be translated to the AI translation processing device, so that the AI translation processing device translates the speech voice data of the specific speaker into the target language and outputs it (S119).
[0245] In addition, the present disclosure can request the AI translation processing device to suspend the translation proxy services of some speakers corresponding to the translation proxy services when a user command is input requesting the suspension of translation of some speakers corresponding to the translation proxy services, and can receive information about the completion of the suspension of the translation proxy services of some speakers and information about the continuation of the translation proxy services of the remaining speakers from the AI translation processing device.
[0246] like Figure 22 As shown, the device-side AI device of the present disclosure may receive a distributed processing request from the AI translation processing device ( S121 ).
[0247] Furthermore, the present disclosure may extract distributed processing information from a distributed processing request when receiving the distributed processing request from the AI translation processing apparatus ( S123 ).
[0248] Here, the present disclosure may extract distributed processing information corresponding to the distributed processing part including AI model information, location information, distributed processing amount information, and input data from the distributed processing request.
[0249] Next, the present disclosure may perform AI translation processing based on the distributed processing information ( S125 ) to generate an AI translation processing result value, and provide the generated AI translation processing result value to the AI translation processing device ( S127 ).
[0250] like Figure 23 As shown, the device-side AI device of the present disclosure may receive a translation proxy service request from the AI translation processing device ( S131 ).
[0251] Also, when a translation proxy service request is received from the AI translation processing device, the present disclosure may measure the AI translation processing amount corresponding to the translation proxy service request ( S133 ).
[0252] Next, the present disclosure may determine whether self-translation processing is possible based on the measured AI translation processing amount ( S135 ).
[0253] Next, if the present disclosure determines that the self-translation process can be performed, the present disclosure may send approval of the translation proxy service request to the AI translation processing device ( S136 ).
[0254] If the present disclosure determines that the self-translation process cannot be performed, the present disclosure may transmit a notification message notifying that the self-processing of the translation proxy service cannot be performed to the AI translation processing device ( S139 ).
[0255] Then, the present disclosure may confirm whether the speaker's utterance voice data and the target language information to be translated are received from the AI translation processing device ( S137 ).
[0256] Next, if the present disclosure receives the speaker's utterance voice data and the target language information to be translated from the AI translation processing device, it is possible to translate the speaker's utterance voice into the target language and output it ( S138 ).
[0257] Here, the present disclosure may suspend the translation proxy services of some speakers upon receiving a request to suspend the translation proxy services of some speakers from the AI translation processing device, and send information about the completion of the suspension of the translation proxy services of some speakers and information about the continuation of the translation proxy services of the remaining speakers to the AI translation processing device.
[0258] Figure 24 and Figure 25 1 is a diagram for explaining a method for providing a multi-party translation service by an AI translation processing device communicatively connected to a device-side AI device according to an embodiment of the present disclosure.
[0259] like Figure 24 As shown, the AI translation processing device of the present invention, which is communicatively connected to the device-side AI device, may include: a communication module, which is communicatively connected to the device-side AI device; a memory, which stores an AI model for AI translation processing; and a processor, which performs AI translation processing in response to a translation agent service request from the device-side AI device.
[0260] The processor of the AI translation processing device may receive a translation proxy service request from the device-side AI device ( S211 ).
[0261] Then, when the processor of the AI translation processing device receives a translation proxy service request from the device-side AI device, the processor of the AI translation processing device may measure the AI translation processing amount corresponding to the translation proxy service request ( S213 ).
[0262] Next, the processor of the AI translation processing device can determine whether self-translation processing is possible based on the measured AI translation processing amount ( S215 ).
[0263] Next, if the processor of the AI translation processing device determines that the self-translation process can be performed, the processor may send approval of the translation proxy service request to the device-side AI device ( S216 ).
[0264] If the processor of the AI translation processing device determines that the self-translation process cannot be performed, the processor may send a notification message to the device-side AI device notifying that the self-translation process of the translation proxy service cannot be performed ( S219 ).
[0265] Then, the processor of the AI translation processing device may check whether the speaker's utterance voice data and the target language information to be translated are received from the device-side AI device ( S217 ).
[0266] Next, upon receiving the speaker's speech data and target language information to be translated from the device-side AI device, the processor of the AI translation processing device may translate the speaker's speech into the target language and output the translated speech ( S218 ).
[0267] Here, the processor of the AI translation processing device can stop the translation proxy service for some speakers when receiving a request to stop the translation proxy service for some speakers from the device-side AI device, and send information about the completion of the stop of the translation proxy service for some speakers and information about the continuation of the translation proxy service for the remaining speakers to the device-side AI device.
[0268] like Figure 25 As shown, the AI translation processing device connected to the device-side AI device of the present disclosure may receive a distributed processing request from the device-side AI device ( S221 ).
[0269] Then, upon receiving the distributed processing request from the device-side AI device, the processor of the AI translation processing device may extract the distributed processing information from the distributed processing request ( S223 ).
[0270] Here, the processor of the AI translation processing device can extract distributed processing information corresponding to the distributed processing part including AI model information, location information, distributed processing amount information and input data from the distributed processing request.
[0271] Next, the processor of the AI translation processing device may perform AI translation processing based on the distributed processing information ( S225 ).
[0272] Next, the processor of the AI translation processing device may generate an AI translation processing result value and provide the generated AI translation processing result value to the device-side AI device ( S227 ).
[0273] Figure 263 is a diagram for explaining a method for providing a multi-party translation service of a device-side AI system according to an embodiment of the present disclosure.
[0274] like Figure 26 As shown, when speech voices are input from multiple speakers (S310), the device-side AI device 500 preprocesses the speech voices (S320), and classifies the preprocessed speech voices by speaker (S330). When a specific speaker is selected from multiple speakers, the speech voice of the selected specific speaker can be extracted from the speech voices classified by speaker (S340).
[0275] Next, the device-side AI device 500 measures the AI translation processing volume for translating the current language into the target language by converting the speech of the specific speaker into text, and determines whether distributed processing is required by checking whether the measured AI translation processing volume exceeds its own processing capacity.
[0276] Next, when the measured AI translation processing volume of the device-side AI device 500 exceeds its own processing capacity, the device-side AI device 500 can select the AI translation processing device 700 for distributed processing.
[0277] Then, the device-side AI device 500 requests a communication connection from the AI translation processing device 700, and the AI translation processing device 700 approves the communication connection request for the device-side AI device 500, so that the AI translation processing device 700 and the device-side AI device 500 can be connected to each other (S350, S360).
[0278] Next, the device-side AI device 500 can request the AI translation processing device 700 to perform distributed processing of the AI translation process ( S370 ).
[0279] Next, when the AI translation processing device 700 receives a distributed processing request from the device-side AI device 500, it extracts distributed processing information from the distributed processing request (S380), performs AI translation processing based on the distributed processing information to generate an AI translation processing result value (S390), and provides the generated AI translation processing result value to the device-side AI device 500 (S400).
[0280] Then, when the device-side AI device 500 receives the first AI translation processing result value from the AI translation processing device 700 , it may provide a final translation result based on the first AI translation processing result value and the second AI translation processing result value processed by itself ( S410 ).
[0281] Then, when a user command requesting a translation proxy service is input, the device-side AI device 500 requests a translation proxy service for the speech voice of a specific speaker from the AI translation processing device 700 (S430), and when approval of the translation proxy service request is received from the AI translation processing device 700 (S440), the device-side AI device 500 is able to send the speech voice data of the specific speaker and the target language information to be translated to the AI translation processing device 700, so that the speech voice of the specific speaker can be translated into the target language and output (S450).
[0282] Next, when the AI translation processing device 700 receives the speaker's speech data and the target language information to be translated from the device-side AI device 500, the AI translation processing device 700 can translate the speaker's speech into the target language and output it (S460).
[0283] In this way, the present disclosure extracts the speech voice of a specific speaker from the speech voices classified by speaker and translates it into the target language when a specific speaker is selected from multiple speakers, thereby enabling real-time and accurate translation of conversations between multiple speakers by speaker recognition.
[0284] In addition, when the amount of AI processing used to translate the speaker's speech exceeds the self-processing capacity, the present invention can improve the processing speed of AI translation processing and the accuracy and service quality of the translation processing result value by selecting an external AI translation processing device and requesting distributed processing of AI translation processing.
[0285] Furthermore, the present disclosure can minimize power consumption and reduce heat generation by distributing AI translation processing with an externally located AI translation processing device, thereby improving performance and lifespan.
[0286] The present disclosure described above can be implemented as computer-readable code on a medium having a program recorded thereon. Computer-readable media include all types of recording devices in which data readable by a computer system is stored. Examples of computer-readable media include hard disk drives (HDDs), solid-state drives (SSDs), silicon disk drives (SDDs), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, optical data storage devices, and the like. In addition, a computer may include a processor of an artificial intelligence device.
Claims
1. A device-side artificial intelligence (AI) device, comprising: an input module configured to receive sensory data associated with a plurality of utterances from a plurality of speakers; as well as a processor configured to perform AI processing to translate an utterance of the plurality of utterances into a target language, Wherein, the processor is configured to: Preprocessing the multiple utterances, classifying each preprocessed utterance of the preprocessed plurality of utterances to a speaker of the plurality of speakers, extracting an utterance of a specific speaker from the classified plurality of utterances, translating the utterance of the specific speaker into the target language, and The utterance of the specific speaker is output in the target language.
2. The device according to claim 1, wherein The operation of translating the speech of the specific speaker includes: determining a current language by converting the speech of the specific speaker into text; determining whether self-translation processing is possible by measuring the amount of AI translation processing used to translate the current language into the target language; and as a result of determining whether self-translation processing is possible, translating the speech of the specific speaker into the target language through self-processing AI translation processing.
3. The device according to claim 2, wherein The operation of determining whether the self-translation process is possible includes comparing the measured AI translation processing amount with the self-processing capability.
4. The device according to claim 3, wherein The operation of determining whether self-translation processing is possible includes determining that self-translation processing is not possible after the measured AI translation processing amount exceeds the self-processing capability, wherein the processor is configured to: Select an AI translation processing device for distributed processing of AI translation processing, Requesting the selected AI translation processing device to perform distributed processing of the AI translation process, and After receiving a first AI translation processing result value from the AI translation processing device, a final translation result value is provided based on the first AI translation processing result value and a self-processed second AI translation processing result value.
5. The device according to claim 3, wherein The operation of requesting the distributed processing of the AI translation processing includes: calculating an excess amount of the measured AI translation processing amount excluding the self-processing capacity, determining that the processing amount of the selected AI translation processing device is greater than the calculated excess amount, and requesting the distributed processing of the AI translation processing for the excess amount from the AI translation processing device.
6. The device according to claim 3, wherein The operation of requesting distributed processing of the AI translation process includes: calculating an excess amount of the measured AI translation processing amount excluding the self-processing capacity, extracting a distributed processing portion of the AI translation process corresponding to the calculated excess amount, and requesting the selected AI translation processing device to perform distributed processing of the AI translation process for the extracted distributed processing portion.
7. The device according to claim 4, wherein The operation of providing the final translation result value includes providing the final translation result value of translating the utterance of the specific speaker into the target language by mapping the first AI translation processing result value received from the AI translation processing device and the self-processed second AI translation processing result value to each other.
8. The device according to claim 1, further comprising a communication module connected to the AI translation processing device via a wired or wireless connection, wherein the AI translation processing device performs translation processing based on the AI model. in, After receiving a distributed processing request from the AI translation processing device, the processor is configured to: extract distributed processing information from the distributed processing request, generate an AI translation processing result value by performing AI translation processing based on the distributed processing information, and provide the generated AI translation processing result value to the AI translation processing device.
9. The device according to claim 1, further comprising a communication module connected to the AI translation processing device via a wired or wireless connection, wherein the AI translation processing device performs translation processing based on the AI model. in, After receiving a translation proxy service request from the AI translation processing device, the processor is configured to: determine whether self-translation processing is possible by measuring the AI translation processing amount corresponding to the translation proxy service request; after determining that the self-translation processing is possible, send approval of the translation proxy service request to the AI translation processing device; and after receiving the speaker's speech voice data and the target language information to be translated from the AI translation processing device, translate the speaker's speech voice into the target language and output it.
10. A method for providing multi-party translation services by a device comprising an artificial intelligence (AI) device, the method comprising the following steps: receiving a plurality of utterances from a plurality of speakers; preprocessing the plurality of utterances; classifying each preprocessed utterance of the preprocessed plurality of utterances to a speaker of the plurality of speakers; extracting an utterance of a specific speaker from the plurality of classified utterances; translating the utterance of the specific speaker into a target language; and The utterance of the specific speaker is output in the target language.
Citation Information
Cited By
Speech processing method and device and related product
CN121306128A