On-device ai apparatus and method for providing multi-party interpretation service

The on-device AI device addresses the challenge of real-time multi-speaker interpretation by preprocessing and classifying speech, and distributes processing when necessary, achieving efficient and accurate interpretation.

JP2025120152APending Publication Date: 2025-08-15LX SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025014189
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-04
Filing Date
2025-01-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Current on-device AI systems struggle to accurately interpret conversations between multiple speakers in real time and distinguish between different speakers, limiting their effectiveness in providing selective interpretation services.

Method used

An on-device AI device preprocesses speech from multiple speakers, classifies it by speaker, and selectively interprets the speech of a specific speaker into a target language, with the option to distribute AI interpretation processing to external devices when needed.

Benefits of technology

Enables real-time, accurate interpretation of multi-speaker conversations by improving processing speed, accuracy, and reducing power consumption and heat generation, while enhancing service quality and device performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025120152000001_ABST
    Figure 2025120152000001_ABST
Patent Text Reader

Abstract

To provide an on-device AI apparatus for identifying conversations among a plurality of speakers on a per-speaker basis and providing real-time interpretation, and a method for providing multi-party interpretation services of the same.SOLUTION: An on-device AI apparatus 500 includes an input unit to which a speech of a speaker is input, and a processor that performs AI processing to interpret the speech of the speaker into a target language. The processor pre-processes the speech when the speech is input from a plurality of speakers, classifies the pre-processed speech on a per-speaker basis, extracts the speech of a selected specific speaker from the speeches classified on the per-speaker basis when the specific speaker is selected from among the plurality of speakers; and interprets the speech of the specific speaker into the target language and outputs it.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an on-device AI device capable of interpreting conversations between multiple speakers in real time and a method for providing such a multi-party interpretation service. [Background technology]

[0002] In general, artificial intelligence is a field of computer engineering and information technology that studies ways to enable computers to think, learn, and develop themselves as human beings do, and it means enabling computers to imitate intelligent human behavior.

[0003] Furthermore, AI does not exist in isolation, but is directly and indirectly connected to many other fields of computer science. In particular, there are currently many active attempts to introduce AI elements into various fields of information technology and utilize them to solve problems in those fields.

[0004] Meanwhile, there has been active research into technologies that use artificial intelligence to recognize and learn about the surrounding situation, provide information desired by the user in a desired format, or have the user's desired operations and functions.

[0005] An electronic device that provides such various operations and functions can be called an artificial intelligence device.

[0006] In recent years, on-device AI, which can process information independently on terminal devices without the need to connect to a server or cloud, has been attracting attention.

[0007] On-device AI does not transmit data to a server or cloud, but instead processes AI calculations independently within the user's own device, which has the advantages of being fast, personal information protection, and cost-effective.

[0008] However, among the services that can currently be provided by on-device AI, there is still a problem in that real-time interpretation services have limitations in terms of their completeness.

[0009] In other words, current on-device AI cannot accurately interpret conversations between multiple speakers in real time, and it also has the problem of not being able to distinguish between the voices of multiple speakers and provide selective interpretation services.

[0010] Therefore, in the future, it will be necessary to develop on-device AI equipment that can identify conversations between multiple speakers and accurately interpret them in real time without the need for additional equipment. Summary of the Invention [Problem to be solved by the invention]

[0011] The present disclosure is directed to solving the above-mentioned problems and other problems.

[0012] The present disclosure aims to provide an on-device AI device and a method for providing a multi-party interpretation service that can identify a conversation between multiple speakers by speaker and accurately interpret it in real time by extracting the speech of the specific speaker from speech sounds classified by speaker when a specific speaker is selected from multiple speakers and interpreting it into a target language. [Means for solving the problem]

[0013] An on-device AI device according to one embodiment of the present disclosure includes an input unit to which a speaker's speech is input, and a processor that performs AI processing to interpret the speaker's speech into a target language. When speech is input from multiple speakers, the processor preprocesses the speech and classifies the preprocessed speech by speaker. When a specific speaker is selected from the multiple speakers, the processor extracts the speech of the selected specific speaker from the speeches classified by speaker, and interprets the speech of the specific speaker into the target language and outputs it.

[0014] An AI interpretation processing device according to one embodiment of the present disclosure is an AI interpretation processing device that is communicatively connected to an on-device AI device, and includes a communication unit communicatively connected to the on-device AI device, a memory that stores an AI model for AI interpretation processing, and a processor that performs AI interpretation processing in response to a request for an interpretation service from the on-device AI device. When the processor receives a request for an interpretation service from the on-device AI device, it measures the amount of AI interpretation processing corresponding to the request for the interpretation service and determines whether or not it can perform its own interpretation processing. If it determines that it can perform its own interpretation processing, it transmits approval for the request for the interpretation service to the on-device AI device. When it receives speaker's speech data and target language information to be interpreted from the on-device AI device, it can interpret the speaker's speech into the target language and output it.

[0015] An on-device AI system according to one embodiment of the present disclosure is an on-device AI system including at least one AI interpretation processing device communicatively connected to an on-device AI device, the on-device AI device performing AI interpretation processing to interpret a speaker's speech into a target language, and at least one AI interpretation processing device performing AI interpretation processing in response to a request for an interpretation service from the on-device AI device, and the on-device AI device pre-processes the speech when speech is input from a plurality of speakers, classifies the pre-processed speech by speaker, and identifies a specific speaker from the plurality of speakers. Once selected, the speech of the selected specific speaker is extracted from the speech sounds classified by speaker, and the speech of the specific speaker is interpreted into the target language and output. When a user command requesting an interpretation service is input, an interpretation service for the speech of the specific speaker is requested from the AI interpretation processing device. When approval for the interpretation service request is received from the AI interpretation processing device, the speech data of the specific speaker and information on the target language to be interpreted can be transmitted to the AI interpretation processing device so that the AI interpretation processing device can interpret the speech of the specific speaker into the target language and output it.

[0016] A method for providing a multi-party interpretation service using an on-device AI device according to one embodiment of the present disclosure may include the steps of receiving input of speech sounds from multiple speakers, preprocessing the speech sounds, classifying the preprocessed speech sounds by speaker, and when a specific speaker is selected from the multiple speakers, extracting the speech sound of the selected specific speaker from the speech sounds classified by speaker, and interpreting the speech sound of the specific speaker into a target language and outputting it. [Effects of the Invention]

[0017] According to one embodiment of the present disclosure, when a specific speaker is selected from multiple speakers, the on-device AI device extracts the speech of the specific speaker from the speech sounds divided by speaker and interprets it into a target language, thereby identifying a conversation between multiple speakers by speaker and accurately interpreting it in real time.

[0018] In addition, the present disclosure can improve the speed of AI interpretation processing, the accuracy of interpretation processing results, and service quality by selecting an external AI interpretation processing device and requesting distributed processing of AI interpretation processing when the amount of AI processing required to interpret a speaker's speech exceeds the amount that can be processed by the device itself.

[0019] In addition, the present disclosure can minimize power consumption and reduce heat generation by distributing AI interpretation processing with an externally located AI interpretation processing device, thereby improving performance and lifespan. [Brief explanation of the drawings]

[0020] [Figure 1] 1 illustrates an artificial intelligence device according to one embodiment of the present disclosure. [Figure 2] 1 illustrates an artificial intelligence server according to one embodiment of the present disclosure. [Figure 3] 1 illustrates an artificial intelligence system according to one embodiment of the present disclosure. [Figure 4] FIG. 1 is a diagram illustrating an on-device AI system according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a diagram illustrating an on-device AI system according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a diagram illustrating an on-device AI system according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a diagram illustrating an on-device AI device according to an embodiment of the present disclosure. [Figure 8] FIG. 1 is a diagram illustrating an on-device AI device according to an embodiment of the present disclosure. [Figure 9] FIG. 1 is a diagram illustrating the speaker identification process of an on-device AI device according to one embodiment of the present disclosure. [Figure 10] FIG. 1 is a diagram illustrating the speaker identification process of an on-device AI device according to one embodiment of the present disclosure. [Figure 11] FIG. 1 is a diagram illustrating the speaker identification process of an on-device AI device according to one embodiment of the present disclosure. [Figure 12] FIG. 10 is a diagram illustrating the speaker selection process of an on-device AI device according to one embodiment of the present disclosure. [Figure 13] FIG. 10 is a diagram illustrating the speaker selection process of an on-device AI device according to one embodiment of the present disclosure. [Figure 14] FIG. 10 is a diagram illustrating the speaker selection process of an on-device AI device according to one embodiment of the present disclosure. [Figure 15] FIG. 10 is a diagram illustrating the speaker selection process of an on-device AI device according to one embodiment of the present disclosure. [Figure 16] FIG. 10 is a diagram illustrating a target language setting process for an on-device AI device according to one embodiment of the present disclosure. [Figure 17] FIG. 10 is a diagram illustrating a process for measuring the amount of AI interpretation processing of an on-device AI device according to one embodiment of the present disclosure. [Figure 18] FIG. 10 is a diagram illustrating a process for measuring the amount of AI interpretation processing of an on-device AI device according to one embodiment of the present disclosure. [Figure 19]FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service using an on-device AI device according to an embodiment of the present disclosure. [Figure 20] FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service using an on-device AI device according to an embodiment of the present disclosure. [Figure 21] FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service using an on-device AI device according to an embodiment of the present disclosure. [Figure 22] FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service using an on-device AI device according to an embodiment of the present disclosure. [Figure 23] FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service using an on-device AI device according to an embodiment of the present disclosure. [Figure 24] FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service by an AI interpretation processing device communicatively connected to an on-device AI device according to one embodiment of the present disclosure. [Figure 25] FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service by an AI interpretation processing device communicatively connected to an on-device AI device according to one embodiment of the present disclosure. [Figure 26] FIG. 1 is a diagram illustrating a method for providing a multi-party interpretation service in an on-device AI system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0021] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Identical or similar components will be designated by the same reference numerals regardless of the reference numerals, and redundant descriptions thereof will be omitted. The suffixes "module" and "unit" used in the following description are added or substituted for components solely for ease of description and do not have any distinct meanings or functions. Furthermore, when describing embodiments of the present disclosure, detailed descriptions of related publicly known technologies will be omitted if they are deemed to obscure the gist of the embodiments disclosed herein. Furthermore, the accompanying drawings are merely intended to facilitate understanding of the embodiments disclosed herein, and the accompanying drawings should not be construed as limiting the technical concept of the present disclosure, and should be understood to include any modifications, equivalents, or alternatives within the concept and technical scope of the present disclosure.

[0022] Terms including ordinal numbers such as first, second, etc. may be used to describe various components, but these components are not limited by such terms. Such terms are used only to distinguish one component from another.

[0023] When a component is said to be "coupled" or "connected" to another component, it should be understood that it may be directly coupled or connected to the other component, but that there may be additional components in between. On the other hand, when a component is said to be "directly coupled" or "directly connected" to another component, it should be understood that there are no additional components in between.

[0024] Also, throughout this specification, the terms neural network, neural network network, and network function may be used interchangeably. A neural network may be composed of a collection of interconnected computational units, generally called "nodes." These "nodes" may also be called "neurons." A neural network is composed of at least two nodes. The nodes (or neurons) that make up a neural network may be interconnected by one or more "links."

[0025] FIG. 1 shows an AI device 100 according to one embodiment of the present disclosure.

[0026] The AI device 100 may be embodied as a fixed or mobile device such as a TV, projector, mobile phone, smartphone, desktop computer, laptop, digital broadcasting terminal, PDA (personal digital assistant), PMP (portable multimedia player), navigation, tablet PC, wearable device, set-top box (STB), DMB receiver, radio, washing machine, refrigerator, digital signage, robot, vehicle, etc.

[0027] Referring to FIG. 1, the AI device 100 may include a communication unit 110, an input unit 120, a learning processor 130, a sensing unit 140, an output unit 150, a memory 170, and a processor 180, etc.

[0028] The communication unit 110 can use wired or wireless communication technology to transmit and receive data to and from external devices such as other AI devices 100a to 100e or the AI server 200. For example, the communication unit 110 can transmit and receive sensor information, user input, learning models, control signals, etc. to and from external devices.

[0029] At this time, the communication technologies used by the communication unit 110 include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth (registered trademark), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, and NFC (Near Field Communication).

[0030] The input unit 120 can acquire various types of data.

[0031] In this case, the input unit 120 may include a camera for inputting a video signal, a microphone for receiving an audio signal, a user input unit for receiving information input from a user, etc. Here, the camera or microphone may be treated as a sensor, and the signal obtained from the camera or microphone may be referred to as sensing data or sensor information.

[0032] The input unit 120 may receive learning data for model learning, input data used when obtaining an output using a learning model, etc. The input unit 120 may also receive raw input data, in which case the processor 180 or the learning processor 130 may extract input features by preprocessing the input data.

[0033] The learning processor 130 can use the training data to train a model configured as an artificial neural network. Here, the trained artificial neural network can be referred to as a training model. The training model can be used to infer a result value for new input data other than the training data, and the inferred value can be used as a basis for making decisions to take certain actions.

[0034] At this time, the learning processor 130 can perform AI processing together with the learning processor 240 of the AI server 200 of FIG.

[0035] In this case, the learning processor 130 may include a memory integrated or embodied in the AI device 100. Alternatively, the learning processor 130 may be embodied using the memory 170, an external memory directly coupled to the AI device 100, or a memory held in an external device.

[0036] The sensing unit 140 can use various sensors to acquire at least one of internal information of the AI device 100, information about the surrounding environment of the AI device 100, and user information.

[0037] In this case, the sensors included in the sensing unit 140 include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a lidar, a radar, and the like.

[0038] The output unit 150 can generate an output related to a visual, an auditory, or a tactile sense.

[0039] In this case, the output unit 150 may include a display unit that outputs visual information, a speaker that outputs auditory information, a haptic module that outputs tactile information, and the like.

[0040] The memory 170 can store data that supports various functions of the AI device 100. For example, the memory 170 can store input data acquired from the input unit 120, learning data, a learning model, a learning history, and the like.

[0041] Processor 180 can determine at least one executable action of AI device 100 based on information determined or generated using data analysis algorithms or machine learning algorithms, and processor 180 can control components of AI device 100 to perform the determined action.

[0042] To that end, processor 180 can request, retrieve, receive, or utilize data from learning processor 130 or memory 170 and control components of AI device 100 to perform an action that is predicted or determined to be preferred among the at least one possible action.

[0043] At this time, if cooperation with an external device is required to perform the determined operation, the processor 180 can generate a control signal for controlling the external device and transmit the generated control signal to the external device.

[0044] The processor 180 can obtain intent information for the user input and determine the user's requirements based on the obtained intent information.

[0045] At this time, the processor 180 can acquire the intention information corresponding to the user input using at least one of an STT (Speech To Text) engine for converting the voice input into a string of characters or an NLP (Natural Language Processing) engine for acquiring the intention information of natural language.

[0046] In this case, at least one of the STT engine or the NLP engine may be configured as an artificial neural network trained at least in part by a machine learning algorithm, and at least one of the STT engine or the NLP engine may be trained by the learning processor 130, the learning processor 240 of the AI server 200, or a distributed processing thereof.

[0047] The processor 180 can collect history information, including the operation details of the AI device 100 or user feedback on the operation, and store the collected history information in the memory 170 or the learning processor 130, or transmit it to an external device such as the AI server 200. The collected history information can be used to update the learning model.

[0048] The processor 180 can control at least some of the components of the AI device 100 to run the application program stored in the memory 170. The processor 180 can also operate two or more components included in the AI device 100 in combination with each other to run the application program.

[0049] FIG. 2 illustrates an AI server 200 according to one embodiment of the present disclosure.

[0050] Referring to Figure 2, the AI server 200 may refer to a device that trains an artificial neural network using a machine learning algorithm or uses the trained artificial neural network. Here, the AI server 200 may be configured with multiple servers to perform distributed processing, and may be defined as a 5G network. In this case, the AI server 200 may be included as part of the AI device 100 and may perform at least a portion of the AI processing together.

[0051] The AI server 200 may include a communication unit 210, a memory 230, a learning processor 240, and a processor 260.

[0052] The communication unit 210 can send and receive data to and from external devices such as the AI device 100.

[0053] The memory 230 may include a model storage unit 231. The model storage unit 231 can store a model (or artificial neural network) 231a that is being trained or has been trained by the learning processor 240.

[0054] The learning processor 240 can train the artificial neural network 231a using the training data. The training model may be installed in the AI server 200 of the artificial neural network, or may be installed in an external device such as the AI device 100.

[0055] The learning model may be implemented by hardware, software, or a combination of hardware and software. When a part or all of the learning model is implemented by software, one or more instructions constituting the learning model may be stored in memory 230.

[0056] The processor 260 can use the learning model to infer outcome values for new input data and generate responses or control instructions based on the inferred outcome values.

[0057] FIG. 3 shows an AI system 1 according to one embodiment of the present invention.

[0058] 3, in the AI system 1, at least one of an AI server 200, a robot 100a, an autonomous vehicle 100b, an XR device 100c, a smartphone 100d, and a home appliance 100e is connected to a cloud network 10. Here, the robot 100a, the autonomous vehicle 100b, the XR device 100c, the smartphone 100d, or the home appliance 100e to which AI technology is applied may be referred to as the AI devices 100a to 100e.

[0059] The cloud network 10 may refer to a network that forms part of a cloud computing infrastructure or exists in a cloud computing infrastructure, and may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network, a 5G network, or the like.

[0060] That is, the devices 100a to 100e, 200 constituting the AI system 1 may be connected to each other through the cloud network 10. In particular, the devices 100a to 100e, 200 may communicate with each other via a base station, or may communicate with each other directly without going through a base station.

[0061] The AI server 200 may include a server that performs AI processing and a server that performs calculations on big data.

[0062] The AI server 200 is connected to at least one of the AI devices constituting the AI system 1, namely, the robot 100a, the autonomous vehicle 100b, the XR device 100c, the smartphone 100d, or the home appliance 100e, via the cloud network 10, and can assist in at least part of the AI processing of the connected AI devices 100a to 100e.

[0063] In this case, the AI server 200 can train the artificial neural network using a machine learning algorithm on behalf of the AI devices 100a to 100e, and can directly store or transmit the learned model to the AI devices 100a to 100e.

[0064] In this case, the AI server 200 receives input data from the AI devices 100a to 100e, uses a learning model to infer a result value for the received input data, and generates a response or control command based on the inferred result value and transmits it to the AI devices 100a to 100e.

[0065] Alternatively, the AI devices 100a to 100e can directly use the learning model to infer a result value for input data, and generate a response or control command based on the inferred result value.

[0066] 4 to 6 are diagrams illustrating an on-device AI system according to an embodiment of the present disclosure.

[0067] As shown in FIG. 4, the on-device AI system of the present disclosure may include an on-device AI device 500 that performs AI interpretation processing to interpret speech from multiple speakers 600 into a target language.

[0068] Here, the on-device AI device 500 is an artificial intelligence device capable of performing on-device AI processing, and may include any of stationary devices such as a personal computer (PC), network TV, hybrid broadcast broadband TV (HBTV), smart TV, and internet protocol TV (IPTV), as well as mobile or handheld devices such as a smartphone, tablet PC, notebook, personal digital assistant (PDA), smart watch, smart glasses, and robot.

[0069] When speech sounds are input from multiple speakers 600, the on-device AI device 500 preprocesses the speech sounds, classifies the preprocessed speech sounds by speaker, and when a specific speaker is selected from the multiple speakers 600, extracts the speech sound of the selected specific speaker from the speech sounds classified by speaker, and interprets the speech sound of the specific speaker into a target language and outputs it.

[0070] In some cases, as shown in Figures 5 and 6, the on-device AI system of the present disclosure may include an on-device AI device 500 that performs AI interpretation processing to interpret the speech of multiple speakers 600 into a target language, and at least one AI interpretation processing device 700 that performs AI interpretation processing in response to an interpretation service request or a distributed processing request from the on-device AI device 500.

[0071] Here, the AI interpretation processing device 700 may be the same device as the on-device AI device 500, or may be a different device.

[0072] As an example, as shown in FIG. 5, the AI interpretation processing device 700 may include at least one of a personal computer (PC) that performs interpretation processing based on a pre-trained AI model and outputs the interpretation processing results, a standing device including a network TV, a hybrid broadcast broadband TV (HBTV), a smart TV, an internet protocol TV (IPTV), etc., and a mobile or handheld device including a smartphone, a tablet PC, a notebook computer, a personal digital assistant (PDA), a robot, etc.

[0073] As another example, as shown in FIG. 6, the AI interpretation processing device 700 may include a battery pack 710, earphones 720, smart glasses 730, a smart watch 740, etc., which store at least one AI model and perform interpretation processing based on a pre-trained AI model, but do not output the results of the interpretation processing, and may further include various peripheral devices such as an external memory, a smart health band, an adapter, a GPS (Global Positioning System), a PDA (Personal Data Assistance), a barcode reader, a character recognition device, a voice recognition device, etc.

[0074] Here, the AI interpretation processing device 700 may be communicatively connected to the on-device AI device 500 via wired or wireless communication.

[0075] In one embodiment, when a user command requesting an interpretation service is input, the on-device AI device 500 requests the AI interpretation processing device 700 to provide an interpretation service for the speech of a specific speaker, and when it receives approval for the interpretation service request from the AI interpretation processing device 700, it can transmit the speech data of the specific speaker and the target language information to be interpreted to the AI interpretation processing device 700 so that the AI interpretation processing device 700 can interpret the speech of the specific speaker into the target language and output it.

[0076] Here, when a user command requesting an interpretation service is input, the on-device AI device 500 checks whether interpretation of the speech of a specific speaker is currently being performed, and if interpretation of the speech of a specific speaker is currently being performed, it interrupts the interpretation of the speech of the specific speaker and can obtain speech data of the specific speaker input after the point in time when interpretation was interrupted.

[0077] When the AI interpretation processing device 700 receives a request for an interpretation service from the on-device AI device 500, it measures the amount of AI interpretation processing corresponding to the request for the interpretation service and determines whether it can perform its own interpretation processing. If it determines that it can perform its own interpretation processing, it transmits approval for the request for the interpretation service to the on-device AI device 500. When it receives the speech data of the speaker 600 and the target language information to be interpreted from the on-device AI device 500, it can interpret the speech of the speaker 600 into the target language and output it.

[0078] In another embodiment, when a user command is input to request the suspension of interpretation for some speakers in accordance with the interpretation service, the on-device AI device 500 can request the AI interpretation processing device 700 to suspend the interpretation service for some speakers, and can receive from the AI interpretation processing device 700 information indicating the completion of the suspension of the interpretation service for some speakers as well as information indicating the continuation of the interpretation service for the remaining speakers.

[0079] Here, when the AI interpretation processing device 700 receives a request to suspend the interpretation service for some speakers from the on-device AI device 500, it suspends the interpretation service for some speakers and transmits to the on-device AI device 500 information on the completion of the suspension of the interpretation service for some speakers and information on the continuation of the interpretation service for the remaining speakers.

[0080] In yet another embodiment, the on-device AI device 500 measures the amount of AI interpretation processing required to convert the speaker's speech into text and interpret the current language into the target language, and if the measured amount of AI interpretation processing exceeds the amount that the on-device AI device 500 can process itself, it requests the AI interpretation processing device 700 to process the AI interpretation processing in a distributed manner, and upon receiving a first AI interpretation processing result value from the AI interpretation processing device 700, it can provide a final interpretation result value based on the first AI interpretation processing result value and the second AI interpretation processing result value that it processed itself.

[0081] Here, when the AI interpretation processing device 700 receives a distributed processing request from the on-device AI device 500, it extracts distributed processing information from the distributed processing request, performs AI interpretation processing based on the distributed processing information to generate an AI interpretation processing result value, and provides the generated AI interpretation processing result value to the on-device AI device 500.

[0082] At this time, the AI interpretation processing device 700 can extract distributed processing information including AI model information, location information, distributed processing amount information, and input data corresponding to the distributed processing part from the distributed processing request.

[0083] In this way, when the AI interpretation processing device 700 is a device that performs interpretation processing based on a pre-trained AI model and outputs the interpretation processing execution result, as shown in Figure 5, the on-device AI device 500 can request an interpretation agency service from the AI interpretation processing device 700, or can request the AI interpretation processing device 700 to perform distributed processing.

[0084] In addition, as shown in Figure 6, if the AI interpretation processing device 700 is a device that only performs interpretation processing based on a pre-trained AI model and does not output the results of the interpretation processing, the on-device AI device 500 can request the AI interpretation processing device 700 to perform distributed processing of the AI interpretation processing.

[0085] Meanwhile, the on-device AI device 500 stores at least one AI model and can provide various service results using the pre-trained AI model in response to the speaker's 600 commands.

[0086] Here, the AI model is a deep neural network (DNN), which may be a neural network that includes multiple hidden layers in addition to an input layer and an output layer.

[0087] AI models can understand the latent structures of data such as photos, text, video, audio, music, etc.

[0088] 7 and 8 are diagrams illustrating an on-device AI device according to an embodiment of the present disclosure.

[0089] As shown in FIG. 7, the on-device AI device 500 of the present disclosure may include an input unit 510 to which a speaker's speech is input, a processor 520 that performs AI processing to interpret the speaker's speech into a target language, and a memory 530 in which at least one AI model 540 is stored.

[0090] The processor 520 of the present disclosure can perform interpretation processing based on a pre-trained AI model.

[0091] When speech sounds from multiple speakers are input, the processor 520 preprocesses the speech sounds, classifies the preprocessed speech sounds by speaker, and when a specific speaker is selected from the multiple speakers, extracts the speech sound of the selected specific speaker from the speech sounds classified by speaker, and interprets the speech sound of the specific speaker into a target language and outputs it.

[0092] When preprocessing the speech, the processor 520 may perform preprocessing by analyzing frequencies corresponding to the speech and removing noise frequencies when the speech of a speaker is input.

[0093] Here, when analyzing the frequencies corresponding to the speech sound, the processor 520 checks whether there is a specific frequency that is outside the human voice frequency range within the frequencies corresponding to the speech sound, and if there is a specific frequency, it can recognize the specific frequency as a noise frequency.

[0094] Optionally, processor 520 may input the speaker's speech into a pre-trained noise classification model to classify and remove noise frequencies.

[0095] When classifying the preprocessed speech by speaker, the processor 520 extracts speech features of the preprocessed speech, identifies the speaker of the preprocessed speech based on the extracted speech features, and classifies the preprocessed speech by the identified speaker.

[0096] Here, the processor 520 inputs the preprocessed speech into a pre-trained feature extraction model to extract speech features of the preprocessed speech, and inputs the extracted speech features into a pre-trained speaker recognition model to identify the speaker of the preprocessed speech.

[0097] In some cases, once the processor 520 has extracted the speech features of the preprocessed speech, it may select the speaker with the most similar voice from a pre-registered speaker voice catalog based on the extracted speech features, and match the preprocessed speech to the selected speaker to classify it by speaker.

[0098] In addition, when extracting the voice features of the preprocessed speech, the processor 520 checks whether the preprocessed speech is a mixed voice in which the speech of multiple speakers is mixed, and if the speech is a mixed voice, it performs speaker separation on the mixed voice to separate the mixed voice into individual voices and extracts voice features for each of the separated individual voices.

[0099] In some cases, when extracting speech features from the preprocessed speech, processor 520 may check whether the preprocessed speech is continuous speech in which speech from multiple speakers is connected, and if the speech is continuous speech, perform speaker diarization on the continuous speech to separate the continuous speech into speaker-unit speeches, group the separated speaker-unit speeches by speaker, and extract speech features for each of the speaker-unit speeches grouped by speaker.

[0100] Then, when the processor 520 selects a specific speaker, if the speech voice is classified by speaker, it generates and provides a speaker list corresponding to the speech voice, and if a user input is received selecting at least one speaker included in the speaker list, it can select the speaker selected by the user input as the specific speaker.

[0101] Next, when selecting a specific speaker, the processor 520 analyzes the amount of speech data per speaker for a predetermined period of time after the speech voice is divided by speaker, and can select a specific speaker from multiple speakers based on the amount of speech voice data per speaker.

[0102] Here, the processor 520 can compare the amount of speech voice data for each speaker with a preset reference data amount, and select a speaker having a speech voice data amount equal to or greater than the reference data amount as a specific speaker.

[0103] Furthermore, when there are a plurality of speakers having an amount of speech voice data equal to or greater than the reference data amount, the processor 520 can select the speaker having the largest amount of speech voice data as the specific speaker.

[0104] Here, when there are multiple speakers with speech voice data amounts equal to or greater than the standard data amount, processor 520 checks whether a standard number of persons for specific speakers has been set in advance, and if a standard number of persons for specific speakers has been set in advance, if the number of speakers with speech voice data amounts equal to or greater than the standard number is equal to or greater than the standard number, processor 520 selects as many specific speakers as the number of speakers with speech voice data amounts equal to or greater than the standard number in order of the amount of speech voice data, and if the number of speakers with speech voice data amounts equal to or greater than the standard number is less than the standard number, processor 520 can select as many specific speakers as the number of speakers with speech voice data amounts equal to or greater than the standard data amount.

[0105] At this time, when the processor 520 presets the reference number of personnel, the processor 520 can preset the reference number of personnel based on the amount of interpretation that can be processed for the speech of the speaker.

[0106] In some cases, when selecting a specific speaker, the processor 520 may convert the speech of each speaker for a predetermined period of time into text after dividing the speech into segments by speaker, analyze the converted text to extract common related keywords for each speaker, and select a specific speaker from among multiple speakers based on the common related keywords for each speaker.

[0107] Here, when extracting common related keywords, the processor 520 may extract common related keywords including the conference topic keyword and its related keywords from the text corresponding to the speech of each speaker, and group the extracted common related keywords by speaker.

[0108] The processor 520 may also compare the number of common related keywords for each speaker with a preset reference keyword number, and select a speaker having a common related keyword number equal to or greater than the reference keyword number as a specific speaker.

[0109] Here, if there are a plurality of speakers who have a number of common related keywords equal to or greater than the reference keyword number, the processor 520 may select the speaker with the largest number of common related keywords as the specific speaker.

[0110] In addition, when there are multiple speakers who have a number of common related keywords that is equal to or greater than the reference keyword number, processor 520 checks whether a reference number of persons for a specific speaker has been set in advance. If a reference number of persons for a specific speaker has been set in advance, if the number of speakers who have a number of common related keywords that is equal to or greater than the reference keyword number is equal to or greater than the reference number, processor 520 selects as many specific speakers as the reference number in order of the number of common related keywords that is greatest. If the number of speakers who have a number of common related keywords that is equal to or greater than the reference keyword number is less than the reference number, processor 520 can select as many specific speakers as the number of speakers who have a number of common related keywords that is equal to or greater than the reference keyword number.

[0111] Here, when the processor 520 preliminarily sets the reference number of personnel, the processor 520 can preliminarily set the reference number of personnel based on the amount of interpretation that can be processed for the speech of the speaker.

[0112] Next, when processor 520 extracts the speech of a specific speaker, if a specific speaker is selected from multiple speakers, it can leave only the speech of the selected specific speaker from the speech sounds divided by speaker and remove the speech sounds of the remaining speakers other than the specific speaker.

[0113] Here, when removing the speech of the remaining speakers, the processor 520 can remove the speech of the remaining speakers while leaving only the speech of other speakers that continues before or after the time of the speech of the specific speaker.

[0114] In addition, when interpreting the speech of a specific speaker, the processor 520 can store in the database only the speech of other speakers that occurs before or after the time of the speech of the specific speaker in order to extract reference analogous keywords.

[0115] Next, when interpreting the speech of a specific speaker, the processor 520 converts the speech of the specific speaker into text and checks the current language. If the target language is set, the processor 520 measures the amount of AI interpretation processing required to interpret the current language into the target language and determines whether self-interpretation processing is possible. If it determines that self-interpretation processing is possible, the processor 520 can process the AI interpretation processing itself to interpret the speech of the specific speaker into the target language.

[0116] Here, when setting the target language, if the current language is confirmed, the processor 520 checks whether the target language to be interpreted is set, and if the target language is not set, it generates and provides a target language list window, and if it receives user input to select a specific language through the target language list window, it can set the selected specific language as the target language.

[0117] In some cases, the processor 520 may generate and provide a user notification requesting the user to set the target language if the target language is not set.

[0118] For example, the processor 520 may generate and output a user notification in at least one of a text format and a sound format.

[0119] In addition, when determining whether or not to perform interpretation processing, processor 520 measures the amount of AI interpretation processing required to interpret the current language into the target language, and if the measured amount of AI interpretation processing is less than or equal to the amount that can be processed by itself, it can determine that interpretation processing is possible.

[0120] In some cases, processor 520 determines that the measured AI interpretation processing volume exceeds its own processable volume and therefore cannot perform the interpretation processing itself, selects an AI interpretation processing device for distributed processing of the AI interpretation processing, requests the selected AI interpretation processing device to process the AI interpretation processing in a distributed manner, and upon receiving the first AI interpretation processing result value from the AI interpretation processing device, provides a final interpretation result value based on the first AI interpretation processing result value and the second AI interpretation processing result value processed by itself.

[0121] Here, when processor 520 selects an AI interpretation processing device for AI interpretation processing distributed processing, it checks whether there is a communication connection with an external device, and if there is a communication connection with the external device, it obtains identification information of the external device from the external device and checks whether the external device is an AI interpretation processing device based on the identification information, and if the external device is an AI interpretation processing device, it can select the external device as an AI interpretation processing device for AI interpretation processing distributed processing.

[0122] As shown in FIG. 8, the present disclosure may further include a communication unit 550 that is communicatively connected to an external device via a wired or wireless connection, and the processor 520 can check whether or not a communication connection with the external device is established via the communication unit 550.

[0123] When the processor 520 determines whether an external device is an AI interpretation processing device, if an AI model for interpretation processing is stored in the external device, the processor 520 can recognize the external device as an AI interpretation processing device.

[0124] As an example, the AI interpretation processing device may include at least one of a personal computer (PC) that performs interpretation processing based on an AI model and outputs the interpretation processing execution results, a standing device including a network TV, a hybrid broadcast broadband TV (HBTV), a smart TV, or an internet protocol TV (IPTV), and a mobile device or handheld device including a smartphone, a tablet PC, a notebook, a personal digital assistant (PDA), or a robot.

[0125] As another example, the AI interpretation processing device may include at least one of a battery pack, earphones, external memory, smart watch, smart glasses, smart health band, adapter, GPS (Global Positioning System), PDA (Personal Data Assistance), barcode reader, character recognition device, and voice recognition device that only perform interpretation processing based on an AI model and do not output the results of the interpretation processing.

[0126] When selecting an AI interpretation processing device, if there are multiple external devices recognized as AI interpretation processing devices, the processor 520 can select all of the multiple external devices as AI interpretation processing devices and prioritize the multiple external devices based on the processing performance indexes for the multiple external devices selected as AI interpretation processing devices.

[0127] Here, the processor 520 may assign the highest priority to the external device having the highest processing performance index among the plurality of external devices, and assign the lowest priority to the external device having the lowest processing performance index.

[0128] Therefore, as shown in FIG. 8, if the measured amount of AI interpretation processing exceeds the amount that the processor 520 can process, the processor 520 checks the amount that the external devices can process for the excess amount, and if multiple external devices are needed for the excess amount, it can request distributed processing of the AI interpretation processing for the excess amount according to the priority assigned to the external devices.

[0129] Here, when the processor 520 requests distributed processing of the AI interpretation processing for the excess amount, it can request a plurality of external devices to process different excess amounts.

[0130] That is, when requesting the processing of different excess amounts, the processor 520 can request the external device with a higher priority to process the first excess amount, and can request the external device with a lower priority to process the remaining second excess amount, excluding the first excess amount, of the total excess amount.

[0131] In some cases, when the processor 520 requests distributed processing of the AI interpretation processing for the excess amount, it may request multiple external devices to process the same excess amount.

[0132] Here, the processor 520 can distribute the total excess amount equally among multiple external devices, request an external device with a higher priority to process a first excess amount, and request an external device with a lower priority to process a second excess amount equal to the first excess amount.

[0133] In addition, when processor 520 requests distributed processing of AI interpretation processing, it calculates the excess amount of the measured AI interpretation processing volume other than its own processable volume, checks the processable volume of the selected AI interpretation processing device, and if the processable volume of the AI interpretation processing device is greater than the excess volume, it can request distributed processing of the AI interpretation processing volume for the excess volume from the AI interpretation processing device.

[0134] Here, if the processable amount of the AI interpretation processing device is less than the excess amount, the processor 520 can further select another AI interpretation processing device and request the multiple AI interpretation processing devices to process the excess amount in a distributed manner.

[0135] In addition, when processor 520 requests distributed processing of AI interpretation processing, it calculates the excess amount of the measured AI interpretation processing amount beyond the amount that it can process itself, extracts the distributed processing portion of the AI interpretation processing corresponding to the excess amount, and requests the selected AI interpretation processing device to perform distributed processing of the AI interpretation processing for the extracted distributed processing portion.

[0136] Here, when processor 520 extracts the distributed processing portion corresponding to the excess amount, it analyzes the AI model for interpretation processing to check whether there are branch points connecting one higher-level operator to multiple lower-level operators and junction points where multiple higher-level operators join one lower-level operator.If branch points and junction points exist, it checks whether there is at least one parallel processing portion based on the branch points and junction points, and can extract the parallel processing portion as the distributed processing portion.

[0137] In some cases, when processor 520 extracts the distributed processing portion corresponding to the excess amount, if there are multiple AI models for interpretation processing, it may check whether any of the multiple AI models can be processed in parallel with each other, and if there are AI models that can be processed in parallel, it may extract the processing portion performed by the AI models that can be processed in parallel as the distributed processing portion.

[0138] Here, when the processor 520 requests distributed processing of the AI interpretation processing, it can provide a distributed processing request including AI model information, location information, distributed processing amount information, and input data corresponding to the distributed processing part to the AI interpretation processing device.

[0139] In addition, when providing the final interpretation result value, the processor 520 can map the first AI interpretation processing result value received from the AI interpretation processing device and the second AI interpretation processing result value processed by itself to each other to provide the final interpretation result value that interprets the speech of a specific speaker into a target language.

[0140] When the processor 520 interprets and outputs the speech of a specific speaker, it can output the interpretation result in at least one of a first method of outputting the interpretation result as audio from a speaker, and a second method of outputting the interpretation result as text through a display.

[0141] Meanwhile, as shown in FIG. 8, the present disclosure may further include a communication unit 550 that is connected to an AI interpretation processing device that performs interpretation processing based on an AI model via wired or wireless communication. When a user command requesting an interpretation service is input, the processor 520 checks whether there is a communication connection with the AI interpretation processing device, and when there is a communication connection with the AI interpretation processing device, it requests the AI interpretation processing device to provide an interpretation service for the speech of a specific speaker. When it receives approval for the interpretation service request from the AI interpretation processing device, it can transmit the speech data of the specific speaker and the target language information to be interpreted to the AI interpretation processing device so that the AI interpretation processing device can translate the speech of the specific speaker into the target language and output it.

[0142] Here, when a user command requesting an interpretation service is input, the processor 520 checks whether interpretation of the speech of a specific speaker is currently being performed, and if interpretation of the speech of a specific speaker is currently being performed, it interrupts the interpretation of the speech of the specific speaker and can obtain speech data of the specific speaker input after the point in time when interpretation was interrupted.

[0143] For example, when a user command is input to request the suspension of interpretation for some speakers corresponding to the interpretation service, the processor 520 can request the AI interpretation processing device to suspend the interpretation service for some speakers, and can receive from the AI interpretation processing device information indicating the completion of the suspension of the interpretation service for some speakers as well as information indicating the continuation of the interpretation service for the remaining speakers.

[0144] In another case, when processor 520 receives a distributed processing request from an AI interpretation processing device, it can extract distributed processing information from the distributed processing request, perform AI interpretation processing based on the distributed processing information to generate an AI interpretation processing result value, and provide the generated AI interpretation processing result value to the AI interpretation processing device.

[0145] Here, when extracting the distributed processing information, the processor 520 can extract the distributed processing information including AI model information, location information, distributed processing amount information, and input data corresponding to the distributed processing portion from the distributed processing request.

[0146] In yet another case, when processor 520 receives a request for an interpretation service from the AI interpretation processing device, it measures the amount of AI interpretation processing corresponding to the interpretation service request and determines whether it can process the interpretation itself. If it determines that it can process the interpretation itself, it transmits approval for the interpretation service request to the AI interpretation processing device. When it receives the speaker's speech data and target language information to be interpreted from the AI interpretation processing device, it can interpret the speaker's speech into the target language and output it.

[0147] Here, when the processor 520 receives a request to suspend the interpretation service for some speakers from the AI interpretation processing device, it suspends the interpretation service for some speakers and transmits to the AI interpretation processing device information indicating the completion of the suspension of the interpretation service for some speakers and information indicating the continuation of the interpretation service for the remaining speakers.

[0148] In this way, when a specific speaker is selected from multiple speakers, the on-device AI device disclosed herein can extract the speech of the specific speaker from the speech sounds divided by speaker and interpret it into the target language, thereby identifying the conversation between multiple speakers by speaker and accurately interpreting it in real time.

[0149] In addition, the present disclosure can improve the speed of AI interpretation processing, the accuracy of interpretation processing results, and service quality by selecting an external AI interpretation processing device and requesting distributed processing of AI interpretation processing when the amount of AI processing required to interpret a speaker's speech exceeds the amount that can be processed by the device itself.

[0150] In addition, the present disclosure can minimize power consumption and reduce heat generation by distributing AI interpretation processing with an externally located AI interpretation processing device, thereby improving performance and lifespan.

[0151] 9 to 11 are diagrams for explaining a speaker identification process of an on-device AI device according to one embodiment of the present disclosure.

[0152] As shown in FIG. 9, when speech sounds are input from speakers, the present disclosure preprocesses the speech sounds and can identify the preprocessed speech sounds for each speaker.

[0153] The present disclosure may include a preprocessing unit 810 that preprocesses speech, a speech feature extraction unit 820 that extracts features of the preprocessed speech, and a speaker identification unit 830 that identifies a speaker for the speech based on the speech features.

[0154] Here, when a speech voice of a speaker is input, the pre-processing unit 810 can perform pre-processing by analyzing frequencies corresponding to the speech voice and removing noise frequencies.

[0155] For example, when analyzing frequencies corresponding to speech, the pre-processing unit 810 checks whether there are specific frequencies within the frequencies corresponding to speech that are outside the human voice frequency range, and if there are specific frequencies, it can recognize the specific frequencies as noise frequencies.

[0156] In some cases, the pre-processing unit 810 may input the speaker's speech into a pre-trained noise classification model to classify and remove noise frequencies.

[0157] Then, the speech feature extraction unit 820 extracts speech features of the preprocessed speech, and the speaker identification unit 830 identifies the speaker of the preprocessed speech based on the extracted speech features, and classifies the preprocessed speech according to the identified speaker.

[0158] Here, the speech feature extraction unit 820 inputs the preprocessed speech into a pre-trained feature extraction model to extract speech features of the preprocessed speech, and the speaker identification unit 830 inputs the extracted speech features into a pre-trained speaker recognition model to identify the speaker of the preprocessed speech.

[0159] In some cases, when the speaker identification unit 830 extracts the voice features of the pre-processed speech, it can select the speaker with the most similar voice from a pre-registered speaker voice list based on the extracted voice features, and match the pre-processed speech to the selected speaker to classify it by speaker.

[0160] Furthermore, as shown in FIG. 10, the present disclosure can confirm whether the preprocessed speech audio is a mixed audio in which speech audio from multiple speakers is mixed, and if the speech audio is a mixed audio, perform speaker separation on the mixed audio to separate the mixed audio into individual audio, and extract audio features for each of the separated individual audio.

[0161] In some cases, as shown in FIG. 11 , the present disclosure can confirm whether the preprocessed speech audio is continuous audio in which speech audio from multiple speakers is connected, and if the speech audio is continuous audio, perform speaker diarization on the continuous audio to separate the continuous audio into speaker-unit audio, group the separated speaker-unit audio by speaker, and extract audio features for each of the speaker-unit audio grouped by speaker.

[0162] 12 to 15 are diagrams for explaining the speaker selection process of an on-device AI device according to an embodiment of the present disclosure.

[0163] As shown in FIG. 12, when selecting a specific speaker, the present disclosure classifies the speech audio by speaker (842), generates a speaker list corresponding to the speech audio (844), and when user input is received selecting at least one speaker included in the speaker list, the speaker selected by the user input can be selected as the specific speaker (846).

[0164] As shown in FIG. 13, the present disclosure can generate a speaker inventory window 910 and output it on the screen of the on-device AI device 500.

[0165] Here, the speaker list window 910 may include various items such as a speaker selection item, a speech voice provision item for each speaker, and a speech voice data amount for each speaker.

[0166] Furthermore, when the present disclosure provides the speaker list window 910, it can provide a speaker selection request notification message 920 together with the speaker list window 910.

[0167] Here, the present disclosure can generate and output the speaker selection request notification message 920 in at least one of a text format and a sound format.

[0168] As shown in FIG. 14, when selecting a specific speaker, the present disclosure classifies speech sounds by speaker (852), analyzes the amount of speech sound data by speaker for a predetermined period of time (854), and selects a specific speaker from multiple speakers based on the amount of speech sound data by speaker (856, 858).

[0169] Here, the present disclosure can compare the amount of speech voice data for each speaker with a preset reference data amount, and select a speaker having a speech voice data amount equal to or greater than the reference data amount as a specific speaker (856).

[0170] Furthermore, in the present disclosure, when there are multiple speakers with speech sound data amounts equal to or greater than the reference data amount, the speaker with the largest amount of speech sound data can be selected as the specific speaker.

[0171] Here, the present disclosure checks whether a standard number of speakers for specific speakers has been set in advance when there are multiple speakers with speech data volumes equal to or greater than the standard data volume, and if a standard number of speakers for specific speakers has been set in advance, selects as many specific speakers as the standard number in order of the amount of speech data largest when the number of speakers with speech data volumes equal to or greater than the standard data volume is equal to or greater than the standard number, and if the number of speakers with speech data volumes equal to or greater than the standard data volume is less than the standard number, selects as many specific speakers as the number of speakers with speech data volumes equal to or greater than the standard data volume (858).

[0172] In this case, when the reference number of personnel is set in advance, the present disclosure can set the reference number of personnel in advance based on the amount of interpretation processable for the speech of the speaker.

[0173] As shown in FIG. 15, when selecting a specific speaker, the present disclosure classifies speech sounds by speaker (862), converts the speech sounds of each speaker for a predetermined period of time into text (864), analyzes the converted text to extract common related keywords for each speaker (866), and can select a specific speaker from multiple speakers based on the common related keywords for each speaker (868).

[0174] Here, when extracting common related keywords, the present disclosure extracts common related keywords including conference topic keywords and their related keywords from text corresponding to the speech of each speaker, and can group the extracted common related keywords by speaker.

[0175] In addition, the present disclosure may compare the number of common related keywords for each speaker with a preset reference keyword number, and select a speaker having a common related keyword number equal to or greater than the reference keyword number as a specific speaker (868).

[0176] Here, in the present disclosure, when there are a plurality of speakers who have a number of common related keywords equal to or greater than the reference keyword number, the speaker with the largest number of common related keywords can be selected as the specific speaker.

[0177] Furthermore, the present disclosure checks whether a standard number of persons for specific speakers has been set in advance when there are multiple speakers who have a number of common related keywords that is equal to or greater than the standard keyword number, and if a standard number of persons for specific speakers has been set in advance, selects as many specific speakers as the standard number in order of the number of common related keywords that is greatest when the number of speakers who have a number of common related keywords that is equal to or greater than the standard keyword number is equal to or greater than the standard number, and if the number of speakers who have a number of common related keywords that is equal to or greater than the standard keyword number is less than the standard number, selects as many specific speakers as the number of speakers who have a number of common related keywords that is equal to or greater than the standard keyword number.

[0178] Here, in the present disclosure, when the reference number of personnel is set in advance, the reference number of personnel can be set in advance based on the amount of interpretation processable for the speech of the speaker.

[0179] FIG. 16 is a diagram illustrating a target language setting process of an on-device AI device according to one embodiment of the present disclosure.

[0180] When interpreting the speech of a specific speaker, the present disclosure converts the speech of the specific speaker into text, checks the current language from the converted text, and once the current language is confirmed, it can check whether the target language to be interpreted is set.

[0181] Here, the present disclosure can generate an interpretation language list window 1010 and output it on the screen of the on-device AI device 500 if the target language is not set, as shown in FIG.

[0182] In this case, the interpretation language list window 1010 may include various items such as a current language item and a target language selection item.

[0183] In addition, when the present disclosure provides the interpretation language list window 1010, it can generate and provide a user notification 1020 requesting the user to set a target language together with the interpretation language list window 1010.

[0184] For example, the present disclosure may generate and output the user notification 1020 in at least one of a text format and a sound format.

[0185] 17 and 18 are diagrams illustrating a process of measuring the amount of AI interpretation processing of an on-device AI device according to one embodiment of the present disclosure.

[0186] The present disclosure measures the amount of AI interpretation processing required to convert a speaker's speech into text and interpret the current language into a target language, and if the measured amount of AI interpretation processing exceeds the amount that can be processed by the AI interpretation processing device itself, requests the AI interpretation processing device to perform distributed processing.When the first AI interpretation processing result value is received from the AI interpretation processing device, the device can provide a final interpretation result value based on the first AI interpretation processing result value and the second AI interpretation processing result value processed by the AI interpretation processing device itself.

[0187] Here, when measuring the amount of AI interpretation processing, the present disclosure can measure the amount of AI interpretation processing for performing interpretation based on the amount of calculation of the AI model 540 for interpretation.

[0188] As shown in FIG. 17, in the present disclosure, when there is one AI model 540 for interpretation, the amount of AI interpretation processing can be measured based on the amount of calculations processed by one AI model 640.

[0189] That is, the present disclosure can analyze an AI model 540 to determine whether there are branch points 542 connecting one higher-level operator to multiple lower-level operators, and whether there are junction points 544 connecting multiple higher-level operators to one lower-level operator.

[0190] Then, in the present disclosure, if a branch point 542 and a junction point 544 exist, it is confirmed whether at least one parallel processing portion 546 exists based on the branch point 542 and the junction point 544, and if a parallel processing portion 546 exists, the amount of AI interpretation processing can be measured based on a first amount of calculation for the parallel processing portion 546 and a second amount of calculation for the remaining portion other than the parallel processing portion.

[0191] Next, the present disclosure can determine whether to perform distributed processing for the parallel processing portion 546 of the total AI interpretation processing amount for one AI model when the AI interpretation processing amount for the interpretation exceeds the amount that can be processed by itself.

[0192] As another example, as shown in FIG. 18, in the present disclosure, when there are multiple AI models 540, the amount of AI interpretation processing can be measured based on the total amount of calculations processed by the multiple AI models 540.

[0193] Here, the present disclosure analyzes an AI model group to check whether there is a branch point 548 connecting one higher-level AI model to multiple lower-level AI models and a junction point 549 connecting multiple higher-level AI models to one lower-level AI model; if there is a branch point 548 and a junction point 549, it checks whether there is at least one parallel processing part based on the branch point 548 and the junction point 549; if there is a parallel processing part, it can measure the AI interpretation processing amount based on the first calculation amount for the parallel processing part and the second calculation amount for the remaining part other than the parallel processing part.

[0194] Next, the present disclosure can determine whether to perform distributed processing for the parallel processing portion of the total AI interpretation processing amount for an AI model group including multiple AI models when the AI interpretation processing amount for the interpretation exceeds the amount that can be processed by the AI model group itself.

[0195] In addition, the present disclosure checks for each AI model in an AI model group whether there are branch points connecting one higher-level operator to multiple lower-level operators and confluence points connecting multiple higher-level operators to one lower-level operator, and if there are branch points and confluence points, it checks whether there is at least one parallel processing part based on the branch points and confluence points and determines whether to perform distributed processing for the parallel processing part.

[0196] Furthermore, in the present disclosure, when requesting distributed processing of AI interpretation processing, it is also possible to request distributed processing of the parallel processing portion between branch points and junction points within each AI model from the AI interpretation processing device, and it is also possible to request distributed processing of the entire processing portion performed by the parallel model between branch points and junction points within an AI model group from the AI interpretation processing device.

[0197] 19 to 23 are diagrams for explaining a method for providing a multi-party interpretation service by an on-device AI device according to an embodiment of the present disclosure.

[0198] As shown in FIG. 19, the on-device AI device of the present disclosure can receive input of speech from multiple speakers (S10) and preprocess the speech (S20).

[0199] Here, when a speech sound from a speaker is input, the present disclosure can perform preprocessing by analyzing frequencies corresponding to the speech sound and removing noise frequencies.

[0200] Then, the present disclosure can classify the pre-processed speech sounds by speaker (S30).

[0201] Here, the present disclosure extracts speech features from the preprocessed speech, identifies the speaker of the preprocessed speech based on the extracted speech features, and classifies the preprocessed speech according to the identified speaker.

[0202] Next, the present disclosure can confirm whether a specific speaker has been selected from among a plurality of speakers (S40).

[0203] Here, when the speech voice is classified by speaker, the present disclosure generates and provides a speaker list corresponding to the speech voice, and when a user input is received to select at least one speaker included in the speaker list, the speaker selected by the user input can be selected as a specific speaker.

[0204] In some cases, the present disclosure may also divide speech sounds into speakers, convert the speech sounds of each speaker for a predetermined period of time into text, analyze the converted text to extract common related keywords for each speaker, and select a specific speaker from multiple speakers based on the common related keywords for each speaker.

[0205] In another case, the present disclosure can classify speech sounds by speaker, convert the speech sounds of each speaker for a predetermined period of time into text, analyze the converted text to extract common related keywords for each speaker, and select a specific speaker from multiple speakers based on the common related keywords for each speaker.

[0206] Next, in the present disclosure, when a specific speaker is selected from among a plurality of speakers, the speech of the selected specific speaker can be extracted from the speeches classified by speaker (S50).

[0207] Here, in the present disclosure, when a specific speaker is selected from among multiple speakers, only the speech of the selected specific speaker is retained from the speech sounds divided by speaker, and the speech sounds of the remaining speakers other than the specific speaker can be removed.

[0208] Then, in the present disclosure, once the speech of the specific speaker is extracted, the speech of the specific speaker can be interpreted into the target language and output (S60).

[0209] Here, the present disclosure converts the speech of a specific speaker into text to check the current language, and when the target language is set, measures the amount of AI interpretation processing required to interpret the current language into the target language and determines whether self-interpretation processing is possible. If it is determined that self-interpretation processing is possible, the AI interpretation processing is performed by itself to interpret the speech of the specific speaker into the target language.

[0210] Furthermore, in the present disclosure, if a specific speaker is not selected from among a plurality of speakers, the speech of all speakers can be interpreted into the target language and output (S70).

[0211] As shown in FIG. 20, the on-device AI device of the present disclosure can make a distributed processing request based on the AI interpretation processing amount for translating the current language into the target language.

[0212] As shown in FIG. 20, the present disclosure converts the speech of a specific speaker into text to confirm the current language, and once the target language is set, measures the amount of AI interpretation processing required to translate the current language into the target language (S61).

[0213] Then, the present disclosure can check whether the measured AI interpretation processing amount exceeds its own processable amount (S62).

[0214] Next, the present disclosure determines that the AI interpretation processing is impossible if the measured AI interpretation processing amount exceeds the amount that can be processed by the AI interpretation processing device itself, and can select an AI interpretation processing device for distributed processing of the AI interpretation processing (S63).

[0215] Here, the present disclosure checks whether there is a communication connection with an external device, and when the external device is connected to the external device, obtains identification information of the external device from the external device, and checks whether the external device is an AI interpretation processing device based on the identification information. If the external device is an AI interpretation processing device, the external device can be selected as the AI interpretation processing device for the AI interpretation processing distributed processing.

[0216] Next, the present disclosure can request the selected AI interpretation processing device to perform distributed processing of the AI interpretation processing (S64).

[0217] Here, the present disclosure calculates the excess amount of the measured AI interpretation processing volume other than the amount that can be processed by the AI interpretation processing device itself, checks the processing volume of the selected AI interpretation processing device, and if the processing volume of the AI interpretation processing device is greater than the excess amount, requests the AI interpretation processing device to perform distributed processing of the AI interpretation processing for the excess amount.

[0218] In this case, if the processing capacity of the AI interpretation processing device is less than the excess amount, the present disclosure can further select another AI interpretation processing device and request the multiple AI interpretation processing devices to process the excess amount in a distributed manner.

[0219] In addition, the present disclosure calculates the excess amount of the measured AI interpretation processing volume beyond the amount that can be processed by the device itself, extracts the distributed processing portion of the AI interpretation processing corresponding to the excess amount, and requests the selected AI interpretation processing device to process the extracted distributed processing portion in a distributed manner.

[0220] Then, the present disclosure can receive a first AI interpretation processing result value from the AI interpretation processing device (S65).

[0221] Next, the present disclosure can provide a final interpretation result value based on the first AI interpretation processing result value and the self-processed second AI interpretation processing result value (S66).

[0222] Here, the present disclosure can provide a final interpretation result value that interprets the speech of a specific speaker into a target language by mapping the first AI interpretation processing result value received from the AI interpretation processing device and the second AI interpretation processing result value processed by the AI interpretation processing device to each other.

[0223] Furthermore, when interpreting and outputting the speech of a specific speaker, the present disclosure can output the interpretation result in at least one of a first method in which the interpretation result is output as audio from a speaker, and a second method in which the interpretation result is output as text through a display.

[0224] Meanwhile, in the present disclosure, if the measured amount of AI interpretation processing is less than the amount that can be processed by itself, it determines whether or not self-interpretation processing is possible (S67), and if it determines that self-interpretation processing is possible, it processes the AI interpretation processing by itself and provides the AI interpretation processing result value (S68).

[0225] As shown in FIG. 21, the on-device AI device of the present disclosure can receive a user command requesting an interpretation service (S111).

[0226] Next, when a user command requesting an interpretation service is input, the present disclosure can check whether or not there is a communication connection with the AI interpretation processing device (S113).

[0227] Here, when a user command requesting an interpretation service is input, the present disclosure checks whether interpretation of the speech of a specific speaker is currently being performed, and if interpretation of the speech of a specific speaker is currently being performed, interrupts the interpretation of the speech of the specific speaker, and can obtain speech data of the specific speaker that will be input after the point in time when interpretation was interrupted.

[0228] Next, when the present disclosure is connected to the AI interpretation processing device, it can request the AI interpretation processing device to provide an interpretation service for the speech of a specific speaker (S115).

[0229] Then, the present disclosure can receive approval for the request for the interpretation service from the AI interpretation processing device (S117).

[0230] Next, when the present disclosure receives approval for the request for interpretation services from the AI interpretation processing device, it can transmit the speech data of the specified speaker and the target language information to be interpreted to the AI interpretation processing device so that the AI interpretation processing device can interpret the speech of the specified speaker into the target language and output it (S119).

[0231] In addition, when a user command is input to request the suspension of interpretation for some speakers in accordance with the interpretation service, the present disclosure requests the AI interpretation processing device to suspend the interpretation service for some speakers, and receives from the AI interpretation processing device information indicating the completion of the suspension of the interpretation service for some speakers as well as information indicating the continuation of the interpretation service for the remaining speakers.

[0232] As shown in FIG. 22, the on-device AI device of the present disclosure can receive a distributed processing request from an AI interpretation processing device (S121).

[0233] Then, when the present disclosure receives a distributed processing request from the AI interpretation processing device, it can extract distributed processing information from the distributed processing request (S123).

[0234] Here, the present disclosure can extract distributed processing information including AI model information, location information, distributed processing amount information, and input data corresponding to the distributed processing portion from the distributed processing request (S125).

[0235] Next, the present disclosure can perform AI interpretation processing based on the distributed processing information to generate an AI interpretation processing result value, and provide the generated AI interpretation processing result value to the AI interpretation processing device (S127).

[0236] As shown in FIG. 23, the on-device AI device of the present disclosure can receive a request for an interpretation service from an AI interpretation processing device (S131).

[0237] Then, when the present disclosure receives a request for an interpretation service from the AI interpretation processing device, it can measure the amount of AI interpretation processing corresponding to the interpretation service request (S133).

[0238] Next, the present disclosure can determine whether or not to perform self-interpretation processing based on the measured amount of AI interpretation processing (S135).

[0239] Next, if the present disclosure determines that self-interpretation processing is possible, it can transmit approval for the request for the interpretation service to the AI interpretation processing device (S136).

[0240] If the present disclosure determines that the self-interpretation process is not possible, it can transmit a notification message to the AI interpretation processing device informing the AI interpretation service that the self-interpretation process is not possible (S139).

[0241] Then, the present disclosure can confirm whether the speaker's speech data and the target language information to be interpreted are received from the AI interpretation processing device (S137).

[0242] Next, when the present disclosure receives the speaker's speech data and the target language information to be interpreted from the AI interpretation processing device, it can interpret the speaker's speech into the target language and output it (S138).

[0243] Here, when the present disclosure receives a request from the AI interpretation processing device to suspend the interpretation service for some speakers, it can suspend the interpretation service for some speakers and transmit to the AI interpretation processing device information on the completion of the suspension of the interpretation service for some speakers, along with information on the continuation of the interpretation service for the remaining speakers.

[0244] 24 and 25 are diagrams illustrating a method for providing a multi-party interpretation service by an AI interpretation processing device communicatively connected to an on-device AI device according to one embodiment of the present disclosure.

[0245] As shown in FIG. 24, the AI interpretation processing device communicatively connected to the on-device AI device of the present disclosure may include a communication unit communicatively connected to the on-device AI device, a memory for storing an AI model for AI interpretation processing, and a processor for performing AI interpretation processing in response to a request for an interpretation service from the on-device AI device.

[0246] The processor of the AI interpretation processing device may receive a request for an interpretation service from the on-device AI device (S211).

[0247] Then, when the processor of the AI interpretation processing device receives a request for an interpretation service from the device AI device, it can measure the amount of AI interpretation processing corresponding to the request for the interpretation service (S213).

[0248] Next, the processor of the AI interpretation processing device can determine whether to perform its own interpretation process based on the measured amount of AI interpretation processing (S215).

[0249] Next, if the processor of the AI interpretation processing device determines that it can perform the interpretation process itself, it can transmit approval for the request for the interpretation service to the on-device AI device (S216).

[0250] If the processor of the AI interpretation processing device determines that it cannot process the interpretation itself, it can transmit a notification message to the on-device AI device informing it that it cannot process the interpretation service itself (S219).

[0251] Then, the processor of the AI interpretation processing device can check whether the speaker's speech data and the target language information to be interpreted are received from the on-device AI device (S217).

[0252] Next, when the processor of the AI interpretation processing device receives the speaker's speech data and the target language information to be interpreted from the on-device AI device, it can interpret the speaker's speech into the target language and output it (S218).

[0253] Here, when the processor of the AI interpretation processing device receives a request from the on-device AI device to suspend the interpretation service for some speakers, it suspends the interpretation service for some speakers and transmits to the on-device AI device information on the completion of the suspension of the interpretation service for some speakers and information on the continuation of the interpretation service for the remaining speakers.

[0254] As shown in FIG. 25, an AI interpretation processing device communicatively connected to an on-device AI device of the present disclosure can receive a distributed processing request from the on-device AI device (S221).

[0255] Then, when the processor of the AI interpretation processing device receives a distributed processing request from the on-device AI device, it can extract distributed processing information from the distributed processing request (S223).

[0256] Here, the processor of the AI interpretation processing device can extract distributed processing information including AI model information, location information, distributed processing amount information, and input data corresponding to the distributed processing part from the distributed processing request.

[0257] Next, the processor of the AI interpretation processing device can perform AI interpretation processing based on the distributed processing information (S225).

[0258] Next, the processor of the AI interpretation processing device can generate an AI interpretation processing result value and provide the generated AI interpretation processing result value to the on-device AI device (S227).

[0259] FIG. 26 is a diagram illustrating a method for providing a multi-party interpretation service in an on-device AI system according to an embodiment of the present disclosure.

[0260] As shown in Figure 26, when speech sounds are input from multiple speakers (S310), the on-device AI device 500 preprocesses the speech sounds (S320) and classifies the preprocessed speech sounds by speaker (S330).When a specific speaker is selected from the multiple speakers, the on-device AI device 500 can extract the speech sounds of the selected specific speaker from the speech sounds classified by speaker (S340).

[0261] Next, the on-device AI device 500 measures the amount of AI interpretation processing required to convert the speech of a specific speaker into text and interpret the current language into the target language, and determines whether distributed processing is possible by checking whether the measured AI interpretation processing amount exceeds its own processing capacity.

[0262] Next, if the measured AI interpretation processing amount exceeds the amount that the on-device AI device 500 can process, the on-device AI device 500 can select the AI interpretation processing device 700 for distributed processing.

[0263] Then, the on-device AI device 500 requests a communication connection from the AI interpretation processing device 700, and the AI interpretation processing device 700 approves the communication connection in response to the communication connection request from the on-device AI device 500, so that the AI interpretation processing device 700 and the on-device AI device 500 can be communicatively connected to each other (S350, S360).

[0264] Next, the on-device AI device 500 can request the AI interpretation processing device 700 to perform distributed processing of the AI interpretation processing (S370).

[0265] Next, when the AI interpretation processing device 700 receives a distributed processing request from the on-device AI device 500, it extracts distributed processing information from the distributed processing request (S380), performs AI interpretation processing based on the distributed processing information to generate an AI interpretation processing result value (S390), and can provide the generated AI interpretation processing result value to the on-device AI device 500 (S400).

[0266] Then, when the on-device AI device 500 receives the first AI interpretation processing result value from the AI interpretation processing device 700, it can provide a final interpretation result based on the first AI interpretation processing result value and the second AI interpretation processing result value processed by itself (S410).

[0267] Next, when a user command requesting an interpretation service is input, the on-device AI device 500 requests the AI interpretation processing device 700 to provide an interpretation service for the speech of a specific speaker (S430). When the on-device AI device 500 receives approval for the interpretation service request from the AI interpretation processing device 700 (S440), the on-device AI device 500 can transmit the speech data of the specific speaker and the target language information to be interpreted to the AI interpretation processing device 700 so that the AI interpretation processing device 700 can translate the speech of the specific speaker into the target language and output it (S450).

[0268] Next, when the AI interpretation processing device 700 receives the speaker's speech data and the target language information to be interpreted from the on-device AI device 500, it can interpret the speaker's speech into the target language and output it (S460).

[0269] In this way, the present disclosure enables a conversation between multiple speakers to be identified by speaker and accurately interpreted in real time by extracting the speech of the specific speaker from the speech sounds divided by speaker and translating it into the target language when a specific speaker is selected from multiple speakers.

[0270] In addition, the present disclosure can improve the speed of AI interpretation processing, the accuracy of interpretation processing results, and service quality by selecting an external AI interpretation processing device and requesting distributed processing of AI interpretation processing when the amount of AI processing required to interpret a speaker's speech exceeds the amount that can be processed by the device itself.

[0271] In addition, the present disclosure can minimize power consumption and reduce heat generation by distributing AI interpretation processing with an externally located AI interpretation processing device, thereby improving performance and lifespan.

[0272] The present disclosure described above can be embodied as computer-readable code on a medium having a program recorded thereon. The computer-readable medium includes any type of storage device that stores data readable by a computer system. Examples of computer-readable media include hard disk drives (HDDs), solid-state disks (SSDs), silicon disk drives (SDDs), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer may also include a processor of an artificial intelligence device.

Claims

1. an input module configured to receive sensing data associated with a plurality of utterances from a plurality of speakers; a processor configured to perform AI processing to interpret any of the plurality of utterances into a target language; and Equipped with The processor: preprocessing the plurality of utterances; classifying the preprocessed utterances in association with each of the plurality of speakers; extracting an utterance of a specific speaker from among the plurality of speakers from the classified utterances; interpreting the utterance of the particular speaker into the target language; configured to output the utterance of the specific speaker in the target language. An on-device AI device characterized by:

2. Interpreting the utterance of the particular speaker includes: identifying a current language by converting the speech of the particular speaker into text; Determine whether self-interpretation processing is possible by measuring the amount of AI interpretation processing required to interpret the current language into the target language; and interpreting the utterance of the specific speaker into the target language based on the result of determining that the system itself is capable of performing the interpretation process. The on-device AI device according to claim 1 .

3. Determining whether self-interpretation processing is possible is Comparing the measured AI interpretation processing amount with its own processing capacity; The on-device AI device according to claim 2 .

4. Determining whether self-interpretation processing is possible includes: determining that the self-interpretation processing is impossible after the measured amount of AI interpretation processing exceeds the self-processing capacity; The processor: Select an AI interpretation processing device for distributed processing of AI interpretation processing; Request the distributed processing of the AI interpretation processing to the selected AI interpretation processing device. configured to provide a final interpretation result value based on the first AI interpretation processing result value and its own processed second AI interpretation processing result value after receiving a first AI interpretation processing result value from the AI interpretation processing device; The on-device AI device according to claim 3 .

5. Requesting the distributed processing of the AI interpretation processing includes: Calculate the surplus amount of the measured AI interpretation processing amount that exceeds the processing capacity of the device itself; Determine that the processing capacity of the selected AI interpretation processing device is greater than the calculated surplus capacity; and requesting the AI interpretation processing device to perform the distributed processing of the AI interpretation processing for the surplus amount. The on-device AI device according to claim 3 .

6. Requesting the distributed processing of the AI interpretation processing includes: Calculate the surplus amount of the measured AI interpretation processing amount that exceeds the amount that can be processed by itself, Extracting a distributed processing portion corresponding to the surplus amount from the entire processing content of the AI interpretation processing; and requesting the selected AI interpretation processing device to perform the distributed processing of the AI interpretation processing for the extracted distributed processing portion. The on-device AI device according to claim 3 .

7. Providing the final interpretation result value includes: providing a final interpretation result value for interpreting the speech of the specific speaker into the target language by mapping the first AI interpretation processing result value received from the AI interpretation processing device and the second AI interpretation processing result value processed by the AI interpretation processing device to each other; The on-device AI device according to claim 3 .

8. Further provided is a communication module that can be connected by wire or wirelessly to an AI interpretation processing device that performs interpretation processing based on an AI model; When a distributed processing request is received from the AI interpretation processing device, The processor: extracting distributed processing information from the distributed processing request; Execute AI interpretation processing based on the extracted distributed processing information to generate an AI interpretation processing result value; and providing the generated AI interpreter processing result value to the AI interpreter processing device. The on-device AI device according to claim 1 .

9. Further provided is a communication module that can be connected by wire or wirelessly to an AI interpretation processing device that performs interpretation processing based on an AI model; When a request for an interpretation service is received from the AI interpretation processing device, The processor: Measure the amount of AI interpretation processing corresponding to the request for the interpretation service and determine whether or not to process the interpretation by itself; If it is determined that the self-interpretation processing is possible, an approval for the request for the interpretation service is sent to the AI interpretation processing device; Thereafter, when the speech data of the speaker and target language information to be interpreted are received from the AI interpretation processing device, the speech data of the speaker is interpreted into the target language and output. The on-device AI device according to claim 1 .

10. A method for providing a multi-party interpretation service using an apparatus equipped with on-device AI, comprising: receiving a plurality of utterances from a plurality of speakers; preprocessing the plurality of utterances; classifying the preprocessed utterances in association with each of the speakers; extracting an utterance of a specific speaker from the plurality of speakers from the classified utterances; interpreting the speech of the particular speaker into a target language; outputting the utterance of the particular speaker in the target language; A method for providing a multi-party interpretation service, comprising: