Electronic device and method for speech recognition

The use of a generative AI model to restore missing frames in audio-visual multimodal speech recognition addresses accuracy issues due to user movement or environmental factors, improving speech recognition reliability.

WO2025264031A1PCT designated stage Publication Date: 2025-12-26SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008582
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-15
Filing Date
2025-06-20
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing speech recognition technologies struggle with accuracy when users' movements or environmental factors cause missing image or audio frames, leading to incomplete voice recognition.

Method used

An electronic device and method that utilizes a generative artificial intelligence model to restore missing image and audio frames in audio-visual multimodal speech recognition, synchronizing frames to improve accuracy by using learned machine learning models.

Benefits of technology

Enhances voice recognition accuracy by restoring missing frames, ensuring reliable speech recognition even with user movement or environmental interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008582_26122025_PF_FP_ABST
    Figure KR2025008582_26122025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment of the present disclosure may: obtain, during a first time period, visual data associated with a lip image of a speaker by means of at least one camera of an electronic device; obtain, during the first time period, audio data associated with the speaker's speech through at least one microphone of the electronic device; identify at least one missing frame on the basis of first frames corresponding to the visual data and second frames corresponding to the audio data; generate, on the basis of a trained machine learning model, at least one frame corresponding to the at least one missing frame by using the first frames and the second frames; and perform a speech recognition operation on the basis of the first frames, the second frames, or the at least one generated frame. Various other embodiments are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method for speech recognition

[0001] The present disclosure relates to an electronic device and method for voice recognition.

[0002] Speech recognition can refer to the technology that converts a user's voice input into text on an electronic device. The electronic device can segment voice data collected through a microphone into frames, predict the word corresponding to each frame, and output it as text.

[0003] Audio-visual multimodal speech recognition can refer to a technology that recognizes speech based on the user's voice and images in an electronic device. For example, audio-visual multimodal speech recognition can improve the accuracy and reliability of speech recognition by utilizing visual data, such as an image of the user's lips, along with the user's voice data.

[0004] The above information may be provided as background information to aid in understanding this document. None of the above is claimed to be prior art related to this document or can be used to determine prior art.

[0005] One embodiment of the present disclosure may provide an electronic device and method for voice recognition.

[0006] One embodiment of the present disclosure may provide an electronic device and method for restoring missing image frames and / or missing audio frames based on a generative artificial intelligence model in audio-visual multimodal speech recognition.

[0007] One embodiment of the present disclosure can provide an electronic device and method for synchronizing image frames and audio frames based on restored frames.

[0008] One embodiment of the present disclosure can provide an electronic device and method capable of providing more accurate voice recognition results even when a user's voice or lip image is missing due to the user's movement or environmental factors.

[0009] An electronic device (101) (301) according to one embodiment of the present disclosure may include at least one camera (180) (310), at least one microphone (150) (320), a memory (130) (330), and at least one processor (120) (340) including a processing circuit. The memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to perform at least one operation. The at least one operation may include an operation of acquiring visual data (904) associated with a lip image of a speaker through the at least one camera (180) (310) during a first time period. The at least one operation may include acquiring audio data (902) associated with the speaker's voice through the at least one microphone (150) (320) during the first time period. The at least one operation may include identifying at least one missing frame based on first frames corresponding to the visual data (904) and second frames corresponding to the audio data (902). The at least one operation may include generating (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712). The at least one operation may include performing a voice recognition operation based on the first frames, the second frames, or the at least one generated frame.

[0010] An electronic device according to one embodiment of the present disclosure may include at least one camera (180) (310), at least one microphone (150) (320), a communication circuit (190), a memory (130) (330), and at least one processor (120) (340) including a processing circuit. The memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to perform at least one operation. The at least one operation may include an operation of acquiring visual data (904) associated with a speaker's lips through the at least one camera (180) (310) during a first time period. The at least one operation may include an operation of acquiring audio data (902) associated with the speaker's voice through the at least one microphone (150) (320) during the first time period. The at least one operation may include an operation of controlling the communication circuit (190) to transmit first frames corresponding to the acquired visual data (904) and second frames corresponding to the acquired audio data (902) to a server. The at least one operation may include an operation of receiving a speech recognition result from the server through the communication circuit (190). The speech recognition result may include a result of identifying at least one missing frame based on the first frames and the second frames, generating (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712), and performing a speech recognition operation based on the first frames, the second frames, or the at least one generated frame.

[0011] A voice recognition method of an electronic device according to one embodiment of the present disclosure may include an operation (1702) of acquiring visual data (704) (904) associated with a lip image of a speaker through at least one camera (180) (310) of the electronic device (101) (301) during a first time period. The voice recognition method may include an operation (1704) of acquiring audio data (702) (902) associated with a voice of the speaker through at least one microphone (150) (320) of the electronic device (101) (301) during the first time period. The voice recognition method may include an operation (1706) of identifying at least one missing frame based on first frames corresponding to the visual data (904) and second frames corresponding to the audio data (902). The speech recognition method may include an operation (1708) of generating (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712). The speech recognition method may include an operation (1710) of performing a speech recognition operation based on the first frames, the second frames, or the at least one generated frame.

[0012] A voice recognition method of an electronic device according to one embodiment of the present disclosure may include an operation (2602) of acquiring visual data (904) associated with a speaker's lips through at least one camera (180) (310) of the electronic device (101) (301) during a first time period. The voice recognition method may include an operation (2604) of acquiring audio data (902) associated with the speaker's voice through at least one microphone (150) (320) of the electronic device (101) (301) during the first time period. The voice recognition method may include an operation (2606) of transmitting first frames corresponding to the acquired visual data (904) and second frames corresponding to the acquired audio data (902) to a server. The voice recognition method may include an operation (2608) of receiving a voice recognition result from the server. The speech recognition result may include a result of identifying at least one missing frame based on the first frames and the second frames, generating (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712), and performing a speech recognition operation based on the first frames, the second frames, or the at least one generated frame.

[0013] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0014] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0015] Figure 2 is an example of a generative artificial intelligence system.

[0016] FIG. 3 is a schematic block diagram of an electronic device according to one embodiment.

[0017] FIG. 4a is a diagram for explaining a voice recognition operation according to one embodiment.

[0018] FIG. 4b is a diagram for explaining a multimodal speech recognition operation based on restoration of audio data according to one embodiment.

[0019] FIG. 4c is a diagram for explaining a multimodal speech recognition operation based on restoration of visual data according to one embodiment.

[0020] FIG. 5A is a diagram illustrating a first type of electronic device according to one embodiment.

[0021] FIG. 5b is a diagram illustrating a second type of electronic device according to one embodiment.

[0022] FIG. 5c is a diagram illustrating a third type of electronic device according to one embodiment.

[0023] FIG. 6A is a diagram illustrating a situation in which lip recognition is possible in an electronic device according to one embodiment.

[0024] FIG. 6b is a diagram illustrating a situation in which lip recognition is impossible in an electronic device according to one embodiment.

[0025] Fig. 7 is a block diagram of a multimodal speech recognition device according to one embodiment.

[0026] FIG. 8a is a diagram illustrating a learning operation of a multimodal synchronization module for image frame restoration according to one embodiment.

[0027] FIG. 8b is a diagram illustrating a learning operation of a multimodal synchronization module for audio frame restoration according to one embodiment.

[0028] FIG. 9a is a diagram illustrating a voice recognition operation when a multimodal synchronization module according to one embodiment is not used.

[0029] FIG. 9b is a diagram illustrating a voice recognition operation when a multimodal synchronization module according to one embodiment is used.

[0030] FIG. 10 is a diagram for explaining the learning operation of a multimodal ASR module according to one embodiment.

[0031] FIG. 11 is a flowchart schematically illustrating the operation of an electronic device according to one embodiment.

[0032] FIG. 12 is a diagram illustrating the results of multimodal speech recognition performed in an electronic device according to one embodiment.

[0033] FIG. 13 is a diagram illustrating a UI indicating a voice or image input status provided by an electronic device according to one embodiment.

[0034] FIG. 14 is a diagram illustrating an operation of an electronic device according to one embodiment of the present invention to output a notification message indicating a failure in voice or lip recognition.

[0035] FIG. 15 is a diagram illustrating an example of a multimodal voice recognition operation being used in an electronic device according to one embodiment.

[0036] FIG. 16 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a voice recognition result.

[0037] FIG. 17 is a flowchart illustrating a voice recognition operation of an electronic device according to one embodiment.

[0038] FIG. 18 is a flowchart illustrating an operation of an electronic device according to one embodiment to identify at least one missing frame.

[0039] FIG. 19 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to identify at least one missing frame through timeline-based sorting.

[0040] FIG. 20 is a flowchart illustrating an operation of an electronic device according to one embodiment to restore a missing frame.

[0041] FIG. 21 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a message based on the number of at least one missing frame.

[0042] FIG. 22 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a result of a voice recognition operation.

[0043] FIG. 23 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a voice recognition result through a server.

[0044] FIG. 24 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a voice recognition result through an on-device operation.

[0045] FIG. 25 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a frame input status.

[0046] FIG. 26 is a flowchart illustrating an operation of an electronic device receiving a voice recognition result from a server according to one embodiment.

[0047] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0048] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.

[0049] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0050] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0051] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0052] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0053] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0054] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0055] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0056] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0057] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0058] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0059] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0060] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0061] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0062] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0063] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0064] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0065] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0066] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0067] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0068] According to various embodiments, the antenna module (197) may generate a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0069] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0070] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0071] Figure 2 is an example of a generative artificial intelligence system.

[0072] Referring to FIG. 2, the user query / response interface (210) can receive a user's input. The user's input may be in the form of natural language, images, and / or videos. Additionally, context information (212) may also be transmitted when the user's input is transmitted. The context information (212) may include various additional information at the time of the user input. For example, there is additional information such as information on the application currently being used by the user or information on the user's location. Furthermore, the user's input may be in the form of a mixture of the aforementioned natural language, images, sounds, or context information (212). Furthermore, the user's input may also be in a non-natural language form, such as selecting a menu. The user query / response interface (210) can output the results of the generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user query / response interface (210) can output the results of the generative artificial intelligence system to the user. The output can be in natural language form, in the form of specific content, or in the form of an action requested by the user.

[0073] The AI ​​(artificial intelligence) framework (220) can receive user input and coordinate and control each component necessary to perform the user's intention based on the user's query.

[0074] User input received from the user query / response interface (210) can be transmitted to the prompt design component (222). The prompt design component (222) can be used to generate a prompt suitable for inputting the user input into a large language model (LLM) or a large multimodal model (LMM). The prompt design component (222) can be an AI component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design component (222) can access user preference data (214), a prompt library (216), and a knowledge component including prompt examples based on the user input to generate a prompt, and transmit the generated prompt to the LLM or LMM.

[0075] The API / plug-in management component (224) can communicate with external information when there is a request for additional information when passing user input as input to the generative model. The API / plug-in management component (224) can establish a channel for communication with the outside of the AI ​​interface through the API, and can enable access to various data sources (e.g., knowledge repositories (240)) through the established channel. In addition, if the API / plug-in management component (224) needs to perform an action that performs the user input as a final result rather than an intermediate result in an application or service, it can request the application / service component (250) through the API to perform the action. Information obtained from the outside can be used to generate a prompt in the prompt design component (222) together with the user input, or can be passed as input to the generative model.

[0076] The refiner component (226) can fine-tune the output from the generative model. For example, the refiner component (226) can verify that the content generated through the LLM and / or LMM is not irrelevant, does not contain biased content, or does not contain harmful content. Furthermore, the refiner component (226) can determine the degree to which the content matches the user's desired result and, if necessary, perform additional processing. The refiner component (226) can additionally configure and provide hints to the user to avoid undesirable output.

[0077] A generative AI model (230) may generally refer to an artificial intelligence neural network that generates new types of data based on user input information. The generative AI model (230) may include an image-generating model and / or a language-generating model. Representative models for generating images include a generative adversarial network (GAN) and a variational autoencoder (VAE), and examples include a diffusion-based generative model that uses a VAE and a transformer structure. A language-generating model is a model trained to statistically output the most appropriate output based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there is also an LMM that can recognize various types of data input, such as text, images, or voice, and generate new data corresponding to them.

[0078] FIG. 3 is a schematic block diagram of an electronic device according to one embodiment.

[0079] Referring to FIG. 3, the electronic device (301) may include at least one camera (310), at least one microphone (320), memory (330), and processor (340).

[0080] According to one embodiment, the electronic device (301) may include additional components (e.g., a display, a speaker, or communication circuitry) in addition to the illustrated components, or may omit at least one of the illustrated components.

[0081] According to one embodiment, the electronic device (301) may be implemented identically or similarly to the electronic device (101) of FIG. 1.

[0082] According to one embodiment, at least one camera (310), at least one microphone (320), memory (330), and processor (340) may be implemented identically or similarly to the camera module (180), input module (150), memory (130), and processor (120) of FIG. 1, respectively.

[0083] According to one embodiment, the electronic device (301) may be any one of a smart phone, a tablet, a computing device (e.g., a personal computer (PC) or a laptop), or a wearable device (e.g., a smart watch, a head mounted display (HMD), or a smart ring), but is not limited thereto and may be implemented as various types of electronic devices.

[0084] At least one camera (310) can capture an image of the outside. In one embodiment, the at least one camera (310) can be used to capture an image of the user's lips while the user's (or speaker's or talker's) voice is input to the electronic device (301).

[0085] At least one microphone (320) can receive sound from an external source. In one embodiment, at least one microphone (320) can be used to input a user's voice and can be activated for substantially the same period of time as at least one camera (310).

[0086] The memory (330) can store at least one program. For example, the memory (330) can store at least one program for processing and controlling the processor (340), and can store input and / or output data (e.g., audio data, visual data, or voice recognition results). The memory (330) can also store at least one AI model.

[0087] The processor (340) can control the overall operation of the electronic device (301). The processor (340) can execute calculations or data processing related to control and / or communication of at least one other component of the electronic device (301). For example, the processor (340) can be electrically connected to at least one camera (310), at least one microphone (320), and a memory (330), and can execute instructions of a program stored in the memory (330).

[0088] The processor (340) may include a processing circuit that executes instructions of a program stored in the memory (330). The processor (340) may include at least one of a central processing unit (CPU), a neural processing unit (NPU), a graphics processing unit (GPU), a micro processing unit (MPU), a micro controller unit (MCU), an application processor (AP), a communication processor (CP), a system on chip (SoC), or an integrated circuit (IC), a sensor hub, a supplementary processor, a communication processor, an application processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), and may have multiple cores.

[0089] The processor (340) can control the operations of the electronic device (101) (301) by executing instructions stored in the memory (330). For example, the processor (340) can correspond to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors.

[0090] An electronic device (101) (301) according to one embodiment of the present disclosure comprises: at least one camera (180) (310); at least one microphone (150) (320); a memory (130) (330); And at least one processor (120) (340) including a processing circuit, wherein the memory (130) (330) is configured to, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: acquire visual data (904) associated with a lip image of a speaker through the at least one camera (180) (310) during a first time period, acquire audio data (902) associated with a voice of the speaker through the at least one microphone (150) (320) during the first time period, identify at least one missing frame based on first frames corresponding to the visual data (904) and second frames corresponding to the audio data (902), and, based on a learned machine learning model (712), use the first frames and the second frames to identify at least one missing frame. At least one frame corresponding to the missing frame can be generated (906), and instructions for causing a speech recognition operation to be performed based on the first frames, the second frames, or the at least one generated frame can be stored.

[0091] According to one embodiment, the memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: determine whether a first time corresponding to the number of the first frames and a second time corresponding to the number of the second frames are the same, and, based on the first time and the second time not being the same, identify the at least one missing frame associated with frames corresponding to a shorter time among the first frames or the second frames.

[0092] According to one embodiment, the memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: determine whether a first time corresponding to the number of the first frames and a second time corresponding to the number of the second frames are the same; align the first frames and the second frames based on a time line corresponding to the first time period based on the fact that the first time and the second time are not the same; perform synchronization between the first frames and the second frames based on the alignment; identify at least one unsynchronized frame among the first frames or the second frames; and identify the at least one missing frame based on the at least one unsynchronized frame.

[0093] According to one embodiment, the memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: extract a first feature from source frames (710) corresponding to a longer time among the first frames or the second frames, based on the first time and the second time being not the same; extract a second feature from target frames (708) corresponding to a shorter time among the first frames or the second frames; fuse the first feature and the second feature using the machine learning model; and generate the at least one frame based on the fused feature.

[0094] According to one embodiment, the electronic device (101) (301) further includes a display (160); and a speaker (155), wherein the memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: output a message (1402) (1408) (1410) through the display (160) or the speaker (155) instructing to re-capture the speaker's lips or re-input the speaker's voice based on the number of the at least one missing frame being greater than or equal to a threshold value.

[0095] According to one embodiment, the electronic device (101)(301) comprises a display (160); And further comprising a speaker (155), wherein the memory (130) (330) may store instructions that, when individually or wholly executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: generate a first result (1502) (1508) of performing the voice recognition operation based on the first frames or the second frames, generate a second result (1506) (1512) of performing the voice recognition operation based on the first frames, the second frames, or the at least one generated frame as an updated result of the first result (1502) (1508), and output the first result (1502) (1508) and the second result (1506) (1512) through the display (160) or the speaker (155) at the same or different times.

[0096] According to one embodiment, the electronic device (101) (301) further includes a communication circuit (190); a display (160); and a speaker (155), and the memory (130) (330) may store instructions that, when individually or wholly executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: control the communication circuit (190) to transmit first information including the first frames, the second frames, or the at least one generated frame to a server; in response to transmitting the first information, receive second information including a voice recognition result from the server through the communication circuit (190); and output the voice recognition result through the display (160) or the speaker (155) based on the received second information.

[0097] According to one embodiment, the electronic device (101) (301) further includes a display (160); and a speaker (155), and the memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: perform the voice recognition operation based on the learned voice recognition model (726) using the first frames, the second frames, or the at least one generated frame; obtain a voice recognition result based on the voice recognition operation; and output the obtained voice recognition result through the display (160) or the speaker (155).

[0098] According to one embodiment, the electronic device (101) (301) further includes a display (160), and the memory (130) (330) may store instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: output information (1302) (1312) (1322) indicating a frame input state based on the number of the first frames, to the display (160), based on the acquisition of the visual data (904), and output information (1304) (1314) (1324) indicating a frame input state based on the number of the second frames, to the display (160), based on the acquisition of the audio data (902).

[0099] An electronic device according to one embodiment of the present disclosure comprises: at least one camera (180) (310); at least one microphone (150) (320); a communication circuit (190); a memory (130) (330); and at least one processor (120) (340) including a processing circuit, wherein the memory (130) (330) controls the communication circuit (190) to control instructions that, when individually or collectively executed by the at least one processor (120) (340), cause the electronic device (101) (301) to: acquire visual data (904) associated with the lips of a speaker through the at least one camera (180) (310) during a first time period, acquire audio data (902) associated with the voice of the speaker through the at least one microphone (150) (320) during the first time period, transmit first frames corresponding to the acquired visual data (904) and second frames corresponding to the acquired audio data (902) to a server, and receive a voice recognition result from the server through the communication circuit (190). The voice recognition result may include a result of identifying at least one missing frame based on the first frames and the second frames, generating (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712), and performing a voice recognition operation based on the first frames, the second frames, or the at least one generated frame.

[0100] FIG. 4a is a diagram for explaining a voice recognition operation according to one embodiment.

[0101] Referring to FIG. 4a, a user's voice (402) input into an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3) can be output as text (408) based on an acoustic model (404) and a language model (406).

[0102] According to one embodiment, a speech signal generated based on a voice input may be input to an acoustic model (404). The acoustic model (404) may analyze the speech signal to extract acoustic features and predict the speech signal on a syllable-by-syllable basis based on the extracted acoustic features. The language model (406) may convert the predicted syllable-by-syllable sequence into words or sentences and output the conversion result as text (408).

[0103] According to one embodiment, when voice input is performed in a noisy environment, or when the user's pronunciation is inaccurate or the user's voice is soft, the sequence predicted by the acoustic model (404) in units of syllables may contain errors. The sequence containing errors may be corrected by the language model (406). For example, the sequence containing errors may be corrected into frequently used words (or sentences) based on the learned data of the language model (406). If the corrected word is linguistically sound or contextually sound, the corrected word may be output from the language model (406). Since the output word is corrected based on statistical data, it may not match the user's voice (402). If the output word does not match the user's voice (402), a voice recognition error may occur. For example, a voice recognition error as shown in the following [Table 1] may occur.

[0104] Reference text (correct information) Output text (prediction information) Error result 1. Just like clothes get wet in a drizzle, love also gets wet in a drizzle. Like a clothes expert, love also gets wet in a drizzle -> expert (recognition error) 2. I plan to enjoy my afternoon tea quietly at the park today. I plan to enjoy my afternoon side dish quietly at the park today. Tea is -> side dish is (recognition error) 3. The cost of enjoying golf is not cheap, but it is a hobby enjoyed by many people. The cost of enjoying golf is not cheap, but it is a hobby enjoyed by many people. Cost is -> rain (recognition error) 4. The highest temperature this summer is 3 degrees Celsius. The highest temperature this summer is 3 degrees Celsius. Degrees -> City (recognition error)

[0105] Referring to [Table 1], the reference text may represent correct answer information as text corresponding to the input user's voice (402). The output text may represent text (408) output based on an acoustic model (404) and a language model (406), and may represent information predicted from the input user's voice (402).

[0106] Referring to Case 1, the spoken word "Love is like getting wet in a drizzle" may be misrecognized as "Expert" in the text output, as in "Love is like getting wet in a drizzle, like being a clothing expert." Since "Wet" and "Expert" are terms with linguistic and contextual meanings, misrecognition can occur even when a language model (406) is used.

[0107] Referring to the second case, the speech of "I plan to enjoy a quiet afternoon tea at the park today" may result in the text being output as "I plan to enjoy a quiet afternoon tea at the park today," with "반차는" being incorrectly recognized as "반찬는." In the reference text of the second case, "반차" indicates a temporal meaning, while in the output text of the second case, "반찬" indicates a food-related meaning. Therefore, texts with different meanings may be output due to the user's incorrect pronunciation.

[0108] Referring to Case 3, the spoken phrase "Golf is not cheap, but it's a hobby enjoyed by many people" may be misrecognized as "cost" and output as "cost." This is because the phrase "cost" is misinterpreted as "cost." During voice input, noise or the user's low voice may result in the output text missing some words (or phonemes).

[0109] Referring to Case 4, the spoken phrase "This summer's maximum temperature is 3 degrees Celsius" may be misrecognized as "City" (city), resulting in the output text "This summer's maximum temperature is 3 degrees Celsius." Similar to the other cases described above, some words may be misrecognized in the output text based on the user's pronunciation, voice volume, or surrounding environment.

[0110] To reduce speech recognition errors as described above and improve speech recognition performance, audio-visual multimodal speech recognition (hereinafter referred to as "multimodal speech recognition") may be utilized. Multimodal speech recognition may be performed based on audio data associated with a user's voice and visual data associated with the user's lips (e.g., data associated with a lip image (or lip image)). When multimodal speech recognition is utilized, missing portions (or damaged or lost portions) of audio data may be restored based on visual data, or missing portions of visual data may be restored based on audio data, thereby providing more accurate speech recognition results.

[0111] FIG. 4b is a diagram for explaining a multimodal speech recognition operation based on restoration of audio data according to one embodiment.

[0112] Referring to FIG. 4b, in an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3), audio data (or audio signal) (452) associated with a speaker's voice can be obtained through at least one microphone (e.g., the input module (150) of FIG. 1 or at least one microphone (320) of FIG. 3).

[0113] According to one embodiment, at least one of the frames included in the audio data (452) may be a missing frame (454). The at least one missing frame (454) in the audio data (452) may represent a frame that does not include a speaker's voice and / or a frame in which a speaker's voice cannot be identified due to noise exceeding a threshold, voices of multiple speakers, or a soft voice of the speaker. The remaining frames in the audio data (452) excluding the at least one missing frame (454) may represent frames in which a speaker's voice is included or in which a speaker's voice is identified.

[0114] In an electronic device, visual data (456) associated with a speaker's lips can be acquired through at least one camera (e.g., the camera module (180) of FIG. 1 or at least one camera (310) of FIG. 3). The visual data (456) can be acquired over substantially the same time period as the audio data (452), and can include frames in which the movement (or shape) of the lips over time is captured.

[0115] In one embodiment, the frames included in the visual data (456) may be non-missing frames. The non-missing frames in the visual data (456) may represent frames that include the speaker's lips or in which the speaker's lips are identifiable. For convenience of explanation, the non-missing frame(s) will be referred to as normal frame(s) hereinafter.

[0116] According to one embodiment, audio data (452) including at least one missing frame (454) and visual data (456) including normal frames can be output to a multimodal synchronization module (458). The multimodal synchronization module (458) can include at least one AI model. The at least one AI model can be a trained machine learning model, which can include a generative AI model capable of restoring at least one missing frame. The generative AI model can be trained to restore at least one missing frame associated with the audio data (452) or the visual data (456).

[0117] According to one embodiment, the operation of restoring at least one missing frame may include the operation of generating at least one frame corresponding to the at least one missing frame, or the operation of generating at least one frame to be included in the position of the at least one missing frame on the time axis.

[0118] According to one embodiment, the multimodal synchronization module (458) can restore at least one missing frame (454) of audio data (452) using visual data (456) based on a generative AI model. For example, the visual data (456) can be used to recognize speech based on lip reading technology, and thus can be used to generate at least one audio frame (462) corresponding to the at least one missing frame (454) (or at least one audio frame (462) to be included at a position on the time axis of the at least one missing frame (454)).

[0119] According to one embodiment, the multimodal synchronization module (458) can generate at least one audio frame (462) and output audio data (460) including at least one audio frame (462). The audio data (460) can be output to a multimodal automatic speech recognition (ASR) module (464). Since the visual data (456) includes normal frames, it can be output to the multimodal ASR module (464) without a restoration operation. The frames included in the audio data (460) and the visual data (456) can be synchronized and output to the multimodal ASR module (464).

[0120] According to one embodiment, the multimodal ASR module (464) can extract acoustic features, such as an audio spectrogram, from audio data (460) and visual features, such as the shape of a speaker's mouth, from visual data (456).

[0121] According to one embodiment, the multimodal ASR module (464) can combine extracted acoustic features and visual features and generate text based on a multimodal ASR model trained using the combined features.

[0122] In one embodiment, text generated by the multimodal ASR module (464) may be passed to a language model (466). The language model (466) may be used to check the grammar or meaning of the text or to supplement the context. Text (468) that passes through the language model (466) may include text that is easy to understand and contextually meaningful.

[0123] FIG. 4c is a diagram for explaining a multimodal speech recognition operation based on restoration of visual data according to one embodiment.

[0124] Referring to FIG. 4c, in an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3), audio data (or audio signal) (472) associated with a speaker's voice can be obtained through at least one microphone (e.g., the input module (150) of FIG. 1 or at least one microphone (320) of FIG. 3).

[0125] According to one embodiment, the audio data (472) may include normal frames. In the audio data (472), normal frames may represent frames that contain a speaker's voice or in which a speaker's voice is identifiable.

[0126] According to one embodiment, an electronic device may acquire visual data (474) associated with a speaker's lips via at least one camera (e.g., a camera module (180) of FIG. 1 or at least one camera (310) of FIG. 3). The visual data (474) may be acquired over substantially the same time period as the audio data (472), and may include frames in which lip movement (or lip shape) over time is captured.

[0127] According to one embodiment, at least one of the frames included in the visual data (474) may be a missing frame (476). At least one missing frame (476) in the visual data (474) may represent a frame in which the speaker's lips are not included or in which the speaker's lips are not clearly identifiable.

[0128] According to one embodiment, audio data (472) including normal frames and visual data (474) including at least one missing frame (476) may be output to a multimodal synchronization module (478). The multimodal synchronization module (478) may be identical to or similar to the multimodal synchronization module (458) of FIG. 4B.

[0129] According to one embodiment, the multimodal synchronization module (478) can restore at least one missing frame (476) of the visual data (474) using the audio data (472) based on the generative AI model. For example, the multimodal synchronization module (458) can generate at least one image frame (482) corresponding to the at least one missing frame (476) (or at least one image frame (482) to be included at a position on the time axis of the at least one missing frame (476)). The multimodal synchronization module (478) can generate the at least one image frame (482) and output visual data (480) including the at least one image frame (482).

[0130] Visual data (480) can be output to a multimodal ASR module (484). Since audio data (472) contains normal frames, it can be output to the multimodal ASR module (484) without a restoration operation. The multimodal ASR module (484) can be the same as or similar to the multimodal ASR module (464) of FIG. 4b.

[0131] According to one embodiment, the multimodal ASR module (484) can extract acoustic features, such as an audio spectrogram, from audio data (472) and visual features, such as a speaker's mouth shape, from visual data (480). The multimodal ASR module (484) can combine the extracted acoustic features and visual features and generate text based on a multimodal ASR model trained using the combined features.

[0132] In one embodiment, text generated by the multimodal ASR module (484) may be passed to a language model (486). The language model (486) may be identical to or similar to the language model (466) of FIG. 4B . The language model (486) may be used to check the grammar or meaning of the text or to supplement the context. Text (488) that passes through the language model (486) may include text that is easy to understand and contextually meaningful.

[0133] As described with reference to FIG. 4b or FIG. 4c, in a multimodal speech recognition operation, missing portions of audio data or missing portions of visual data can be restored based on a generative AI model, so speech recognition errors such as those shown in [Table 1] can be prevented.

[0134] According to one embodiment, the electronic device (301) may be various types of electronic devices. Examples of various types are described below with reference to FIGS. 5A to 5C.

[0135] FIG. 5A is a diagram illustrating a first type of electronic device according to one embodiment.

[0136] Referring to FIG. 5A, the electronic device (301) may be a first type of electronic device. The first type of electronic device may include a mobile terminal such as a smartphone or tablet.

[0137] According to one embodiment, when the electronic device (301) is a first type of electronic device, it may include the following configuration. The front of the electronic device (301) may include a display (503) (e.g., the display module (160) of FIG. 1) and at least one camera (507) (e.g., the camera module (180) of FIG. 1 or at least one camera (310) of FIG. 3). The bottom of the electronic device (301) may include at least one microphone (not shown) (e.g., the input module (150) of FIG. 1 or at least one microphone (320) of FIG. 3).

[0138] When multimodal voice recognition is used, the electronic device (301) can activate at least one camera and at least one microphone. The electronic device (301) can display a user interface (UI) or a graphical user interface (GUI) so that a user can recognize that at least one camera (507) and at least one microphone are activated for multimodal voice recognition. For example, the electronic device (301) can indicate that at least one camera (507) is activated by providing a shaded mark or a circled mark around at least one camera (507). For example, the electronic device (301) can indicate that at least one microphone is activated by displaying at least one display object (e.g., a microphone icon or an icon indicating that voice recognition is in progress) (505) at a set position on the display (503) (e.g., a position adjacent to at least one microphone or a position at the bottom of the display (503).

[0139] At least one camera (507) and at least one microphone may be activated at substantially the same time for multimodal speech recognition. In one embodiment, at least one camera (507) and at least one microphone may be activated at substantially the same time based on a button input or touch input to initiate multimodal data input. Based on a button input or touch input to initiate multimodal data input, information for capturing the user's face and receiving voice input may be output to the display (503) or as audio information.

[0140] In one embodiment, at least one camera (507) and at least one microphone may be deactivated at substantially the same time based on a button input or touch input to terminate multimodal data input. Information indicating that filming has ended and voice input has ended based on a button input or touch input to terminate multimodal data input may be output to the display (503) or as audio information.

[0141] The electronic device (301) can acquire an image including a user's face and receive the user's voice while at least one camera (507) and at least one microphone are activated.

[0142] According to one embodiment, an image (e.g., an image including a user's face or mouth) acquired by at least one activated camera (507) may or may not be displayed on the display (503).

[0143] In one embodiment, when a user's voice is input through at least one activated microphone, a message or graphic may be displayed indicating that the user's voice is being input.

[0144] According to one embodiment, the electronic device (301) may detect a portion corresponding to the user's face from an image acquired through at least one camera (507). When the portion corresponding to the user's face is detected, the electronic device (301) may detect a portion of the detected portion that includes lips as a lip image. The electronic device (301) may perform a multimodal voice recognition operation based on visual data associated with the detected lip image and audio data associated with the user's voice.

[0145] FIG. 5b is a diagram illustrating a second type of electronic device according to one embodiment.

[0146] Referring to FIG. 5B, the electronic device (301) may be a second type of electronic device. The second type of electronic device may include a mobile terminal similar to the first type of electronic device (301) of FIG. 5A, or may include a foldable electronic device or a flexible electronic device whose display can be folded or unfolded.

[0147] According to one embodiment, when the electronic device (301) is a second type of electronic device, it may include the following configuration. The electronic device (301) may include a flexible display on the inner surface and a display (513) (e.g., the display module (160) of FIG. 1) on the outer surface. According to one embodiment, the electronic device (301) may perform a multimodal voice recognition operation through the display (513) included on the outer surface without unfolding the display.

[0148] When multimodal voice recognition is used, the electronic device (301) may activate at least one camera (517) and at least one microphone (not shown). The electronic device (301) may display a UI or GUI so that a user can recognize that at least one camera (517) (e.g., the camera module (180) of FIG. 1 or at least one camera (310) of FIG. 3) and at least one microphone (e.g., the input module (150) of FIG. 1 or at least one microphone (320) of FIG. 3) are activated for multimodal voice recognition. For example, the electronic device (301) may display at least one display object (e.g., a camera icon) (518) on the display (513) indicating that at least one camera (517) is activated. For example, the electronic device (301) may indicate that at least one microphone is activated by displaying at least one display object (e.g., a microphone icon or an icon indicating that voice recognition is in progress) (515) at a set location on the display (513) (e.g., a location adjacent to at least one microphone or a location below the display (513)).

[0149] The electronic device (301) may activate at least one camera (517) and at least one microphone at substantially the same time for multimodal voice recognition. The electronic device (301) may acquire an image including a user's face and receive the user's voice while the at least one camera (517) and at least one microphone are activated. According to one embodiment, an image (e.g., an image including the user's face or mouth) acquired by the at least one activated camera (517) may or may not be displayed on the display (513). According to one embodiment, when the user's voice is input through the at least one activated microphone, a message or graphic indicating that the user's voice is being input may be displayed.

[0150] According to one embodiment, the electronic device (301) can detect a portion corresponding to the user's face from an image acquired through at least one camera (517). When the portion corresponding to the user's face is detected, the electronic device (301) can detect a portion of the detected portion that includes lips as a lip image. The electronic device (301) can perform a multimodal voice recognition operation based on visual data associated with the detected lip image and audio data associated with the user's voice.

[0151] FIG. 5c is a diagram illustrating a third type of electronic device according to one embodiment.

[0152] Referring to FIG. 5c, the electronic device (301) may be a third type of electronic device. The third type of electronic device may include a wearable electronic device. According to one embodiment, the electronic device (301) may be a smartwatch, but is not limited thereto, and may also be an electronic device such as smart glasses, a head-mounted display (HMD), or a smart ring.

[0153] When the electronic device (301) is a third type of electronic device, it may include the following configuration. When multimodal voice recognition is used, the electronic device (301) may activate at least one camera (527) (e.g., the camera module (180) of FIG. 1 or at least one camera (310) of FIG. 3) and at least one microphone (not shown) (e.g., the input module (150) of FIG. 1 or at least one microphone (320) of FIG. 3). The electronic device (301) may display a UI or GUI so that a user can recognize that at least one camera (527) and at least one microphone are activated for multimodal voice recognition. For example, the electronic device (301) may indicate that at least one camera (527) is activated by providing a shaded mark or a circling mark around at least one camera (527). For example, the electronic device (301) may indicate that at least one microphone is activated by displaying at least one display object (e.g., a microphone icon or an icon indicating that voice recognition is in progress) (525) at a set location (e.g., a location adjacent to at least one microphone or a location below the display (523)) on the display (523) (e.g., the display module (160) of FIG. 1).

[0154] The electronic device (301) can activate at least one camera (527) and at least one microphone at substantially the same time for multimodal speech recognition.

[0155] The electronic device (301) can acquire an image including the user's face (528) or mouth and receive the user's voice while at least one camera (527) and at least one microphone are activated. According to one embodiment, the image (e.g., an image including the user's face (528) or mouth) acquired by the at least one activated camera (527) may or may not be displayed on the display (523). According to one embodiment, when the user's voice is input through the at least one activated microphone, a message or graphic indicating that the user's voice is being input may be displayed.

[0156] According to one embodiment, the electronic device (301) can detect a portion corresponding to the user's face (528) from an image acquired through at least one camera (527). When the portion corresponding to the user's face (528) is detected, the electronic device (301) can detect a portion (529) containing lips from the detected portion as a lip image. The electronic device (301) can perform a multimodal voice recognition operation based on visual data associated with the detected lip image and audio data associated with the user's voice.

[0157] At least one camera mentioned below may correspond to, for example, any one of the camera module (180) of FIG. 1, at least one camera (310) of FIG. 3, or at least one camera (507) of FIG. 5a, at least one camera (517) of FIG. 5b, or at least one camera (527) of FIG. 5c.

[0158] At least one microphone mentioned below may correspond to, for example, the input module (150) of FIG. 1 or at least one microphone (320) of FIG. 3.

[0159] FIG. 6A is a diagram illustrating a situation in which lip recognition is possible in an electronic device according to one embodiment.

[0160] Referring to FIG. 6A, the electronic device (301) can acquire an image through at least one camera. According to one embodiment, capturing may be performed while the lips of the speaking user (602) are included in the field of view (FOV) (603) of at least one camera. As a result of capturing, the electronic device (301) can acquire an image including the lips of the speaking user (602) and perform lip recognition from the acquired image.

[0161] FIG. 6b is a diagram illustrating a situation in which lip recognition is impossible in an electronic device according to one embodiment.

[0162] Referring to FIG. 6B, the electronic device (301) can acquire an image through at least one camera. In one embodiment, the capture can be performed in a state where the lips of the speaking user (602) are not included in the FOV (603) of at least one camera.

[0163] In one embodiment, when the user (602) briefly lifts his / her face while speaking, the FOV (603) of at least one camera may include the chin of the user (602) instead of his / her lips, as shown in (a) of FIG. 6b.

[0164] In one embodiment, when the user (602) turns his / her face to the side while speaking, the FOV (603) of at least one camera may include the cheek instead of the user's (602) lips, as shown in (b) of FIG. 6B.

[0165] According to one embodiment, when the position of the electronic device (301) changes upward during the user's (602) speech, the FOV (603) of at least one camera may include the user's (602) eyes instead of his / her lips, as shown in (c) of FIG. 6b.

[0166] As a result of performing a photographing operation in a situation as shown in (a), (b), or (c) of FIG. 6B, the electronic device (301) may obtain an image that does not include the lips of the speaking user (602). The electronic device (301) may be unable to perform lip recognition based on the fact that the lips of the speaking user (602) are not included in the obtained image.

[0167] In one embodiment, during a time period during which multimodal data input is performed, an image that is not lip-recognizable may be acquired. At least one frame corresponding to the acquired image may be identified as at least one missing frame (e.g., at least one missing frame (476) of FIG. 4C) in the visual data (e.g., visual data (474) of FIG. 4C).

[0168] Fig. 7 is a block diagram of a multimodal speech recognition device according to one embodiment.

[0169] Referring to FIG. 7, the multimodal speech recognition device (701) may include a modality frame selection module (706), a multimodal synchronization module (712), or a multimodal ASR module (726).

[0170] According to one embodiment, first modality data (702) and second modality data (704) may be input to a modality frame selection module (706). The first modality data (702) and the second modality data (704) may be different types of data and may be acquired during substantially the same time period. The first modality data (702) may include first frames acquired during time T (or time period T), and the second modality data (704) may include second frames acquired during time T. The first frames may include either audio frames or image frames, and the second frames may include the other of the audio frames or image frames. The audio frames may be associated with a speaker's voice, and the image frames may be associated with an image of the speaker's lips.

[0171] According to one embodiment, the modality frame selection module (706) can determine source frames (710) and target frames (708) based on the number of first frames included in the first modality data (702) and second frames included in the second modality data (704).

[0172] In one embodiment, the modality frame selection module (706) may determine whether a first time corresponding to the number of first frames and a second time corresponding to the number of second frames are substantially the same. Based on the fact that the first time and the second time are not substantially the same, the modality frame selection module (706) may identify at least one missing frame. The at least one missing frame may be associated with frames corresponding to a shorter time among the first frames or the second frames.

[0173] According to one embodiment, at least one missing frame can be identified by alignment and / or synchronization between the first frames and the second frames along a timeline based on time T. For example, the modality frame selection module (706) can align the first frames and the second frames along the timeline and synchronize the aligned first frames and the second frames. If, as a result of the synchronization, at least one unsynchronized frame among the first frames or the second frames is identified, the modality frame selection module (706) can identify the at least one missing frame based on the at least one unsynchronized frame.

[0174] According to one embodiment, the modality frame selection module (706) can determine whether a first frame reception rate associated with the number of first frames and a second frame reception rate associated with the number of second frames are substantially the same. The first frame reception rate and the second frame reception rate can be determined based on the number of image frames received per second (e.g., frames per second (fps)) or the number of audio frames received per second (e.g., sampling rate).

[0175] According to one embodiment, the modality frame selection module (706) may calculate the first frame reception ratio as a ratio (e.g., 80%) of the total number of frames to be acquired for 10 seconds (T=10) (e.g., 300) and the number of image frames actually acquired (or the number of image frames including lip images) (e.g., 240), based on the fact that the first frames are image frames and the number of image frames received per second is 30 (e.g., 30 fps).

[0176] According to one embodiment, the modality frame selection module (706) may calculate the second frame reception ratio as a ratio (e.g., 99%) of the total number of frames to be acquired for 10 seconds (T=10) (e.g., 1000) and the number of image frames (or the number of image frames containing the speaker's voice) actually acquired (e.g., 990) based on the fact that the second frames are audio frames and the number of audio frames received per second is 100 (e.g., sampling rate 100 Hz).

[0177] According to one embodiment, the modality frame selection module (706) may determine frames corresponding to a larger ratio of the first frame reception ratio and the second frame reception ratio (e.g., the second frames) as source frames (710), and determine frames corresponding to a smaller ratio of the first frames and the second frames (e.g., the first frames) as target frames (708).

[0178] According to one embodiment, the modality frame selection module (706) can output source frames (710) and target frames (708) to a multimodal synchronization module (712). The multimodal synchronization module (712) can include a first encoder (714), a second encoder (716), a fusion module (718), a first decoder (720), or a synchronization model (722).

[0179] In one embodiment, the first encoder (714) may compress the target frames (708) to generate first embedding information. For example, the first encoder (714) may extract a first feature (e.g., an acoustic feature such as an audio spectrogram or a visual feature associated with lip movement or shape) from the target frames (708) and generate the first embedding information (or first embedding vector) based on the extracted first feature.

[0180] In one embodiment, the second encoder (716) may compress the source frames (710) to generate second embedding information. For example, the second encoder (716) may extract second features (e.g., visual features associated with lip movement or shape, or acoustic features such as an audio spectrogram) from the source frames (710) and generate second embedding information (or second embedding vector) based on the extracted second features.

[0181] According to one embodiment, the fusion module (718) can fuse (or combine or integrate) the first embedding information output from the first encoder (714) and the second embedding information output from the second encoder (716) to generate fused embedding information (or fused embedding vector). The first embedding information can include the first feature extracted from the target frames (708), and the second embedding information can include the second feature extracted from the source frames (710). The fused embedding information can include the fused result of the first feature and the second feature.

[0182] According to one embodiment, the first decoder (720) may infer and restore at least one missing frame using the fused embedding information based on a loss function or a restoration model. For example, the first decoder (720) may generate at least one frame corresponding to the at least one missing frame, or generate at least one frame to be included in the position of the at least one missing frame on the time axis. The first decoder (720) may output target frames (724) including the at least one generated frame as updated target frames.

[0183] According to one embodiment, the synchronization model (722) can be used for frame synchronization between source frames (710) and updated target frames (724). The source frames (710) and the updated target frames (724) can be synchronized through the synchronization model (722) and then generated as a sequence of a set time length. For example, the source frames (710) can be generated as a first sequence including frames corresponding to time T, and the updated target frames (724) can be generated as a second sequence including frames corresponding to time T. The first sequence and the second sequence can be input to the multimodal ASR module (726).

[0184] In one embodiment, the multimodal ASR module (726) may include a multimodal ASR model. The multimodal ASR module (726) may perform speech recognition based on the first sequence and the second sequence based on the multimodal ASR model. In one embodiment, the multimodal ASR module (726) may include or be connected to a language model. The language model may linguistically modify or supplement the output result of the multimodal ASR module (726) to output natural text appropriate to the context.

[0185] According to one embodiment, the multimodal speech recognition device (701) may be included in the electronic device (301). Based on the inclusion of the multimodal speech recognition device (701) in the electronic device (301), the above-described operations may be performed as on-device operations.

[0186] According to one embodiment, at least one component included in the multimodal speech recognition device (701) may be included in the electronic device (301) or may be included in the server. For example, the modality frame selection module (706) may be included in the electronic device (301), and the multimodal synchronization module (712) and / or the multimodal ASR module (726) may be included in the server. For example, the modality frame selection module (706) and / or the multimodal synchronization module (712) may be included in the electronic device (301), and the multimodal ASR module (726) may be included in the server. The operation of each component may be performed substantially the same as described above, except that communication is performed between the electronic device (301) and the server.

[0187] FIG. 8a is a diagram illustrating a learning operation of a multimodal synchronization module for image frame restoration according to one embodiment.

[0188] Referring to FIG. 8A, image frames included in visual data may be determined as target frames (802), and audio frames included in audio data may be determined as source frames (804). In the multimodal synchronization module (712), learning to restore at least one missing image frame based on audio frames (e.g., learning to generate at least one image frame corresponding to at least one missing image frame, or learning to generate at least one image frame to be included in the position of at least one missing image frame on the time axis) may be performed.

[0189] According to one embodiment, the target frames (802) are image frames (e.g., {a1,…,a) for a time T. m}) may include source frames (804) and audio frames (e.g. {l1,…,l) for a time T. n}) can be included. The number of image frames can be m, and the number of audio frames can be n. m and n can be natural numbers greater than or equal to 0, and can be the same or different.

[0190] In one embodiment, m and n may correspond to the number of frames that can be acquired during a time period T based on the number of frames received per second (e.g., fps or sampling rate). For example, if fps is 30 and the sampling rate is 100, m and n for 10 seconds (e.g., T=10) may be 300 and 1000, respectively.

[0191] According to one embodiment, image frames and audio frames may be used as input data for learning to restore at least one missing image frame based on audio frames. At least one of the image frames may be set as a missing image frame, and the audio frames may be set not to include at least one missing audio frame. Information about the at least one missing image frame may be set to include specific information or a value (e.g., 0) and used for learning.

[0192] According to one embodiment, for learning to restore at least one missing image frame, image reference frames (or image correct frames) (810), which are output data, may be used. The image reference frames (810) may include m image frames (or m normal image frames) that do not include at least one missing image frame.

[0193] According to one embodiment, the following learning may be performed in the multimodal synchronization module (712) based on input data (e.g., image frames and audio frames) and output data (e.g., image reference frames (810)).

[0194] According to one embodiment, image frames, which are target frames (802), may be input to a first encoder (714). The first encoder (714) may compress the image frames to extract a first feature (e.g., a visual feature associated with lip movement or lip shape). The first encoder (714) may generate first embedding information based on the extracted first feature.

[0195] According to one embodiment, audio frames, which are source frames (804), may be input to a second encoder (716). The second encoder (716) may compress the audio frames to extract second features (e.g., acoustic features such as an audio spectrogram). The second encoder (716) may generate second embedding information based on the extracted second features.

[0196] According to one embodiment, the first embedding information generated by the first encoder (714) and the second embedding information generated by the second encoder (716) can be fused by the fusion module (718).

[0197] According to one embodiment, the fusion module (718) may include a plurality of fusion conformers as attention modules for fusion of a first feature included in the first embedding information and a second feature included in the second embedding information. For example, the fusion module (718) may include a first fusion conformer (717-1), a second fusion conformer (717-2), a third fusion conformer (717-3), a fourth fusion conformer (717-4), or a fifth fusion conformer (717-5).

[0198] In one embodiment, a plurality of fusion conformers may perform hierarchical or step-wise learning associated with the fusion of each image frame and each audio frame. For example, a first fusion conformer (717-1) may learn the fusion of the first feature and the second feature based on the input first embedding information and the input second embedding information, and a second fusion conformer (717-2) may perform more detailed learning based on the learning result of the first fusion conformer (717-1). Similarly, a third fusion conformer (717-3), a fourth fusion conformer (717-4), or a fifth fusion conformer (717-5) may perform more detailed learning based on the previous learning result in a hierarchical or step-wise manner.

[0199] According to one embodiment, the fusion module (718) may include a plurality of FCs as fully connected layers. For example, the fusion module (718) may include a first FC (719-1), a second FC (719-2), or a third FC (719-3). The first FC (719-1), the second FC (719-2), or the third FC (719-3) may increase the accuracy of the fusion result (or fused feature) by transforming the output of at least one of the plurality of fusion conformers (e.g., at least one of the third fusion conformer (717-3), the fourth fusion conformer (717-4), or the fifth fusion conformer (717-5)) based on different weight values. The weight values ​​of each of the first FC (719-1), the second FC (719-2), or the third FC (719-3) may be determined or adjusted to reduce the difference between the image data predicted by the fusion result and the output data (e.g., image reference frames (810)).

[0200] According to one embodiment, the fusion result output from the fusion module (718) may be input to the first decoder (720). The first decoder (720) may predict at least one lip image based on the fusion result, and may restore at least one missing image frame based on the prediction result.

[0201] According to one embodiment, the first decoder (720) can predict at least one lip image based on the prediction and / or reconstruction model (808) and generate at least one image frame including the at least one predicted lip image. According to one embodiment, the at least one generated image frame can be a result of reconstruction of at least one missing image frame, and can correspond to the at least one missing image frame or be included at a position of the at least one missing image frame on the time axis.

[0202] In one embodiment, an image loss function (806) may be utilized in connection with the prediction and / or reconstruction model (808). The image loss function (806) may be utilized to compute a difference between a lip image included in at least one generated image frame and at least one lip image included in image reference frames (810). The prediction and / or reconstruction model (808) may be trained such that the output value of the image loss function (806) is minimized.

[0203] According to one embodiment, the first decoder (720) can update the image frames to include at least one generated image frame. The updated image frames can be synchronized with the audio frames based on a synchronization model (722). According to one embodiment, the synchronization is performed using a loss function (L) associated with the synchronization. sync ) can be performed so that the output value is minimized, or so that the temporal characteristics between the input data and the output data correspond.

[0204] FIG. 8b is a diagram illustrating a learning operation of a multimodal synchronization module for audio frame restoration according to one embodiment.

[0205] Referring to FIG. 8B, audio frames included in audio data may be determined as target frames (832), and image frames included in visual data may be determined as source frames (834). In the multimodal synchronization module (712), learning to restore at least one missing audio frame based on the image frames (e.g., learning to generate at least one audio frame corresponding to at least one missing audio frame, or learning to generate at least one audio frame to be included in the position of at least one missing audio frame on the time axis) may be performed.

[0206] According to one embodiment, the target frames (832) are audio frames (e.g., {l1,…,l) for a time T. n}) may include source frames (834) and the source frames (834) may include image frames (e.g., {a1,…,a) for a time T. m}) may be included.

[0207] According to one embodiment, audio frames and image frames may be used as input data for learning to restore at least one missing audio frame based on image frames. At least one of the audio frames may be set as a missing audio frame, and the image frames may be set not to include at least one missing image frame. Information about the at least one missing audio frame may be set to include specific information or a value (e.g., 0) and used for learning.

[0208] According to one embodiment, for learning to restore at least one missing audio frame, audio reference frames (or audio correct frames) (840), which are output data, may be used. The audio reference frames (840) may be n audio frames (or n normal audio frames) that do not include at least one missing audio frame.

[0209] According to one embodiment, the following learning may be performed in the multimodal synchronization module (712) based on input data (e.g., image frames and audio frames) and output data (e.g., audio reference frames (840)).

[0210] According to one embodiment, audio frames, which are target frames (832), may be input to a first encoder (714). The first encoder (714) may compress the audio frames to extract a first feature (e.g., an acoustic feature such as an audio spectrogram). The first encoder (714) may generate first embedding information based on the extracted first feature.

[0211] According to one embodiment, image frames, which are source frames (834), may be input to a second encoder (716). The second encoder (716) may compress the image frames to extract second features (e.g., visual features associated with lip movement or lip shape). The second encoder (716) may generate second embedding information based on the extracted second features.

[0212] According to one embodiment, the first embedding information generated by the first encoder (714) and the second embedding information generated by the second encoder (716) can be fused by the fusion module (718).

[0213] According to one embodiment, a plurality of fusion conformers included in the fusion module (718) can perform hierarchical or step-by-step learning associated with the fusion of each audio frame and each image frame. The first FC (719-1), the second FC (719-2), or the third FC (719-3) included in the fusion module (718) can increase the accuracy of the fusion result (or fused feature) by transforming the output of at least one of the plurality of fusion conformers (e.g., at least one of the third fusion conformer (717-3), the fourth fusion conformer (717-4), or the fifth fusion conformer (717-5)) based on different weight values. The weight values ​​of each of the first FC (719-1), the second FC (719-2), or the third FC (719-3) can be determined or adjusted to reduce a difference between audio data predicted by the fusion result and output data (e.g., audio reference frames (840)).

[0214] According to one embodiment, the fusion result output from the fusion module (718) may be input to the first decoder (720). The first decoder (720) may predict at least one audio signal (e.g., audio spectrogram) associated with speech based on the fusion result, and may restore at least one missing audio frame based on the prediction result.

[0215] According to one embodiment, the first decoder (720) can predict at least one audio signal based on the prediction and / or reconstruction model (838) and generate at least one audio frame including the at least one predicted audio signal. According to one embodiment, the at least one generated audio frame can be a result of reconstruction of at least one missing audio frame, and can correspond to the at least one missing audio frame or be included at a position of the at least one missing audio frame on the time axis.

[0216] In one embodiment, an audio loss function (836) may be utilized in connection with a prediction and / or reconstruction model (838). The audio loss function (836) may be utilized to compute a difference between an audio signal included in at least one generated audio frame and an audio signal included in audio reference frames (840) that are output data. The prediction and / or reconstruction model (838) may be trained such that an output value of the audio loss function (836) is minimized.

[0217] According to one embodiment, the first decoder (720) can update the audio frames to include at least one generated audio frame. The updated audio frames can be synchronized with the image frames based on a synchronization model (722). According to one embodiment, the synchronization is performed using a loss function (L) associated with the synchronization. sync ) can be performed so that the output value is minimized, or so that the temporal characteristics between the input data and the output data correspond.

[0218] FIGS. 9A and 9B are diagrams for comparing voice recognition operations when a multimodal synchronization module according to one embodiment is used and when it is not used.

[0219] FIG. 9a is a diagram illustrating a voice recognition operation when a multimodal synchronization module according to one embodiment is not used.

[0220] Referring to FIG. 9A, audio data (902) may be input as a voice modality, and visual data (904) may be input as a lip image modality. The audio data (902) and the visual data (904) may be input during substantially the same time period, and each may include at least one missing frame. The at least one missing frame may be indicated as a frame in which either the voice or the lip image is missing. In one embodiment, the visual data (904) may include more missing frames than the audio data (902).

[0221] When the multimodal synchronization module (712) is not used, audio data (902) and visual data (904) can be input to the multimodal ASR module (726). The multimodal ASR module (726) can perform a speech recognition operation based on the audio data (902) and the visual data (904). However, the multimodal ASR module (726) may output unclear speech recognition results because each of the audio data (902) and the visual data (904) includes missing frames. For example, the multimodal ASR module (726) may output text information with some content lost as a speech recognition result (e.g., I'm on the subway now...thank goodness I'm on it now (908)).

[0222] FIG. 9b is a diagram illustrating a voice recognition operation when a multimodal synchronization module according to one embodiment is used.

[0223] Referring to FIG. 9B, audio data (902) and visual data (904) can be input to a multimodal synchronization module (712). The multimodal synchronization module (712) can operate as a learned generative AI model, and thus can generate (906) missing frames of each of the audio data (902) and visual data (904).

[0224] In one embodiment, the missing image frames (e.g., image frames 3, 4, 5, or 8) and the missing audio frames (e.g., audio frame 6 or 7) may not correspond in time. Based on this, the image frames to be included in the positions of the missing image frames in time can be generated based on the corresponding audio frames in time (e.g., audio frame 3, 4, 5, or 8), and the audio frames to be included in the positions of the missing audio frames in time can be generated based on the corresponding image frames in time (e.g., image frame 6 or 7).

[0225] In one embodiment, audio data (902) may be updated to include generated audio frames, and image data (904) may be updated to include generated image frames. The updated audio data and the updated image data may be input to the multimodal ASR module (726).

[0226] According to one embodiment, the multimodal ASR module (726) can perform a speech recognition operation based on updated audio data and updated image data. The multimodal ASR module (726) can output a clear speech recognition result based on the updated audio data and updated image data. For example, the multimodal ASR module (726) can output text information (e.g., "I'm taking the subway now. It's a good train. I'm on it now (910)") as a speech recognition result.

[0227] FIG. 10 is a diagram for explaining the learning operation of a multimodal ASR module according to one embodiment.

[0228] Referring to FIG. 10, the multimodal ASR module (726) may include a conformal encoder (1008) and a decoder (1010).

[0229] According to one embodiment, audio data and visual data may be input to the multimodal ASR module (726) in the form of a sequence. The sequence may be generated by sequentially arranging frames corresponding to a set number or a set time. According to one embodiment, an audio sequence (1002) including audio frames and a visual sequence (1004) including image frames may be input to the multimodal ASR module (726).

[0230] According to one embodiment, the audio sequence (1002) and the visual sequence (1004) may be embedded through a conformer encoder (1008). For example, the encoder (1008) may convert each of the audio sequence (1002) and the visual sequence (1004) into low-dimensional data based on a neural network such as a convolutional neural network (CNN) and / or a residual network (ResNet), and may extract features from the converted data.

[0231] According to one embodiment, the conformer encoder (1008) can perform an audio-visual fusion operation (1006) that fuses (or combines or integrates) features extracted from each of an audio sequence (1002) and a visual sequence (1004). The conformer encoder (1008) can output the features fused by the audio-visual fusion operation (1006) to the decoder (1010).

[0232] According to one embodiment, the decoder (1010) can predict and output text (e.g., "Hello, I'm a library") based on the fused features. According to one embodiment, the decoder (1010) can perform the prediction or output of the text based on at least one token. According to one embodiment, the decoder (1010) can predict or output a start token (e.g., "Hello, I'm a library") that indicates the beginning of a sentence. <s>) can start a prediction or output operation of the text based on the predicted end of the sentence. According to one embodiment, the decoder (101) ends the prediction or operation of the text when the end of the sentence is predicted, and an end token (e.g.,< / s> ) can be generated. For example, the decoder (1010) can generate “ <s> When “Hello, I am a library” is entered, it indicates that the sentence ends with “Hello, I am a library”.< / s> can be printed. <s> and / or< / s>Tokens such as are used to clearly distinguish the boundaries of text or sentences and may not be used as text for users.

[0233] According to one embodiment, in the multimodal ASR module (726), learning for speech recognition may be performed based on fused features output from the conformer encoder (1008), predicted text from the decoder (1010), or correct text.

[0234] FIG. 11 is a flowchart schematically illustrating the operation of an electronic device according to one embodiment.

[0235] According to one embodiment, operations 1102 to 1118 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1 or processor (340) of FIG. 3) of an electronic device (e.g., electronic device (101) of FIG. 1 or electronic device (301) of FIG. 3).

[0236] The operations illustrated in FIG. 11 are not limited to the order illustrated and may be performed in various orders. In one embodiment, at least some of the operations illustrated in FIG. 11 may be omitted, or more operations may be performed than those illustrated in FIG. 11.

[0237] Referring to FIG. 11, in operation 1102, the electronic device may receive an image through at least one camera during a first time period.

[0238] According to one embodiment, in operation 1104, the electronic device may perform a face detection operation based on an input image. According to one embodiment, the electronic device may determine location information of a face in the image based on at least one camera. The at least one camera may include a face detection module for the face detection operation. According to one embodiment, the face detection module may perform face recognition based on a small-size deep learning model or a pattern recognition method. When a face is detected through the face detection module, the electronic device may store location information of the face. According to one embodiment, the location information of the face may include coordinate information (e.g., Pi=(Pi_x, Pi_y, Pi_w, Pi_h)) of an area (e.g., a bounding box) including a face in the image.

[0239] In one embodiment, if a face is not detected in an image, the electronic device may output a notification via the display or speaker indicating that a face is not detected. If a face is not detected for a set period of time, the electronic device may turn off at least one camera.

[0240] According to one embodiment, in operation 1106, the electronic device may perform a lip detection operation based on the face detection. According to one embodiment, the electronic device may estimate lips based on coordinate information stored through the face detection. Once the lips are estimated, the electronic device may perform an image preprocessing operation (e.g., cropping and / or resizing) to obtain an image including the estimated lips.

[0241] In one embodiment, at operation 1108, the electronic device may receive a speaker's voice input through at least one microphone during a first time period. In one embodiment, operation 1108 may be performed in parallel with or substantially at the same time as operation 1102.

[0242] According to one embodiment, the electronic device may obtain image frames associated with a lip image based on operations 1102 to 1106, and may obtain audio frames associated with a speaker's voice based on operation 1108.

[0243] According to one embodiment, in operation 1110, the electronic device can identify a first time corresponding to the image frames and a second time corresponding to the audio frames based on which the image frames and the audio frames were acquired.

[0244] According to one embodiment, in operation 1112, the electronic device may determine whether the first time and the second time are substantially the same. Based on the fact that the first time and the second time are not substantially the same, the electronic device may perform operation 1114, and based on the fact that the first time and the second time are substantially the same, the electronic device may perform operation 1116.

[0245] According to one embodiment, at operation 1114, the electronic device may perform a multimodal restoration and / or synchronization operation based on the first time and the second time being substantially not the same. According to one embodiment, the multimodal restoration and / or synchronization operation may include a restoration operation of at least one missing frame performed by the multimodal synchronization module (712) (e.g., generating at least one frame corresponding to the at least one missing frame, or generating at least one frame to be included at a time axis location of the at least one missing frame), and / or an operation of synchronizing image frames and audio frames with the at least one restored frame.

[0246] According to one embodiment, at operation 1116, the electronic device can perform a multimodal ASR operation based on input image frames and audio frames based on the first time and the second time being substantially the same.

[0247] According to one embodiment, the electronic device can perform a multimodal ASR operation based on image frames and audio frames output by a multimodal restoration and / or synchronization operation based on the first time and the second time being substantially not the same.

[0248] According to one embodiment, at operation 1118, the electronic device may output text based on the multimodal ASR operation.

[0249] According to one embodiment, the electronic device may perform operation 1118 and perform a natural language understanding (NLU) operation or a text-to-speech (TTS) operation. According to one embodiment, the NLU operation may include an operation of interpreting the meaning (or intent) of output text and generating text in a natural context based on the interpreted meaning, and the TTS operation may include an operation of converting text generated by the NLU operation into an audio signal.

[0250] FIG. 12 is a diagram illustrating the results of multimodal speech recognition performed in an electronic device according to one embodiment.

[0251] Referring to FIG. 12, a user can input a facial image and voice (e.g., “Just like clothes get wet in a drizzle, love grows by accumulating small attention and consideration”) using an electronic device (301) such as a smartwatch (e.g., the second type of electronic device (301) of FIG. 5b). The electronic device (301) can obtain visual data (1202) and audio data (1204) based on the input facial image and voice.

[0252] In one embodiment, the visual data (1202) may include at least one missing image frame that makes lip recognition impossible due to the user's movement. The audio data (1204) may not include at least one missing audio frame, but may contain noise, which may cause errors in speech recognition. For example, when a multimodal speech recognition operation based on the visual data (1202) and the audio data (1204) is performed, the electronic device (301) may output text (e.g., "Love is like a clothes expert in a drizzle...") (1208) that includes misrecognized content.

[0253] According to one embodiment, since at least one audio frame corresponding to at least one missing image frame in the audio data (1204) is not missing, at least one image frame can be generated through the generative AI model of the multimodal synchronization module (712). The at least one generated image frame can correspond to the at least one missing frame. The at least one generated image frame can be included in the visual data (120) to generate updated visual data (1206).

[0254] When the updated visual data (1206) is used for a multimodal speech recognition operation, a clearer speech recognition result can be output compared to when the updated visual data (1206) is not used. For example, when a multimodal speech recognition operation is performed based on the updated visual data (1206) and audio data (1204), the electronic device (301) can output a text (e.g., “Love also gets wet in a drizzle…”) (1210) containing content corresponding to the input speech.

[0255] FIG. 13 is a diagram illustrating a UI indicating a voice or image input status provided by an electronic device according to one embodiment.

[0256] Referring to FIG. 13, the electronic device (301) may output information indicating a voice input status on the display based on the number of audio frames acquired based on voice input. The electronic device (301) may output information indicating an image input status on the display based on the number of image frames acquired based on lip image input.

[0257] According to one embodiment, information indicating a voice input state or an image input state may include graphic information such as an equalizer, a bar, or a circle, or graph information indicating the number of frames acquired over time.

[0258] Figures 13 (a) to (c) show examples in which a voice input state or an image input state is indicated as graphic information in the form of an equalizer based on the number of frames.

[0259] Referring to (a) of FIG. 13, the image frame equalizer (1302) may include a first graphic indicator associated with the image frames. The first graphic indicator may have a length based on the number of image frames and may include a graphical representation in the form of a bar having a position based on the input time of each of the image frames.

[0260] The audio frame equalizer (1304) may include a second graphic indicator associated with the audio frames. The second graphic indicator may have a length based on the number of audio frames and may include a graphical representation in the form of a bar having a position based on the input point of each of the audio frames.

[0261] According to one embodiment, based on the first graphic indicator and the second graphic indicator having the same length, it can be identified that the image frames and audio frames are input in substantially the same number.

[0262] According to one embodiment, based on the first graphic indicator and the second graphic indicator indicating input points at the same location, the image frames and the audio frames can be identified as being input at substantially the same point in time.

[0263] According to one embodiment, based on the image frame equalizer (1302) and the audio frame equalizer (1304), it can be identified that the image frames and the audio frames have substantially the same time length and that there is no missing frame in each of the image frames and the audio frames. Based on the identification result, the electronic device (301) can perform a multimodal speech recognition operation based on the image frames and the audio frames.

[0264] Referring to (b) of FIG. 13, the first graphic indicator included in the image frame equalizer (1312) may not indicate consecutive input points. Non-consecutive input points may mean that lip images are not input consecutively, or that at least one image frame is missing.

[0265] The second graphic indicator included in the audio frame equalizer (1314) may indicate consecutive input points. Consecutive input points may mean that there is no missing audio frame at least.

[0266] According to one embodiment, based on the image frame equalizer (1312) and the audio frame equalizer (1314), it can be identified that the image frames and the audio frames have different time lengths and that at least one of the image frames is a missing frame. The electronic device (301) can perform a multimodal restoration and / or synchronization operation based on the identification result, and perform a multimodal speech recognition operation based on the updated image frames and audio frames after performing the operation.

[0267] Referring to (c) of FIG. 13, the first graphic indicator included in the image frame equalizer (1322) may have a shorter length than the second graphic indicator included in the audio frame equalizer (1324). The fact that the first graphic indicator has a shorter length than the second graphic indicator may mean that the number of missing frames in the image frames is greater than that in the audio frames.

[0268] According to one embodiment, based on the image frame equalizer (1322) and the audio frame equalizer (1324), it can be identified that the image frames and the audio frames have different time lengths and that at least one of the image frames is a missing frame. The electronic device (301) can perform a multimodal restoration and / or synchronization operation based on the identification result, and perform a multimodal speech recognition operation based on the updated image frames and audio frames after performing the operation.

[0269] FIG. 14 is a diagram illustrating an operation of an electronic device according to one embodiment of the present invention to output a notification message indicating a failure in voice or lip recognition.

[0270] Referring to (a) of FIG. 14, based on a failure in voice recognition through at least one microphone in the electronic device (301), a notification message may be output via the display or speaker. For example, the electronic device (301) may induce re-performing of voice input by outputting a notification message (1402) on the display, “Recognition failed. Please look at the screen and input again.”

[0271] Referring to (b) of FIG. 14, based on the failure of lip recognition through at least one camera in the electronic device (301), a notification message may be output through the display or speaker. For example, the electronic device (301) may output a notification message through the speaker that prompts the user to take a photo of the lips, such as a notification message (1408) of “Look at the camera and input voice” or a notification message (1410) of “Turn your wrist.”

[0272] FIG. 15 is a diagram illustrating an example of a multimodal voice recognition operation being used in an electronic device according to one embodiment.

[0273] Referring to FIG. 15, the electronic device (301) can utilize multimodal voice recognition operations in the text-to-call function. According to one embodiment, the text-to-call function may include a function for conversing with other users through multimodal voice input. The electronic device (301) can perform voice recognition operations based on voice and lip images input by the user, and output the voice recognition results as text on the display.

[0274] In one embodiment, a voice misrecognition may occur based on noise in the user input, a low voice, or missing lip images due to the user's movements. The electronic device (301) may output a first text generated as a result of the voice misrecognition (e.g., "Sorry, I'm on hold right now, there are a lot of people here") (1502) on the display.

[0275] According to one embodiment, the electronic device (301) may perform sentence correction via the multimodal synchronization module (712) (1504). For example, the electronic device (301) may perform the speech recognition operation again after restoring at least one missing frame in the audio frames and / or image frames that caused the speech misrecognition via the multimodal synchronization module (712). The electronic device (301) may output the second text (e.g., “Sorry, I’m running now, there are too many people here”) (1506) generated as a result of the re-performed speech recognition on the display. The second text (1506) may be indicated as the corrected or updated text of the first text (1502). For example, the electronic device (301) may display information (e.g., “corrected text”) indicating that the second text (1506) is the corrected or updated text of the first text (1502) along with the second text (1506).

[0276] In one embodiment, the first text (1502) and the second text (1506) may be displayed simultaneously, the first text (1502) may be displayed first, and then the second text (1506) may be displayed, or the display of the first text (1502) may be omitted and the second text (1506) may be displayed.

[0277] In one embodiment, sentence correction via the multimodal synchronization module (712) may be performed repeatedly. Based on the occurrence of a voice misrecognition result again, the electronic device (301) may display third text generated as a result of the voice misrecognition result (e.g., "Here is a city around here" (1508).

[0278] According to one embodiment, the electronic device (301) can perform sentence correction via the multimodal synchronization module (712) (1510). For example, the electronic device (301) can perform the speech recognition operation again after restoring at least one missing frame within the audio frames and / or image frames that caused the speech misrecognition via the multimodal synchronization module (712). The electronic device (301) can output the fourth text (e.g., "It's noisy around here") (1512) generated as the result of the re-performed speech recognition to the display.

[0279] FIG. 16 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a voice recognition result.

[0280] According to one embodiment, operations 1602 to 1612 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1 or processor (340) of FIG. 3) of an electronic device (e.g., electronic device (101) of FIG. 1 or electronic device (301) of FIG. 3).

[0281] Referring to FIG. 16, in operation 1602, an electronic device may receive voice and image input. The electronic device may obtain audio frames associated with the input voice. The electronic device may recognize lips from the input image and obtain image frames associated with the recognized lips.

[0282] According to one embodiment, at operation 1604, the electronic device can perform a multimodal ASR operation based on the acquired audio frames and image frames.

[0283] According to one embodiment, at operation 1606, the electronic device may output a first result (e.g., first text (1502) or third text (1508) of FIG. 15) based on the multimodal ASR operation.

[0284] According to one embodiment, at operation 1608, the electronic device may perform a multimodal restoration and / or synchronization operation. According to one embodiment, the multimodal restoration and / or synchronization operation may include a restoration operation of at least one missing frame performed by a multimodal synchronization module (712) (e.g., generating at least one frame corresponding to the at least one missing frame, or generating at least one frame to be included in the position of the at least one missing frame on the time axis), and / or an operation of synchronizing image frames and audio frames with the at least one restored frame. Based on the multimodal restoration and / or synchronization operation, at least one of the image frames or the audio frames may be updated to include the at least one restored frame.

[0285] According to one embodiment, at operation 1610, the electronic device may perform a multimodal ASR operation based on image frames and audio frames output as a result of the multimodal restoration and / or synchronization operation.

[0286] According to one embodiment, at operation 1612, the electronic device may output a second result (e.g., the second text (1506) or the fourth text (1512) of FIG. 15) based on the multimodal ASR operation.

[0287] According to one embodiment, the operations illustrated in FIG. 16 are not limited to the illustrated order and may be performed in various orders. According to one embodiment, at least some of the operations illustrated in FIG. 16 may be omitted, or more operations may be performed than those illustrated in FIG. 16. For example, after operations 1604 and 1606 are performed, operations 1608 to 1612 may be performed, or either operations 1604 and 1606 or operations 1608 to 1612 may not be performed.

[0288] The operation of the electronic device is described below with reference to FIGS. 17 to 26.

[0289] According to one embodiment, the operations illustrated in FIGS. 17 to 26 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1 or processor (340) of FIG. 3) of an electronic device (e.g., electronic device (101) of FIG. 1 or electronic device (301) of FIG. 3).

[0290] The operations illustrated in FIGS. 17 to 26 are not limited to the illustrated order and may be performed in various orders. According to one embodiment, at least some of the operations illustrated in FIGS. 17 to 26 may be omitted, or more operations may be performed than those illustrated in FIGS. 17 to 26.

[0291] FIG. 17 is a flowchart illustrating a voice recognition operation of an electronic device according to one embodiment.

[0292] Referring to FIG. 17, in operation 1702, the electronic device (101) (301) may acquire visual data (904) associated with a speaker's lip image through at least one camera (180) (310) of the electronic device (101) (301) during a first time period.

[0293] According to one embodiment, in operation 1704, the electronic device (101) (301) may obtain audio data (902) associated with the speaker's voice through at least one microphone (150) (320) of the electronic device (101) (301) during a first time period.

[0294] According to one embodiment, in operation 1706, the electronic device (101) (301) can identify at least one missing frame based on the first frames corresponding to the visual data (904) and the second frames corresponding to the audio data (902).

[0295] According to one embodiment, in operation 1708, the electronic device (101) (301) may generate (906) at least one frame corresponding to at least one missing frame (or at least one frame to be included in the position of the at least one missing frame on the time axis) using the first frames and the second frames based on the learned machine learning model (712).

[0296] According to one embodiment, in operation 1710, the electronic device (101) (301) may perform a voice recognition operation based on the first frames, the second frames, or at least one generated frame.

[0297] FIG. 18 is a flowchart illustrating an operation of an electronic device according to one embodiment to identify at least one missing frame.

[0298] According to one embodiment, the operations illustrated in FIG. 18 may be operations associated with operation 1706 of FIG. 17.

[0299] Referring to FIG. 18, in operation 1802, the electronic device (101) (301) can determine whether a first time corresponding to the number of first frames and a second time corresponding to the number of second frames are substantially the same.

[0300] According to one embodiment, in operation 1804, the electronic device (101) (301) may identify at least one missing frame associated with frames corresponding to a shorter time among the first frames or the second frames, based on the first time and the second time being substantially not the same.

[0301] FIG. 19 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to identify at least one missing frame through timeline-based sorting.

[0302] According to one embodiment, the operations illustrated in FIG. 19 may be operations associated with operation 1706 of FIG. 17.

[0303] Referring to FIG. 19, in operation 1902, the electronic device (101) (301) can determine whether a first time corresponding to the number of first frames and a second time corresponding to the number of second frames are substantially the same.

[0304] According to one embodiment, in operation 1904, the electronic device (101) (301) may align the first frames and the second frames based on a timeline corresponding to the first time interval, based on the first time and the second time being not substantially the same.

[0305] According to one embodiment, in operation 1906, the electronic device (101) (301) may perform synchronization between the first frames and the second frames based on alignment.

[0306] According to one embodiment, at operation 1908, the electronic device (101) (301) may identify at least one frame that is out of synchronization among the first frames or the second frames.

[0307] According to one embodiment, at operation 1910, the electronic device (101) (301) can identify at least one missing frame based on at least one out-of-synchronization frame.

[0308] FIG. 20 is a flowchart illustrating an operation of an electronic device according to one embodiment to restore a missing frame.

[0309] According to one embodiment, the operations illustrated in FIG. 20 may be operations associated with operation 1708 of FIG. 17.

[0310] Referring to FIG. 20, in operation 2002, the electronic device (101) (301) can determine whether a first time corresponding to the number of first frames and a second time corresponding to the number of second frames are substantially the same.

[0311] According to one embodiment, in operation 2004, the electronic device (101) (301) may extract a first feature from source frames (710) corresponding to a longer time among the first frames or the second frames, based on the first time and the second time being not substantially the same.

[0312] According to one embodiment, in operation 2006, the electronic device (101) (301) may extract a second feature from target frames (708) corresponding to a shorter time among the first frames or the second frames.

[0313] According to one embodiment, in operation 2008, the electronic device (101) (301) may fuse the first feature and the second feature using a machine learning model and generate at least one frame based on the fused feature. According to one embodiment, the at least one frame may correspond to at least one missing frame or may be included in the position of at least one missing frame on the time axis.

[0314] FIG. 21 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a message based on the number of at least one missing frame.

[0315] Referring to FIG. 21, in operation 2102, the electronic device (101) (301) can determine whether the number of at least one missing frame is greater than or equal to a threshold value.

[0316] According to one embodiment, in operation 2104, the electronic device (101) (301) may output a message (1402) (1408) (1410) instructing to re-capture the speaker's lips or re-input the speaker's voice through the display (160) or speaker (155) of the electronic device (101) (301) based on the number of at least one missing frame being greater than or equal to a threshold value.

[0317] FIG. 22 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a result of a voice recognition operation.

[0318] According to one embodiment, the operations illustrated in FIG. 22 may be operations associated with operation 1710 of FIG. 17.

[0319] Referring to FIG. 22, in operation 2202, the electronic device (101) (301) may generate a first result (1502) (1508) of performing a voice recognition operation based on the first frames or the second frames.

[0320] According to one embodiment, in operation 2204, the electronic device (101) (301) may generate a second result (1506) (1512) of performing a voice recognition operation based on the first frames, the second frames, or at least one generated frame as an updated result of the first result (1502) (1508).

[0321] According to one embodiment, in operation 2206, the electronic device (101) (301) may output the first result (1502) (1508) and the second result (1506) (1512) at the same or different times through the display (160) or speaker (155) of the electronic device (101) (301).

[0322] FIG. 23 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a voice recognition result through a server.

[0323] According to one embodiment, the operations illustrated in FIG. 23 may be operations associated with operation 1710 of FIG. 17.

[0324] Referring to FIG. 23, in operation 2302, the electronic device (101) (301) may transmit first information including first frames, second frames, or at least one generated frame to the server.

[0325] According to one embodiment, in operation 2304, the electronic device (101) (301) may, in response to transmitting the first information, receive second information including a voice recognition result from the server.

[0326] According to one embodiment, in operation 2306, the electronic device (101) (301) may output a voice recognition result through the display (160) or speaker (155) of the electronic device (101) (301) based on the received second information.

[0327] FIG. 24 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a voice recognition result through an on-device operation.

[0328] According to one embodiment, the operations illustrated in FIG. 24 may be operations associated with operation 1710 of FIG. 17.

[0329] Referring to FIG. 24, in operation 2402, the electronic device (101) (301) may perform a voice recognition operation using the first frames, the second frames, or at least one generated frame based on the learned voice recognition model (726).

[0330] According to one embodiment, in operation 2404, the electronic device (101) (301) can obtain a voice recognition result based on a voice recognition operation.

[0331] According to one embodiment, in operation 2406, the electronic device (101) (301) can output the acquired voice recognition result through the display (160) or speaker (155) of the electronic device (101) (301).

[0332] FIG. 25 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present invention to output a frame input status.

[0333] Referring to FIG. 25, in operation 2502, the electronic device (101) (301) may output information (1302) (1312) (1322) indicating a frame input status based on the number of first frames, based on the visual data (904) obtained, to the display (160) of the electronic device (101) (301).

[0334] According to one embodiment, in operation 2504, the electronic device (101) (301) may output information (1304) (1314) (1324) indicating a frame input status based on the number of second frames, to the display (160), based on the audio data (902) obtained.

[0335] FIG. 26 is a flowchart illustrating an operation of an electronic device receiving a voice recognition result from a server according to one embodiment.

[0336] Referring to FIG. 26, in operation 2602, the electronic device (101) (301) may acquire visual data (904) associated with the speaker's lips through at least one camera (180) (310) of the electronic device (101) (301) during a first time period.

[0337] According to one embodiment, in operation 2604, the electronic device (101) (301) may obtain audio data (902) associated with the speaker's voice through at least one microphone (150) (320) of the electronic device (101) (301) during a first time period.

[0338] According to one embodiment, in operation 2606, the electronic device (101) (301) may transmit first frames corresponding to the acquired visual data (904) and second frames corresponding to the acquired audio data (902) to the server.

[0339] According to one embodiment, in operation 2608, the electronic device (101) (301) may receive a voice recognition result from the server. According to one embodiment, the voice recognition result may include a result of identifying at least one missing frame based on the first frames and the second frames, generating (906) at least one frame corresponding to the at least one missing frame (or at least one frame to be included in the position of the at least one missing frame on the time axis) using the first frames and the second frames based on the learned machine learning model (712), and performing a voice recognition operation based on the first frames, the second frames, or the at least one generated frame.

[0340] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0341] According to one embodiment, a voice recognition method of an electronic device (101) (301) comprises: an operation (1702) of acquiring visual data (904) associated with a lip image of a speaker through at least one camera (180) (310) of the electronic device (101) (301) during a first time period; an operation (1704) of acquiring audio data (902) associated with a voice of the speaker through at least one microphone (150) (320) of the electronic device (101) (301) during the first time period; an operation (1706) of identifying at least one missing frame based on first frames corresponding to the visual data (904) and second frames corresponding to the audio data (902); an operation (1708) of generating at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712) (906); And it may be a voice recognition method including an operation (1710) of performing voice recognition based on the first frames, the second frames, or the at least one generated frame.

[0342] According to one embodiment, the operation (1706) of identifying at least one missing frame comprises:

[0343] The method may include an operation (1802) of determining whether a first time corresponding to the first frames and a second time corresponding to the second frames are the same; and / or an operation (1804) of identifying at least one missing frame associated with frames corresponding to a shorter time among the first frames or the second frames, based on the first time and the second time not being the same.

[0344] According to one embodiment, the operation of identifying at least one missing frame (1706) may be a speech recognition method including: an operation of determining whether a first time corresponding to the first frames and a second time corresponding to the second frames are the same (1902); and an operation of aligning the first frames and the second frames based on a time line corresponding to the first time period based on the first time and the second time not being the same (1904); an operation of performing synchronization between the first frames and the second frames based on the alignment (1906); an operation of identifying at least one unsynchronized frame among the first frames or the second frames (1908); and / or an operation of identifying the at least one missing frame based on the at least one unsynchronized frame (1910).

[0345] According to one embodiment, the operation (1708) of generating (906) at least one frame may be a speech recognition method including: extracting (2004) a first feature from source frames (710) corresponding to a longer time among the first frames or the second frames, based on the first time and the second time being not the same; extracting (2006) a second feature from target frames (708) corresponding to a shorter time among the first frames or the second frames; and / or fusing the first feature and the second feature using the machine learning model and generating the at least one frame based on the fused feature (2008).

[0346] According to one embodiment, the voice recognition method may further include an operation (2104) of outputting a message (1402) (1408) (1410) instructing to re-photograph the speaker's lips or re-input the speaker's voice through a display (160) or a speaker (155) of the electronic device (101) (301) based on the number of the at least one missing frame being greater than or equal to a threshold value.

[0347] According to one embodiment, the operation (1710) of performing a voice recognition operation may be a voice recognition method including the operation (2202) of generating a first result (1502) (1508) of performing the voice recognition operation based on the first frames or the second frames; the operation (2204) of generating a second result (1506) (1512) of performing the voice recognition operation based on the first frames, the second frames, or the at least one generated frame as an updated result of the first result (1502) (1508); and / or the operation (2206) of outputting the first result (1502) (1508) and the second result (1506) (1512) through a display (160) or a speaker (155) of an electronic device (101) (301) at the same or different times.

[0348] According to one embodiment, the operation of performing a voice recognition operation may be a voice recognition method including an operation (2302) of transmitting first information including first frames, second frames, or at least one generated frame to a server; an operation (2304) of receiving second information including a voice recognition result from the server in response to transmitting the first information; and an operation (2306) of outputting the voice recognition result through a display (160) or a speaker (155) of an electronic device (101) (301) based on the received second information.

[0349] According to one embodiment, the operation of performing a voice recognition operation may be a voice recognition method including an operation (2402) of performing a voice recognition operation using first frames, second frames, or at least one generated frame based on a learned voice recognition model (726); an operation (2404) of obtaining a voice recognition result based on the voice recognition operation; and an operation (2406) of outputting the obtained voice recognition result through a display (160) or a speaker (155) of an electronic device (101) (301).

[0350] A method of an electronic device according to one embodiment of the present disclosure may further include: an operation (2502) of outputting information (1302) (1312) (1322) indicating a frame input state based on the number of first frames to a display (160) of an electronic device (101) (301) based on the acquisition of visual data (904); and an operation (2504) of outputting information (1304) (1314) (1324) indicating a frame input state based on the number of second frames to the display (160) based on the acquisition of audio data (902).

[0351] A method of an electronic device according to one embodiment of the present disclosure comprises: an operation (2602) of acquiring visual data (904) associated with lips of a speaker through at least one camera (180) (310) of an electronic device (101) (301) during a first time period; an operation (2604) of acquiring audio data (902) associated with voice of the speaker through at least one microphone (150) (320) of the electronic device (101) (301) during a first time period; an operation (2606) of transmitting first frames corresponding to the acquired visual data (904) and second frames corresponding to the acquired audio data (902) to a server; And may include an operation (2608) of receiving a voice recognition result from the server, wherein the voice recognition result may include a result of identifying at least one missing frame based on the first frames and the second frames, generating at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on the learned machine learning model (712) (906), and performing a voice recognition operation based on the first frames, the second frames, or the at least one generated frame.

[0352] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0353] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0354] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0355] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0356] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separately arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices (101,301), At least one camera (180,310); At least one microphone (150,320); memory (130,330); and At least one processor (120,340) comprising a processing circuit, The above memory (130, 330), when executed individually or collectively by the at least one processor (120, 340), causes the electronic device (101, 301) to: During the first time period, visual data (904) associated with the speaker's lip image is acquired through at least one camera (180,310), Acquire audio data (902) associated with the speaker's voice through at least one microphone (150,320) during the first time period, Obtain first frames corresponding to the above visual data (904) and second frames corresponding to the above audio data (902), Based on a comparison of a first time corresponding to the first frames and a second time corresponding to the second frames, identifying at least one missing frame among the first frames or the second frames, An electronic device that generates (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712), and stores instructions that cause a voice recognition operation to be performed based on the first frames, the second frames, or the at least one generated frame.

2. In the first paragraph, the memory (130, 330), when individually or collectively executed by the at least one processor (120, 340), causes the electronic device (101, 301) to: Determine whether the first time corresponding to the first frames and the second time corresponding to the second frames are the same, An electronic device storing instructions that cause the at least one missing frame to be identified based on the first time and the second time not being the same, wherein the at least one missing frame is associated with frames corresponding to a shorter time among the first frames or the second frames.

3. In any one of the first or second paragraphs, the memory (130, 330), when individually or collectively executed by the at least one processor (120, 340), causes the electronic device (101, 301) to: Determine whether the first time corresponding to the first frames and the second time corresponding to the second frames are the same, Based on the fact that the first time and the second time are not the same, the first frames and the second frames are aligned based on a time line corresponding to the first time interval, Perform synchronization between the first frames and the second frames based on the above alignment, Identify at least one frame that is out of sync among the first frames or the second frames, An electronic device storing instructions that cause the at least one missing frame to be identified based on the at least one out-of-synchronization frame.

4. In any one of the second or third paragraphs, the memory (130, 330), when individually or collectively executed by the at least one processor (120, 340), causes the electronic device (101, 301) to: Based on the fact that the first time and the second time are not the same, a first feature is extracted from source frames (710) corresponding to a longer time among the first frames or the second frames, Extracting a second feature from target frames (708) corresponding to a shorter time among the first frames or the second frames, An electronic device storing commands that cause the first feature and the second feature to be fused using the machine learning model and to generate the at least one frame based on the fused feature.

5. In any one of paragraphs 1 to 4, display (160); and Includes additional speakers (155), The above memory (130, 330), when executed individually or as a whole by the at least one processor (120, 340), causes the electronic device (101, 301) to: An electronic device storing commands that cause a message (1402, 1408, 1410) to be output through at least one of the display (160) or the speaker (155) instructing to re-photograph the speaker's lips or re-input the speaker's voice based on the number of at least one missing frame being greater than or equal to a threshold value.

6. In any one of paragraphs 1 to 5, display (160); and Includes additional speakers (155), The above memory (130, 330), when executed individually or as a whole by the at least one processor (120, 340), causes the electronic device (101, 301) to: Generating a first result (1502, 1508) of performing the speech recognition operation based on at least one of the first frames or the second frames, Based on the first frames, the second frames, or the at least one generated frame, a second result (1506, 1512) of performing the speech recognition operation is generated as an updated result of the first result (1502, 1508), An electronic device that stores commands that cause the first result (1502, 1508) and the second result (1506, 1512) to be output through the display (160) or the speaker (155) at the same or different times.

7. In any one of paragraphs 1 to 6, Communication circuit (190); display (160); and Includes additional speakers (155), The above memory (130, 330), when executed individually or as a whole by the at least one processor (120, 340), causes the electronic device (101, 301) to: Control the communication circuit (190) to transmit first information including the first frames, the second frames, or the at least one generated frame to the server, In response to transmitting the first information, second information including a voice recognition result is received from the server through the communication circuit (190), An electronic device storing commands that cause the voice recognition result to be output through at least one of the display (160) or the speaker (155) based on the received second information.

8. In any one of paragraphs 1 to 6, display (160); and Includes additional speakers (155), The above memory (130) (330), when executed individually or as a whole by the at least one processor (120, 340), causes the electronic device (101, 301) to: Based on the learned speech recognition model (726), the speech recognition operation is performed using the first frames, the second frames, or the at least one generated frame, Obtaining a voice recognition result based on the above voice recognition operation, An electronic device that stores commands that cause the acquired voice recognition result to be output through at least one of the display (160) or the speaker (155).

9. In any one of paragraphs 1 to 8, It further includes a display (160), The above memory (130, 330), when executed individually or as a whole by the at least one processor (120, 340), causes the electronic device (101, 301) to: Based on the acquisition of the above visual data (904), information (1302) (1312, 1322) indicating a frame input status based on the number of the first frames is output to the display (160), An electronic device storing commands that cause information (1304, 1314, 1324) indicating a frame input status based on the number of second frames to be output to the display (160) based on the acquisition of the above audio data (902).

10. In the electronic device (101)(301), At least one camera (180,310); At least one microphone (150,320); Communication circuit (190); memory (130)(330); and At least one processor (120,340) comprising a processing circuit, The above memory (130, 330), when executed individually or collectively by the at least one processor (120, 340), causes the electronic device (101, 301) to: During the first time period, visual data (904) related to the speaker's lips are acquired through at least one camera (180,310), Acquire audio data (902) associated with the speaker's voice through at least one microphone (150,320) during the first time period, Control the communication circuit (190) to transmit first frames corresponding to the acquired visual data (904) and second frames corresponding to the acquired audio data (902) to the server, Store instructions that cause the voice recognition result to be received from the server through the above communication circuit (190), The above voice recognition results are: An electronic device comprising a result of identifying at least one missing frame based on the first frames and the second frames, generating (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on a learned machine learning model (712), and performing a voice recognition operation based on the first frames, the second frames, or the at least one generated frame.

11. In the method of electronic device (101,301), An operation (1702) of acquiring visual data (904) associated with a speaker's lip image through at least one camera (180, 310) of the electronic device (101, 301) during a first time period; An operation (1704) of obtaining audio data (902) associated with the speaker's voice through at least one microphone (150, 320) of the electronic device (101, 301) during the first time period; An operation of obtaining first frames corresponding to the above visual data (904) and second frames corresponding to the above audio data (902); An operation (1706) of identifying a missing frame of at least one of the first frames or the second frames based on a comparison of a first time corresponding to the first frames and a second time corresponding to the second frames; An operation (1708) of generating (906) at least one frame corresponding to the at least one missing frame using the first frames and the second frames based on the learned machine learning model (712); and A method comprising an operation (1710) of performing speech recognition based on the first frames, the second frames, or the at least one generated frame.

12. In paragraph 11, The operation (1706) of identifying at least one missing frame is: An operation (1802) for determining whether the first time corresponding to the first frames and the second time corresponding to the second frames are the same; and A method comprising the operation (1804) of identifying at least one missing frame, wherein the at least one missing frame is associated with frames corresponding to a shorter time among the first frames or the second frames, based on the first time and the second time not being the same.

13. In either of paragraphs 11 or 12, The operation (1706) of identifying at least one missing frame is: An operation (1902) for determining whether the first time corresponding to the first frames and the second time corresponding to the second frames are the same; and An operation (1904) of aligning the first frames and the second frames based on a time line corresponding to the first time interval, based on the first time and the second time not being the same; An operation (1906) of performing synchronization between the first frames and the second frames based on the above alignment; An operation (1908) of identifying at least one frame that is out of synchronization among the first frames or the second frames; and A method comprising the operation (1910) of identifying at least one missing frame based on at least one out-of-synchronization frame.

14. In either of paragraphs 12 or 13, The operation (1708) of generating (906) at least one frame is: An operation (2004) of extracting a first feature from source frames (710) corresponding to a longer time among the first frames or the second frames, based on the first time and the second time being not the same; An operation (2006) of extracting a second feature from target frames (708) corresponding to a shorter time among the first frames or the second frames; and A method comprising an operation (2008) of fuse the first feature and the second feature using the machine learning model and generating the at least one frame based on the fused feature.

15. In any one of paragraphs 11 to 14, A method further comprising the operation (2104) of outputting a message (1402, 1408, 1410) instructing to re-photograph the speaker's lips or re-input the speaker's voice through the display (160) or speaker (155) of the electronic device (101, 301) based on the number of at least one missing frame being greater than or equal to a threshold value.

Citation Information

Patent Citations

  • Mobile terminal and method for controlling the terminal

    KR1020150041279A

  • Semiconductor memory device

    KR1020250146747A

  • Gas Sensor with Structurally Improved Sensing Accuracy and Its Manufacturing Method

    KR102759567B1

  • Media insertion system

    US20200236420A1

  • KR20230015235A