Information processing system and information processing method

The information processing system addresses the high equipment and cost challenges of LLMs by integrating video and audio processing to enhance user interaction and safety through location-based acoustic information generation and hazard detection.

WO2026104955A1PCT designated stage Publication Date: 2026-05-21SEMICON ENERGY LAB CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SEMICON ENERGY LAB CO LTD
Filing Date
2025-11-07
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

The challenge of building and operating large language models (LLMs) like GPT-4 is the high equipment and cost requirements, making it difficult for users to utilize them effectively.

Method used

An information processing system comprising components that create and process video and audio information using object detection, speech recognition, and speech generation models to generate and modify acoustic information based on location and attribute information, enhancing user interaction and safety.

Benefits of technology

The system provides enhanced convenience, reliability, and usefulness by improving user interaction with nearby speakers and detecting potential dangers without time lags, and alerting to distant hazards through notification sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025061369_21052026_PF_FP_ABST
    Figure IB2025061369_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a novel information processing system that has excellent convenience, usefulness, and reliability. The information processing system is composed of three components. The first component has a function of creating video information and first acoustic information and transmitting said information to the third component. The second component has an object detection model, a voice recognition model, etc., generates various types of information from the received information, and transmits the various types of information to the third component. The third component creates second acoustic information from the received information and transmits the second acoustic information to the first component. By this system, various kinds of information can be provided to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method

[0001] One aspect of the present invention relates to an information processing system, an information processing method, or a semiconductor device.

[0002] Note that one aspect of the present invention is not limited to the above technical field. The technical field of one aspect of the invention disclosed in this specification and the like relates to an object, a method, or a manufacturing method. Or, one aspect of the present invention relates to a process, a machine, a manufacture, or a composition of matter. Therefore, more specifically, examples of the technical field of one aspect of the invention disclosed in this specification include an information processing device, a semiconductor device, a storage device, a driving method thereof, or a manufacturing method thereof.

[0003] In recent years, the development of language models using neural networks has been actively carried out, and in particular, large language models (LLMs) have attracted attention. A large language model is a natural language processing model learned using a large amount of data. With a large language model, for example, a dialogue model that answers user instructions can be realized. In Non-Patent Document 1, GPT-4 (Generative Pre-trained Transformer 4) (registered trademark) is disclosed as a large language model, and ChatGPT is disclosed as a dialogue model.

[0004] By using a large language model, the capabilities of natural language processing models have been greatly improved. On the other hand, due to the enlargement of language models, it is difficult for users to build and operate language models themselves in terms of equipment and costs. Therefore, using an external service that provides a language model has become one form of using a language model.

[0005] Summary of ChatGPT / GPT-4 Research and Perspective Towards the Future of Large Language Models, Yiheng Liu et al. (Submitted on 4 Apr 2023, [online], Internet <URL: https: / / arxiv.org / abs / 2304.01852>

[0006] One aspect of the present invention aims to provide a novel information processing system that is superior in convenience, usefulness, or reliability. Alternatively, it aims to provide a novel information processing method that is superior in convenience, usefulness, or reliability. Alternatively, it aims to provide a novel information processing system, a novel information processing method, or a novel semiconductor device.

[0007] Furthermore, the description of these problems does not preclude the existence of other problems. Moreover, one aspect of the present invention does not need to solve all of these problems. Other problems will naturally become apparent from the description in the specification, drawings, and claims, and it is possible to extract other problems from the description in the specification, drawings, and claims.

[0008] (1) One aspect of the present invention is an information processing system having a first component, a second component, and a third component.

[0009] The first component includes a function for creating video information and first audio information and transmitting them to the third component, and a function for receiving and providing second audio information.

[0010] The second component has the function of receiving video information and first audio information and transmitting sound source information, voice information, and synthesized voice to the third component. The second component also has the function of processing using an object detection model, a speech recognition model, and a speech generation model.

[0011] The object detection model has the function of generating sound source information from video information. This sound source information includes location information and attribute information.

[0012] The speech recognition model has the function of separating speech information from the first acoustic information. The speech generation model also has the function of generating synthesized speech to supplement the speech information.

[0013] The third component has the function of receiving video information, first audio information, voice information, synthesized voice, and sound source information, and the function of transmitting the video information and the first audio information to the second component. The third component also has the function of creating second audio information from the first audio information based on the sound source information and transmitting it to the first component.

[0014] This makes it possible to determine a means for creating second acoustic information based on location information or attribute information. Furthermore, the second acoustic information created by this means can be provided, for example, to a user. As a result, a novel information processing system with superior convenience, usefulness, or reliability can be provided.

[0015] (2) Another aspect of the present invention is the above-described information processing system, wherein a third component has a function to modify the first acoustic information and use it for the second acoustic information when the location information is located within a predetermined range and the attribute information is classified as a conversation partner or a danger.

[0016] This allows, for example, the volume of the voice of a person speaking nearby to be increased. It also allows the user to become aware of potential dangers approaching. Furthermore, the user can converse without experiencing a time lag between their speech and the other person's. Finally, the user can become aware of potential dangers approaching them without experiencing a time lag. As a result, a novel information processing system with superior convenience, usefulness, and reliability can be provided.

[0017] (3) Another aspect of the present invention is the above-described information processing system, in which a third component has the function of using speech information for second acoustic information when the location information is located outside a predetermined range and the attribute information is classified as a speaker.

[0018] Furthermore, when the audio information includes unclear or missing parts, the third component uses synthesized speech for the second audio information.

[0019] This allows, for example, the separation of a speaker's utterance, even if they are located far from the user, from ambient noise and presenting it to the user. Furthermore, a speech recognition model can be used to accurately separate speech information from the first acoustic information. Additionally, even if there is a slight time lag between the speaker's utterance and the second acoustic information when the speaker is located far from the user, it is less likely to be perceived as unnatural. As a result, a novel information processing system with superior convenience, usefulness, and reliability can be provided.

[0020] (4) Another aspect of the present invention is the above-described information processing system, wherein the third component has a function to create a notification sound based on attribute information.

[0021] Furthermore, when the location information is outside a predetermined range and the attribute information is classified as dangerous, the third component uses a notification sound as the second acoustic information.

[0022] This allows, for example, the system to alert users to potential hazards located far from them using a notification sound. As a result, it is possible to provide a novel information processing system that is superior in terms of convenience, usefulness, and reliability.

[0023] (5) One aspect of the present invention is an information processing method comprising the first to fifth steps.

[0024] In the first step, the first component creates video information and first audio information and transmits them to the second component.

[0025] In the second step, the second component receives video information and first audio information and transmits them to the third component.

[0026] In the third step, the third component receives video information and first audio information and generates sound source information using an object detection model.

[0027] In the fourth step, the third component transmits the sound source information to the second component.

[0028] In the fifth step, the second component creates second acoustic information from the first acoustic information based on the sound source information and transmits it to the first component. The sound source information includes location information and attribute information.

[0029] This makes it possible to determine a means for creating second acoustic information based on location information or attribute information. Furthermore, the second acoustic information created by this means can be provided, for example, to a user. As a result, a novel information processing method with superior convenience, usefulness, or reliability can be provided.

[0030] (6) Another aspect of the present invention is the above-described information processing method, wherein in the fifth step, when the location information is located within a predetermined range and the attribute information is classified as a conversation partner or a danger, the second component modifies the first acoustic information and uses it for the second acoustic information.

[0031] This allows, for example, the volume of the voice of a person speaking nearby to be increased. It also allows the user to become aware of potential dangers approaching them. Furthermore, the user can converse without experiencing a time lag between their speech and the other person's. Finally, the user can become aware of potential dangers approaching them without experiencing a time lag. As a result, a novel information processing method with superior convenience, usefulness, and reliability can be provided.

[0032] (7) Another aspect of the present invention is the above-described information processing method, wherein in the fifth step, when the location information is located outside a predetermined range and the attribute information is classified as a speaker, the second component uses a speech recognition model to separate speech information from the first acoustic information and use it for the second acoustic information.

[0033] (8) Another aspect of the present invention is the above-described information processing method, wherein in the fifth step, when the audio information includes unclear or missing parts, the second component uses a speech generation model to generate synthesized speech from the audio information in which the unclear or missing parts have been filled in, and uses this synthesized speech for the second acoustic information.

[0034] This allows, for example, the separation of a speaker's utterance, even if they are located far from the user, from ambient noise and presenting it to the user. Furthermore, a speech recognition model can be used to accurately separate speech information from the first acoustic information. Additionally, even if there is a slight time lag between the speaker's utterance and the second acoustic information when the speaker is located far from the user, it is less likely to be perceived as unnatural. As a result, a novel information processing method with superior convenience, usefulness, and reliability can be provided.

[0035] (9) Another aspect of the present invention is the above-described information processing method, wherein in the fifth step, when the location information is located outside a predetermined range and the attribute information is classified as dangerous, the second component creates a notification sound based on the attribute information and uses it for the second acoustic information.

[0036] This allows, for example, the system to alert users to potential hazards located far from them using a notification sound. As a result, it is possible to provide a novel information processing method that is superior in terms of convenience, usefulness, and reliability.

[0037] One aspect of the present invention can provide a novel information processing system that is superior in convenience, usefulness, or reliability. Alternatively, it can provide a novel information processing method that is superior in convenience, usefulness, or reliability. Alternatively, it can provide a novel information processing system, a novel information processing method, or a novel semiconductor device.

[0038] Furthermore, the description of these effects does not preclude the existence of other effects. Moreover, one aspect of the present invention does not necessarily have to possess all of these effects. Other effects will naturally become apparent from the description in the specification, drawings, and claims, and it is possible to extract other effects from the description in the specification, drawings, and claims.

[0039] Figure 1 is a diagram illustrating the configuration of an information processing system according to an embodiment. Figure 2A is a front view of an information processing device that can be used in the information processing system according to an embodiment, Figure 2B is a side view of an information processing device that can be used in the information processing system according to an embodiment, and Figure 2C is a top view of an information processing device that can be used in the information processing system according to an embodiment. Figure 3 is a diagram illustrating the configuration of an information processing system according to an embodiment. Figures 4A and 4B are diagrams illustrating the configuration of components used in the information processing system according to an embodiment. Figure 5 is a diagram illustrating the configuration of an information processing device used in the information processing system according to an embodiment. Figure 6 is a diagram illustrating an information processing method according to an embodiment. Figure 7 is a diagram illustrating an information processing method according to an embodiment.

[0040] An information processing system according to one aspect of the present invention comprises a first component, a second component, and a third component. The first component has the function of creating video information and first audio information and transmitting them to the third component, and the function of receiving and providing second audio information. The second component has the function of receiving video information and first audio information and transmitting sound source information, voice information, and synthesized voice to the third component. The second component also has the function of performing processing using an object detection model, a speech recognition model, and a speech generation model. The object detection model has the function of generating sound source information from video information. The sound source information includes location information and attribute information. The speech recognition model has the function of separating voice information from first audio information. The speech generation model also has the function of generating synthesized voice that supplements the voice information. The third component has the function of receiving video information, first audio information, voice information, synthesized voice, and sound source information, and the function of transmitting video information and first audio information to the second component. The third component has the function of creating second acoustic information from the first acoustic information based on the sound source information and sending it to the first component.

[0041] This makes it possible to determine a means for creating second acoustic information based on location information or attribute information. Furthermore, the second acoustic information created by this means can be provided, for example, to a user. As a result, a novel information processing system with superior convenience, usefulness, or reliability can be provided.

[0042] Embodiments will be described in detail with reference to the drawings. However, it will be readily apparent to those skilled in the art that the present invention is not limited to the following description, and that its form and details can be modified in various ways without departing from the spirit and scope of the present invention. Accordingly, the present invention is not to be interpreted as being limited to the contents of the embodiments shown below. In the configuration of the invention described below, the same reference numerals are used in common across different drawings for the same parts or parts having similar functions, and repeated descriptions are omitted.

[0043] In this specification and the like, ordinal numbers such as "first", "second", etc. are attached to avoid confusion of components, and do not limit the number of components or the order of components (for example, the process order or the stacking order). Also, even for terms without ordinal numbers in this specification and the like, ordinal numbers may be attached in the claims to avoid confusion of components. Even for terms with ordinal numbers in this specification and the like, different ordinal numbers may be attached in the claims. Even for terms with ordinal numbers in this specification and the like, ordinal numbers may be omitted in the claims.

[0044] In the drawings attached to this specification, components are classified by function and shown as independent blocks in a block diagram. However, it is difficult to completely separate actual components by function, and one component may be related to multiple functions.

[0045] (Embodiment 1) In this embodiment, an information processing system according to an aspect of the present invention will be described with reference to FIGS. 1 to 5.

[0046] FIG. 1 is an external view for explaining the configuration of an information processing system according to an aspect of the present invention.

[0047] FIG. 2A is a front view of an information processing apparatus that can be used in the information processing system according to the embodiment, FIG. 2B is a side view of the information processing apparatus shown in FIG. 2A, and FIG. 2C is a top view of the information processing apparatus shown in FIG. 2A.

[0048] FIG. 3 is a block diagram for explaining the configuration of an information processing system according to an aspect of the present invention.

[0049] FIG. 4A is a diagram for explaining the configuration of components used in an information processing system according to an aspect of the present invention, and FIG. 4B is a diagram for explaining the configuration of components different from those in FIG. 4A.

[0050] FIG. 5 is a diagram for explaining the configuration of an information processing apparatus that can be used in an information processing system according to an aspect of the present invention.

[0051] <Example of Information Processing System Configuration 1> The information processing system described in this embodiment has component 110, component 130, and component 120 (see Figure 3).

[0052] <Component 110> Component 110 has the function of collecting sound and video and creating video information VdI and sound information AcI_1. Component 110 also has the function of transmitting video information VdI and sound information AcI_1 to component 120. Component 110 also has the function of receiving sound information AcI_2 and providing it, for example, to a user 99 of the information processing system.

[0053] An example of an information processing device that can be used in an information processing system according to one embodiment of the present invention is shown in Figures 2A to 2C. The information processing device 700 has an information processing device 700L and an information processing device 700R (see Figure 2A). A user of the information processing device 700 can, for example, wear the information processing device 700L on their left ear. They can also wear the information processing device 700R on their right ear. They can also wear the information processing devices 700 on both ears. Furthermore, each of the information processing devices 700R and 700L is equipped with an acoustic module, a camera module, and a speaker module. The acoustic module, camera module, and speaker module perform the functions provided by component 110.

[0054] The information processing device 700L includes an acoustic module 110Mc, a camera module 110Cm, and a speaker module 110Sp (see Figure 2B). The acoustic module 110Mc of the information processing device 700L collects sound from the left side of the user of the information processing system, and the camera module 110Cm of the information processing device 700L collects video from the left side of the user. The information processing device 700R has the same configuration as the information processing device 700L, and the acoustic module 110Mc of the information processing device 700R collects sound from the right side of the user of the information processing system, and the camera module 110Cm of the information processing device 700R collects video from the right side of the user.

[0055] The camera module 110Cm can, for example, collect video footage of the surroundings of a user of an information processing system and create video information VdI. Preferably, the camera module 110Cm is equipped with an ultra-wide-angle lens or a fisheye lens. Specifically, a camera with a field of view exceeding 180 degrees can be used. Alternatively, multiple camera modules can be used, for example, a camera module for shooting forward, a camera module for shooting to the side, and a camera module for shooting behind. This allows, for example, capturing a panoramic view of the environment where the user of the information processing system is located. It also allows detecting dangers approaching from behind.

[0056] The acoustic module 110Mc can, for example, collect sounds from the surroundings of a user of an information processing system and create acoustic information AcI_1.

[0057] The speaker module 110Sp can reproduce the acoustic information AcI_2 and provide sound to, for example, a user of an information processing system.

[0058] Furthermore, each of the information processing devices 700R and 700L is equipped with a system board that performs the functions of component 120.

[0059] <Component 130> Component 130 has the function of receiving video information VdI and sound information AcI_1 and transmitting sound source information SSI, voice information VoI, and synthesized voice SV to component 120 (see Figure 4A). Component 130 also has the function of performing processing using an object detection model ODM, a speech recognition model VRM, and a speech generation model VGM. For example, a wristwatch-type information processing device 130a, a portable information processing device 130b, a server 130c, etc., can be used as component 130 (see Figure 1).

[0060] 《Object Detection Model ODM》 The object detection model ODM has the function of generating sound source information SSI from video information VdI. The sound source information SSI includes the location information PI and the attribute information Attr of the sound source, which are included in the acoustic information AcI_1.

[0061] For example, video footage captured using the camera module 110Cm can be used as video information VdI. Specifically, video footage of the surroundings of the user of the information processing system can be used as video information VdI. Furthermore, the object detection model ODM can use the video information VdI to determine the distance from the camera module 110Cm to the sound source and the direction of the sound source, and use this information as position information PI. In addition, attribute information Attr can be used to indicate that the sound source is, for example, a person speaking, a speaker at a lecture or a reader at a reading event, an animal, a car, or a natural phenomenon such as thunder.

[0062] 《VRM Speech Recognition Model》 The VRM speech recognition model has the function of separating speech information VoI from acoustic information AcI_1.

[0063] Furthermore, the VRM speech recognition model can separate human voices into speech information VoI from various sounds contained in acoustic information AcI_1. The VRM speech recognition model also has the function of identifying speakers from human voices. Specifically, it can identify speakers using voiceprints registered in a database. It can also identify speakers located in areas that cannot be captured by the camera module 110Cm. Additionally, it can selectively separate human voices identified by voiceprints.

[0064] Furthermore, the VRM speech recognition model can grasp the context from the conversation contained in the VoI (Voice Information). For example, a large-scale language model can be used in the VRM speech recognition model.

[0065] 《VGM Speech Generation Model》 The VGM speech generation model has the function of generating synthesized speech (SV) to supplement speech information (VoI).

[0066] For example, when the audio information VoI contains unclear or missing portions, the speech generation model VGM can fill in those unclear or missing portions based on the context grasped by the speech recognition model VRM. Furthermore, the speech generation model VGM has the function of generating synthesized speech SV by translating the audio information VoI into a different language. In other words, one embodiment of the information processing system of the present invention can be used as a simultaneous translation device.

[0067] <Example 1 of Component 120> Component 120 has the function of receiving video information VdI, sound information AcI_1, voice information VoI, synthesized voice SV, and sound source information SSI, and the function of transmitting video information VdI and sound information AcI_1 to component 130 (see Figure 4B).

[0068] Component 120 includes a selection switch SS and an amplifier Amp. The selection switch SS has the function of creating acoustic information AcI_2 from acoustic information AcI_1 based on sound source information SSI and transmitting it to component 110. For example, depending on the location or type of sound source, the volume or frequency characteristics of acoustic information AcI_1 can be changed and used in acoustic information AcI_2. Alternatively, for example, depending on the location or type of sound source, information can be added to acoustic information AcI_1 and used in acoustic information AcI_2.

[0069] This makes it possible to provide, for example, the user with acoustic information AcI_2, which is created by means determined based on location information PI or attribute information Attr. As a result, a novel information processing system with superior convenience, usefulness, or reliability can be provided.

[0070] <Example 2 of Component 120> For example, when the location information PI included in the sound source information SSI is located within a predetermined range SA, and the attribute information Attr is classified as the conversation partner CP or the dangerous Dng, component 120 modifies the acoustic information AcI_1 and uses it for acoustic information AcI_2. For example, an amplifier Amp can be used to increase the volume of acoustic information AcI_1 and use it for acoustic information AcI_2. Alternatively, an equalizer can be used to modify the sound quality of acoustic information AcI_1 and use it for acoustic information AcI_2. Alternatively, the noise of acoustic information AcI_1 can be reduced. Alternatively, the volume of acoustic information AcI_1 can be reduced. Note that when acoustic information AcI_2 is created using component 120, the processing time can be shortened compared to when using the speech recognition model VRM or speech generation model VGM provided by component 130. Furthermore, for example, the delay time from the timing of the utterance of the conversation partner CP to the provision of acoustic information AcI_2 to the user of the information processing system according to one embodiment of the present invention can be shortened compared to the case where acoustic information AcI_2 is generated using the speech recognition model VRM or the speech generation model VGM.

[0071] This allows, for example, the volume of the voice of a conversation partner (CP) near the user to be increased, or the low frequencies of the conversation partner's voice to be suppressed, or the mid and high frequencies of the conversation partner's voice to be emphasized, or the quality of the conversation partner's voice to be altered to make it easier to hear. It also allows, for example, the user to be alerted to approaching dangers (Dng). Furthermore, for example, the user can converse without feeling a time lag between their speech and the conversation partner's speech. Furthermore, for example, the user can be alerted to approaching dangers (Dng) without feeling a time lag. As a result, a novel information processing system with superior convenience, usefulness, and reliability can be provided.

[0072] <Example 3 of Component 120> Also, for example, when the location information PI included in the sound source information SSI is located outside the predetermined range SA, and the attribute information Attr is classified as speaker Spkr, component 120 uses the speech information VoI for the acoustic information AcI_2.

[0073] Furthermore, for example, when the speech information VoI includes unclear or missing parts, component 120 uses synthesized speech SV for the acoustic information AcI_2. Note that if the location information PI is located outside the predetermined range SA, even if a delay occurs between the timing of the speaker Spkr's utterance and the provision of the acoustic information AcI_2 to a user of the information processing system according to one embodiment of the present invention, the user is unlikely to perceive any discomfort. Also, even if a delay occurs when generating the acoustic information AcI_2 using the speech recognition model VRM or the speech generation model VGM, the user is unlikely to perceive any discomfort.

[0074] This allows, for example, the speech of a speaker Spkr located at a distance from the user to be separated from ambient noise and provided to the user. Furthermore, the speech recognition model VRM can be used to accurately separate speech information VoI from acoustic information AcI_1. Also, even if there is a slight time lag between the speaker Spkr's speech and the acoustic information AcI_2 when the speaker Spkr is at a distance from the user, it is less likely to be perceived as unnatural. Additionally, the speech recognition model VRM and the speech generation model VGM can be used to clarify unclear parts of the speech information VoI. Furthermore, missing parts of the speech information VoI can be supplemented. As a result, a novel information processing system with superior convenience, usefulness, and reliability can be provided.

[0075] <Example 4 of Component 120> Component 120 has a function to create notification sound NS based on attribute information Attr. Specifically, component 120 includes a sampling device SmpD that creates notification sound NS. For example, the sound of a car driving, a dog barking, etc., can be recorded in the sampling device SmpD and played back based on the attribute information Attr.

[0076] For example, when the location information PI included in the sound source information SSI is located outside a predetermined range SA, and the attribute information Attr is classified as dangerous Dng, component 120 uses the notification sound NS for the acoustic information AcI_2.

[0077] This allows, for example, the user to be alerted to a dangerous location (Dng) far away using a notification sound (NS). As a result, a novel information processing system with superior convenience, usefulness, and reliability can be provided.

[0078] <Example of Information Processing System Configuration 2> The information processing system described in this embodiment has component 110, component 120, and component 130 (see Figure 3).

[0079] For example, an information processing system according to one aspect of the present invention can be configured with an information processing device that performs the functions of component 110 and component 120, and an information processing device that performs the functions of component 130 (see Figure 1). The number of information processing devices that constitute the information processing system according to one aspect of the present invention is one or more. Furthermore, for example, an information processing system according to one aspect of the present invention can be configured by connecting multiple information processing devices using a network 51.

[0080] By configuring an information processing system according to one aspect of the present invention using multiple information processing devices, the load related to information processing can be distributed.

[0081] <Configuration Example 1 of the Information Processing Device> Configuration Example 1 of the information processing device described in this embodiment can be used for component 110 and component 120. Configuration Example 1 of the information processing device can also be called a terminal device or a wearable device. For example, earphones, hearing aids, augmented reality (AR) devices, or cross reality (XR) devices can be used for component 110.

[0082] Configuration Example 1 of the information processing device can collect information from the surroundings of a user of an information processing system according to one aspect of the present invention. For example, it can collect acoustic and visual information. Furthermore, Configuration Example 1 of the information processing device can add the acoustic information of the virtual space output by the information processing system according to one aspect of the present invention to the real environment in real time, giving the user a sense of hearing things as if they were real.

[0083] In component 110, for example, application software operates. Users of an information processing system according to one embodiment of the present invention can access the information processing system via the application software. This allows them to enjoy services using the information processing system according to one embodiment of the present invention.

[0084] <Example of Information Processing Device Configuration 2> For example, the example of information processing device configuration 2 described in this embodiment can be used for component 130. For example, a portable information processing device, a server computer, a supercomputer, or other large computer can be used for component 130.

[0085] Furthermore, it is preferable that the second example of the information processing device configuration has the functionality of a parallel computer. By using it as a parallel computer, for example, it can perform large-scale calculations necessary for AI learning and inference.

[0086] Furthermore, the second example of the information processing device configuration can perform processing using an AI-based natural language model. In particular, it can execute processing using a general-purpose language model capable of performing various natural language processing tasks.

[0087] For example, processing can be performed using natural language models such as GPT-3®, GPT-3.5, GPT-4®, LaMDA, Llama2, and Llama3. Processing using GPT-4® is particularly preferable. For instance, being able to perform processing using a large-scale language model, compared to conventional natural language models, can enable more natural document generation or dialogue.

[0088] Furthermore, a provider of services using an information processing system according to one embodiment of the present invention is not necessarily required to own the Information Processing Device Configuration Example 2. For example, a service provider may utilize a portion of the services provided by other businesses using the Information Processing Device Configuration Example 2.

[0089] <Example of Network 51 Configuration> A network 51 that can be used in an information processing system according to one embodiment of the present invention can connect multiple information processing devices. As a result, the connected multiple information processing devices can transmit and receive data from each other. In addition, the load related to information processing can be distributed.

[0090] When performing wireless communication, communication protocols or technologies such as 4G, 5G, 6G, or specifications standardized by IEEE, such as Wi-Fi® and Bluetooth®, may be used.

[0091] For example, a local network can be used as network 51. Alternatively, an intranet or extranet can be used as network 51. Furthermore, PAN (Personal Area Network), LAN (Local Area Network), CAN (Campus Area Network), MAN (Metropolitan Area Network), WAN (Wide Area Network), GAN (Global Area Network), etc., can be used as network 51.

[0092] Furthermore, for example, a global network can be used for network 51. Specifically, the Internet, which is the foundation of the World Wide Web (WWW), can be used.

[0093] Furthermore, a person who provides services using an information processing system according to one aspect of the present invention can, for example, provide services using an information processing method according to one aspect of the present invention via a network 51.

[0094] Furthermore, if an information processing system according to one embodiment of the present invention is built within a local network, the possibility of confidential information leakage can be reduced compared to using the Internet.

[0095] <Example of Information Processing Device Configuration 3> An information processing device 20 that can be used in an information processing system according to one embodiment of the present invention has, for example, an input unit 21, a storage unit 22, a processing unit 23, an output unit 24, and a transmission line 25 (see Figure 5).

[0096] In the drawings attached to this specification, the components are classified by function and shown as independent blocks in the block diagram. However, in reality, it is difficult to completely separate the components by function, and one component may be involved in multiple functions. For example, a part of the processing unit 23 may function as the input unit 21. Also, one function may be involved in multiple components. For example, the processing performed by the processing unit 23 may be executed by different information processing devices depending on the processing.

[0097] Input Unit 21 The input unit 21 can receive data from outside the information processing device. For example, the input unit 21 can receive data via the acoustic module 110Mc, the camera module 110Cm, or the network 51. Specifically, a device equipped with a communication port can be used.

[0098] The input unit 21 supplies the received data to either or both of the storage unit 22 and the processing unit 23 via the transmission line 25.

[0099] <Storage Unit 22> The storage unit 22 has the function of storing the program executed by the processing unit 23. The storage unit 22 may also have the function of storing data generated by the processing unit 23 (for example, calculation results, analysis results, inference results), data received by the input unit 21, etc.

[0100] The storage unit 22 may have a database. The information processing device may also have a database separate from the storage unit 22. The information processing device may have the function to retrieve data from a database located outside the storage unit 22, outside the information processing device itself, or outside the information processing system. Furthermore, the information processing device may have the function to retrieve data from both its own database and an external database.

[0101] Storage can be used in the storage unit 22. Additionally, a database recording the paths of files stored on the file server can be used in the storage unit 22.

[0102] The memory unit 22 has at least one of volatile memory and non-volatile memory. Examples of volatile memory include DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). Examples of non-volatile memory include ReRAM (Resistive Random Access Memory), PRAM (Phase Change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory), and flash memory. Furthermore, the storage unit 22 may have at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark). The storage unit 22 may also have a recording media drive. Examples of recording media drives include hard disk drives (HDDs) and solid state drives (SSDs).

[0103] NOSRAM is an abbreviation for "Nonvolatile Oxide Semiconductor Random Access Memory (RAM)". NOSRAM is a memory in which the memory cell is a 2-transistor type (2T) or 3-transistor type (3T) gain cell, and the transistors are transistors that use metal oxide in the channel formation region (also called OS transistors). OS transistors have an extremely small current flowing between the source and drain when off, i.e., a leakage current. By using the characteristic of extremely low leakage current to hold charge according to the data in the memory cell, NOSRAM can be used as a non-volatile memory. In particular, because NOSRAM can read the stored data without destroying it (non-destructive read), it is suitable for arithmetic processing that involves repeating data read operations a large amount. Since the data capacity of NOSRAM can be increased by stacking them, it can be used as a large-scale cache memory, main memory, or storage memory to improve the performance of semiconductor devices.

[0104] DOSRAM is an abbreviation for "Dynamic Oxide Semiconductor RAM," and refers to RAM with a 1T (transistor) 1C (capacitance) type memory cell. DOSRAM is a DRAM formed using OS transistors, and it is a memory that temporarily stores information sent from the outside. DOSRAM is a memory that takes advantage of the small off-current of OS transistors.

[0105] In this specification, "metal oxide" refers to an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also called oxide semiconductors or simply OS), etc. For example, when a metal oxide is used in the semiconductor layer of a transistor, that metal oxide may be referred to as an oxide semiconductor.

[0106] The metal oxide in the channel-forming region preferably contains indium (In). When the metal oxide in the channel-forming region contains indium, the carrier mobility (electron mobility) of the OS transistor increases. For example, indium oxide (InOx) or indium gallium zinc oxide (In-Ga-Zn oxide, also written as "IGZO") can be used in the channel-forming region. Furthermore, the metal oxide in the channel-forming region is preferably an oxide semiconductor containing element M. Element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements applicable to element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). However, in some cases, element M may be a combination of multiple elements as mentioned above. Element M is, for example, an element with a high bond energy with oxygen. For example, an element with a higher bond energy with oxygen than indium. Furthermore, the metal oxide containing the channel-forming region is preferably a metal oxide containing zinc (Zn). Metal oxides containing zinc may be more prone to crystallization.

[0107] The metal oxides present in the channel-forming regions are not limited to indium-containing metal oxides. For example, the metal oxides present in the channel-forming regions may be zinc-tin oxides, gallium-tin oxides, or other metal oxides that do not contain indium but contain zinc, gallium, or tin.

[0108] Processing Unit 23: The processing unit 23 has the function of performing calculations, analyses, and inferences using data supplied from either or both of the input unit 21 and the storage unit 22. The processing unit 23 can supply the generated data (for example, calculation results, analysis results, and inference results) to either or both of the storage unit 22 and the output unit 24.

[0109] The processing unit 23 has the function of acquiring data from the storage unit 22. The processing unit 23 may also have the function of recording or registering data in the storage unit 22.

[0110] The processing unit 23 may, for example, have an arithmetic circuit. The processing unit 23 may, for example, have a central processing unit (CPU). The processing unit 23 may also have a graphics processing unit (GPU). The processing unit 23 may also have a neural network processing unit (NPU).

[0111] The processing unit 23 may have a microprocessor such as a DSP (Digital Signal Processor). The microprocessor can be implemented using a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) or FPAA (Field Programmable Analog Array). The processing unit 23 may also have a quantum processor. The processing unit 23 can perform various data processing and program control by interpreting and executing instructions from various programs using the processor. Programs that can be executed by the processor are stored in at least one of the processor's memory area and storage unit 22.

[0112] The processing unit 23 may have main memory. The main memory may include at least one of volatile memory such as RAM and non-volatile memory such as ROM (Read Only Memory). Furthermore, the main memory may include at least one of the above-mentioned NOSRAM and DOSRAM.

[0113] For RAM, for example, DRAM or SRAM is used, and a memory space is virtually allocated and used as the workspace for the processing unit 23. The operating system, application programs, program modules, program data, and lookup tables stored in the storage unit 22 are loaded into RAM for execution. These data, programs, and program modules loaded into RAM are directly accessed and manipulated by the processing unit 23.

[0114] ROM can store the BIOS (Basic Input / Output System) and firmware, etc., which do not require rewriting. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), and EPROM (Erasable Programmable Read Only Memory). Examples of EPROMs include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory), which allows data to be erased by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), and flash memory.

[0115] The processing unit 23 may have either or both an OS transistor and a transistor having silicon in its channel formation region (Si transistor).

[0116] The processing unit 23 preferably has an OS transistor. Because the OS transistor has an extremely small off-current, by using the OS transistor as a switch to hold the charge (data) that has flowed into a capacitive element that functions as a memory element, the data retention period can be ensured over a long period. By using this characteristic in at least one of the registers and cache memory of the processing unit, the processing unit can be operated only when necessary, and in other cases, the information of the previous processing is saved to the memory element, thereby turning off the processing unit. In other words, normally-off computing becomes possible, and the power consumption of the information processing system can be reduced.

[0117] It is preferable that the information processing device uses AI for at least some of its processing.

[0118] Information processing devices preferably utilize artificial neural networks (ANNs, also simply referred to as neural networks). Neural networks are implemented using circuits (hardware) or programs (software).

[0119] In this specification, the term "neural network" refers to any model that mimics the neural network of living organisms, determines the strength of connections between neurons through learning, and possesses problem-solving capabilities. A neural network has an input layer, an intermediate layer (hidden layer), and an output layer.

[0120] In this specification and other documents, when discussing neural networks, the process of determining the connection strength (also called weight coefficient) between neurons from existing information is sometimes referred to as "learning."

[0121] In this specification and other documents, the process of constructing a neural network using connection strengths obtained through learning and deriving new conclusions from it may be referred to as "inference."

[0122] Output Unit 24 The output unit 24 can output at least one of the calculation results, analysis results, and inference results from the processing unit 23 to the outside of the information processing device. For example, the output unit 24 can transmit data via the speaker module 110Sp or the network 51. Specifically, a communication port or a device with communication functionality can be used. Alternatively, a device with communication functionality may be used for both the input unit 21 and the output unit 24.

[0123] 《Transmission Line 25》 The transmission line 25 has the function of transmitting data. Data can be transmitted and received between the input unit 21, the storage unit 22, the processing unit 23, and the output unit 24 via the transmission line 25. Specifically, an external bus can be used as the transmission line 25.

[0124] This embodiment can be appropriately combined with other embodiments shown in this specification.

[0125] (Embodiment 2) In this embodiment, an information processing method according to one aspect of the present invention will be described with reference to Figures 6 and 7.

[0126] Figure 6 is a flowchart illustrating an information processing method according to one embodiment of the present invention.

[0127] Figure 7 is a flowchart illustrating an information processing method according to one embodiment of the present invention, as shown in step S5 of Figure 6.

[0128] <Example 1 of Information Processing Method> An information processing method according to one aspect of the present invention is an information processing method comprising a first step S1 to a fifth step S5 (see Figure 6).

[0129] <Step S1> In step S1, component 110 creates video information VdI and audio information AcI_1 and transmits them to component 120.

[0130] <Step S2> In step S2, component 120 receives video information VdI and audio information AcI_1 and transmits them to component 130.

[0131] <Step S3> In step S3, component 130 receives video information VdI and sound information AcI_1, and generates sound source information SSI using the object detection model ODM.

[0132] <Step S4> In step S4, component 130 transmits sound source information SSI to component 120.

[0133] <Example 1 of Step S5> In Step S5, component 120 creates acoustic information AcI_2 from acoustic information AcI_1 based on sound source information SSI and transmits it to component 110. The sound source information SSI includes location information PI and attribute information Attr.

[0134] This makes it possible to provide, for example, the user with acoustic information AcI_2, which is created by means determined based on location information PI or attribute information Attr. As a result, a novel information processing method with superior convenience, usefulness, or reliability can be provided.

[0135] <Example 2 of Step S5> In Step S5, when the location information PI is located within a predetermined range SA and the attribute information Attr is classified as the conversation partner CP or the danger Dng, component 120 modifies the acoustic information AcI_1 and uses it for acoustic information AcI_2 (see Figure 7). Note that when the attribute information Attr is, for example, the sound Other of the environment where the user of the information processing device is located, component 120 can use information that does not contain sound (No Sound) for acoustic information AcI_2. Alternatively, information with the phase inverted from acoustic information AcI_1 can be used for acoustic information AcI_2.

[0136] This allows, for example, the volume of the voice of a conversation partner (CP) near the user to be increased, or the low frequencies of the conversation partner's voice to be suppressed, or the mid-range and high frequencies of the conversation partner's voice to be emphasized, or the quality of the conversation partner's voice to be changed to make it easier to hear. It also allows, for example, the user to be alerted to approaching dangers (Dng), or to have a conversation without feeling a time lag between the user and the conversation partner's speech, or to be alerted to approaching dangers (Dng) without feeling a time lag, or to reduce noise that may bother the user. As a result, a novel information processing method with superior convenience, usefulness, or reliability can be provided.

[0137] <Example 3 of Step S5> Also, in Step S5, when the location information PI is located outside the predetermined range SA and the attribute information Attr is classified as speaker Spkr, the component 120 uses the speech recognition model VRM to separate the speech information VoI from the acoustic information AcI_1 and use it for the acoustic information AcI_2 (see Figure 7).

[0138] <Example 4 of Step S5> Also, in Step S5, when the speech information VoI includes unclear or missing parts (missing), component 120 uses the speech generation model VGM to generate a synthesized speech SV from the speech information VoI with the unclear or missing parts filled in, and uses it for the acoustic information AcI_2.

[0139] This allows, for example, the speech of a speaker Spkr located at a distance from the user to be separated from ambient noise and provided to the user. Furthermore, the speech recognition model VRM can be used to accurately separate speech information VoI from acoustic information AcI_1. In addition, even if there is a slight time lag between the speaker Spkr's speech and the acoustic information AcI_2 when the speaker Spkr is located at a distance from the user, it is less likely to be perceived as unnatural. As a result, a novel information processing method with superior convenience, usefulness, and reliability can be provided.

[0140] <Example 5 of Step S5> Also, in Step S5, when the location information PI is located outside the predetermined range SA and the attribute information Attr is classified as dangerous Dng, the component 120 creates a notification sound NS based on the attribute information Attr and uses it for the acoustic information AcI_2 (see Figure 7). For example, when lightning is detected, a thunderclap can be created and used for the notification sound NS.

[0141] This allows, for example, the user to be alerted to a dangerous location (Dng) far from them using a notification sound (NS). As a result, a novel information processing method with superior convenience, usefulness, and reliability can be provided.

[0142] This embodiment can be appropriately combined with other embodiments shown in this specification.

[0143] AcI_1: Acoustic information, AcI_2: Acoustic information, Dng: Danger, NS: Notification sound, ODM: Object detection model, Othr: Sound, SmpD: Sampling device, Spkr: Speaker, SSI: Sound source information, SV: Synthesized speech, VdI: Video information, VGM: Speech generation model, VoI: Speech information, VRM: Speech recognition model, 20: Information processing device, 21: Input unit, 22: Memory unit, 23: Processing unit, 24: Output unit, 25: Transmission line, 51: Network, 110: Component, 110Cm: Camera module, 110Mc: Acoustic module, 110Sp: Speaker module, 120: Component, 130: Component, 130a: Wristwatch-type information processing device, 130b: Portable information processing device, 130c: Server, 700: Information processing device, 700L: Information processing device, 700R: Information processing device

Claims

The first component and The second component, It has a third component, The first component includes a function for creating video information and first audio information and transmitting it to the third component, and a function for receiving second audio information and providing the second audio information. The second component has the function of receiving the video information and the first audio information and transmitting the sound source information, voice information and synthesized voice to the third component. The second component includes a function for processing using an object detection model, a speech recognition model, and a speech generation model, The object detection model includes a function to generate sound source information from the video information, The aforementioned sound source information includes location information and attribute information, The speech recognition model includes a function for separating the speech information from the first acoustic information, The voice generation model includes a function to generate synthesized speech that supplements the voice information, The third component includes a function for receiving the video information, the first audio information, the voice information, the synthesized voice, and the sound source information, and a function for transmitting the video information and the first audio information to the second component. The third component is an information processing system that has the function of creating second acoustic information from the first acoustic information based on the sound source information and transmitting it to the first component.   When the location information is located within a predetermined range, and the attribute information is classified as a conversation partner or a danger, The information processing system according to claim 1, wherein the third component has a function for modifying the first acoustic information and using it for the second acoustic information.   When the location information is located outside a predetermined range, and the attribute information is classified as a speaker, The third component has a function to use the audio information for the second acoustic information, When the aforementioned audio information includes unclear or missing portions, The information processing system according to claim 1, wherein the third component has a function of using the synthesized speech for the second acoustic information.   The third component has a function to create a notification sound based on the attribute information, When the location information is located outside a predetermined range, and the attribute information is classified as dangerous, The information processing system according to claim 1, wherein the third component has a function of using the notification sound for the second acoustic information.   An information processing method comprising the first to fifth steps, In the first step described above, the first component creates video information and first audio information and transmits them to the second component. In the second step, the second component receives the video information and the first audio information and transmits them to the third component. In the third step, the third component receives the video information and the first audio information, and generates sound source information using an object detection model. In the fourth step, the third component transmits the sound source information to the second component. In the fifth step, the second component creates second acoustic information from the first acoustic information based on the sound source information and transmits it to the first component. The aforementioned sound source information includes location information and attribute information, and is an information processing method.   The information processing method according to claim 5, In the fifth step described above, When the location information is located within a predetermined range, and the attribute information is classified as a conversation partner or a danger, The second component is an information processing method that modifies the first acoustic information and uses it for the second acoustic information.   The information processing method according to claim 5, In the fifth step described above, When the location information is located outside a predetermined range, and the attribute information is classified as a speaker, The second component is an information processing method that uses a speech recognition model to separate speech information from the first acoustic information and use it for the second acoustic information.   The information processing method according to claim 7, In the fifth step described above, When the aforementioned audio information includes unclear or missing portions, The second component is an information processing method that uses a speech generation model to generate synthesized speech from the speech information, in which unclear or missing parts are supplemented, and uses this synthesized speech for the second acoustic information.   The information processing method according to claim 5, In the fifth step described above, When the location information is located outside a predetermined range, and the attribute information is classified as dangerous, The second component is an information processing method that creates a notification sound based on the attribute information and uses it in the second acoustic information.