Recording sounds separated from a mixed sound stream on a personal device

AR glasses with multiple microphones and machine learning algorithms facilitate sound localization and separation, enabling users to isolate and record specific audio sources from mixed environments.

JP7811071B2Active Publication Date: 2026-02-04INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2023541566
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-28
Filing Date
2022-02-18
Publication Date
2026-02-04
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

Existing technologies struggle to isolate and record specific audio sources from a mixed sound environment, particularly in environments with multiple simultaneous sound sources, due to challenges in sound localization and separation.

Method used

A method and system using augmented reality (AR) glasses with multiple microphones and machine learning algorithms to separate and categorize sounds, allowing users to select and record specific sounds through icons displayed on the device interface.

Benefits of technology

Enables effective sound localization and separation, allowing users to selectively record and modify audio sources, enhancing user control over sound selection and recording.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811071000001
    Figure 0007811071000001
  • Figure 0007811071000002
    Figure 0007811071000002
  • Figure 0007811071000003
    Figure 0007811071000003
Patent Text Reader

Abstract

The method provides for one or more processors to receive, on a personal device, a mixture of sounds in a sound stream from a plurality of sound sources. The one or more processors identify one or more sounds of the mixture from the plurality of sound sources based on a sound separation technique. The one or more processors display icons on a user interface of the personal device, each corresponding to a classification of the one or more sounds identified from the plurality of sound sources. The one or more processors receive a selection of a sound from the mixture of sounds based on an action by a user of the personal device of selecting an icon displayed on the user interface of the personal device, and the one or more processors record the sound from the mixture of sounds selected by the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to the field of audio source recording, and more particularly to isolating audio sources of interest for selection of recording targets. [Background technology]

[0002] Many environments contain multiple sound sources that appear to blend together and mix into an indistinguishable collective stream of sound. Collective sound streams can exist indoors, such as at large social gatherings where multiple simultaneous conversations are taking place. Collective sound streams can also exist outdoors, and can include a combination of natural sounds such as wind, birds, and rain, and man-made sounds such as people playing, talking, and cars driving by. Multiple sound sources can blend together to appear as a unified background sound.

[0003] In general, sound source and sound discrimination are affected by the simultaneous occurrence of multiple sounds. Determining the location, azimuth, height, and distance of a sound source in three dimensions. Determining the location of a sound source is based on three types of cues: two binaural cues (interaural time difference and interaural level difference) and one monaural spectral cue (head-related transfer function). Sound localization is based on binaural cues (interaural differences), differences in sound arrival at two detectors such as human ears or dual microphones (i.e., differences in sound arrival time or intensity at the left and right ears), or monaural spectral cues (e.g., frequency-dependent sound patterns).

[0004] Augmented reality glasses are often used to include features and functionality that apply to the actual peripheral vision. In some cases, the augmented reality glasses can add images or indicators to the viewing screen that are displayed in addition to the peripheral vision in the direction the augmented reality (AR) glasses are pointed. In other cases, the AR glasses include information associated with the direction of the peripheral vision, which may be in the form of text, symbols, or audio playback. Summary of the Invention

[0005] Embodiments of the present invention disclose a method, computer program product, and system for selectively recording one or more sounds separated from a multi-sound environment. The method includes one or more processors receiving a mixture of sounds in a sound stream from multiple sound sources on a personal device. The one or more processors identify one or more sounds from the mixture of sounds from the multiple sound sources based on a sound separation technique. The one or more processors display icons on a user interface of the personal device, each corresponding to a classification of the one or more sounds identified from the multiple sound sources. The one or more processors receive a selection of a sound from the mixture of sounds based on an action by a user of AR glasses selecting an icon displayed on the user interface of the personal device, and the one or more processors record the sound from the mixture of sounds selected by the user. [Brief explanation of the drawings]

[0006] [Figure 1] 1 is a functional block diagram illustrating a distributed data processing environment, according to one embodiment of the present invention. [Figure 2] 10A-10C illustrate example sound category icons displayed within the user's personal device area according to one embodiment of the present invention. [Figure 3]2 is a flowchart illustrating the operational steps of a sound selection program operating in the distributed data processing environment of FIG. 1 in accordance with an embodiment of the present invention. [Figure 4] 4 is a block diagram of components of a computing system including a computing device configured to operatively execute the sound selection program of FIG. 3, according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0007] Embodiments of the present invention recognize that sounds from various sources may occur simultaneously and the current difficulty in identifying and distinguishing one sound from one source from other respective sounds. Embodiments further recognize the difficulty in determining the direction and proximity of sounds that occur together, such as birdsong, multiple vehicles, human voices, and natural sounds such as wind and rain. Embodiments also recognize that sound localization can be determined and sound separation can be achieved by applying recurring patterns, temporal regularities, and time-frequency decomposition techniques and algorithms that focus on feature composition discrepancies within mixed sounds.

[0008] Embodiments of the present invention provide a computer program product and a computer system for determining the localization of detected sound sources that form a sound mixture, enabling user-selected sound separation and recording on a personal device. In some embodiments, the personal device is a pair of augmented reality (AR) glasses configured with two or more microphones, a wireless network connection, and resources for running a sound selection program. In other embodiments, the personal device may be a smartphone or other smart device configured to receive a sound stream of the sound mixture and capable of running a sound selection program.

[0009] In some embodiments, the detected mixed sounds are separated and categorized by icons displayed to the user of the personal device. The user can select an icon representing the type and source of the sound and listen to and record the selected sound. In some embodiments, a new recording can be made by adding the raw separated sound selected for recording to a previously recorded sound. In some embodiments, a user of a personal device configured to separate sounds and present separate sound sources on the display of the personal device can modify characteristics of one or more sounds to be recorded, which may include attributes such as adjusting the volume or changing the pitch of the sound. Control of recording parameters is selectable from the personal device. In some embodiments in which a collaborative broadcast is received by multiple users, each user can select a separate sound from the broadcast on their respective personal device and record the selected sound. In still other embodiments, the user of the personal device is presented with icons in preferred locations representing sounds of interest to the user based on interest history or direct input by the user.

[0010] In embodiments of the present invention, two or more microphones on a personal device receive a sound stream containing multiple sounds from distinct sound sources. The sounds are separated using effective sound selection techniques and algorithms, such as time-frequency offset decomposition and temporal regularity of sound repetition. The separated sounds are then classified into categories established by training an artificial intelligence (AI) model using machine learning. Training involves applying supervised learning techniques to individual sounds and clustering similar sounds into classification categories. In some embodiments, the categories may be further drilled down to one or more sub-levels of categories. The separated sound categories include corresponding icons, which are presented to the user on a display component of the user's personal device, such as a smartphone display screen, a display on the inner portion of AR glasses, or a smartwatch display.

[0011] In some embodiments, the display of the icon includes a directional indicator of the sound source. In other embodiments, the relative distance may be indicated, for example, by the length of the directional indicator. The direction of the separated sound is determined by measuring the time delay of the received sound between the separated microphones. In some embodiments, an auxiliary array of microphones may be connected to the personal device to improve the accuracy of the sound localization detection with respect to direction and distance.

[0012] In embodiments of the present invention, a user of a personal device selects a sound to record by viewing an isolated sound icon on the display and performing an action to select the icon associated with the sound. The selection action of the AR glasses as a personal device may include detecting an eye focus direction toward the icon displayed on the inner surface of the AR glasses display and performing a blinking pattern. Optionally, the selection action of the AR glasses may be a hand gesture directed toward the location of the selected sound icon displayed on the AR glasses display. In some embodiments, the user is presented with an option to record characteristics of the selected sound. For example, the user may select to enhance the volume attribute of the selected sound, or, if multiple recordings of the selected sound are made, the user may increase the volume of one isolated sound and decrease the volume of the other recorded sounds.

[0013] The present invention will now be described in detail with reference to the drawings. Figure 1 is a functional block diagram illustrating a distributed computing environment, generally designated 100, in accordance with one embodiment of the present invention. Figure 1 is intended only as an example of one implementation and is not intended to suggest any limitations on the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made by one skilled in the art without departing from the scope of the present invention as defined by the claims.

[0014] The distributed computing environment 100 includes a computing device 110 and augmented reality (AR) glasses 120 interconnected via a network 150. The distributed computing environment 100 includes a sound stream, which represents a sound stream including a mixture of sounds from multiple sound sources, called a mixed sound 130. The network 150 may be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a virtual local area network (VLAN), or any combination that may include wired, wireless, or optical connections. In general, the network 150 may be any combination of connections and protocols that support communication and data transmission.

[0015] Computing device 110 includes a user interface 115 and a sound selection program 300, which is further shown to include sound recording functionality 117. In some embodiments, computing device 110 is a separate device communicatively connected to AR glasses 120 via network 150 (as shown in FIG. 1 ) and provides sound selection and recording functionality as well as memory storage. In other embodiments, computing device 110 is an integrated component (not shown) of AR glasses 120.

[0016] In some embodiments, computing device 110 may be a blade server, a web server, a laptop computer, a desktop computer, a standalone mobile computing device, a smartphone, a tablet computer, or another electronic device or computing system capable of receiving, transmitting, and processing data. In other embodiments, computing device 110 may be a wearable item or may be included in a wearable item for a user, such as AR glasses. In still other embodiments, computing device 110 may be a computing device that interacts with applications and services hosted and operating in a cloud computing environment. In another embodiment, computing device 110 may be a netbook computer, a personal digital assistant (PDA), or other programmable electronic device capable of communicating with and receiving data from other devices (shown and not shown) in distributed computing environment 100 via network 150 and performing the operations of resource prediction program 300. Alternatively, in some embodiments, computing device 110 may be communicatively connected to a remotely operating sound selection program 300. Computing device 110 may include internal and external hardware components, which are illustrated in more detail in FIG.

[0017] User interface 115 provides an interface for accessing features and functionality of computing device 110. In some embodiments of the present invention, user interface 115 provides access for operating and selecting options of sound selection program 300 and may also support initiating and selecting options for recording function 117 or other applications, features, and functionality (not shown) of computing device 110. In some embodiments, user interface 115 provides display input and output functionality for computing device 110. In other embodiments, user interface 115 is a component of AR glasses 120, such as display area 125, that provides display output and enables selection of options and functionality associated with sound selection program 300 running on computing device 110.

[0018] The user interface 115 supports access to alerts, notifications, and provides access to forms of communication. In one embodiment, the user interface 115 may be a graphical user interface (GUI) or web user interface (WUI) that can receive user input and display text, documents, web browser windows, user options, application interfaces, and instructions for operation, including information (e.g., graphics, text, and sounds) that a program presents to the user, as well as control sequences used by the user to control the program. In another embodiment, the user interface 115 may also include mobile application software that provides a respective interface to the features and functionality of the computing device 110. The user interface 115 enables users of the computing device 110 and the AR glasses 120 to receive input, see, hear, respond, access applications, view online conversations, and perform available functions.

[0019] The sound selection program 300 is an application for detecting and selecting one or more sounds from a sound stream containing a mixture of sounds from multiple sound sources and recording the selected separated sounds. In embodiments of the present invention, the sound selection program 300 operates from a user's personal device configured to receive sound input from two or more separated microphones that enable sound direction detection. In some embodiments, the user's personal device may be a suitably configured smartphone that includes two or more microphones positioned to detect sound source direction, such as on either side or both ends of the smartphone. In other embodiments, the user's personal device is a wearable item, such as AR glasses, that includes the functional capabilities of the computing device 110 and is capable of operating the sound selection program 300 and the recording function 117.

[0020] 1 illustrates computing device 110 as separate from AR glasses 120 to illustrate the roles of sound selection program 300 and recording function 117, it is recognized that in some embodiments of the present invention, AR glasses 120 include the computer functionality of computing device 110 and operatively execute sound selection program 300 and recording function 117. To clearly and concisely convey features of embodiments of the present invention, a user's personal device will be referred to herein by reference to AR glasses, such as AR glasses 120. Furthermore, it should be noted that embodiments of the present invention are not limited to AR glasses as the personal device that executes the operational steps of sound selection program 300 and recording function 117.

[0021] The sound selection program 300 includes machine learning techniques for recognizing and categorizing sounds from sound sources contained in a mixed sound stream of multiple sounds. In one embodiment, the sound selection program 300 is trained by submitting and identifying multiple types of sounds and further trained to detect the submitted and identified sounds within the simultaneous mixed sound of the sound stream. In some embodiments, sounds are clustered into sound categories, each of which is associated with an icon to enable and facilitate selection of the detected sounds. In some embodiments, categories may be drilled down into subcategories. The sound selection program 300 receives a sound stream containing multiple sounds from each sound source. In some embodiments, the sound selection program 300 determines the direction of the source of each sound and identifies the category of each sound. The sound selection program 300 displays an icon corresponding to the separated sound category on a user interface display, such as the display area 125 of the AR glasses 120. The sound selection program 300 displays an icon and a directional pointer for sounds detected and separated from the mixed sound stream of multiple sounds.

[0022] In some embodiments, the sound selection program 300 receives a selection of icons corresponding to the isolated sounds from a user of the AR glasses 120. In some embodiments, the icon selection presents the user with an option to review the recording of the selected sound and may include options for modifying the characteristics of the sound as it is recorded, such as attributes to increase or decrease the volume of the sound. In some embodiments, the user selects one or more icons for simultaneous recording, including options for modifying the characteristics of the recorded sound.

[0023] The recording function 117 is a module of the sound selection program 300 that provides functionality for recording selected sounds and applying selected characteristics to the recording. In some embodiments, the recording function 117 includes functionality for storing recordings and recalling previously stored recordings. In some embodiments, the recording function 117 may allow a user to copy a previously recorded sound as an option presented to the user and to mix a recording of another sound with a copy of a previously recorded sound. In embodiments in which a broadcast containing a mixture of multiple sounds is received by multiple users, the sound selection program 300 allows each user to select and record a separate sound from the broadcast of the mixture of multiple sounds.

[0024] The AR glasses 120 are augmented reality glasses and are shown in an exemplary configuration including a power source 122, a microphone 124, a display area 125, a processing and memory component 126, wireless communication 127, and an audio speaker 128. The AR glasses are shown as being wirelessly connected to the computing device 110. In some embodiments, the computing functionality, the sound selection program 300, and the sound recording functionality 117 are included in the AR glasses 120 (not shown). In some embodiments, the AR glasses 120 include the sound selection program 300 and operate the sound selection program 300 to display icons of sounds detected from a sound stream of mixed sounds on the display area 125. In some embodiments, a user of the AR glasses 120 selects an icon corresponding to a classification of the detected sound by focusing their eyes on the icon displayed in the display area 125 and performing a blinking action detected by the camera functionality (not shown) of the AR glasses 120. In other embodiments, a user of the AR glasses 120 selects an icon displayed in the display area 125 by performing a hand gesture that matches the display of the selected icon.

[0025] The power supply 122 is a component of the AR glasses 120, shown as an example earpiece of the AR glasses 120. The power supply 122 provides power to the processing and display functions of the AR glasses 120. The microphones 124 are shown as a pair of microphones located on opposing temple arms of the AR glasses 120. The microphones 124 receive a sound stream that may include a mixture of sounds from multiple sound sources. The microphones 124 are positioned to enable determination of the direction of the sound source. The memory component 126 is shown as an example component of the AR glasses 120, including a primary volatile memory and a storage memory for storing recorded selected sounds. The memory component 126 supports processing of sounds from the sound stream received through the microphones 124 and operation of the sound selection program 300. In some embodiments in which the AR glasses 120 are separate from but communicatively connected to the computing device 110, the wireless communication 127 enables wireless connection of the AR glasses 120 to the computing device 110 via the network 150. The audio speaker 128 provides an audio output to the user of the AR glasses 120 of sounds processed by the sound selection program 300, which are separated and delivered to the audio speaker 128 based on selections made by the user.

[0026] The sound mixture 130 is a sound stream that includes a mixture of multiple sounds from multiple respective sound sources. The sound mixture 130 is shown as including a mixture of bird sounds 140, car sounds 142, playground sounds 144, and additional sounds including wind sounds 146. The sound mixture 130 is received by the microphone 124 of the AR glasses 120 and processed by the sound selection program 300 to present icons corresponding to the sounds separated from the sound mixture 130 for selection by a user of the AR glasses 120.

[0027] FIG. 2 illustrates an example of sound category icons displayed in an area of ​​a user's personal device according to one embodiment of the present invention. FIG. 2 includes a display area 210, a vehicle icon 220 and corresponding directional pointer 222, a person icon 230 and directional pointer 232, a nature icon 240 and directional pointer 242, a playground icon 250 and directional pointer 252, a wind icon 260 and directional pointer 265, and a selection indicator 270. Each icon in display area 210 represents a sound that has been separated and identified from a mixture of sounds in a sound stream received by the user's personal device, such as AR glasses 120 (FIG. 1). Each corresponding directional pointer indicates the detected direction of the sound source.

[0028] The icons displayed in display area 210 represent sound categories, which in some embodiments of the present invention are assigned during training of sound selection program 300. Examples of sound streams represented by icons displayed in display area 210 include sounds from automobile traffic detected in the direction of direction pointer 222 and represented by vehicle icon 220. In some embodiments, sounds emanating from vehicles such as trucks, buses, motorcycles, trains, and bicycles are represented by vehicle icon 220. Similarly, sounds emitted by people who may be talking, singing, shouting, coughing, etc. are represented by person icon 230, and the detected direction of the person sound is pointed to by direction pointer 232. In some embodiments, sounds from various birds, dogs, cats, or other animals are represented by nature icon 240, and the detected direction of the sound is pointed to by direction pointer 242. In some embodiments, playground icon 250 and corresponding directional pointer 252 represent sounds detected from a playground or sporting event area, and detection of sounds from blowing wind is represented by wind icon 260 and directional pointer 265 indicating that the sound is not unidirectional.

[0029] Selection indicator 270 represents a user's instruction to select an icon displayed in display area 210 of the user's personal device. In some embodiments, the user's personal device presents user icons representing various sounds separated from a mixture of sounds in a sound stream, allowing the user to select a sound and record it. In one embodiment, the user's eyes turn and focus toward one of the icons presented on display area 210. The user maintains eye focus while performing a selection action, such as blinking multiple times. Selection indicator 270 provides user confirmation feedback of the selection made. If the user determines that the selected icon is not the intended one, the user can turn their eyes toward undo icon 280 to remove the current selection and make a different selection.

[0030] 3 is a flowchart illustrating the operational steps of a sound selection program 300 operating in the distributed computing environment 100 of FIG. 1, according to an embodiment of the present invention. An embodiment of the present invention includes a user's personal device on which the sound selection program 300 operates. The sound selection program 300 enables the user's appropriately configured personal device to receive mixed sounds in a sound stream, perform sound separation functions, categorize the separated sounds, and present icons corresponding to the classifications on a display component of the user's device. The sound separation program 300 enables the user to select sounds, perform recordings of the separated sounds, and adjust characteristics of the sound recordings.

[0031] In some embodiments, the user's personal device is a smartphone or smart device (i.e., a smartwatch) configured to receive a mixed sound stream and perform sound separation and sound localization functions on the received sounds. In other embodiments, the user's personal device is a pair of AR glasses configured with the computing capabilities and features to operate the sound selection program 300 and record the selected separated sounds. In other embodiments, the sound selection program 300 is contained in and operates from another wearable device. For purposes of clearly describing the functions and steps of the sound selection program 300, the user's personal device will be referred to as a suitably configured pair of AR glasses, recognizing that the user's personal device is not limited to AR glasses.

[0032] The sound selection program 300 receives a mixed sound from multiple sound sources (step 310). The sound selection program 300 receives the mixed sound from a microphone connected to the user's AR glasses. The mixed sound includes multiple sounds from each of the multiple sound sources that are perceived as blended into a single sound stream. The microphones that detect the sound streams are located on the AR glasses to enable sound localization, including the direction and possibly the distance of the sound sources, to be determined based on the relative amplitudes of the sound signals.

[0033] For example, the sound selection program 300 receives a detected sound stream from a microphone located on the arm of the AR glasses, such as microphone 124 of the AR glasses 120 of Figure 1. The sound selection program 300 receives the sound stream that is determined to include multiple distinct sounds from multiple respective sound sources.

[0034] The sound selection program 300 performs separation of the mixed sounds using sound separation techniques (step 320). The sound selection program 300 applies sound localization and sound separation techniques and algorithms to the received sound stream to perform separation of multiple sounds from multiple sound sources. In some embodiments, the sound separation techniques utilize time-frequency methods that detect temporal regularities in the mixed audio input. In some embodiments, the sound separation techniques include temporal coherence, emitting coherently modulated signatures as patterns of sound sources. In some embodiments, sound localization determines the direction of the sound separated from the mixed sound by determining the time delay in receiving sound signals between two or more microphones of the AR glasses.

[0035] For example, the sound selection program 300 applies sound separation techniques to the received sound stream to determine at least four component sounds contained within the sound stream. In some embodiments, once the sound source is separated from other sounds in the sound mix, the sound selection program determines the direction of the sound and associates that direction with a magnetic compass direction so that the relative direction can be displayed regardless of which way the user wearing the AR glasses is facing.

[0036] The sound selection program 300 identifies one or more sounds separated from the sound mixture (step 330). In some embodiments of the present invention, the sound selection program 300 includes machine learning training to identify sound types separated from the sound mixture and assign the identified sounds to classification categories. In exemplary embodiments, the sound selection program 300 is trained using supervised learning techniques involving a variety of previously recorded sounds presented at various volume levels, presented individually, and then presented with additional background and interfering sounds. In some embodiments, the training of the sound selection program 300 enables sound recognition of specific sound sources or "sound types" (i.e., motor vehicles). In some embodiments, the training of the sound separation program 300 includes speech recognition, which enables the sound selection program 300 to distinguish between separate speakers and, in some cases, identify speakers with sufficient training.

[0037] For example, the sound separation program 300 undergoes machine learning training using, among other things, bird sounds, car sounds, and people talking, as well as sounds from playgrounds. The training results in the recognition of these sounds or sounds that closely resemble these sounds. After separating a sound from a mixture of sounds in the received sound stream, the sound separation program 300 determines that the separated sound most closely matches a bird sound.

[0038] The sound separation program 300 assigns the identified sounds to classification categories represented by corresponding icons (step 340). After identifying sounds separated from the mixed sounds in the sound stream, the sound separation program 300 determines the category that most closely matches the identified sounds and assigns the identified sounds to the category represented by the corresponding icon. In some embodiments, during machine learning training of the sound separation program 300 operating in the AR glasses, the category into which the sounds fall and the corresponding icon are selected and input by a user.

[0039] For example, after the sound selection program 300 identifies a sound isolated from a received sound stream as a bird sound, it classifies the bird sound into a category of nature sounds represented by a corresponding tree image icon.

[0040] The sound selection program 300 displays a set of icons corresponding to the classification of the identified sounds (step 350). The identified sounds of the sound stream mixture are associated with respective icons corresponding to classification categories and presented to the user on the display area of ​​the AR glasses. The sound selection program 300 draws an icon on the display area of ​​the AR glasses for each separated sound of the sound stream. In some embodiments, as more separated sounds are identified and assigned category icons, the sound selection program 300 may display a limited number of icons at a time on the display area of ​​the AR glasses, with paging selection to display the next set of icons to be considered by the user. In some embodiments, the displayed icon also includes a directional pointer indicating the direction in which the separated sound is detected.

[0041] For example, after identifying a separated sound from the received sound mixture as a bird sound that is categorized as a nature sound and represented by a corresponding tree icon, the sound selection program 300 presents the tree icon on a display area of ​​the AR glasses worn by the user. By presenting the tree icon, the sound selection program 300 allows the user of the AR glasses to select the icon corresponding to the bird sound.

[0042] The sound selection program 300 records sounds separated from the mixed sound based on the user's selection (step 370). The user of the AR glasses is presented with a set of icons on a display area of ​​the AR glasses, each icon corresponding to a different sound separated from the received sound stream. The user of the AR glasses selects an icon from the displayed set of icons and begins recording the separated sound associated with the selected icon. In some embodiments, the sound selection program 300 includes a hierarchical structure of icons representing categories and subcategories of separated sounds in a social gathering, for example, where a first icon represents the category of "human voices" and a subcategory may include three different icons representing a group of three people having a conversation.

[0043] For example, the sound selection program 300 presents a set of icons on the display area 125 of the AR glasses 120 ( FIG. 1 ), each icon representing a separate sound isolated from a sound stream. A user of the AR glasses 120 looks at the set of icons and directs their eyes toward a tree icon associated with a category of nature sounds. The user performs a selection action, such as blinking quickly multiple times while keeping their eyes focused on the tree icon, and the sound selection program 300 presents a confirmation message to start recording the isolated sounds associated with the tree icons on the display area 125 of the AR glasses 120. The user performs a selection action to confirm the recording, and the sound selection program 300 starts recording the isolated sounds associated with the tree icons.

[0044] In some embodiments, the sound selection program 300 continues to learn as the user of the AR glasses performs multiple recording acts, and displays icons associated with the sound categories the user of the AR glasses prefers in more prominent, preferred positions on the display area of ​​the AR glasses. In some embodiments, isolated sounds that are not identified by the sound selection program 300 are assigned an icon corresponding to an "unknown" status, such as a question mark, providing the user with an opportunity to categorize the sound and associate an existing icon or assign a new icon to the sound.

[0045] After recording the selected separated sounds, the sound selection program 300 ends.

[0046] FIG. 4 illustrates a block diagram of components of a computing system 400 including a computing device 405 configured to include or operatively connect to the components shown in FIG. 1 and capable of operatively executing the sound selection program 300 of FIG. 2, according to one embodiment of the present invention.

[0047] Computing device 405 includes components and functional capabilities similar to those of computing device 110 (FIG. 1) in accordance with the illustrative embodiment of the invention. It should be understood that FIG. 4 is intended as an illustration of one implementation only and does not imply any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made.

[0048] The computing device 405 includes a communications fabric 402 that provides communication between a computer processor 404, memory 406, persistent storage 408, a communications unit 410, and an input / output (I / O) interface 412. The communications fabric 402 may be implemented using any architecture designed to pass data and / or control information between a processor (such as a microprocessor, communications, and network processor), system memory, peripheral devices, and any other hardware components in the system. For example, the communications fabric 402 may be implemented using one or more buses.

[0049] Memory 406, cache memory 416, and persistent storage 408 are computer-readable storage media. In this embodiment, memory 406 includes random access memory (RAM) 414. In general, memory 406 may include any suitable volatile or non-volatile computer-readable storage media.

[0050] In one embodiment, the sound selection program 300 is stored in persistent storage 408 for execution by one or more of the respective computer processors 404 via one or more of the memories 406. In this embodiment, persistent storage 408 includes a magnetic hard disk drive. Alternatively, or in addition to a magnetic hard disk drive, persistent storage 408 may include a solid-state hard drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or other computer-readable storage medium capable of storing program instructions or digital information.

[0051] The media used by persistent storage 408 may also be removable. For example, a removable hard drive may be used for persistent storage 408. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive to transfer data to another computer-readable storage medium that is also part of persistent storage 408.

[0052] In these examples, communications unit 410 provides for communication with other data processing systems or devices, including resources of distributed data processing environment 100. In these examples, communications unit 410 includes one or more network interface cards. Communications unit 410 may provide communications using either or both physical and wireless communications links. Sound selection program 300 may be downloaded to persistent storage 408 via communications unit 410.

[0053] The I / O interface 412 allows for the input and output of data with other devices that may be connected to the computing system 400. For example, the I / O interface 412 may provide a connection to an external device 418, such as a keyboard, keypad, touch screen, or any other suitable input device or combination thereof. The external device 418 may also include portable computer-readable storage media, such as thumb drives, portable optical or magnetic disks, and memory cards. Software and data used to implement embodiments of the present invention, such as the sound selection program 300, may be stored on such portable computer-readable storage media and loaded into the persistent storage 408 via the I / O interface 412. The interface 412 also connects to a display 420.

[0054] Display 420 provides a mechanism for displaying data to a user and may be, for example, a computer monitor.

[0055] The programs described herein are identified based on the application in which they are implemented in particular embodiments of the invention. However, it should be understood that any particular program nomenclature herein is used merely for convenience, and thus the invention should not be limited to use solely with the particular application identified and / or implied by such nomenclature.

[0056] The present invention may be a system, method, or computer program product, or a combination thereof, at any possible level of integration of technical details. The computer program product may include a computer-readable storage medium (or multiple computer-readable storage media) having computer-readable program instructions for causing a processor to implement aspects of the present invention.

[0057] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or raised structures in grooves with instructions recorded on them, and any suitable combination of the above. As used herein, computer-readable storage media should not be construed as being ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses through fiber optic cable), or electrical signals transmitted over electrical wires.

[0058] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or downloaded to an external computer or storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, fiber optic transmission cables, wireless transmission cables, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.

[0059] Computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, C++, and procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer as a standalone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to carry out aspects of the present invention, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry.

[0060] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0061] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, produce means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable medium, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions implementing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, and can direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner.

[0062] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to create a computer-implemented process that causes the computer, other programmable apparatus, or other device to perform a series of operational steps, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0063] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may actually be accomplished as a single step, or may be executed concurrently, substantially concurrently, partially, or fully in a time-overlapping manner, depending on the functionality involved, or the blocks may even be executed in the reverse order. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or actions or a combination of dedicated hardware and computer instructions.

Claims

1. 1. A method for selectively recording one or more sounds from a multi-sound environment by computerized information processing, comprising: receiving, by one or more processors, a mixture of sounds from a plurality of sound sources; performing, by the one or more processors, separation of the mixed sounds by applying one or more sound separation techniques; identifying, by the one or more processors, one or more sounds of the mixture of sounds from the multiple sound sources based on training using machine learning techniques; generating, by the one or more processors, a set of icons each corresponding to a classification assigned to each identified sound of the mixed sound from the plurality of sound sources; displaying, by the one or more processors, the set of icons and a directional icon indicating a direction of a source of the sound on a viewing screen on a user interface of a personal device of the user, wherein the personal device of the user is augmented reality (AR) glasses configured to record sounds and store the recorded sounds; receiving, by the one or more processors, a selection of a first icon from the generated set of icons associated with the sound based on a user action of selecting the first icon; recording, by the one or more processors, the sound from the mixed sound from the plurality of sound sources associated with the first icon selected by the user; A method comprising:

2. receiving, by the one or more processors, the mixed sound using two or more microphones; The method of claim 1 further comprising:

3. receiving, by the one or more processors, the mixed sound using two or more microphones; displaying, by the one or more processors, the set of icons on a user interface of a personal device of the user, wherein the personal device of the user is a smart device configured to record sounds and store the recorded sounds; 3. The method of claim 1 or 2, further comprising:

4. 4. The method of claim 1, wherein a set of parameters associated with recording the sound from the sound mixture selected by the action of the user selecting the first icon of the displayed set of icons associated with the sound mixture is controlled by the user selecting a displayed option.

5. The method of any one of claims 1 to 4, wherein the icons in the set of icons correspond to category classifications.

6. 6. The method of claim 1, wherein the user modifies characteristics of the sound from the sound mixture while recording the sound, and the characteristics of the sound include volume and pitch attributes of the recording.

7. The method of any one of claims 1 to 6, wherein the machine learning technique comprises supervised learning, in which a plurality of distinct audio sounds are delivered along with identities of the distinct audio sounds.

8. 1. A computer program for selectively recording one or more sounds from a multi-sound environment, comprising: receiving a mixed sound from a plurality of sound sources; performing separation of the mixed sounds by applying one or more sound separation techniques; identifying one or more sounds of the mixture from the multiple sound sources based on training using machine learning techniques; generating a set of icons each corresponding to a classification assigned to each identified sound of the mixed sounds from the plurality of sound sources; displaying the set of icons and a direction icon indicating the direction of the source of the sound on a viewing screen on a user interface of a personal device of the user, wherein the personal device of the user is augmented reality (AR) glasses configured to record sounds and store the recorded sounds; receiving a selection of a first icon from the generated set of icons associated with the sound based on a user action of selecting the first icon; recording the sound from the mixed sound from the plurality of sound sources associated with the first icon selected by the user; A computer program for executing the above.

9. On the computer, receiving the mixed sound by two or more microphones; 9. The computer program product of claim 8, further comprising:

10. On the computer, receiving the mixed sound with two or more microphones; displaying the set of icons on a user interface of a personal device of the user, wherein the personal device of the user is a smart device configured to record sounds and store the recorded sounds; 10. The computer program product according to claim 8, further comprising:

11. 1. A computer system for selectively recording one or more sounds from a multi-sound environment, said computer system comprising: one or more computer processors; one or more computer-readable storage media; and program instructions stored on the one or more computer-readable storage media, the program instructions comprising: program instructions for receiving a mixed sound from a plurality of sound sources; - program instructions for performing separation of the mixed sounds by applying one or more sound separation techniques; program instructions for identifying one or more sounds of the mixture of sounds from the multiple sound sources based on training using machine learning techniques; program instructions for generating a set of icons each corresponding to a classification assigned to each identified sound of the mixed sounds from the plurality of sound sources; program instructions for displaying the set of icons and a directional icon indicating a direction of a source of the sound on a viewing screen on a user interface of a personal device of a user, wherein the personal device of the user is augmented reality (AR) glasses configured to record sounds and store the recorded sounds; program instructions for receiving a selection of a first icon from the generated set of icons associated with the sound based on a user action of selecting the first icon; program instructions for recording the sound from the mixed sound from the plurality of sound sources associated with the first icon selected by the user; 1. A computer system comprising:

12. program instructions for receiving the mixed sound with two or more microphones; 12. The computer system of claim 11, further comprising:

13. program instructions for receiving the mixed sound with two or more microphones; program instructions for displaying the set of icons on a user interface of a personal device of the user, wherein the personal device of the user is a smart device configured to record sounds and store the recorded sounds; and 13. The computer system of claim 11 or 12, further comprising:

14. 14. The computer system of claim 11, wherein program instructions for modifying characteristics of the sound from the mixed sound while recording the sound are based on program instructions of selected options received from the user, and wherein the characteristics of the sound include volume and pitch attributes of the recording.

15. the program instructions for identifying the one or more sounds based on training using machine learning techniques, Applying machine learning techniques, including supervised learning, in which a plurality of distinct audio sounds are delivered along with an identification of said distinct audio sounds.

15. A computer system according to any one of claims 11 to 14, comprising programming instructions for:

16. A computer-readable recording medium storing the computer program according to any one of claims 8 to 10.

Citation Information

Patent Citations

  • Voice reproducing device and recording medium

    JP2001222300A

  • Recording apparatus and recording method

    JP2010178124A

  • Acoustic signal processing apparatus and reproducing device

    JP2010187363A

  • Imaging controller, imaging control method, program for imaging control method, and imaging apparatus

    JP2013106298A

  • Image display correction based on sensor input for transmissive myopia displays.

    JP2015504616A