Audio augmented reality system
By automatically retrieving and presenting relevant information through audio input, the problem of users having difficulty obtaining information when they cannot visually confirm the source of the sound is solved, and the automation and real-time performance of information in the audio augmented reality system are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2017-06-22
- Publication Date
- 2026-08-04
AI Technical Summary
In existing audio augmented reality systems, it is difficult for users to automatically retrieve and present relevant information through non-textual audio input, especially when the source of the sound cannot be visually confirmed.
The system captures sound waveforms from audio input on the device, automatically formulates queries and submits them to an online search engine, uses machine learning algorithms to improve the relevance of results, optimizes queries based on user feedback, and presents audio and visual information in real time.
It enables automatic retrieval and real-time presentation of relevant information even without text-based audio input, improving the relevance and accuracy of the information and enhancing the user experience.
Smart Images

Figure CN116859327B_ABST
Abstract
Description
[0001] Divisional Application Instructions
[0002] This application is a divisional application of Chinese invention patent application No. 201780037903.3, entitled "Audio Augmented Reality System", which was filed internationally on June 22, 2017, entered the Chinese national phase on December 18, 2018. Technical Field
[0003] The embodiments of this application relate to audio augmented reality systems. Background Technology
[0004] With the emergence of technologies for processing environmental input and transmitting information in real time, augmented reality systems will become increasingly prevalent in consumer, commercial, academic, and research settings. In audio augmented reality systems, real-time information can be presented to the user through one or more audio channels (e.g., headphones, speakers, or other audio devices). To improve the performance of audio augmented reality systems, it is desirable to provide technologies that increase the relevance and accuracy of the presented real-time information. Attached Figure Description
[0005] Figure 1 The illustration depicts a first scenario illustrating various aspects of this disclosure.
[0006] Figure 2 An exemplary sequence of functional blocks illustrating certain aspects of this disclosure is shown.
[0007] Figure 3 An exemplary embodiment of the operation performed by an audio input and / or output device locally available to the user is illustrated.
[0008] Figure 4 The illustration shows other aspects of an audio augmented reality system used to recover and retrieve relevant information.
[0009] Figure 5 An illustrative correspondence between exemplary formulaic queries and exemplary search results is shown.
[0010] Figure 6 The illustration shows another aspect of an augmented reality system used to visually display information related to received digital sound waveforms.
[0011] Figure 7 The illustration shows an exemplary application of the technology described in the audio augmented reality system above to a specific scenario.
[0012] Figure 8 An exemplary embodiment of the method according to this disclosure is illustrated.
[0013] Figure 9An exemplary embodiment of the apparatus according to the present disclosure is illustrated.
[0014] Figure 10 An exemplary embodiment of the device according to this disclosure is illustrated. Detailed Implementation
[0015] The various aspects of the technology described in this paper generally relate to techniques for searching and retrieving online information in response to queries that include digital audio waveforms. Specifically, the query is submitted to an online engine and may include multiple digital audio waveforms. One or more online results related to the formulated query are retrieved and presented to the user in real time in audio and / or visual formats. Based on user feedback, machine learning algorithms can be used over time to improve the relevance of the online results.
[0016] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of an exemplary apparatus "used as an example, instance, or illustration" and should not be construed as being more preferred or advantageous than other exemplary aspects. The detailed description includes specific details for the purpose of providing a thorough understanding of exemplary aspects of the invention. It will be apparent to those skilled in the art that exemplary aspects of the invention can be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form in order to avoid obscuring the novelty of the exemplary aspects presented herein.
[0017] Figure 1 The illustration depicts a first scenario 100 illustrating various aspects of this disclosure. Note that scenario 100 is shown for illustrative purposes only and is not intended to limit the scope of this disclosure to, for example, any particular type of audio signal that can be processed, a device for capturing or outputting audio input, a particular field of knowledge, search results, any type of information shown or suggested, or any illustrative scenario such as birdwatching or any other particular scenario.
[0018] exist Figure 1 The image shows multiple devices, including active earbuds 120 (including a left earbud 120a and a right earbud 120b), a smartphone 130, a smartwatch 140, a laptop computer 150, etc. A user 110 is schematically depicted listening to the audio output of the active earbuds 120. In an exemplary embodiment, each of the (left and right) active earbuds 120 may include a built-in microprocessor (not shown). The earbuds 120 may be configured to sample live audio and use the built-in microprocessor to process the sampled audio and further generate an audio output that modifies, enhances, or otherwise amplifies the live audio heard by the user in real time.
[0019] It should be understood that any device shown may be equipped with the ability to generate audio output for user 110 and / or receive audio input from user 110's environment. For example, in order to receive audio input, active earphone 120 may be provided with a built-in microphone or other type of sound sensor. Figure 1 (Not shown in the image) Smartphone 130 may include microphone 132, smartwatch 140 may include microphone 142, laptop computer 150 may include microphone 151, and so on.
[0020] In the first illustrative scenario 100, user 110 may be walking while possessing any or all of devices 120, 130, 140, and / or 150. User 110 may happen to encounter bird 160 singing bird song 162. User 110 may perceive bird 160 through his or her visual and / or audio-sensory perception (i.e., sight and / or sound). In this scenario, user 110 may wish to obtain additional information about bird 160 and / or bird song 162, such as the bird's identity and other information, the bird's position relative to user 110 (e.g., if only bird song 162 is heard, but bird 160 is not visible), etc. Note that the example of observing a bird is described for illustrative purposes only and is not intended to limit the scope of this disclosure to any particular type of sound or information that can be processed. In alternative exemplary embodiments, any sound waveform can be adapted, including but not limited to music (e.g., identifiers of music genres, bands, performers, etc.), speech (e.g., identifiers of speakers, natural language understanding, translation, etc.), artificial (e.g., identifiers of sirens, emergency calls, etc.), or natural sounds. Such alternative exemplary embodiments are contemplated within the scope of this disclosure.
[0021] Furthermore, devices 120, 130, 140, and 150 do not necessarily need to be owned by user 110. For example, while earphone 120 may be owned by and near user 110, laptop computer 150 may not belong to the user and / or may typically be located immediately inside or outside of user 110. According to the techniques of this disclosure, devices can generally be placed in the same general environment as the user, for example, allowing each device to provide useful input regarding a specific sound perceived by user 110.
[0022] It should be understood that any of devices 120, 130, 140, and 150 may have the ability to connect to a local network or the World Wide Web, while user 110 is observing bird 160 or listening to birdsong 162. User 110 may utilize this connectivity to, for example, access a network or the Web to retrieve desired information about bird 160 or birdsong 162. In an exemplary embodiment, user 110 may verbally express or otherwise input a query, and any of devices 120, 130, 140, and 150 may submit a formulated query to one or more databases located on such a network or the World Wide Web to retrieve relevant information. In an exemplary embodiment, such a database may correspond to a search engine, such as an Internet search engine.
[0023] However, it should be understood that in certain scenarios, if user 110 lacks expertise on the subject, they will find it difficult to adequately formulate a query to obtain the desired information from an online search engine, even if such a search engine is accessible via devices such as 130, 140, and 150. For example, if user 110 has seen bird 160 and identified certain colors or other characteristics of the bird, they can formulate a suitable text query for the search engine to identify bird 160. However, if user 110 has only heard birdsong 162 but has not seen bird 160, they will find it difficult to formulate a suitable text query. User 110 may also encounter similar difficulties when presented with other types of sounds (e.g., unfamiliar or barely audible language spoken by a human speaker, unfamiliar music that user 110 expects to identify, etc.).
[0024] Therefore, it is desirable to provide a system that can automatically retrieve and present information related to sounds perceived by a user in his or her environment without requiring the user to explicitly formulate queries for such information.
[0025] In an exemplary embodiment, one or more devices in a user environment may receive audio input corresponding to a sound perceived by the user. For example, any or all of devices 120, 130, 140, and 150 may have audio input capabilities and may use their corresponding audio input mechanisms (e.g., the built-in microphone of active earphone 120, microphone 132 of smartphone 130, etc.) to capture bird calls 162. The received audio input may be transmitted from the receiving device to a central device, which may automatically formulate a query based on the received sound waveform and submit such a query to an online search engine (also referred to herein as an "online engine"). Based on the formulated query, the online engine may use the techniques described below to retrieve information identifying the bird 160 and specific characteristics of the bird call 162 received by the device.
[0026] The retrieved information can then be presented to user 110 through one or more presentation modalities, including, for example, synthesized speech audio via earpiece 120, and / or audio output by a speaker (not shown) present on any of devices 130, 140, 150, and / or visual presentation on any of devices 130, 140, 150 with an adaptive display. For example, as shown on the display of smartphone 130, a graphic and text 132 identifying the bird 160, along with other detailed textual descriptions 134, can be displayed.
[0027] The techniques for implementing systems with the capabilities described above are further described below. Figure 2 An exemplary sequence of functional blocks 200 illustrating certain aspects of this disclosure is shown in the figure. Note that... Figure 2 The examples are shown for illustrative purposes only and are not intended to limit the scope of this disclosure. For example, sequence 200 does not need to be performed by a single device, and the described operations can be distributed across devices. Furthermore, in alternative exemplary embodiments, any block in sequence 200 may be modified, omitted, or rearranged in different sequences. Such alternative exemplary embodiments are contemplated within the scope of this disclosure.
[0028] exist Figure 2 In block 210, an audio waveform is received by one or more devices. In an exemplary embodiment, such a device may include any device with audio input and audio digitization capabilities that can communicate with other devices. For example, in scenario 100, such a device may include any or all of devices 120, 130, 140, and 150.
[0029] At box 220, the digital audio waveform is processed to recover and / or retrieve relevant information. In an exemplary embodiment, the digital audio waveform may be processed in conjunction with other input data, such as parameters related to the user profile, for example, the user's usage patterns to which subsequent information will be presented, the device's geographic location determined by the Global Positioning System (GPS) and / or other technologies, and other parameters.
[0030] In an exemplary embodiment, the processing at block 220 may include associating one or more digital sound waveforms with an online repository of sounds or sound models to identify one or more features of the sound waveforms. For example, in exemplary scenario 100 where user 110 hears birdsong 162, the sound waveforms received by each device may correspond to, for example, a first audio version of birdsong 162 received by earbud 120, a second audio version of birdsong 162 received by smartphone 130, a third audio version of birdsong 162 received by smartwatch 140, and so on.
[0031] In an exemplary embodiment, the digital waveform may be transmitted to a single processing unit, such as running on any of devices 120, 130, 140, or 150. In an alternative exemplary embodiment, the digital audio waveform may be transmitted to an online engine, such as those further described below, for example, directly or via an intermediate server or processor running on any of devices 120, 130, 140, 150, or any other device. In an exemplary embodiment, one or more digital audio waveforms may be included in a digital audio-enabled query against an online engine, and relevant information may be recovered and / or retrieved from, for example, the World Wide Web using online search engine technologies.
[0032] It should be understood that relevant information can correspond to any type of information that an online search engine classifies as relevant to a query. For example, relevant information may include an identifier of the characteristics of a sound waveform (e.g., "The bird song you are listening to is sung by a goldfinch"), other relevant information (e.g., "Goldfinches inhabit certain areas of Northern California during the summer"), the geographical origin of the received sound waveform (e.g., "The song of a goldfinch originates from 100 feet northwest"), and triangulations of sounds received from multiple devices, such as devices 120, 130, 140, 150, etc., as further described below.
[0033] At box 230, output sound waveforms and / or visual data can be synthesized to present the results of the processing at box 220 to the user. In an exemplary embodiment, the output sound waveform may include an artificially synthesized version of the information to be presented, such as, "The birdsong you are listening to is sung by a goldfinch...". In an exemplary embodiment, the visual data may include relevant text or graphic data to be presented to the user on a device with a display. The sound waveforms and / or visual data may be synthesized, for example, by an online engine described below, or such data may be synthesized locally by a device available to the user, etc.
[0034] At box 240, a user's local sound generator can be used to output synthesized sound waveforms, and / or a visual display of the user's local device can be used to output synthesized visual data. In an exemplary embodiment, an active earphone 120 can be used to output synthesized sound waveforms. For example, in scenario 100, assuming user 110 hears a bird song 162 sung by bird 160 in real time, the active earphone 120 can output synthesized text-to-speech rendering of information related to the bird song 162, such as, "The bird song you are listening to is sung by a goldfinch located 100 feet northwest of your current location," etc.
[0035] Figure 3An exemplary embodiment 300 illustrating operations performed by an audio input and / or output device locally available to the user is illustrated. In this exemplary embodiment, such a device may correspond to, for example, an active earphone 120, or generally to any of the devices 130, 140, and 150. Note that... Figure 3 The techniques described herein are shown for illustrative purposes only and are not intended to limit the scope of this disclosure to any particular implementation of the techniques described herein.
[0036] exist Figure 3 In this context, the input sound waveform from the user's environment is represented by a curve 301a corresponding to the sound pressure time curve of the waveform. The sound waveform 301a is received by an audio input and / or output device 310, which has a front-end stage corresponding to the sound transducer / digital converter block 320.
[0037] Box 320 converts the sound waveform 301a into a digital sound waveform 320a.
[0038] Box 322 performs operations that result in the recovery or retrieval of relevant information from the digital audio waveform 320a. Specifically, box 322 may send the received digital audio waveform to the central processing unit (CPU). Figure 3 (Not shown in the text) or an online engine, and if the device 310 is capable of presenting information to the user, it may optionally receive relevant information 322a from the central processing unit or the online engine. See below for reference. Figure 4 Specific exemplary operations performed by box 322 (e.g., in conjunction with other modules with which box 322 communicate) are further described.
[0039] In an exemplary embodiment, device 310 may optionally include block 324 for synthesizing sound based on information retrieved from block 322. Block 324 may include, for example, a text-to-speech module for locally synthesizing artificial speech waveforms from the information for presentation to a user. In an alternative exemplary embodiment, block 324 may be omitted, and text-to-speech synthesis of the information may be performed remotely from device 310, for example, by an online engine. In this case, the retrieved information 322a may be understood to already contain synthesized sound information to be presented. Such alternative exemplary embodiments are contemplated within the scope of this disclosure.
[0040] At box 326, speaker 326 generates audio output 301b based on synthesized sound information received, for example, from box 322 or from box 324. Audio output 301b may correspond to an output sound waveform played back to the user.
[0041] Given the description above, it should be understood that the user of device 310 can simultaneously perceive audio from two sources: an input sound waveform 301a from the user's "real" (external device 310) environment and an output sound waveform 301b from the speaker 326 of device 310. In this sense, the output sound waveform 301b can be understood as "overlapping" 305 or "enhancing" the input sound waveform 301a.
[0042] Figure 4 Further aspects of an audio augmented reality system for retrieving and recovering relevant information are illustrated, for example, the operations described in reference boxes 220 and 322 above are described in more detail. Note that... Figure 4 This description is for illustrative purposes only and is not intended to limit the scope of this disclosure to any particular implementation or functional division of the described boxes. In some exemplary embodiments, Figure 4 One or more functional blocks or modules shown (e.g., computer 420 and any device 310.n) can be integrated into a single module. Instead, as shown, functions performed by a single module can be divided across multiple modules. Such alternative exemplary embodiments are contemplated within the scope of this disclosure.
[0043] exist Figure 4 The diagram illustrates multiple audio input / output devices 310.1 to 310.N, and each device may have the same features as those described in the reference above. Figure 3 The architecture of the described device 310 is similar. Specifically, for each device 310.n (where n represents a general index from 1 to N), a block 322, previously described above, is depicted that performs operations resulting in the recovery or retrieval of relevant information from the corresponding digital audio waveform 320a. Specifically, block 322.1, corresponding to block 322 for the first audio input / output device 310.1, performs digital audio processing functions and includes a communication receive and transmit (RX / TX) module 410.1. Similarly, block 322.n, corresponding to block 322 for the nth audio input / output device 310.n, performs its own digital audio processing functions and includes a communication receive and transmit (RX / TX) module 410.n, and so on. Communication between blocks 322.n and 422 can be performed via channel 322.na.
[0044] In an exemplary embodiment, block 322.1 may correspond to block 322 for earbud 120, block 322.2 may correspond to block 322 for smartphone 130, and so on. In an exemplary embodiment where N equals 1, only one block 322.1 may exist in the system. Such alternative exemplary embodiments are contemplated within the scope of this disclosure.
[0045] exist Figure 4In this embodiment, each module 410.n corresponding to block 322.n communicates with one or more other entities remote from that corresponding device 310.n, for example, via a wireless or wired channel. In an exemplary embodiment, communication may occur between each module 410.n and the communication module 422 of the computer 420. The computer 420 may correspond to a central processing unit for processing audio input and / or other input signals from devices 310.1 to 310.N. Specifically, the computer 420 may include a multi-channel signal processing module 425 for jointly processing audio input signals and / or other data received from blocks 322.1 to 322.N.
[0046] In an exemplary embodiment, the multi-channel signal processing module 425 may include an information extraction / retrieval box 428. Box 428 may extract information from multiple received audio input signals and / or other data. Box 428 may include a query formulation box 428.1, which formulates a query 428.1a based on received digital sound waveforms and / or other data. Box 428 may further include a result retrieval box 428.2, which retrieves results in response to a query 428.1a from the online engine 430.
[0047] In an exemplary embodiment, block 428.1 is configured to formulate query 428.1a by concatenating multiple digital sound waveforms. In this sense, formulating query 428.1a also represents a query that enables digital sound, i.e., a query that includes digital sound waveforms as one or more query search terms. For example, referring to scenario 100, query 428.1a may include multiple digital sound waveforms as query search terms, where each digital sound waveform is encapsulated as a standard audio file (such as mp3, wav, etc.). Each digital sound waveform may correspond to a sound waveform received by one of devices 120, 130, 140, or 150. In exemplary scenario 100, where bird call 162 is received by each of devices 120, 130, 140, and 150, then formulating query 428.1a may include up to four digital sound waveforms corresponding to versions of bird call 162 received by each of the four devices. In an alternative exemplary embodiment, any number of digital sound waveforms may be concatenated by block 428.1 to generate formulating query 428.1a.
[0048] When processing queries that enable digital sound, the online engine 430 can be configured to retrieve and rank online results based on the similarity or correspondence between the online results and one or more digital sound waveforms included in the query. In an exemplary embodiment, the relevance of the digital sound waveforms to sound records in the online database can be determined at least in part based on sound pattern recognition and matching techniques, and can utilize techniques known in the fields of speech recognition, sound recognition, and pattern recognition. For example, one or more relevance measures between the recorded sound and candidate sounds can be calculated. In an exemplary embodiment, this calculation can be further informed by knowledge of other parameters included in the formulaic query 428.1a described above.
[0049] In an exemplary embodiment, other data included in the formula query 428.1a may include, for example, annotations for each digital sound waveform, which have data identifying the device that captured the sound waveform and / or describing the environment in which the sound waveform was captured. For example, the version of bird call 162 captured by smartphone 130 may be annotated with data identifying the hardware model / version number of smartphone 130 and the location data of smartphone 130 (e.g., derived from the GPS component of smartphone 130), the relative location data of smartphone 130 with other devices 120, 140, 150, etc., the speed of smartphone 130, the ambient temperature measured by the temperature sensor of smartphone 130, etc. When included as part of formula query 428.1a, such data can be used by an online engine to more accurately identify bird call 162 and retrieve more relevant information.
[0050] In an exemplary embodiment, the formula query 428.1a may further include data other than the audio waveform and data describing such waveform. For example, such data may include parameters such as the user's user profile and / or usage patterns, the geographical location of the device determined by the Global Positioning System (GPS) and / or other technologies, the device's location relative to each other, and other parameters.
[0051] To facilitate the identification and matching of submitted query sounds with relevant online results, the online engine 430 may maintain a sound index 434. The index 434 may include, for example, a list of online-accessible sound models and / or sound categories considered relevant and / or useful in satisfying search queries containing sound files.
[0052] In an exemplary embodiment, query formulation box 428.1 may record information received from device 310.n (e.g., audio and non-audio) to help evaluate and predict query formulations that may be useful to the user. In an exemplary embodiment, box 428.1 may include an optional machine learning module (not shown) that learns to map inputs received from device 310.n to relevant query formulations with increasing accuracy over time.
[0053] A formulaic query 428.1a is submitted from computer 420 to online engine 430, for example, via a wired or wireless connection. In an exemplary embodiment, online engine 430 may be an online search engine accessible via the Internet. Online engine 430 may retrieve relevant results 430a in response to query 428.1a. Subsequently, results 430a may be transmitted back to computer 420 by online engine 430, and computer 420 may then transmit the results back to any of devices 120, 130, 140, and 150.
[0054] In an exemplary embodiment, a user may specifically specify one or more sounds to be included in a search query. For example, while listening to birdsong 162, user 110 may explicitly instruct the system (e.g., via voice command, gesture, text input, etc.) that a query will be formulated and submitted based on the received sound input, for example, immediately after hearing a sound of interest or within a predetermined time. In an exemplary embodiment, this explicit instruction may automatically trigger block 428.1a to formulate the query. In an exemplary embodiment, user 110 may further explicitly specify that all or part of the query string will be included in the formulated query.
[0055] In an alternative exemplary embodiment, user 110 does not need to explicitly instruct the query to be formulated and submitted based on the received voice input. In such an exemplary embodiment, an optional machine learning module (not shown) can "learn" appropriate trigger points to automatically formulate the machine-generated query 428.1a based on the received accumulated data.
[0056] In an exemplary embodiment, the online engine 430 may include a machine learning module 432, which learns to map query 428.1a to relevant results with increasing accuracy over time. Module 432 may employ techniques derived from machine learning, such as neural networks, logistic regression, decision trees, etc. In an exemplary embodiment, channels 322.1a to 322.Na may transmit certain training information useful for training the machine learning module 432 of the engine 430 to the engine 430. For example, user identity may be transmitted to the machine learning module 432. Previously received audio waveforms and / or search results corresponding to such audio waveforms may also be transmitted to module 432. Such received data can be used by the online engine 430 to train the machine learning module 432 to better process and provide query 428.1a.
[0057] As an illustrative example, user 110 in scenario 100 may have a corresponding user identity, for example, associated with the user alias "anne123". The user alias anne123 may be associated with a corresponding user profile, for example, identifying previous search history, user preferences, etc. Assuming that such information can be used to train the machine learning module 432 of search engine 430, search engine 430 can advantageously provide more relevant and accurate results for submitted queries.
[0058] For example, in response to a query submitted by anne123 that includes a digital sound waveform derived from bird song 162, search engine 430 may rank certain search results related to “Goldfinch” higher based on knowledge derived from the user profile that user anne123 resides in a specific geographical vicinity. Note that the foregoing discussion is provided for illustrative purposes only and is not intended to limit the scope of this disclosure to any particular type of information or techniques for processing and / or determining patterns in such information that can be adopted by machine learning module 432.
[0059] Figure 5 An illustrative correspondence is shown between exemplary formulaic query 428.1a and exemplary search result 430a. Note that... Figure 5 The information shown is for illustrative purposes only and is not intended to limit the scope of this disclosure to any particular type of sound, technical field, query formula, query field, length or size of the formulaic query, number or type of results, etc. Note that any information shown may be omitted from any particular search query or any exemplary embodiment, depending on the specific configuration of the system. Furthermore, additional query fields not shown can be readily included in any particular search query or any exemplary embodiment and can be readily adapted using the techniques of this disclosure. Such alternative exemplary embodiments are contemplated within the scope of this disclosure.
[0060] Specifically, exemplary formulaic query 428.1a includes, for example, Figure 5 The left side 501 shows several fields. Specifically, the exemplary formula query 428.1a includes a first field 510, which includes a digital sound waveform 510b encapsulated as an mp3 audio file 510a. The first field 510 also includes additional attributes represented as device 1 attributes 510c, which correspond to other parameters of the device used to capture the digital sound waveform 510b, including device type, sampling rate, geographic location, etc. Note that specific attributes are described herein for illustrative purposes only and are not intended to limit the scope of this disclosure. The exemplary formula query 428.1a further includes a second field 511 and a third field 512, having similar fields including the sound waveform encapsulated as an audio file, as well as additional device attributes.
[0061] Query 428.1a also includes additional parameter fields 513 that can help the online engine 430 retrieve more relevant search results. For example, field 513 may specify the identity of the user (e.g., to whom the retrieved information is to be presented), such a user's profile, such a user's previous search history, ambient temperature (e.g., measured by one or more devices), etc.
[0062] After submitting query 428.1a to online engine 430, query result 430a can be provided in response to query 428.1a. An example query result 430a is shown in... Figure 5 551 on the right. Note that... Figure 5 The exemplary query result 430a is shown for illustrative purposes only and is not intended to limit the scope of this disclosure.
[0063] Specifically, the exemplary query result 430a includes one or more visual results 560, including, for example, a graphic 561 and text 562 describing the result 560. For example, for in Figure 5 The exemplary digital sound waveform and other parameters submitted in illustrative query 428.1, a graphic 561 of a goldfinch, and corresponding text 562 describing the goldfinch's behavior are shown. In an exemplary embodiment, the visual result 560 may be displayed on a user's local device with visual display capabilities, for example, as referenced below. Figure 6 Further description.
[0064] Query result 430a may further or alternatively include one or more results containing content not intended to be visually displayed to the user. For example, query result 430a may include audio result 563, which includes a digital sound waveform representing the speech of text 564 related to search query 428.1a. Audio result 563 may be a computer-generated text-to-speech reproduction of the corresponding text, or it may be read by a human speaker, etc. In an exemplary embodiment, any audio result in query result 430a may be played back using a user's local device (e.g., user 110's local earphone 120, etc.).
[0065] Audio result 563 may further or alternatively include personalized audio result 565 corresponding to a digital sound waveform customized for the user. For example, in the exemplary embodiment shown, the user's favorite song 566 (e.g., determined by user profile parameters submitted in query 428.1a or elsewhere) may be mixed with a goldfinch song 568 (e.g., any digital sound waveform from waveform 510b submitted in query 428.1a, or a bird song extracted from the digital sound waveform associated with audio result 563 or any other result in query result 430a) 567.
[0066] In an exemplary embodiment, user feedback can be received in the audio augmented reality system to train a machine learning algorithm running in the online engine 430 to retrieve results with increased relevance to the formulaic query. For example, when presented with either a visual result 560 or an audio result 563 (including a personalized audio result 565), the user 110 can choose one of the presented results to retrieve additional information related to the result. For example, when viewing text 562 in the visual result 560, the user 110 can express interest in learning more about the goldfinch migration by submitting another query for “goldfinch migration” to the online engine 430, for example, via an available device, or by otherwise indicating that result 562 is considered relevant by the user. Alternatively, when listening to the synthesized speech presentation 564 of the audio result 563, the user 110 can express interest in the synthesized audio information by, for example, increasing the volume of the audio output or otherwise submitting additional queries related to the retrieval result (e.g., via voice commands or multiple entries of additional text). Upon receiving user feedback indicating a positive relevance to the search results, the online engine 430 can further adjust and / or train a basic machine learning algorithm, such as that executed by the machine learning module 432, to retrieve relevant results in response to a formulaic query.
[0067] Figure 6 The illustration depicts another aspect of an augmented reality system used to visually display information related to received digital sound waveforms. Note that... Figure 6 This disclosure is described for illustrative purposes only and is not intended to limit the scope of this disclosure to any particular implementation or functional division of the described box.
[0068] exist Figure 6 In, for example, with reference Figure 4 The computer 420 described herein is further communicatively coupled to one or more devices 610.1 to 610.M, each having a visual display, any of which is referred to herein as 610.m. Specifically, any of the devices 610.1 to 610.M may correspond to any of the devices 130, 140, and 150 mentioned earlier above, provided that such devices have a visual display, such as smartphone 130, smartwatch 140, laptop computer 150, etc. Alternatively, any of the devices 610.1 to 610.M may be a standalone device that does not have audio input capability or otherwise does not receive any audio input for use by the augmented reality system. Each device 610.m includes a communication RX / TX box 620 for communicating, for example, directly or via another intermediate device (not shown) such as a router, server, etc., with box 422 of computer 420.
[0069] In an exemplary embodiment, computer 420 includes a visual information presentation box 630 coupled to a result retrieval box 428.2. Specifically, retrieval results 430a may be formatted or otherwise collected for visual presentation and display by box 630, which transmits the formatted and / or collected results to devices 610.1 to 610.M via communication box 422 for visual display. For example, in the case where device 610.1 corresponds to a laptop computer with a display, box 630 may be based on... Figure 5 The visual result 560 shown is used to format one or more search results, and then such formatted results are sent to a laptop computer for display.
[0070] Figure 7 The illustration shows an exemplary application of the techniques described above with reference to the audio augmented reality system to a specific scenario 100. Note that... Figure 7 The information is shown for illustrative purposes only and is not intended to limit the scope of this disclosure to any particular scenario shown.
[0071] exist Figure 7 In the diagram, at frame 210.1, multiple devices 120, 130, and 140 receive sound waveforms, for example, corresponding to birdsong 162. Devices 120, 130, and 140 digitize the received sound waveforms into digital sound waveforms.
[0072] At block 220.1, a digital audio waveform is sent to the central processing unit by any or all of devices 120, 130, and 140 for remote processing. Note that the central processing unit may be separate from devices 120, 130, and 140, or it may be implemented on one or more of devices 120, 130, and 140. In an exemplary embodiment, the central processing unit may perform operations as described in the reference... Figure 4 The functions described in the computer 420.
[0073] In an exemplary embodiment, for example, as described in reference box 428.1 above, and also as Figure 7 As shown, query formulation can be performed by computer 420. Specifically, at box 428.1, computer 420 can formulate a query and submit it to online engine 430 to perform an online search. The query may include digital sound waveforms and other data, such as those referenced above. Figure 5 As stated above.
[0074] like Figure 7 As shown, the exemplary online engine 430.1 can receive a formulated and submitted query from box 428.1. The online engine 430.1 may include a machine learning module 432.1 configured to map query 428.1a to relevant results with increasing accuracy.
[0075] In the specific scenario shown, module 432.1 is specifically configured to use acoustic triangulation techniques to estimate the origin of the sound waveforms received by devices 120, 130, and 140. Specifically, assuming the same bird call 162 generates three different sound waveforms corresponding to sound waveforms received at three separate devices, digital sound waveforms can be used to perform triangulation to determine the location of the bird 160 relative to the devices and therefore relative to the user.
[0076] For example, sound triangulation can explain the relative delay of the bird call 162 within each digital sound waveform (e.g., assuming each device is equipped with an accurate time reference that can be derived from a GPS signal), the frequency shift of the received sound due to the movement of the source (e.g., bird 160) or devices 120, 130, 140, etc.
[0077] Based on the sound triangulation described above, the machine learning module 432.1 can be configured to triangulate the source of the bird call 162, and thus the location of the bird 160 relative to the user. The machine learning module 432.1 can be further configured to extract a standard version of the bird call 162 from the multiple received versions, for example, by taking into account any calculated frequency shifts and delays. This standard version of the bird call 162 can then be associated with sound models or samples that may be available on the World Wide Web (WWW) 440, for example, as referenced by the sound index 434 of the online engine 430.1, as previously mentioned above. Figure 4 As described. After association, birdcall162 can be identified as corresponding to one or more specific bird species. This information can then be further used to extract relevant information as query results, such as... Figure 5 The result shown is 430a.
[0078] Based on the retrieved information, sound synthesis can be performed at box 710, and visual synthesis can be performed at box 712. For example, an exemplary visual result can be as shown in the reference. Figure 5 The results described in 560 are as follows, while exemplary audio results may be described as in reference results 563 or 565. In an exemplary embodiment, blocks 710, 712 may be implemented individually or jointly at the online engine 430.1 or at the computer 420.
[0079] After sound synthesis at box 710, the synthesized sound can be output to the user, for example, via earpiece 120, at box 240.1. In an exemplary embodiment, for example, while user 110 listens to birdsong 162, he or she listens to the output of earpiece 120, the synthesized sound output of earpiece 120 constituting an audio augmented reality, wherein the user receives real-time synthesized audio information related to sounds that are otherwise naturally perceived through the environment.
[0080] Following the visual composition at box 712, the composed visual information can be output to the user, for example, via smartphone 140, at box 240.2. The composed visual information can identify the bird 160 to the user, as well as provide other relevant information.
[0081] Figure 8 An exemplary embodiment of method 800 according to this disclosure is illustrated. Note that... Figure 8 The methods shown are for illustrative purposes only and are not intended to limit the scope of this disclosure to any particular method shown.
[0082] exist Figure 8In block 810, a query including a first digital audio waveform from a first source and a second digital audio waveform from a second source is received. In an exemplary embodiment, the first source and the second source may correspond to different audio input devices having separate locations, for example, referring to... Figure 1 Any two of the multiple devices described, 120, 130, 140, 150, etc. A first digital audio waveform can be recorded by a first source, and a second digital audio waveform can be recorded by a second source.
[0083] At box 820, at least one online result related to both the first and second digital sound waveforms is retrieved. In an exemplary embodiment, the first and second digital sound waveforms correspond to different recordings of the same sound event received from different sources, such as a separate digital sound recording of birdsong 162.
[0084] At box 830, generate a synthesized sound corresponding to at least one online result.
[0085] At box 840, the generated synthesized sound is provided in response to the received query.
[0086] Figure 9 An exemplary embodiment of the apparatus 900 according to this disclosure is shown. Figure 9 In the device 900, there are query processing modules 910 configured to receive queries including a first digital audio waveform from a first source and a second digital audio waveform from a second source, a search engine 920 configured to retrieve at least one online result related to both the first and second digital audio waveforms, a synthesis module 930 configured to generate a synthesized sound corresponding to the at least one online result, and a transmission module 940 configured to provide the generated synthesized sound in response to the received query.
[0087] In an exemplary embodiment, the structures for implementing module 910, search engine 920, module 930, and module 940 may correspond to one or more server computers, for example, operating remotely from devices used to capture the first and second digital audio waveforms and communicating with these devices, for example, via a network connection using the Internet. In an alternative exemplary embodiment, the structures for implementing module 910 and search engine 920 may correspond to one or more server computers, while the structures for implementing modules 930 and 940 may correspond to one or more processors residing on one or more devices used to capture the first and second digital audio waveforms. Specifically, the generation of synthesized sound may be performed at a server and / or local device. Such alternative exemplary embodiments are contemplated within the scope of this disclosure.
[0088] Figure 10An exemplary embodiment of the device 1000 according to this disclosure is shown. Figure 10 In the computing device 1000, a memory 1020 is included, which stores instructions executable by a processor 1010 to perform the following operations: receiving a query including a first digital sound waveform from a first source and a second digital sound waveform from a second source; retrieving at least one online result related to both the first and second digital sound waveforms; generating a synthesized sound corresponding to the at least one online result; and providing the generated synthesized sound in response to the received query.
[0089] In this specification and claims, it should be understood that when an element is referred to as "connected to" or "coupled to" another element, it may be directly connected to or coupled to the other element, or there may be intermediate elements. Conversely, when an element is referred to as "directly connected to" or "directly coupled to" another element, there are no intermediate elements. Furthermore, when an element is referred to as "electrically coupled" to another element, it indicates that there is a low-resistance path between these elements, while when an element is referred to as simply "coupled" to another element, a low-resistance path may or may not exist between these elements.
[0090] The functions described herein can be performed, at least in part, by one or more hardware and / or software logic components. For example, and not as a limitation, illustrative types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.
[0091] While the invention is readily adaptable to various modifications and alternative constructions, certain illustrated embodiments are shown in the accompanying drawings and have been described in detail above. However, it should be understood that the invention is not intended to be limited to the specific forms disclosed, but rather, it is intended to encompass all modifications, alternative constructions, and equivalents falling within the spirit and scope of the invention.
Claims
1. A method for collecting sound waveforms and generating a response synthesized sound output, the method comprising: Receive a query, the query including a first digital audio waveform from a first source and a second digital audio waveform from a second source, captured by at least one microphone, the query also including user preference parameters; Using a network-connected device, retrieve the audio file representing the user's preferences identified based on the preference parameters, as well as at least one online result related to both the first digital audio waveform and the second digital audio waveform; The sound is synthesized using computer sound synthesis hardware, the sound comprising the user’s preferred sound file and an audio mix of at least one online result related to both the first digital sound waveform and the second digital sound waveform; as well as The synthesized sound is provided in response to the received query.
2. The method according to claim 1, wherein the received query comprises at least three digital audio waveforms, each audio waveform being received by a device at a different location, and the retrieval of the at least one online result comprises: The location of the source object is calculated based on the at least three digital audio waveforms, the calculation including determining the relative delay between the at least three digital audio waveforms and determining the relative location of the different positions of the device.
3. The method of claim 2, wherein synthesizing sound includes generating a computed text-to-speech reproduction of the location.
4. The method according to claim 1, wherein retrieving the at least one online result comprises: Each of the first and second digital audio waveforms is associated with at least one online audio file, and each audio file has corresponding identification information. The at least one online result includes identification information corresponding to the audio file that is most highly correlated with the first digital audio waveform and the second digital audio waveform.
5. The method according to claim 1, wherein the user's preference parameter includes the user's preferred music type, the preferred sound file includes a music file, and the music file contains music of the preferred music type.
6. The method of claim 5, wherein the at least one online result comprises a sound file corresponding to an object characterized by the first digital sound waveform and the second digital sound waveform, and the synthesized sound comprises an audio mix of the sound file and the music file corresponding to the object.
7. The method according to claim 1, further comprising: Generate a synthetic visual output corresponding to the at least one online result; as well as The generated synthetic visual output is provided in response to the received query.
8. The method according to claim 1, further comprising: Receive user approval for the provided synthesized sound; as well as The machine learning algorithm used to retrieve at least one online result related to the received query is updated based on the user-approved received instruction.
9. The method of claim 1, wherein the first digital sound waveform and the second digital sound waveform comprise at least one digital sound waveform captured by an active earpiece, the active earpiece being further configured to receive and reproduce the synthesized sound provided in response to the received query.
10. The method of claim 1, wherein the at least one digital audio waveform comprises: At least one digital audio waveform received by a smartphone; as well as At least one digital sound waveform received by a smartwatch.
11. An apparatus for collecting sound waveforms and generating a response synthesized sound output, the apparatus comprising: A query processing module is configured to receive a query, the query including a first digital audio waveform from a first source and a second digital audio waveform from a second source, captured by at least one microphone, and the query also includes user preference parameters; The search engine is configured to use a network-connected device to retrieve audio files identified based on the user's preferences according to the preference parameters, as well as at least one online result related to both the first digital audio waveform and the second digital audio waveform; A computer sound synthesis module is configured to generate synthesized sound, the synthesized sound comprising the user’s preferred sound file and an audio mix of at least one online result related to both the first digital sound waveform and the second digital sound waveform; as well as The transmission module is configured to provide the synthesized sound in response to the received query.
12. The apparatus of claim 11, wherein the received query comprises at least three digital audio waveforms, each audio waveform being received by a device at a different location, and the search engine is further configured to: The location of the source object is calculated based on the at least three digital audio waveforms, the calculation including determining the relative delay between the at least three digital audio waveforms and determining the relative location of the different positions of the device.
13. The apparatus of claim 12, wherein the synthesis module is further configured to generate computed text-to-speech reproduction of the location.
14. The apparatus of claim 11, wherein the search engine is further configured to retrieve the at least one online result by associating each of the first digital sound waveform and the second digital sound waveform with at least one online sound file, each sound file having corresponding identification information; the at least one online result includes identification information corresponding to the sound file most highly correlated with the first digital sound waveform and the second digital sound waveform.
15. The apparatus of claim 11, wherein the user's preference parameter includes the user's preferred music type, the preferred sound file includes a music file containing music of the preferred music type.
16. The apparatus of claim 15, wherein the at least one online result comprises a sound file corresponding to an object characterized by the first digital sound waveform and the second digital sound waveform, and the synthesized sound comprises an audio mix of the sound file and the music file corresponding to the object.
17. The apparatus of claim 11, wherein the first digital sound waveform and the second digital sound waveform comprise at least one digital sound waveform captured by an active earpiece, the active earpiece being further configured to receive and reproduce the synthesized sound provided in response to the received query.
18. The apparatus of claim 11, wherein the at least one digital audio waveform comprises: At least one digital audio waveform received by a smartphone; as well as At least one digital sound waveform received by a smartwatch.
19. A computing device for collecting sound waveforms and generating a response synthesized sound output, the computing device comprising a memory holding instructions executable by a processor to: Receive a query, the query including a first digital audio waveform from a first source and a second digital audio waveform from a second source, captured by at least one microphone, the query also including user preference parameters; Using a network-connected device, retrieve the audio file representing the user's preferences identified based on the preference parameters, as well as at least one online result related to both the first digital audio waveform and the second digital audio waveform; A computer sound synthesis module is used to generate a synthesized sound, the synthesized sound comprising the user’s preferred sound file and an audio mix of at least one online result related to both the first digital sound waveform and the second digital sound waveform; as well as The synthesized sound is provided in response to the received query.
20. The device of claim 19, wherein the memory also stores instructions executable by the processor to: Receive user approval instructions for the provided synthesized sound; and The machine learning algorithm used to retrieve at least one online result related to the received query is updated based on the user-approved received instruction.