Presentation of sound associated with a selected target object external to the device
By detecting the selection of external target objects in the Internet of Vehicles system and starting the communication channel with the associated devices, receiving and decoding audio packets, multi-party conversations between vehicles and spatial presentation of audio signals is realized, and the problem of audio signal transmission delay and reliability in the Internet of Vehicles system is solved, and the safety and user experience of autonomous driving cars are improved.
Patent Information
- Application Number
- CN201980082469.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-19
- Filing Date
- 2019-12-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-12-20
AI Technical Summary
The prior art is difficult to effectively realize the presentation of external target objects associated with sound in the Internet of Vehicles system, especially in multi-vehicle communication scenarios, and how to accurately share and present sensor data and audio signals is a challenge.
By detecting the selection of an external target object in the first device and starting a communication channel with the associated second device, receiving and decoding an audio packet, an audio signal is output based on the selected target object. This technology uses C-V2X or V2X communication system to realize direct audio data transmission between vehicles.
It realizes multi-party conversations between vehicles and spatial presentation of audio signals, improves the safety of autonomous vehicles and the user experience of infotainment systems, and solves the problems of audio signal transmission delay and reliability in traditional technologies.
Smart Images

Figure CN113196795B_ABST
Abstract
Description
[0001] Priority Claim under 35 U.S.C.§119
[0002] This patent application claims priority to non - provisional application No. 16 / 720,639, entitled "RENDERING OF SOUNDS ASSOCIATED WITH SELECTED TARGET OBJECTS EXTERNAL TO A DEVICE", filed on December 19, 2019, and provisional application No. 62 / 783,887, entitled "RENDERING OF SOUNDS ASSOCIATED WITH SELECTED TARGET OBJECTS EXTERNAL TO A DEVICE", filed on December 21, 2018. These applications are assigned to the assignee of this application and are hereby incorporated herein by reference in their entirety. Technical Field
[0003] This application relates to rendering sounds associated with selected target objects external to a first device. Background Art
[0004] The following generally relates to wireless communication and, more specifically, to vehicle - to - everything (V2X) control channel design.
[0005] Wireless communication systems are widely deployed to provide various types of communication content, such as voice, video, packet data, messaging, broadcasts, etc. These systems are capable of supporting communication with multiple users by sharing available system resources (e.g., time, frequency, and power). Examples of such multi - access systems include code - division multiple - access (CDMA) systems, time - division multiple - access (TDMA) systems, frequency - division multiple - access (FDMA) systems, and orthogonal frequency - division multiple - access (OFDMA) systems (e.g., Long - Term Evolution (LTE) systems or New Radio (NR) systems).
[0006] A wireless multi - access communication system may include multiple base stations or network access nodes, each of which simultaneously supports communication for multiple communication devices, which may also be referred to as user equipment (UE). Additionally, a wireless communication system may include a network that supports communication - based vehicles. For example, vehicle - to - vehicle (V2V) and vehicle - to - infrastructure (V2I) communications are wireless technologies that enable the exchange of data between a vehicle and its surrounding environment. V2V and V2I are collectively referred to as vehicle - to - everything (V2X). V2X uses communication wireless links for fast - moving objects such as vehicles. Recently, the emergence of cellular vehicle - to - everything (C - V2X) has distinguished it from WLAN - based V2X.
[0007] The 5G Automotive Association (5GAA) has promoted C-V2X. C-V2X was initially defined in LTE Release 14 and is designed to operate in multiple modes: (a) vehicle-to-vehicle (V2V); (b) vehicle-to-infrastructure (V2I); and (c) vehicle-to-network (V2N). In 3GPP Release 15, C-V2X includes support for both V2V and traditional cellular network-based communications, and the functionality is extended to support the 5G air interface standard. The PC5 interface in C-V2X allows direct communication between vehicles and other devices (via the "sidelink channel") without using a base station.
[0008] A vehicle-based communication network can provide always-on telematics, in which a UE such as a vehicle UE (v-UE) communicates directly with the network (V2N), pedestrian UEs (V2P), infrastructure devices (V2I), and other v-UEs (e.g., via the network). A vehicle-based communication network can support a safe, always-connected driving experience by providing intelligent connectivity in which traffic signals / timings, real-time traffic and routes, safety alerts for pedestrians / cyclists, collision avoidance information, etc. are exchanged.
[0009] However, such networks that support vehicle-based communications can also be associated with various requirements, such as communication requirements, safety and privacy requirements, etc. Other example requirements can include, but are not limited to, the requirement for reduced latency, the requirement for higher reliability, etc. For example, vehicle-based communications can include transmitting sensor data that can support autonomous vehicles. The sensor data can also be used between vehicles to improve the safety of autonomous vehicles.
[0010] V2X and C-V2X allow for a variety of applications, including the applications described in this disclosure. SUMMARY
[0011] Generally, this disclosure describes techniques related to presenting a sound associated with a selected target object external to a first device. In one example, this disclosure describes a first device for initiating communication with a second device, the first device including one or more processors configured to detect a selection of at least one target object external to the first device and initiate a communication channel between the first device and a second device associated with at least one target object external to the first device. The one or more processors can be configured to receive an audio packet from the second device in response to a selection of at least one target object external to the first device; decode the received audio packet from the second device to generate an audio signal; and output the audio signal based on the selection of at least one target object external to the first device. The first device can also include a memory coupled to the one or more processors and configured to store the audio packet.
[0012] In one example, the present disclosure describes a method for initiating communication with a second device, the method comprising detecting a selection of at least one target object external to the first device; initiating a communication channel between the first device and a second device associated with at least one target object external to the first device; and in response to the selection of at least one target object external to the device, receiving an audio packet from the second device. The method further comprises decoding the audio packet received from the second device to generate an audio signal; and outputting the audio signal based on the selection of at least one target object external to the first device.
[0013] In one example, the present disclosure describes an apparatus, the apparatus comprising means for detecting a selection of at least one target object external to the first device; and means for initiating a communication channel between the first device and a second device associated with at least one target object external to the first device. The apparatus further comprises means for receiving an audio packet from the second device in response to the selection of at least one target object external to the device. The apparatus may further comprise means for decoding the audio packet received from the second device to generate an audio signal; and means for outputting the audio signal based on the selection of at least one target object external to the first device.
[0014] In one example, the present disclosure describes an apparatus, the apparatus comprising means for detecting a selection of at least one target object external to the first device; and means for initiating a communication channel between the first device and a second device associated with at least one target object external to the first device. The apparatus further comprises means for receiving an audio packet from the second device in response to the selection of at least one target object external to the device. The apparatus may further comprise means for decoding the audio packet received from the second device to generate an audio signal; and means for outputting the audio signal based on the selection of at least one target object external to the first device.
[0015] In one example, the present disclosure describes a non-transitory computer-readable medium storing computer-executable code that can be executed by one or more processors to detect a selection of at least one target object external to the first device and initiate a communication channel between the first device and a second device associated with at least one target object external to the first device. The code, when executed, can cause one or more processors to receive an audio packet from the second device in response to the selection of at least one target object external to the device; decode the audio packet received from the second device to generate an audio signal. The code, when executed, can cause one or more processors to output the audio signal based on the selection of at least one target object external to the first device.
[0016] Details of one or more examples of the present disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the various aspects of the technology will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1a A conceptual diagram showing a first device detecting a selection of another device and communicating with the other device (e.g., a second device).
[0018] Figure 1b A conceptual diagram showing a first device that can detect a selection of another device and communicate with the other device (e.g., a second device) with the assistance of a tracker, where audio communication can be spatialized.
[0019] Figure 1c A conceptual diagram showing different vehicles transmitting and receiving wireless connections according to the techniques described in the present disclosure.
[0020] Figure 1d A conceptual diagram showing different vehicles transmitting and receiving wireless connections using a cache server in the vehicle or a memory in the vehicle.
[0021] Figure 2 A flowchart showing a process in which a first device initiates communication with a second device according to the techniques described in the present disclosure.
[0022] Figure 3 A conceptual diagram showing a first vehicle having different components on or in the first vehicle operating according to the techniques described in the present disclosure.
[0023] Figure 4a A block diagram showing a first device having different components on or in the first device operating according to the techniques described in the present disclosure.
[0024] Figure 4b A block diagram showing a first device having different components on or in the first device operating according to the techniques described in the present disclosure.
[0025] Figure 5 A conceptual diagram showing transforming world coordinates to pixel coordinates according to the techniques described in the present disclosure.
[0026] Figure 6a A conceptual diagram showing an embodiment of estimating the distance and angle of a remote vehicle / passenger (e.g., a second vehicle).
[0027] Figure 6b A conceptual diagram showing the estimation of distance and angle in the x - y plane of a remote device.
[0028] Figure 6cConceptual diagram showing the estimation of distance and angle in the y-z plane of a remote device.
[0029] Figure 7a Shows an embodiment of an audio spatializer according to the techniques described in the present disclosure.
[0030] Figure 7b Shows an embodiment of an audio spatializer including a decoder according to the techniques described in the present disclosure.
[0031] Figure 8 Shows an embodiment in which the positions of persons in a first vehicle and a selected (remote) vehicle can be in the same coordinate system. Detailed Description
[0032] Certain wireless communication systems can be used to transmit data associated with high reliability and low latency. A non-limiting example of such data includes C-V2X and V2X communications. For example, autonomous vehicles can rely on wireless communication. Autonomous vehicles can include some sensors, such as, for example, light detection and ranging (LIDAR), radio detection and ranging (RADAR), cameras, etc., which are line-of-sight sensors. However, C-V2X and V2X communications can include line-of-sight and non-line-of-sight wireless communications. Current C-V2X and V2X communications are examples of using non-line-of-sight wireless communications to handle communications between vehicles that are approaching a common intersection but are not in each other's line of sight. C-V2X and V2X communications can be used to share sensor information between vehicles. This and other communication scenarios present certain considerations. For example, for a particular location or geographical area, several vehicles may sense the same information (e.g., an obstacle or a pedestrian). This raises questions such as which vehicle should broadcast such information (e.g., sensor data), how to share such information (e.g., which channel configuration provides reduced latency and improved reliability), etc.
[0033] The C-V2X communication system can have logical channels and transport channels. The logical channels and transport channels can be part of the uplink and downlink data transmission between a first device (e.g., a headset or a vehicle) and a base station or another intermediate node in the network. Those of ordinary skill in the art can recognize that the logical channels can include different types of control channels, e.g., xBCCH, xCCH, xDCCH. When the first device downloads broadcast system control information from another entity (e.g., a server or a base station), the xBCCH type of channel can be used. The xCCCH control channel can be used to send control information between the first device (e.g., a vehicle, a mobile device or a headset) and the network (e.g., a node in the network base station). When the first device (e.g., a vehicle, a mobile device or a headset) does not have a radio resource control connection with the network, the xCCCH control channel can be used. The xDCCH control channel includes control information between the first device and the network and is used by the first device having a radio resource control connection with the network. The xDCCH is also bidirectional, i.e., the control information can be sent and received by the first device and the network.
[0034] Generally, some of the information bits conveyed in the different types of control channels mentioned above can provide an indication of the location of the data channel (or resource). Since the data may span multiple subcarriers (depending on the amount of data being transmitted) and the size of the control channel is currently fixed, this can introduce a time / frequency transient or gap between the control channel and the corresponding data channel. This results in unused frequency / time resources of the control channel. It may be possible to utilize the unused frequency / time resources for other purposes of transmitting media between vehicles or between devices. It may also be possible to create new channels in the V2X or C-V2X system, specifically, for exchanging media between vehicles or between devices.
[0035] As mentioned above, vehicles use many advancements from other fields to improve their safety, infotainment systems, and overall user experience.
[0036] For example, object detection algorithms combined with sensors (e.g., RADAR, LIDAR or computer vision) can be used in vehicles to perform object detection while driving. These objects can include lanes in the road, stop signs, other vehicles or pedestrians. Some V2X and C-V2X use cases envision a cooperative V2X system warning the vehicle or the vehicle driver when a collision may occur between the vehicle and another object (e.g., a car, a bicycle or a person). Due to the relatively nascent nature of the V2X and C-V2X systems, many improvements have not been envisioned.
[0037] One area of improvement is communication between people in different vehicles. Although a person in one vehicle can communicate with another person in a different vehicle, the communication is done by making a phone call. The initiator of the phone call knows what phone number to dial to communicate with the other person and then dials it.
[0038] The present disclosure contemplates improvement in such a way that the device initiates a target object selection for a selected target object based on using a direct channel communication or a peer-to-peer connection, V2X, or C-V2X communication system, allowing communication or an auditory experience with other people or other devices.
[0039] For example, a first device for communicating with a second device may include one or more processors configured to detect a selection of at least one target object external to the first device and initiate a channel for communication between the first device and the second device associated with at least one target object external to the first device. Whether the selection of at least one target object external to the first device is performed first or the channel for communication between the first device and the second device associated with at least one target object external to the first device is initiated may not be important. It may depend on such background or circumstances, whether the channel has been established, and whether the initiation of the communication channel occurs or the initiation of the communication channel is based on the detection of the selection of at least one target object external to the first device.
[0040] For example, a channel for communication between the first device and the second device may have been established before the selection of at least one target object external to the device is detected. The channel for communication between the first device and the second device is initiated in response to the detection of the selection.
[0041] In addition, one or more processors in the first device may be configured to receive an audio packet from the second device as a result of a channel for communication between at least one target object external to the first device and the second device. Subsequently, after receiving the audio packet, one or more processors may be configured to decode the audio packet received from the second device to generate an audio signal; and output the audio signal based on the selection of at least one target object external to the first device. The first device and the second device may be a first vehicle and a second vehicle. The present disclosure has different examples illustrating vehicles, but many of the described techniques are also applicable to other devices. That is, the two devices may be headsets, including: mixed reality headsets, head-mounted displays, virtual reality (VR) headsets, augmented reality (AR) headsets, etc.
[0042] The audio signal may be reproduced by one or more speakers coupled to the first device. If the first device is a vehicle, the speakers may be in the vehicle's cabin. If the first device is a headset, the speakers may reproduce a binauralized version of the audio signal.
[0043] Based on the selection of the target object, communication can be performed between one or more target objects and the first device using C-V2X or V2X systems, or other communication systems. The second device (i.e., the headset or the vehicle) can have one or more people speaking or playing music associated with the second device. The voice or music emitted from inside the second vehicle or from the second headset can be compressed using an audio / voice codec and audio packets are generated. The audio / voice codec can be two separate codecs, such as an audio codec, or can be a voice codec. Alternatively, one codec can have the ability to compress both audio and voice.
[0044] Additional techniques and background are described herein with reference to the accompanying drawings.
[0045] Figure 1a A conceptual diagram of a first device that can communicate with another device (e.g., a second device) is shown. The conceptual diagram also includes the detection of the selection of another device within the first device. For example, the first device can be a first vehicle 303a that is capable of communicating with a second vehicle via a V2X or C-V2X communication system. The first vehicle 303a can include different components or a person 111 as shown by the circle 103 above. If the first vehicle 303a is self-driving, the person 111 may be driving, or the person 111 may not be driving. The person 111 can see other vehicles traveling on the road through the mirror 127 or window 132 of the first vehicle 303a and wishes to hear the type of music being played on the radio inside another vehicle. In some configurations of the first vehicle 303a, the camera 124 of the first vehicle 303a can assist the person 111 in seeing other vehicles, which may be challenging to see through the mirror 127 or window 132.
[0046] The person 111 can select at least one target object outside the vehicle, or if the person 111 is wearing a headset, the at least one target object is outside the headset. The target object can be the vehicle itself, i.e., the second vehicle can be the target object. Alternatively, the target object can be another person. The selection can be the result of an image detection algorithm encoded in instructions executed by a processor in the first vehicle. The image detection algorithm can be assisted by an external camera mounted on the first vehicle. The image detection algorithm can detect different types of vehicles or can detect only faces.
[0047] Additionally, or alternatively, person 111 may speak a descriptor to identify the target vehicle. For example, if the second vehicle is a black Honda Accord, the person may say "Honda Accord", "the black Honda Accord in front of me", "the Accord to my left", etc., and the speech recognition algorithm may be encoded in instructions executed on a processor in the first vehicle to detect and / or identify phrases or keywords (e.g., the make and model of the car). Thus, the first device may include a selection of at least one target object based on the detection of a command signal, which detection is based on keyword detection.
[0048] The processor that executes the instructions for the image detection algorithm may not have to be the same processor that executes the instructions for the speech recognition algorithm. If the processors are not the same, they may work independently or in a coordinated manner, e.g., to assist the other processor with image or speech recognition. One or more processors (which may include the same processors used in image detection or speech recognition), or different processors, may be configured to detect the selection of at least one target object of the first device. That is, one or more processors may be used to detect which target object (e.g., a face or other vehicle or headset) is selected. This selection may initiate communication between the first device and a second device (another vehicle or headset). In some cases, a channel for communication between the first device and the second device may already have been established. In some cases, the image detection algorithm may also incorporate aspects of image recognition, e.g., detecting a vehicle versus detecting "Honda Accord". For simplicity, in this disclosure, unless otherwise explicitly stated, the image detection algorithm may include aspects of image recognition.
[0049] As described above, when two people wish to communicate with each other and speak, one person calls the other by dialing a phone number. Alternatively, two devices may be wirelessly connected to each other, and if both devices are connected to a communication network, each device may register the Internet Protocol (IP) address of the other device. In Figure 1a this case, communication between the first device and the second device may also be established via the respective IP addresses of each device in a V2X, C-V2X communication network or a network that has the ability to directly (e.g., without using a base station) connect the two devices. However, unlike instant messaging, chat, or email, communication between the first device and the second device is initiated based on the selection of a target object associated with the second device or directly based on the selection of the second device itself.
[0050] For example, person 111 in vehicle 303a may see a second vehicle 303b or a different second vehicle 303c and may wish to initiate communication with a person in one of those vehicles based on image detection, image recognition, or speech recognition of the vehicle.
[0051] After the selection of a target object, one or more processors in a first device may be configured to initiate communication including communication based on an IP address. In the case where Person 111 is the driver of a first vehicle, it is unsafe to initiate messaging, email, or chat using the hands via a dialogue window. However, audio user interfaces for speaking without using the hands are becoming increasingly popular, and in the Figure 1a system shown, it may be possible to initiate communication between two devices and speak with another person based on a V2X or C-V2X communication system. The vehicle may communicate using V2V communication or using the sidelink channel of C-V2X. An advantage of the C-V2X system is that the vehicle can send communication signals between vehicles without relying on whether the vehicle is connected to a cellular network.
[0052] When the vehicle is wirelessly connected to a cellular network, the vehicle may also be able to communicate using V2V communication or the sidelink channel.
[0053] It may be possible to include other data in the sidelink channel. For example, audio packets and / or one or more tags of audio content may be received via the sidelink channel. In the case where Person 111 is not driving, either because the vehicle is driving itself or because Person 111 is a passenger, it may also be possible to send instant messages between devices in the sidelink channel. The instant message may be part of a media exchange between a first device and a second device, which may include audio packets.
[0054] A display device 119 is also shown in the upper circle 103. The display device 119 may represent an image or icon of the vehicle. When communication is initiated or during communication between a first vehicle 303a and a second vehicle (e.g., 303b or 303c), the pattern 133 may light up or may flash.
[0055] In addition, after the selection of a target object, as a result of a channel for communication between at least one target object external to the first device and a second device, an audio packet may be received from the second device. For example, the lower circle 163 below includes a processor 167, which may be configured to decode the audio packet received from the second device to generate an audio signal and output the audio signal based on the selection of at least one target object external to the first device. That is, a person may be able to hear what voice or music is being played in a second vehicle (or a headphone device) through the playback of a speaker 169.
[0056] As will be explained later in this disclosure, other modes of selection are also possible, including gesture detection of Person 111 and eye gaze detection of Person 111.
[0057] Figure 1bA conceptual diagram of a first device that can communicate with another device (e.g., a second device) is shown. The conceptual diagram also includes the detection of the selection of another device within the first device with the help of a tracker, and audio communication can be spatialized.
[0058] Figure 1b Has a description similar to the description associated with Figure 1a except that other elements are added. For example, device 119 is not shown in the upper circle 104 above because it is shown in the lower circle 129 below. The upper circle 104 shows the vehicle outside the window 132, the mirror 127, and the interior camera 124, which function as described with respect to Figure 1a as described.
[0059] The lower circle 129 shows the display device 119. In addition to just an icon or image representing the vehicle 133, the display device can also represent an image of a real vehicle that may be a potential selection of a person 111 in the first vehicle 303a. For example, an image of a vehicle captured by one or more external cameras (e.g., Figure 3 310b in, 402 in FIG. 4) is represented on the display device 119. The image of the vehicle can have bounding boxes 137a - 137d enclosing each image of the vehicle. The bounding boxes can assist in the selection of a target object, e.g., one of the vehicles represented on the display device. Additionally, instead of the pattern 133 between the icon and image of the vehicle, from the perspective of the person 111 who selects the second vehicle, there can be a separate pattern 149. Thus, the bounding box 137d can show the selected second vehicle 303b, and the direction of the separate pattern 149 can be illuminated or can also blink to indicate that communication has been initiated or is in progress with the second vehicle 303b.
[0060] Additionally, the processor can include a tracker 151 and a feature extractor (not shown) that can perform feature extraction on the images on the display device 119. The extracted features, either individually or in some configurations in combination with RADAR / LIDAR sensors, can assist in the estimation of the relative position of the selected vehicle (e.g., 303b). In other configurations, the tracker 151 can only assist or operate on the input of the GPS position from the selected vehicle, which can also be sent to the first vehicle 303a through a communication channel in a V2X or C-V2X system.
[0061] For example, the second vehicle 303b or another second vehicle 303c may be invisible to the camera. In such a scenario, each of the vehicles (vehicles 303b and 303c) may have a GPS receiver that detects the position of each vehicle. The position of each vehicle may be received by the first device (e.g., vehicle 303a) via assisted GPS, or directly via the V2X or C-V2X system if the V2X or C-V2X system permits. The received vehicle position may be represented by GPS coordinates as determined by each of one or more GPS satellites 160, or in combination with a base station (e.g., as used in assisted GPS). The first device may calculate its own position relative to the other vehicles (vehicles 303b and 303c) based on knowing its own (the first device's) GPS coordinates via its own GPS receiver. Additionally or alternatively, the first device may calculate its own position based on the user of a RADAR sensor, LIDAR sensor, or camera coupled to the first device. It should be understood that the calculation may also be referred to as an estimation. Thus, the first device may estimate its own position based on a RADAR sensor, LIDAR sensor, camera coupled to the first device, or received GPS coordinates. Additionally, each vehicle or device may know its own position by using assisted GPS, i.e., having a base station or other intermediate structure receive the GPS coordinates and relay them to each vehicle or device.
[0062] In addition, the display device 119 may represent an image of the second device in the relative position of the first device. That is, the outward-facing camera 310b or 402 coordinated with the display device 119 may represent the second device in the relative position of the first device. Thus, the display device 119 may be configured to represent the relative position of the second device. Additionally, the relative position of the second device may be represented as an image of the second device on the display device 119.
[0063] Additionally, the audio engine 155, which may be integrated into one or more processors, may process the decoded audio packets based on the relative position of the devices. The audio engine 155 may be part of an audio spatializer that may be integrated as part of a processor, which may output the audio signal as a three-dimensional spatialized audio signal based on the relative position of the second device as represented on the display device 119.
[0064] As discussed above, the relative position may also be GPS receiver-based, where the GPS receiver may be coupled to the tracker 155 and integrated with one or more processors, and the first device may perform assisted GPS to determine the relative position of the second device. The audio engine 155 may be part of an audio spatializer that may be integrated as part of a processor, which may output the audio signal as a three-dimensional spatialized audio signal based on the relative position determined by the assisted GPS of the second device 161.
[0065] In addition, in some configurations, the externally-facing cameras 310b and 402 can capture devices or vehicles in front of or behind the first vehicle 303a. In such a scenario, it may be desirable to hear the sounds from a vehicle or device behind the first vehicle 303a (or, if a headset, behind the person wearing the headset) that have a different spatial resolution than the sounds heard from those vehicles or devices in front of the first vehicle 303a. Thus, the output of the three-dimensional spatialized audio signal has a different spatial resolution when the second device is in a first position relative to the first device (e.g., in front of the first device) compared to when the second device is in a second position relative to the second device (e.g., behind the first device).
[0066] Additionally, when tracking the relative position of at least one target object external to the first device (e.g., a second device or a second vehicle), one or more processors can be configured to receive an updated estimate of the relative position of the at least one target object external to the first device. Based on the updated estimate, a three-dimensional spatialized audio signal can be output. Thus, the first device can present the three-dimensional spatialized audio signal through the speaker 157. A person in the first vehicle 303a or wearing a headset can hear the sounds received by the second device (e.g., the vehicle 303c that is diagonally in front of the first device) as if the audio is coming from diagonally in front. If the first device is the vehicle 303a, diagonally in front is relative to the potential driver of the vehicle 303a looking out of the window 132 as if he or she were driving the vehicle 303a. If the first device is a headset, diagonally in front is relative to the person wearing the headset looking straight ahead.
[0067] In some scenarios, the audio engine 155 may be able to receive multiple audio streams, i.e., audio / voice packets from multiple devices or vehicles. That is, there may be multiple target objects that are selected. The multiple target objects external to the first device can be vehicles, headsets, or a combination of headsets and vehicles. In such scenarios where there are multiple target objects, the speaker 157 can be configured to present the three-dimensional spatialized audio signal based on the relative position of each of the multiple vehicles (e.g., 303b and 303c) or devices (e.g., headsets). It is also possible that the audio streams can be mixed into a single auditory channel and heard together as if there were a multi-party conversation among at least one person in the second vehicles (e.g., 303b and 303c).
[0068] In some configurations, audio / voice packets can be received from each of multiple vehicles in respective communication channels. That is, a first vehicle 303a can receive audio / voice packets from a second vehicle 303b in one communication channel and also receive audio / voice packets from a different second vehicle 303c in a different communication channel 303c. The audio packets (for simplicity) can represent speech spoken by at least one person in each of the second vehicles.
[0069] In such a scenario, a person in the first vehicle 303a or a passenger in a headset can select two target objects by techniques described elsewhere in this disclosure. For example, a person 111 in the first vehicle 303a can tap on regions enclosed by bounding boxes 137a - 137d on a display device 119 to select at least two vehicles (e.g., 303b and 303c) with which to have a multi-party communication. Optionally, the person 111 can use speech recognition to select at least two vehicles (e.g., 303b and 303c) with which to have a multi-party communication.
[0070] In some configurations, one or more processors can be configured to authenticate each of the people or vehicles of the second vehicles to facilitate a trusted multi-party session between at least one person in the second vehicles (e.g., 303b and 303c) and a person 111 in the first vehicle 303a. If people comfortably store samples of each other's voices in their vehicles, the authentication can be based on speech recognition. Other authentication methods can be possible, including facial or image recognition of people or vehicles in the multi-party session.
[0071] Figure 1c A conceptual diagram shows different vehicles sending and receiving wireless connections according to the techniques described in this disclosure.
[0072] Vehicles can be Figure 1c shown to be directly wirelessly connected or can be wirelessly connected to different access points or nodes that are part of a C-V2V or V2X communication system 176 and capable of sending and receiving data and / or messages.
[0073] Figure 1d A conceptual diagram shows different vehicles sending and receiving wireless connections using a cache server in the vehicle or a memory in the vehicle.
[0074] Instant messages exchanged between a first device and a second device wirelessly connected via a sidelink channel can include data packets and / or audio packets transmitted from one vehicle to another. For example, a second device (e.g., vehicle 303d) can broadcast or send an instant message on the sidelink channel, where the instant message includes metadata1. In some configurations, metadata1 is sent on the sidelink and may not have to be part of the instant message.
[0075] In various embodiments, a vehicle in a C-V2X or V2X communication system 176 may receive an instant message or metadata including one or more tags associated with audio content delivered from a static broadcast station to the vehicle (e.g., vehicles 303a, 303d, 303e) via a content delivery network (CDN). The CDN may transfer data efficiently and quickly between a sender and a receiver. In a distributed network, there are many possible combinations of network links and routers through which packets can be forwarded. The selection of network links and routers provides a fast and reliable content delivery network.
[0076] High-demand content may be stored or cached in a memory location near the network edge where the consumers of the data are located. This is more likely when there is media content being broadcast (e.g., entertainment with many viewers and listeners). Caching in a physical location closer to the media consumers may mean a faster network connection and better content delivery. In a configuration scenario where both the sender and receiver of data are in vehicles moving and the vehicles change positions relative to each other, the role of the CDN can provide an effective way to deliver media content over a sidelink channel. Content cached at the edge of the network closest to the consumer may be stored in a device on the move (e.g., vehicle 303d). Media content (e.g., one or more tags of audio content or metadata) is being sent to other vehicles on the move. If traveling along the road in the same direction, the broadcaster device (e.g., vehicle 303e) and the listener device (e.g., vehicle 303a) may be only within a few miles of each other. So a strong local connection is likely. In contrast, if two vehicles are traveling in opposite directions on the same road, the listener vehicle 303a may fall out of the range of the broadcaster device (e.g., vehicle 303e).
[0077] In a vehicle-to-vehicle communication system, it may be possible to receive radio stations outside the vehicle's range. For example, a vehicle traveling 300 miles between cities will undoubtedly lose the signal from the departure city. However, with a CDN, radio signals may be relayed and rebroadcast from the vehicle beyond the range limits of the radio station signal. A vehicle at a certain radial distance from the broadcast station becomes a cache for the radio station, which allows other vehicles within a certain range to request the stream. That is, the broadcasting vehicle 303e may include a cache server 172 and broadcast metadata 2 in the C-V2X or V2X communication system network 176. The listening vehicle 303a may receive the metadata 2.
[0078] Machine learning algorithms can be used to listen to, parse, understand, and broadcast a driver's listening preferences. Along with the driver's geographical location, information can be collected to determine the most popular content most frequently received by vehicles from other vehicles within each geographical area.
[0079] As can be seen in Figure 1d There can be a first device for receiving metadata from a second device. The first device and the second device can be wirelessly connected via a sidelink channel that is part of a C-V2X or V2x communication system network 172. Once the first device (e.g., vehicle 303d) receives the metadata (e.g., metadata 1 171 or metadata 2 173), the first device can read the metadata and extract one or more tags representing the audio content.
[0080] One or more tags can include the song name, artist name, album name, writer, or International Standard Recording Code. The International Standard Recording Code (ISRC) uniquely identifies sound recordings and music video recordings and is encoded as an ISO 3901 standard.
[0081] The metadata can be indexed and made searchable by my search engine. If the audio content is streamed or broadcast by the second device (e.g., vehicle 303d or 303e), then one or more of the tags can be read by the audio player or, in some cases, by the radio interface to the radio. Additionally, one or more of the audio tags can be represented on a display device. The metadata associated with the audio content can include songs, audio books, tracks from movies, etc.
[0082] The metadata can be structural or descriptive. Structural metadata represents data as a container for the data. Descriptive metadata describes the audio content or some property associated with the audio content (e.g., song, author, creation date, album, etc.).
[0083] After one or more processors of the first device extract one or more tags representing the audio content, the audio content can be identified based on the extracted one or more tags. One or more processors of the first device can be configured to output the audio content.
[0084] In Figure 1dIn this case, the first device can also be part of a group of devices configured to receive one of one or more tags. A device (e.g., vehicle 303a) can be part of a group of devices (e.g., vehicles 303b and 303c as well) configured to receive at least one tag of metadata from another device (e.g., vehicle 303d or 303e). The group of devices can also include other devices (e.g., vehicles 303d and 303e) that send metadata. That is, there can be a group of devices including five devices, where each device is a vehicle (e.g., vehicles 303a, 303b, 303c, 303d, and 303e), or there can be a mix of vehicles and headsets. It can be that the group of devices includes these five devices.
[0085] In one embodiment, the group of devices can be part of a content delivery network (CDN). Additionally or alternatively, a second device (e.g., 303e) in the group of devices can be a respective content delivery network and send one or more tags to the remaining devices in the group.
[0086] Figure 2 A flowchart of process 200 is shown in which a first device initiates communication with a second device based on the techniques described in this disclosure.
[0087] 210, the first device can include one or more processors configured to detect a selection of at least one target object external to the first device. 220, the one or more processors can be configured to initiate a communication channel between the first device and a second device associated with at least one target object external to the first device. 230, the one or more processors can be configured to receive an audio packet from the second device in response to a selection of at least one target object external to the device.
[0088] 240, the one or more processors can be configured to decode the audio packet received from the second device to generate an audio signal. 250, the one or more processors can be configured to output the audio signal based on a selection of at least one target object external to the first device.
[0089] Figure 3 A conceptual diagram of a first vehicle having different components operating according to the techniques described in this disclosure on or in the first vehicle is shown. As Figure 3 shown, person 111 can move within vehicle 303a. The selection of a target object external to vehicle 303a can be directly within the driver's field of view, which can be captured by an eye gaze tracker (i.e., person 111 is looking at the target object) or a gesture detector (person 111 makes a gesture, such as pointing at the target object) coupled to a camera 310a within vehicle 303a. Similarly,
[0090] The first device may include a selection of at least one target object based on the detection of a command signal, where the command signal detection is based on eye gaze detection.
[0091] If the target object is a person outside vehicle 303a, or there are some other recognizable images associated with vehicle 303b, the camera 310b mounted on vehicle 303a may also assist in the selection of the target object itself (e.g., vehicle 303b) or another device associated with the target object.
[0092] One or more antennas 356, optionally coupled with a depth sensor 340, may assist in determining the relative position of the target object with respect to vehicle 303a via a wireless local area network (WLAN) that can be part of a cellular network such as C-V2X, or a coexistence of a cellular network and a Wi-Fi network, or just a Wi-Fi network, or a V2X network.
[0093] It should be noted that the camera 310a installed inside vehicle 303a, or the camera 310b mounted on vehicle 303a, or both cameras 310a and 310b, depending on the available bandwidth, can form a personal area network (PAN) as part of vehicle 303a via one or more antennas 356. Through the PAN, the camera 310a in vehicle 303a or the camera 310b on vehicle 303a may be able to have an indirect wireless connection with the device associated with the target object or the target object itself. Although the external camera 310b is shown near the front of vehicle 303a, vehicle 303a may be able to have one or more external cameras 310b mounted near or in the rear of vehicle 303a to view what device or vehicle is behind vehicle 303a. For example, the second device may be vehicle 303c.
[0094] The external camera 310b may assist in the selection, or as explained above and below, GPS may also assist in determining the location of the second device, such as where the second vehicle 303c is located.
[0095] The relative position of the second device may be represented on the display device 319. The relative position of the second device may be based on receiving the position by one or more antennas 356. In another embodiment, the depth sensor 340 may be used to assist in or determine the location of the second device. Other position detection techniques (e.g., GPS) for detecting the location of the second device or assisting GPS may also be used to determine the relative position of the second device.
[0096] The representation of the relative position of the second device may be presented as a composite image, an icon, or other representation associated with the second device, such that a person in vehicle 303a can make a selection of the second device by eye gaze toward the representation on display device 319 or a gesture (pointing or touching) toward the representation on display device 319.
[0097] The selection can also be made by voice recognition and using one or more microphones 360 located inside vehicle 303a. When the second device communicates with vehicle 3030a, the audio signal can be received by (the first) vehicle 303a through a transceiver installed in or on vehicle 303a and coupled to one or more antennas 356.
[0098] Those of ordinary skill in the art will also understand that as autonomous vehicles continue to improve, the driver of vehicle 303a may not actually manually direct (i.e., “drive”) vehicle 303a. Instead, for some periods of time, vehicle 303a may be self-driving.
[0099] Figure 4a A block diagram 400a of a first device having different components operating according to the techniques described in this disclosure on or in the first device is shown. One or more different components may be integrated in one or more processors of the first device.
[0100] As Figure 4a shown, selecting a target object external to the first device may be based on an eye gaze tracker 404 that detects and tracks where the wearer of the headset is looking or where the person 111 in the first vehicle is looking. When the target object is within the person's field of view, the eye gaze tracker 404 can detect and track the eye gaze and assist in selecting the target object via a target object selector 414. Similarly, a gesture detector 406 coupled to one or more inward-facing cameras 403 within vehicle 303a or mounted on a headset (not shown) can detect gestures, e.g., the direction of pointing to the target object. Additionally, a voice command detector 408 can assist in selecting the target object based on the person 111 saying a phrase as described above (e.g., “the black Honda Accord in front of me”). The output of the voice command detector 408 can be used by the target object selector 414 to select the intended second device, such as vehicle 303b or 303c.
[0101] As previously mentioned, vehicle 303a may optionally have one or more outward-facing cameras 402 mounted near or in the rear of vehicle 303a to view what device or vehicle is behind vehicle 303a. For example, the second device may be vehicle 303c.
[0102] A target object (e.g., a second device) can be represented relative to a first device and based on image features, an image, or both the image and image features, where the image is captured by one or more cameras coupled to the first device.
[0103] One or more outward-facing cameras 402 can assist in the selection of where the second vehicle 303c is located, e.g., behind the vehicle 303a (in other figures).
[0104] It is also possible that, based on one or more transmitter antennas 425 and possibly a depth sensor 340 ( Figure 4a not shown in the figure), or other position detection techniques (e.g., GPS) for detecting the position of the second device, the relative position of the second device can be represented on the display device 410. The representation of the relative position of the second device can appear as a synthetic image, an icon, or other representations associated with the second device, so that a person in the vehicle 303a can select the second device by eye gaze towards the representation on the display device 410 or a gesture (pointing or touching) towards the representation on the display device 410.
[0105] If the selection of the remote device (i.e., the second device) is based on touch, the display device including a representation of at least one target object of the external device (i.e., the first device) can be configured to select at least one target object outside the device based on a change in state of a capacitance sensor or an ultrasonic sensor on the display device.
[0106] One or more transmitter antennas 425 of the first device coupled to one or more processors included in the first device can be configured to send communication data to the second device based on the initiation of a channel for communication between the first device and the second device, which is an external target object relative to the first device. That is, after the selection of the second device, one or more processors can initiate a protocol or other form of communication between the first device and the second device, using C-vVX and / or V-2X communication in the channel for communication between the first device and the second device.
[0107] The selection can also be by voice recognition and use one or more microphones located inside the vehicle 303a ( Figure 4a(not shown in the figure). When the second device communicates with the vehicle 3030a, the audio signal can be received by the (first) vehicle 303a through one or more receiver antennas 430 installed in or on the vehicle 303a. The receiver antenna is coupled to the reception of a transceiver (e.g., a modem capable of V2X or C-V2X communication). That is, one or more receiving antennas 430 coupled to one or more processors can be configured to receive audio packets based on the result of the activation of a channel for communication between at least one target object (e.g., the second device) outside the first device and the first device.
[0108] In addition, the first device may include one or more externally facing cameras 402. If the target object is a person outside the vehicle 303a or some other recognizable image associated with the vehicle 303b, the externally facing camera 402 can be mounted on the vehicle 303a and can also assist in the selection of the target object itself (e.g., the vehicle 303b) or another device associated with the target object. One or more externally facing cameras can be coupled to one or more processors, which include a feature extractor (not shown) that can perform feature extraction on the images on the display device 410. The extracted features alone, or in some configurations in combination with external sensors 422 (e.g., RADAR / LIDAR sensors), can assist in the estimation of the relative position of the second device (e.g., the selected vehicle 303b).
[0109] The output of the extracted features or the external sensors 422 can be input to the determiner 420 of the relative position / orientation of the selected target object. The determiner 420 of the relative position / orientation of the selected target object can be integrated into one or more processors and can be part of a tracker, or in other configurations (as Figure 4a shown) can be separately integrated into one or more processors. In Figure 4a the figure, the tracker 151 is not shown.
[0110] The distance and angle can be provided by the determiner 420 of the relative position / orientation of the selected target object. The distance and angle can be used by the audio spatializer 420 to output a three-dimensional audio signal based on the relative position of the second device. There can be at least two speakers 440 coupled to one or more processors, which are configured to present the three-dimensional spatialized audio signal based on the relative position of the second device, or if there are multiple second devices (e.g., multiple vehicles), then the three-dimensional spatialized audio signal can be presented as described above.
[0111] After the selection of at least one target object external to the first device is performed by the target object selector 414, the command interpreter 416 integrated in one or more processors in the first device activates a channel for communication between the first device and a second device associated with at least one target object external to the first device. In response to the selection of at least one target object external to the first device, an audio packet can be received from the second device.
[0112] The audio packet 432a from the second device can be decoded by the codec 438 to generate an audio signal. The audio signal can be output based on the selection of at least one target object external to the first device. In some scenarios, the audio packet can represent a stream from a cloud associated with a remote device (i.e., the second device) 436a. The codec 438 can decompress the audio packet, and the audio spatializer can operate on the uncompressed audio packet 432b or 436b. In other scenarios, the audio can be spatialized based on the passenger position of the person making the second vehicle selection.
[0113] The transmission of the audio packet by the audio codec to be used can include one or more of the following: MPEG-2 / AAC stereo, MPEG-4 BSAC stereo, Real Audio, SBC Bluetooth, WMA, and WMA 10Pro. Since C-v2X and v2V systems can use a data service channel or a voice channel, the audio packet (which can carry a voice signal) can use one or more of the following codecs to decompress the audio signal: AMR narrowband voice codec (5.15 kbp), AMR wideband voice codec (8.85 Kbps), G.729AB voice codec (8 kbps), GSM-EFR voice codec (12.2 kbps), GSM-FR voice codec (13 kbps), GSM-HR voice codec (5.6 kbps), EVRC-NB, EVRC-WB, enhanced voice service (EVS). A voice codec is sometimes referred to as a vocoder. Before being sent over the air, the vocoder packet is inserted into a larger packet. Voice is sent in the voice channel, although voice can also be sent in the data channel using VOIP (voice-over-IP, voice based on IP). The codec 438 can represent a voice codec, an audio codec, or a combination of functions for decoding a voice packet or an audio packet. Generally, for ease of explanation, the term audio packet also includes the definition of the packet.
[0114] It is also possible that, in one configuration, the spatialization effect can be disabled after the second vehicle is at a certain distance from the first vehicle.
[0115] One or more processors included in a first device may be configured to disable a spatialization effect after a second vehicle is greater than a configurable distance from the first device. The certain distance may be configurable based on distance, such as one-eighth of a mile. The configurable distance may be input as a function of distance measurement or time measurement. The certain distance may be configurable based on time, for example, depending on the speeds of the first vehicle and the second vehicle. For example, instead of indicating that one-eighth of a mile is the distance for which the spatial effect should persist, the distance between them may be measured in terms of time. A vehicle traveling at 50 miles per hour (mph), one-eighth of a mile corresponds to 9 seconds, i.e., 125 mi / 50 mi / hr = 0.0025 hr = 0.0025 * 60 min = 0.15 min = 9 seconds. Thus, in this example, after 9 seconds, the spatial effect will fade out or suddenly stop.
[0116] Figure 4b FIG. 400b is a block diagram of a first device having different components operating in accordance with the techniques described in the present disclosure on or in the first device. One or more different components may be integrated in one or more processors of the first device.
[0117] Block diagram 400b includes a communication interpreter 416 and an rx antenna 430. Via the rx antenna 430, one or more processors may be configured to receive metadata 435 from a second device that is wirelessly connected to the first device via a sidelink channel. One or more processors may store the metadata in a buffer 444. The metadata 435 may be read from the buffer 444. One or more processors may be configured to extract one or more tags representative of audio content. For example, the communication interpreter 416 may send a control signal to a controller 454, and the controller, which may be integrated as part of one or more processors, may control an extractor 460, which may also be integrated as part of one or more processors. The extractor 460 may be configured to extract one or more tags representative of audio content. If one or more tags are not yet in a form where they can be extracted in-place in the buffer 444, they may be written back to the buffer 444 via a bus 445. That is, the extractor 460 may extract one or more tags from the buffer 444, or the extractor 460 may receive the metadata via the bus 445 and then write one or more tags back to the buffer 444 via the bus 445. One of ordinary skill in the art will recognize that the location where one or more tags may be written may be the same buffer 444 or a different memory location in an alternative buffer. However, for ease of explanation, it may still be referred to as buffer 444.
[0118] One or more processors may be configured to identify audio content based on one or more tags. The identification can be done in a variety of ways. For example, one of the tags can identify the name of a song, and the tag identifying the song can be displayed on the display device 410, or one or more processors can store the "song" tag in a memory location (e.g., also in buffer 444), or in an alternative memory location. Based on this identification, one or more processors may output the audio content.
[0119] The output of the audio content can be done in a variety of ways. For example, one or more processors in a first device may be configured to switch to a radio station that is playing the identified audio content based on one or more tags. This can occur by causing the radio interface 458 to receive a control signal from the controller 460. The radio interface 458 may be configured to scan different radio stations on the radio 470 and switch the radio 470 to a radio station that is playing the identified audio content (e.g., song) based on one or more tags.
[0120] In another example, one or more processors may be configured to turn on a media player and cause the media player to play the identified content based on one or more tags. The media player may read from a playlist having tags that may be associated with the one or more received tags. For example, the controller may be configured to compare one or more tags received via metadata and extracted with its own tags to the audio content stored in the memory. The media player may be coupled to the database 448, and the database 448 may store tags associated with the audio content of the media player's playlist. The database 448 may also store a compressed version of the audio content in the form of an audio bitstream that includes audio packets. The audio packets 453 may be sent to the codec 438. The codec 438 may be integrated as part of the media player. It should be observed that the audio packets 453 may be stored in the database 448. It may also be possible to receive the audio packets 432a as described in Figure 4a In addition, it may be possible to receive the audio packets 432a associated with one or more tags that are associated with the audio content received via the rx antenna 430.
[0121] The first device includes one or more processors that may receive metadata from a second device that is wirelessly connected to the first device via a sidelink channel, the one or more processors read the metadata received from the second device to extract one or more tags representing audio content, identify the audio content based on the tags, and then output the audio content.
[0122] A wireless link via a sidelink channel can be part of a C-2VX communication system. The first device and the second device in a C-V2Vx system can both be vehicles, or one of the devices (first or second) can be a headset and the other can be a vehicle (first or second).
[0123] Similarly, a wireless link via a sidelink channel can be part of a V2X or V2V communication system. The first device and the second device in a V2V system can both be vehicles.
[0124] The first device can include one or more processors configured to scan a buffer 444 based on configuration preferences stored on the first device. For example, there can be many metadata sets received from multiple second devices. A person listening to audio content on the first device (whether a vehicle or a headset) may only want to listen to audio content based on configuration preferences (e.g., rock music). The configuration preferences can also include attributes from the second device. For example, the second device itself can have a label identifying itself. For example, a blue BMW. Thus, a person listening to audio content on the first device may want to listen to content from a blue BMW.
[0125] In the same or an alternative embodiment, the first device is coupled to a display device. The coupling can be an integration, e.g., the display device is integrated as part of a headset or part of a vehicle. One or more processors in the first device can be configured to represent one or more labels on the screen of the display device. When the buffer 444 is coupled to the display device 410, one or more labels including the song name, artist, and even a blue BMW can appear on the screen of the display device 410. Thus, a person can see which songs are from a blue BMW.
[0126] As previously discussed with respect to Figure 4a The first device can include a display device configured to represent the relative position of the second device. Similarly, with respect to audio content identified based on one or more tags extracted from metadata received from the second device, the first device can include one or more processors configured to output three-dimensional spatialized audio content. After decoding the audio packet 453 from the database 448 by the codec 438, the three-dimensional spatialized audio content can optionally be generated by the audio spatializer 424. In the same or an alternative embodiment, the audio packet 432a associated with one or more audio tags of the identified audio content can be decoded from the codec 438. The codec 438 can implement with respect to Figure 4aThe described audio codec or speech codec. One or more processors may be configured to output three-dimensional spatialized audio content based on where the relative position of the second device is represented on the display device 410. The output three-dimensional spatialized audio content may be presented by two or more speakers 440 coupled to the first device.
[0127] In some configurations, the output of the audio content may be three-dimensional spatialized audio content based on the relative position of the second device, regardless of whether the position of the second device is represented on the display device 410.
[0128] Additionally, in the same or alternative embodiments, one or more processors may be configured to fade in or fade out audio content associated with one or more tags.
[0129] The fade in or fade out of the audio content associated with one or more tags may be based on a configurable distance of the second device. For example, if the distance of the second device is within 20 meters or within 200 meters, a person listening to the audio content on the first device may expect the audio content to fade in or fade out. Additionally, as described with respect to Figure 4a One or more processors may be configured to disable the spatialization effect after the second device is at a distance greater than a configurable distance from the first device. Thus, there may be a first configurable distance to fade in and fade out the audio content (e.g., within 0 to 200 meters), and a second configurable distance where if the second device is within 200 meters or even farther (e.g., up to 2000 meters), the spatialization effect for a listener listening to the spatialization effect is disabled. As previously mentioned, the configurable distance (the first configurable distance or the second configurable distance) may be a distance measurement or a time measurement.
[0130] As described with respect to Figure 1d The first device may be part of a group of devices. Figure 1d One or more of the illustrated tags 170 or the cache server 172 may also be part of the buffer 444, or may alternatively be drawn as being associated with Figure 4bis adjacent to buffer 444, where metadata 435a can be metadata 1 or metadata 2, depending on whether the second device is a device having one or more tags 170 in the memory (e.g., vehicle 303d), or whether the second device is a device having a caching server 172 (e.g., vehicle 303e). Thus, the fade-in or fade-out of the audio content may also be based on when one of the devices in the group disconnects from the group. For example, the first device may disconnect from the group of devices, and the audio content may fade out. Similarly, when connecting to become part of the group of devices, the audio content may fade in. In both the fade-in and fade-out when a device (e.g., the first device) connects or disconnects from a group of devices, the fade-in or fade-out may also be based on a configurable distance and may be a distance measurement or a time measurement.
[0131] Additionally, the first device and other devices in the group of devices may be part of a content delivery network (CDN), as described above when describing Figure 1d it.
[0132] The first device or the second device may be a separate content delivery network and may send one or more tags to other devices in the group.
[0133] Although one or more outward-facing cameras 402 and a target object selector 414 are depicted in Figure 4b and there are no other components coupled to them in Figure 4a in the same or alternative configurations, it may also receive audio packets associated with one or more tags, where the one or more tags are associated with the audio content received via one or more rx antennas 430.
[0134] Like this, after the target object selector 414 selects at least one target object outside the first device, a command interpreter 416 integrated within one or more processors in the first device initiates a channel for communication between the first device and a second device associated with at least one target object outside the first device. In response to the selection of at least one target object outside the first device, audio packets may be received from the second device.
[0135] One or more tags from the second device may be received in the metadata, which is read from the buffer 444, extracted, and used to identify the audio content. The audio content may be output based on the selection of at least one target object outside the first device. In some scenarios, one or more tags may represent a stream from the cloud associated with a remote device (i.e., the second device).
[0136] Figure 5A conceptual diagram 500 is shown of transforming world coordinates to pixel coordinates according to the techniques described in this disclosure. An external camera (e.g., Figure 3 310b, Figure 4a and Figure 4b 402) can capture an image (e.g., a video frame) and represent an object in three-dimensional (3D) world coordinates [x, y, z] 502. The world coordinates can be transformed to 3D camera coordinates [xc, yc, zc] 504. The 3D camera coordinates 504 can be projected into a 2D xy plane (normal vector perpendicular to the face of the camera (310b, 402)) and in pixel coordinates (x p ,y p ) 506. Those skilled in the art will recognize that this transformation from world coordinates to pixel coordinates is based on using the input rotation matrix [R], translation vector [t] and camera coordinates [x c ,y c ,z c ] to transform the world coordinates [xyz]. For example, the camera coordinates can be expressed as [xc, yc, zc] = [xy z] * [R] + t, where the rotation matrix [R] is a 3×3 matrix and the translation vector is a 1×3 vector.
[0137] The bounding box of the region of interest (ROI) may be represented by pixel coordinates (xP, yP) on the display device 510. There may be a visual indication (e.g., a color change or an icon or composite pointer enhanced within the bounding box 512) to alert the occupants of the vehicle that a target object (e.g., a second vehicle) has been selected to initiate communication therewith.
[0138] Figure 6a A conceptual diagram of one embodiment of an estimation of the distance and angle of a remote vehicle / passenger (e.g., a second vehicle) is shown. The distance may be derived from a bounding box 622d in a video frame. A distance estimator 630 may receive sensor parameters 632a, intrinsic and extrinsic parameters 632d of an external viewing camera (310b, 402), and a size 632b of the bounding box 622d. In some embodiments, there may be a database of vehicle information that includes the sizes 632c of different vehicles and may also contain certain image characteristics that may help identify a vehicle.
[0139] The distance and angle parameters can be estimated at the video frame rate and interpolated to match the audio frame rate. From the database of vehicles, the actual size of the remote vehicle, i.e., width and height, can be obtained. The pixel coordinates of the corners of the bounding box (x p ,y p ) can correspond to a line with a given azimuth and elevation in 3D world coordinates.
[0140] For example, using the lower left and lower right corners of the bounding box and having the width w of the vehicle, the estimated distance d and azimuth angle (θ) 640a can be obtained as Figure 6b shown.
[0141] Figure 6b Fig. shows a conceptual diagram of the estimated distance 640c and angle 640a in the x-y plane of the remote device.
[0142] Figure 6b The point A in can be represented by world coordinates (a, b, c). Figure 6b The point B in can also be represented by world coordinates (x, y, z). The azimuth angle (θ) 640a can be expressed as (θ1 + θ2) / 2. For small angles, the distance d xy *(sinθ1 - sinθ2) is approximately equal to w, which is Figure 6b the width of the remote device in. The world coordinates (x, y, z) and (a, b, c) can be expressed in terms of the width in the x-y plane, for example, using the following formulas:
[0143] x = a
[0144] |y - b| = w
[0145] z = c
[0146] Figure 5 The pixel coordinates described in can be represented as x p = x = a and y p = y = w + / - b.
[0147] Similarly, using the lower left and upper left corners of the bounding box and knowing the height h of the second vehicle 303b and the elevation angle (φ) 640b of the second vehicle 30b, the distance d of the second vehicle can be calculated as shown in Figure 6c Fig. yz .
[0148] Figure 6c Fig. shows a conceptual diagram of the estimated distance and elevation angle 640b in the y-z plane of the remote device.
[0149] Figure 6c The point A in can be represented by world coordinates (a, b, c). Figure 6c The point B in can also be represented by world coordinates (x, y, z). The elevation angle (φ) 640b can be expressed as (φ1 + φ2) / 2. For small angles, the distance d yz *(sinφ1 – sinφ2) is approximately equal to h, which is Figure 6c the height of the remote device in. The world coordinates (x, y, z) and (a, b, c) can be expressed in terms of the height in the y-z plane, for example, using the following formulas:
[0150] x = a
[0151] y = b
[0152] |z - c| = h
[0153] Figure 5 The pixel coordinates described in can be expressed as x p = x = a, and y p = y = b.
[0154] Depending on the position of the sound source, for sounds from the left half, right half, or middle of the remote device 670, the elevation angle 640b and azimuth angle 640a can be further adjusted. For example, if the remote device 670 is a remote vehicle (e.g., the second vehicle), the position of the sound source can depend on whether the driver or a passenger is speaking. For example, the driver side (left) azimuth angle 640a of the remote vehicle can be expressed as (3*θ1 + θ2) / 4. This provides the azimuth angle 640a in the left half of the vehicle as represented in Figure 8 .
[0155] The video frame rate usually does not match the audio frame rate. To compensate for the misalignment of the frame rates in different domains (audio and video), the parameter distances 640c, elevation angle φ, azimuth angles 640a, θ can be interpolated as linear interpolations of the corresponding values from the previous two video frames for each audio frame. Optionally, the value from the nearest video frame can be used (sampled and held). Additionally, these values can be smoothed by taking the median (removing outliers) or average from the past few video frames at the cost of reduced responsiveness.
[0156] Figure 6a The distances 640c, d shown can be d xy or d yz or d xy and d yzSome combination thereof, such as an average value. In some embodiments, it may be desirable to ignore the height difference between the first vehicle and the remote device 670. For example, if the remote device 670 is at the same height as the first vehicle. Another example could be that the listener in the first vehicle is configured to receive spatial audio by projecting the z - component of the sound field emitted from the remote device 670 onto the x - y plane. In other examples, the remote device 670 could be a drone (e.g., flying around playing music), or a device that streams music in a tall building. In such examples, it may be desirable for the angle estimator 630 to output the elevation angle, or for other optional blocks to also operate on it. That is, for the smoothing of the parameter frame rate conversion from video to audio, 640 also operates on the elevation angle and produces a smoother version of the elevation angle. Since the vehicle and / or the remote device will likely move around, the Doppler estimator 650 is responsible for the relative change in the sound frequency. Therefore, it may be desirable for the listener in the first vehicle to additionally hear the sound of the remote device (e.g., a second vehicle) with the Doppler effect. As the remote device 670 approaches or moves away from the first vehicle, the Doppler estimator 650 can increase or decrease the change in the frequency (i.e., pitch) heard by the listener in the first vehicle. As the remote device 670 approaches the first vehicle, since the pressure wave of the sound is compressed by the remote device approaching the first device, the sound reaches the listener at a higher frequency if propagated through the air. In the case where the audio signal (or audio content) is compressed and received as part of a radio frequency signal, there is no Doppler shift perceptible to the human ear. Therefore, the Doppler estimator 650 must compensate and use the distance and angle to produce the Doppler effect. Similarly, when the remote device 670 is moving away from the first vehicle, if propagated through the air, the pressure sound wave of the audio signal (or audio content) will be expanded and result in a lower - pitched sound. The Doppler estimator 650 will compensate for the low - frequency effect, since the audio signal (or audio content) is compressed in the bitstream and is also transmitted by the remote device and received by the first vehicle using radio frequency waves according to the modulation scheme that is part of the air interface of a C - V2X or V - 2VX communication link. Or, if the remote device is not a vehicle, different types of communication links and air interfaces can be used.
[0157] Figure 7a An embodiment of an audio spatializer 724a according to the techniques in the present disclosure is shown. In Figure 7a this, the reconstructed sound field is presented into the speaker feed that is provided to the speakers 440 or the headset or any other audio delivery mechanism. The reconstructed sound field can include spatial effects provided to account for the distance and azimuth / elevation angle of the device (e.g., a remote vehicle or a wearable device) relative to the person 111 in the vehicle 303a (or another wearable device).
[0158] A distance 702a (e.g., from a distance estimator 630, a smoother 650 for video-to-audio parametric frame rate conversion, or a Doppler estimator 660) can be provided to a distance compensator 720. The input to the distance compensator 720 can be an audio signal (or audio content). The audio signal (or audio content) can be the output of a codec 438. The codec 438 can output a pulse code modulated audio signal. The PCM audio signal can be represented in the time domain or the frequency domain. A distance effect can be added as a filtering process, a finite impulse response (FIR), or an infinite impulse response (IIR) with additional attenuation proportional to the distance (e.g., the applied attenuation can be 1 / distance). Optional parameters (gains) can also be applied to increase the gain for intelligibility. Additionally, a reverb filter is an example of a distance simulator filter.
[0159] Another distance cue that can be modeled and added to the audio signal (or audio content) is the Doppler effect described for the Doppler estimator 650 in Figure 6c The relative speed of a remote vehicle is determined by calculating the rate of change of distance per unit time, and the distance and angle are used to provide the Doppler effect as described above.
[0160] An acoustic field rotator 710 can use the output of the distance compensator 720 and the input angle 702b (e.g., azimuth angle 640a, elevation angle 640b, or a combination based on these angles), and can translate the audio from a remote device (e.g., a second vehicle) to the expected azimuth and elevation angles. The input angle 720b can be converted to an output at an audio frame interval rather than a video frame interval by smoothing the parametric frame rate conversion for video-to-audio 650. In Figure 7b Another embodiment is shown that can include an acoustic field rotator 710 that does not depend on distance interdependently. Among other means, the translation can be achieved by using object-based rendering techniques (e.g., vector-based amplitude panning (VBAP), a renderer based on high-fidelity stereo), or by using high-resolution head-related transfer functions (HRTFs) for headphone-based spatialization and rendering.
[0161] Figure 7b An embodiment of an audio spatializer 424 including a decoder used according to the techniques described in the present disclosure is shown. In Figure 7b this, the decoder 724b can utilize the distance 702a information during the decoding process. As Figure 7aAs described, additional distance effects can be applied. The decoder 730 can be configured to ignore the highest frequency bands when decoding for distances greater than a certain threshold. The distance filter may wipe out these higher frequencies and may not be required to maintain the highest fidelity in these frequency bands. Additionally, Doppler shift can be applied in the frequency domain during the decoding process to provide a computationally efficient implementation of the Doppler effect. Reverberation and other distance filtering effects can also be effectively implemented in the frequency domain and integrated with the decoding process themselves. During the decoding process, rendering and / or binauralization can also be applied in the time domain or frequency domain within the decoder to produce a properly panned speaker feed at the output of the decoder.
[0162] The decoder 730 can be a voice decoder, or an audio decoder, or a combined voice / audio decoder capable of decoding audio packets including compressed voice and music. The input to the decoder 730 can be a stream from a cloud server associated with one or more remote devices. That is, there can be multiple streams as input 432b. The cloud server can include streaming of music or other media. The input to the decoder 730 can also be compressed voice and / or music directly from a remote device (e.g., a remote vehicle).
[0163] Figure 8 Embodiment 800 is described where the positions of the first vehicle and the person 111 in the selected (remote) vehicle 810 can be in the same coordinate system. The angles and distances relative to the previously described external camera may need to be readjusted relative to the head position 820 (X', Y', Z') of the person 111 in the first vehicle. The position (X, Y, Z) of the selected remote device (e.g., remote vehicle 303b) and the position (X, Y, Z) 802 of the first vehicle 303a can be calculated from the distance and azimuth / elevation angles as follows. X = d * cos(azimuth), Y = d * sin(azimuth) and Z = d * sin(elevation). The head position 820 from the (first vehicle's) inward-facing camera 188 can be determined and transformed to the same coordinate system as the coordinates of the first vehicle to obtain X', Y' and Z' 820. Given X, Y, Z 802 and X', Y', Z' 820, then the updated distance and angle relative to the person 111 can be determined using the trigonometric relations d = sqrt[(X - X')^2+(Y - Y')^2+(Z - Z')^2] and azimuth = asin[(Y - Y') / d] and elevation = asin[(Z - Z') / d]. These updated d and angles can be used for finer spatialization and distance resolution as well as better accuracy.
[0164] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. These techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device handset, or an integrated circuit device, which have a variety of uses including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include a memory or data storage media, such as, for example, random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and so forth. Additionally or alternatively, these techniques may be realized at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a device having computing capabilities.
[0165] The program code or instructions may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term “processor” may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein may be provided within a special software module or hardware module configured for encoding and decoding, or incorporated into a combined video codec (CODEC).
[0166] The decoding techniques discussed in this document can be examples of embodiments in a video encoding and decoding system. The system includes a source device that provides encoded video data to be decoded by a destination device at a later time. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device can include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, cellular handsets such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and so on. In some cases, the source device and the destination device can be equipped for wireless communication.
[0167] The destination device can receive the encoded video data to be decoded via a computer-readable medium. The computer-readable medium can include any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium can include a communication medium to enable the source device to send the encoded video data directly to the destination device in real time. The encoded video data can be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to the destination device. The communication medium can include any wireless or wired communication media, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium can include routers, switches, base stations, or any other device that can be useful for facilitating communication from the source device to the destination device.
[0168] In some examples, the encoded data may be output from the output interface to a storage device. Similarly, the encoded data may be accessed from the storage device by the input interface. The storage device may include any one of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data. In a further example, the storage device may correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. The destination device may access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing the encoded video data and sending the encoded video data to the target device. Example file servers include a web server (e.g., for a website), an FTP server, a network-attached storage (NAS) device, or a local disk drive. The target device can access the encoded video data through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device can be a streaming transmission, a download transmission, or a combination of both.
[0169] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a group of ICs (e.g., a chipset). Various components, modules, or units are described in the present disclosure to emphasize the functional aspects of the devices configured to perform the disclosed techniques, but need not be implemented by different hardware units. Instead, as described above, the various units may be combined in an encoder / decoder hardware unit with suitable software and / or firmware or provided by some interoperating hardware units that include one or more processors as described above.
[0170] Specific embodiments of the present disclosure are described below with reference to the accompanying drawings. In the description, common features are indicated by common reference numerals throughout the drawings. As used herein, various terms are used only for the purpose of describing specific embodiments and are not intended to be limiting. For example, unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are also intended to include the plural forms. Additionally, it can be further understood that the terms "comprise," "comprises," and "comprising" may be used interchangeably with "include," "includes," or "including." Additionally, it should be understood that the term "wherein" may be used interchangeably with "where." As used herein, "exemplary" may indicate an example, an embodiment, and / or an aspect, and should not be construed as limiting or indicating a preference or a preferred embodiment. As used herein, ordinal terms (e.g., "first," "second," "third," etc.) used to modify elements such as structures, components, operations, etc. do not themselves indicate any priority or order of the element relative to another element, but merely distinguish the element from another element having the same name (but a different ordinal term used). As used herein, the term "group" refers to a grouping of one or more elements, and the term "plurality" refers to a plurality of elements.
[0171] As used herein, "coupled" may include "communicatively coupled," "electrically coupled," or "physically coupled," and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. As an illustrative but non-limiting example, two devices (or components) that are electrically coupled may be included in the same device or different devices and may be connected via an electronic device, one or more connectors, or inductive coupling. In some embodiments, two devices (or components) that are communicatively coupled (e.g., electrically communicatively) may send and receive electrical signals (digital signals or analog signals) directly or indirectly, for example, via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without an intermediate component.
[0172] As used herein, "integrated" may include "manufactured or sold devices". If a user purchases a package that bundles or includes a device as part of the package, then the device may be integrated. In some descriptions, two devices may be coupled, but not necessarily integrated (e.g., different peripheral devices may not be integrated into a command device, but may still be "coupled"). Another example may be that any of the transceivers or antennas described herein may be "coupled" to a processor, but does not have to be part of a package that includes a video device. When using the term "integrated", other examples may be inferred from the context disclosed herein (including this paragraph).
[0173] As used herein, a "wireless" connection between devices may be based on various wireless technologies, such as being "wirelessly connected" based on different cellular communication systems such as V-2X and C-V2X. C-V2X allows direct communication between vehicles and other devices (via "sidelink") without using a base station. In this case, the devices may be "wirelessly connected via sidelink".
[0174] Long Term Evolution (LTE) system, Code Division Multiple Access (CDMA) system, Global System for Mobile Communications (GSM) system, Wireless Local Area Network (WLAN) system, or some other wireless system. The CDMA system may implement Wideband CDMA (WCDMA), 1X, Evolution-Data Optimized (EVDO), Time Division-Synchronous CDMA (TD-SCDMA), or other versions of CDMA. Additionally, two devices may be wirelessly connected based on Bluetooth, Wireless Fidelity (Wi-Fi), or a variant of Wi-Fi (such as Wi-Fi Direct). When two devices are in line of sight, a "wireless connection" may also be based on other wireless technologies, such as ultrasonic, infrared light, pulsed radio frequency electromagnetic energy, structured light, or direction-of-arrival techniques used in signal processing (such as audio signal processing or radio frequency processing).
[0175] As used herein, A "and / or" B may mean either "A and B" or "A or B", or both "A and B" and "A or B" are applicable or acceptable.
[0176] As used herein, a unit may include, for example, dedicated hardwired circuitry, software, and / or firmware in combination with programmable circuitry, or a combination thereof.
[0177] The term "computing device" is generally used herein to refer to any one or all of the following: servers, personal computers, laptop computers, tablet computers, mobile devices, cellular telephones, smartbooks, ultrabooks, handheld computers, personal digital assistants (PDAs), wireless email receivers, Internet-enabled multimedia cellular telephones, global positioning system (GPS) receivers, wireless game controllers, and similar electronic devices that include programmable processors and circuitry for wirelessly sending and / or receiving information.
[0178] Various examples have been described. These and other examples are within the scope of the appended claims.
Claims
1. A first device capable of communicating with a second device, the first device comprising: One or more processors configured to: Detect a selection of at least one target object external to the first device; Initiate a communication channel between the first device and a second device associated with the at least one target object external to the first device; Receive an audio packet from the second device in response to the selection of the at least one target object external to the first device; Decode the audio packet received from the second device to generate an audio signal; Apply a spatialization effect to the audio signal based on the selection of the at least one target object external to the first device; Output the audio signal with the spatialization effect; Disable the spatialization effect on the output of the audio signal after the second device is more than a configurable distance away from the first device; Continue to receive the audio packet and decode the audio packet received from the second device to generate the audio signal; And Output an audio signal without the spatialization effect; And A memory coupled to the one or more processors, configured to store the audio packet before and after applying the spatialization effect, Wherein applying a spatialization effect to the audio signal includes reconstructing the sound field of the audio signal based on the distance and angle of the second device relative to the first device for providing to the speakers of the first device, and Wherein the distance and angle of the second device relative to the first device are obtained by: Capturing a video frame including the at least one target object using one or more cameras coupled to the first device; Estimating the distance and angle of the second device relative to the first device from the bounding box of the video frame at the video frame rate; and Interpolating the distance and angle to match the audio frame rate of the audio signal.
2. The first device according to claim 1, wherein the representation of the at least one target object relative to the first device is based on image-based features, the image, or both the image and the features of the image, wherein the image is captured by one or more cameras coupled to the first device.
3. The first device according to claim 1, further comprising one or more transmitter antennas coupled to the one or more processors, configured to transmit communication data of a communication channel between the first device and the second device associated with the at least one target object external to the first device by the one or more processors to the second device.
4. The first device according to claim 1, further comprising one or more receiver antennas coupled to the one or more processors, configured to receive the audio packet based on the result of a communication channel between the at least one target object external to the first device and the first device.
5. The first device according to claim 1, wherein, The selection of the at least one target object is based on the detection of a command signal, the command signal being based on keyword detection.
6. The first device according to claim 1, further comprising a display device configured to represent the at least one target object outside the first device, and wherein, The selection of the at least one target object external to the first device is based on a change in state of a capacitance sensor or an ultrasonic sensor on the display device.
7. The first device according to claim 1, wherein, The selection of the at least one target object is based on detection of a command signal, the command signal detection being based on eye gaze detection.
8. The first device according to claim 1, wherein, The relative position of the second device is represented on the display device as an image of the second device.
9. The first device according to claim 1, wherein, The output of the audio signal is a three-dimensional spatialized audio signal.
10. The first device according to claim 9, further comprising a display device configured to represent the relative position of the second device, and wherein, The output of the three-dimensional spatialized audio signal is based on where the relative position of the second device is represented on the display device.
11. The first device according to claim 9, further comprising a Global Positioning Satellite (GPS) receiver coupled to the one or more processors, configured to assist the first device in performing assisted GPS to determine the relative position of the second device, and wherein, The output of the three-dimensional spatialized audio signal for the selection of the at least one target object external to the first device is based on the assisted GPS.
12. The first device according to claim 9, further comprising one or more sensors coupled to the one or more processors, configured to assist in estimating the relative position of the second device.
13. The first device according to claim 9, wherein, The one or more processors are configured to output the three-dimensional spatialized audio signal at a different spatial resolution when the second device is in a first position relative to the first device as compared to a second position relative to the second device.
14. The first device according to claim 9, Among them, The one or more processors are configured to receive an updated estimate of the relative position of the at least one target object external to the first device based on tracking of the at least one target object external to the first device, and wherein the one or more processors are configured to output the three-dimensional spatialized audio signal based on the updated estimated relative position of the at least one target object external to the first device.
15. The first device according to claim 14, further comprising two or more speakers coupled to the one or more processors, configured to render the three-dimensional spatialized audio signal based on the updated estimated relative position of the at least one target object.
16. The first device according to claim 1, wherein, The first device is a first vehicle.
17. The first device according to claim 1, wherein one of the target objects is a second vehicle, and wherein, The plurality of target objects among the at least one target object includes a plurality of vehicles external to the first device.
18. The first device according to claim 16, wherein, The one or more processors in the first vehicle are configured to receive the audio packets in respective communication channels from each of the plurality of vehicles, and each of the plurality of vehicles is a second vehicle.
19. The first device according to claim 18, wherein, The audio packets represent speech spoken by at least one person in each of the second vehicles.
20. The first device according to claim 19, wherein, The one or more processors are configured to authenticate each person or each vehicle in the second vehicle to facilitate a trusted multi-party conversation between at least one person in the second vehicle and a person in the first vehicle.
21. The first device according to claim 20, wherein, The configurable distance is a distance measurement or a time measurement.
22. A method comprising communication between a first device and a second device, the method comprising: Detecting a selection of at least one target object external to the first device; Initiating a channel for communication between the first device and a second device associated with the at least one target object external to the first device; Receiving an audio packet from the second device in response to the selection of the at least one target object external to the first device; Decode the audio packet received from the second device to generate an audio signal; Apply a spatialization effect to the audio signal based on the selection of the at least one target object external to the first device; Output the audio signal with the spatialization effect; Disable the spatialization effect on the output of the audio signal after the second device is more than a configurable distance away from the first device; Continue to receive the audio packet and decode the audio packet received from the second device to generate the audio signal; And Output an audio signal without the spatialization effect, where applying the spatialization effect to the audio signal includes reconstructing the sound field of the audio signal based on the distance and angle of the second device relative to the first device for providing to the speakers of the first device, and where the distance and angle of the second device relative to the first device are obtained by: Capturing video frames including the at least one target object using one or more cameras coupled to the first device; Estimating the distance and angle of the second device relative to the first device from the bounding boxes of the video frames at the video frame rate; and Interpolating the distance and angle to match the audio frame rate of the audio signal.
23. The method according to claim 22, wherein, The configurable distance is a distance measurement or a time measurement.
24. The method according to claim 22, wherein, The representation of the at least one target object relative to the first device is based on image features, the image, or both the image and the features of the image.
25. The method according to claim 22, wherein The selection of the at least one target object is based on the detection of a command signal, the command signal being based on keyword detection.
26. An apparatus for communication between a first device and a second device, comprising: Components for detecting the selection of at least one target object external to the first device; Components for initiating a channel for communication between the first device and a second device associated with the at least one target object external to the first device; Components for receiving an audio packet from the second device in response to the selection of at least one target object external to the first device; Components for decoding the audio packet received from the second device to generate an audio signal; Components for applying a spatialization effect to the audio signal based on the selection of the at least one target object external to the first device; Components for outputting the audio signal with the spatialization effect; Components for disabling the spatialization effect on the output of the audio signal after the second device is more than a configurable distance away from the first device; Components for continuing to receive the audio packet and decoding the audio packet received from the second device to generate the audio signal; And Components for outputting an audio signal without the spatialization effect, where the components for applying the spatialization effect to the audio signal include components for reconstructing the sound field of the audio signal based on the distance and angle of the second device relative to the first device for providing to the speakers of the first device, The apparatus further includes: Components for capturing video frames including the at least one target object using one or more cameras coupled to the first device; Components for estimating the distance and angle of the second device relative to the first device from the bounding boxes of the video frames at the video frame rate; and Components for interpolating the distance and angle to match the audio frame rate of the audio signal.
27. A non-transitory computer-readable medium storing computer-executable code, the code being executed by one or more processors to: Detect a selection of at least one target object external to the first device; Initiate a communication channel between the first device and a second device associated with the at least one target object external to the first device; Receive audio packets from the second device in response to the selection of the at least one target object external to the first device; Decode the audio packets received from the second device to produce an audio signal; Apply a spatialization effect to the audio signal based on the selection of the at least one target object external to the first device; Output the audio signal with the spatialization effect; Disable the spatialization effect on the output of the audio signal after the second device is more than a configurable distance away from the first device; Continue to receive the audio packets and decode the audio packets received from the second device to produce the audio signal; And Output an audio signal without the spatialization effect, wherein applying a spatialization effect to the audio signal includes reconstructing the sound field of the audio signal based on the distance and angle of the second device relative to the first device for providing to the speakers of the first device, and wherein the distance and angle of the second device relative to the first device are obtained by: Capturing video frames including the at least one target object using one or more cameras coupled to the first device; Estimating the distance and angle of the second device relative to the first device from the bounding boxes of the video frames at the video frame rate; and Interpolating the distance and angle to match the audio frame rate of the audio signal.
Citation Information
Patent Citations
Positional Audio in a Vehicle-to-Vehicle Network
US20090226001A1