Neck-worn audio devices and apparatuses, training data generation methods, apparatuses, and media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2024-09-27
- Publication Date
- 2026-05-29
AI Technical Summary
In split-type VR display devices, neckband-mounted audio devices often fail to achieve optimal head-related spatial sound fields due to improper user posture or obstruction by hair or clothing.
An array consisting of at least three speakers and a microphone is used, combined with a neural network model to detect the wearing posture in real time, and the strategy of the speakers and microphone is adjusted by a controller to ensure the establishment of the optimal sound field.
It achieves a stable head-related spatial sound field under different wearing postures and occlusion conditions, improving the user experience and immersion, while reducing hardware complexity and device size.
Smart Images

Figure CN122122923A_ABST
Abstract
Description
Neck-mounted audio device and equipment, training data generation method, equipment and medium TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of wearable devices, in particular to a neck-mounted audio device, an equipment comprising the neck-mounted audio device, and further relates to a training data generation method, a computer equipment and a computer readable storage medium for implementing the method. BACKGROUND
[0002] With the development of Virtual Reality (VR) technology, it is increasingly applied to various wearable smart devices, such as VR display devices. Most of the current VR display devices adopt an all-in-one machine, which has the problems of large volume and heavy weight. Therefore, long-term wearing will cause the user to feel fatigue, discomfort and other symptoms. The split-type VR display device moves the unnecessary parts (such as system on chip, speaker, microphone, etc.) in the head-mounted display device out of the head-mounted display device and places them in other devices as split-type devices, such as neck-mounted audio devices, thereby overcoming the problems of large volume and heavy weight of the all-in-one VR display device. However, in the split-type VR display device, the neck-mounted audio device may be difficult to achieve the best head-related spatial sound field due to the inappropriate wearing pose of the user or due to the obstruction of hair, clothes, etc.
[0003] SUMMARY
[0004] According to a first aspect of the present disclosure, a neck-mounted audio device is provided, comprising: a wearing support for wearing the neck-mounted audio device on a user's neck; at least three speakers located on the wearing support, the speakers being configured to emit sound signals to establish a head-related spatial sound field around the user's head; at least three sound receiving devices located on the wearing support, the sound receiving devices being configured to: receive the sound signals, and generate corresponding sound data based on the received sound signals; and a controller located on the wearing support, the controller being configured to: determine, based on the sound data obtained from the at least three sound receiving devices, a wearing pose of the neck-mounted audio device on the user's neck relative to the head by using a neural network model, wherein the neural network model is pre-trained using training data, and the training data records the relationship between various preset wearing poses of the neck-mounted audio device on the user's neck and the head-related spatial sound field established by the neck-mounted audio device.
[0005] According to some example embodiments, the controller is further configured to apply a corresponding head-related transfer function to the neck-mounted audio device based on the determined wearing pose.
[0006] According to some example embodiments, the preset wearing poses further include a preset recommended wearing pose, and the controller is further configured to: determine coordinates of a positioning point in the neck-worn audio device in a head space coordinate system with respect to the user based on the wearing pose; determine recommended coordinates of the positioning point in the neck-worn audio device in the head space coordinate system based on the recommended wearing pose; determine the wearing pose as an improper wearing pose when a difference between the coordinates and the recommended coordinates is greater than a preset difference threshold; and prompt the user when the wearing pose is determined as the improper wearing pose.
[0007] According to some example embodiments, the controller is further configured to: assist the user in adjusting the wearing pose of the neck-worn audio device on the user's neck.
[0008] According to some example embodiments, the controller is further configured to: prompt the user to make a corresponding adjustment to the wearing pose of the neck-worn audio device on the user's neck based on the difference between the coordinates and the recommended coordinates.
[0009] According to some example embodiments, the controller is further configured to: determine an occlusion condition in the at least three loudspeakers and the at least three microphones based on the sound data obtained from the at least three microphones; and prompt the user to make an adjustment to the wearing of the neck-worn audio device when there is an occlusion in the at least three loudspeakers and / or the at least three microphones.
[0010] According to some example embodiments, the at least three loudspeakers are caused to emit sound signals in turn; all the microphones are caused to receive the sound signals, and a sum of intensities of the sound signals received by each microphone for all the loudspeakers is determined as a microphone signal overall intensity; and it is determined that one microphone is occluded when the microphone signal overall intensity of the one microphone is lower than that of the other microphones.
[0011] According to some example embodiments, the controller is further configured to: adjust an audio emission strategy applied to the at least three loudspeakers when there is an occlusion in the at least three loudspeakers and the at least three microphones, so that the microphone signal overall intensities of the microphones are the same.
[0012] According to some example embodiments, the controller is further configured to: determine a failure condition of the at least three loudspeakers based on the sound data obtained from the at least three microphones; and adjust an audio emission strategy applied to the loudspeakers that do not have a failure when it is determined that one of the at least three loudspeakers has a failure, to establish a head-related spatial sound field around the user's head.
[0013] According to some example embodiments, the controller is further configured to: cause the at least three speakers to emit sound signals in turn; cause all sound receivers to receive sound signals, and determine a sum of intensities of sound signals received by all sound receivers for one speaker as a total intensity of speaker signals for the one speaker; and determine that the one speaker is faulty when the total intensity of speaker signals for the one speaker is lower than the total intensities of speaker signals for other speakers.
[0014] According to some example embodiments, the number of the at least three sound receivers is equal to the number of the at least three speakers, and each speaker is arranged in a one-to-one corresponding relationship with one sound receiver.
[0015] According to some example embodiments, the number of the at least three speakers is even, and the controller is configured to: when the neck-mounted audio device is worn on a user’s neck, the at least three speakers are symmetrically arranged on both sides of the user’s head.
[0016] According to some example embodiments, the sound receivers comprise microphones.
[0017] According to a second aspect of the present disclosure, there is provided a split-type virtual reality display device, comprising: a head-mounted display device; and a neck-mounted audio device according to the first aspect of the present disclosure and example embodiments thereof.
[0018] According to a third aspect of the present disclosure, there is provided a training data generation method for training a neural network model in a neck-mounted audio device according to the first aspect of the present disclosure and example embodiments thereof, the training data generation method comprising: setting a plurality of preset wearing poses of the neck-mounted audio device on a user’s neck; generating sound data corresponding to each preset wearing pose; and determining a whole of sound data corresponding to all preset wearing poses as the training data.
[0019] According to some example embodiments, the generating sound data corresponding to each preset wearing pose comprises: selecting one preset wearing pose; causing the at least three speakers to emit sound signals in turn using training audio; in a case where each speaker emits a sound signal, causing the at least three sound receivers to generate corresponding sound data based on received sound signals, respectively; selecting another preset wearing pose, and repeating the above steps of causing speakers to emit sound signals and causing sound receivers to generate sound data until all preset wearing poses have been selected.
[0020] According to a fourth aspect of the present disclosure, there is provided a computer device comprising a processor, a memory, and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of the training data generation method according to the third aspect of the present disclosure and each of the exemplary embodiments thereof.
[0021] According to a fifth aspect of the present disclosure, there is provided a computer readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements the steps of the training data generation method according to the third aspect of the present disclosure and each of the exemplary embodiments thereof. BRIEF DESCRIPTION OF DRAWINGS
[0022] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings; in the drawings:
[0023] Fig. 1 schematically illustrates an application scenario of a neck-mounted audio device according to one exemplary embodiment of the present disclosure;
[0024] Fig. 2 schematically illustrates the structure of a neck-mounted audio device according to one exemplary embodiment of the present disclosure in the form of a block diagram;
[0025] Fig. 3 schematically illustrates the structure of a neural network model that can be used to determine a wearing pose according to one embodiment of the present disclosure;
[0026] Fig. 4 schematically illustrates a situation in which a neck-mounted audio device determines an occlusion condition based on sound data acquired from a sound collecting device according to one embodiment of the present disclosure;
[0027] Fig. 5 schematically illustrates the structure of a split-type VR display device according to one exemplary embodiment of the present disclosure in the form of a block diagram;
[0028] Fig. 6 schematically illustrates a training data generation method according to one embodiment of the present disclosure in the form of a flowchart;
[0029] Fig. 7 schematically illustrates details of the training data generation method shown in Fig. 6 according to one exemplary embodiment of the present disclosure in the form of a flowchart;
[0030] Fig. 8 schematically illustrates the structure of a computer device according to one exemplary embodiment of the present disclosure in the form of a block diagram.
[0031] It should be understood that the accompanying drawings are only schematic and are intended to provide a general illustration of the exemplary embodiments of the present disclosure, and are not a limitation. Additionally, in the drawings, the same or similar features are denoted by the same or similar reference signs. DETAILED DESCRIPTION
[0032] The technical solutions according to the present disclosure will be described below in conjunction with the accompanying drawings, so as to make those skilled in the art fully understand and implement the technical solutions according to the present disclosure.
[0033] Referring to FIG. 1, it schematically shows an application scenario of a neck-mounted audio device according to one exemplary embodiment of the present disclosure. In the application scenario 100, the neck-mounted audio device 110 is worn on the neck of a user 120. The neck-mounted audio device 110 comprises a wearing support 111, loudspeakers 112 and microphones 113. The number of loudspeakers 112 and microphones 113 is four respectively, and they are all located on the wearing support 111 and arranged symmetrically on both sides of the head of the user 120 when the neck-mounted audio device 110 is worn, each microphone 113 is arranged in a one-to-one corresponding relationship with one loudspeaker 112. It should be understood that the number and arrangement of the loudspeakers 112 and microphones 113 shown in FIG. 1 are exemplary and not limiting. According to actual needs, the loudspeakers 112 and microphones 113 can have any other suitable number and arrangement, and the present disclosure does not make any limitation thereto. As shown in FIG. 1, the neck-mounted audio device 110 can be worn on the neck of the user 120 through the wearing support 111, so that in the head-neck space coordinate system XYZ of the user 120 as shown in the figure, the neck-mounted audio device 110 has a specific wearing pose relative to the head of the user 120, i.e. the neck-mounted audio device 110 has a specific position and orientation. As shown in FIG. 1, when the neck-mounted audio device 110 is in a certain wearing pose, the positioning points 114a, 114b on the neck-mounted audio device 110 have specific coordinates in the head-neck space coordinate system XYZ, which can be used as the positioning of the neck-mounted audio device 110 corresponding to the wearing pose. In the embodiment shown in FIG. 1, the two end points 114a, 114b of the end of the wearing support 111 of the neck-mounted audio device 110 located in the inner side bending part are selected as the positioning points, however, according to actual needs, any suitable part of the neck-mounted audio device 110 can be selected as the positioning point, as long as its coordinates can determine the position and orientation of the neck-mounted audio device 110. It should also be understood that the head space coordinate system XYZ shown in FIG. 1 is schematic and not limiting, and according to actual needs, the directions of the coordinate axes of the head space coordinate system XYZ can also be determined otherwise, thereby constructing the required head space coordinate system. In this way, the neck-mounted audio device 110 can establish a head-related spatial sound field around the head of the user 120 with the array of loudspeakers 112, so as to provide high-quality audio playback effect to the user 120. It should be understood that the neck-mounted audio device 110 can be used independently as a separate audio player, or can be used together with, for example, a head-mounted display device, thereby being used as a separate device in a separate VR display device.
[0034] Referring to FIG. 2 in conjunction with FIG. 1, a structure of the neck-mounted audio device according to one example embodiment of the present disclosure is schematically shown in a block diagram. As shown in FIG. 2, the neck-mounted audio device 200 includes a wearing support 210, at least three loudspeakers 220, at least three sound receivers 230, and a controller 240. The wearing support 210 is used to wear the neck-mounted audio device 200 on the neck of a user. The at least three loudspeakers 220 are located on the wearing support 210 and are used to form a loudspeaker array. The at least three loudspeakers 220 can emit sound signals so as to be able to construct a corresponding head-related spatial sound field around the head of the user. The types of the at least three loudspeakers 220 can be full-band types covering low frequencies and high frequencies, but can also be a combination of low-frequency band types and high-frequency band types. The at least three sound receivers 230 are also located on the wearing support 210, which form a sound receiver array and are configured to: receive sound signals emitted by the loudspeakers, and based on the received sound signals, generate corresponding sound data. The sound data generated by the sound receivers 230 records the characteristics of the corresponding head-related spatial sound field, which is associated with the wearing pose of the neck-mounted audio device 200 on the neck of the user. Specifically, the sound signals received by the sound receivers generally include direct sound, early reflection sound, and reverberation. The direct sound reflects the distance relationship between the sound source and the sound receiver, the farther the distance, the weaker the sound, and the greater the time difference, while the early reflection sound and the reverberation reflect the spatial position of the sound source. Based on this principle, when all the loudspeakers emit sound signals in turn and all the sound receivers receive sound signals at the same time, as the wearing pose of the neck-mounted audio device changes, the frequency response of the sound received by each sound receiver also differs (i.e., embodies the respective characteristics of the head-related spatial sound field), and there is a correlation between the difference and the change in the wearing pose. Therefore, based on the sound data generated by all the sound receivers, the characteristics of the corresponding head-related spatial sound field can be determined, and in turn the wearing pose of the neck-mounted audio device 200 on the neck of the user can be determined. It should be understood that the sound receivers 230 can be any suitable sound receivers, including but not limited to microphones, sound sensors, etc., the number of the loudspeakers 220 and the sound receivers 230 can be any suitable number respectively according to actual needs, as long as each is not less than three, and the loudspeakers 220 and the sound receivers 230 can be arranged in any suitable manner on the wearing support 210. For example, the loudspeakers 220 and the sound receivers 230 can have equal numbers, and each loudspeaker 220 is arranged in a one-to-one corresponding relationship adjacent to one sound receiver 230. In other embodiments, the number of loudspeakers 220 can be an even number (e.g., can be four), thereby enabling the loudspeakers 220 to be arranged symmetrically about the head of the user on both sides thereof when the neck-mounted audio device 200 is worn, as shown in FIG. 1.The controller 240 can be located on the wearing support 210, and can be configured to determine, based on sound data acquired from the at least three sound collecting devices 220, a wearing pose of the neck-worn audio device 200 on a neck of a user relative to a head by using a neural network model, wherein the neural network model is pre-trained by using training data recording relationships between various preset wearing poses of the neck-worn audio device 200 on the neck of the user and head-related spatial sound fields established by the neck-worn audio device 200. In this way, the neck-worn audio device according to the present disclosure can track and determine a wearing pose of the neck-worn audio device on the neck of the user in real time based on sound data generated by a sound collecting device array formed by all the sound collecting devices, and then can further take corresponding measures according to the determined wearing pose.
[0035] Referring to FIG. 3, a structure of a neural network model that can be used to determine a wearing pose according to one embodiment of the present disclosure is schematically shown, which can be applied in the controller 240 of the neck-worn audio device 200 shown in FIG. 2. As shown in FIG. 3, the neural network model 300 includes an input layer 310, a hidden layer 320 (which can also be referred to as an intermediate layer), and an output layer 330. The input layer 310 includes n input neurons, where n is an integer greater than or equal to 3. Each input neuron receives an input x i , where subscript i is an integer and satisfies 3≤i≤n. For example, for the neck-worn audio device 200 according to the present disclosure, the number of input neurons can be equal to the number of sound collecting devices 230, and thus sound data generated by one sound collecting device 230 can be provided as the input x i to this input neuron. The hidden layer 320 can include a plurality of intermediate neurons, each of which can receive, as input, an output from each input neuron in the input layer 310, perform a corresponding calculation thereon, and then output a calculation result to the next layer. The output layer 330 includes m output neurons, where m is an integer greater than or equal to 1. Each output neuron can receive, as input, an output from each intermediate neuron in the hidden layer 320, perform a corresponding calculation thereon, and then output a calculation result. Thus, the output layer 330 can provide m outputs y j , where subscript j is an integer and satisfies 1≤j≤m. For the neck-worn audio device 200 according to the present disclosure, each output y jThe determined one of the wearing poses can correspond to. In this way, the controller 240 in the neck-worn audio device 200 can determine the wearing pose of the neck-worn audio device 200 on the user's neck based on the sound data acquired from the at least three microphones 230 using the neural network model 300. With continued reference to FIG. 3, the illustrated neural network model 300 is a fully connected feedforward neural network including one hidden layer, however this is merely exemplary and not limiting. The neural network model 300 can have any suitable structure according to actual needs, for example, the number of included hidden layers can be selected according to actual needs (thus, in some cases more than one hidden layer can be included), the number of included intermediate neurons in each hidden layer can also be selected according to actual needs, and the connections between layers do not necessarily have to be fully connected, but can also be partially connected. In addition, the overall architecture of the neural network can also have other types of structures according to actual needs, for example, a recurrent neural network or a convolutional neural network can be used, etc. The present disclosure does not make any limitation on the features of the neural network model used by the neck-worn audio device in the above aspects.
[0036] It should be understood that the neural network model used by the neck-worn audio device according to the present disclosure can be pre-trained using pre-generated training data, where the training data records the relationship between various preset wearing poses of the neck-worn audio device on the user's neck and the head-related spatial sound field established by the neck-worn audio device. In this way, the neck-worn audio device according to the present disclosure does not need to use an inertial measurement unit, a pose estimation sensor, a monocular depth camera, etc. scheme, but only uses the array composed of microphones to detect the head-related spatial sound field constructed by the neck-worn audio device, to track and determine the wearing pose of the neck-worn audio device on the user's neck in real time, thereby greatly reducing the complexity of the hardware structure, reducing the size and weight of the device, and also reducing the manufacturing cost of the device.
[0037] With reference back to FIG. 2 and in conjunction with FIG. 3, in one example embodiment, the controller 240 in the neck-worn audio device 200 can also be configured to apply a corresponding head-related transfer function to the neck-worn audio device 200 based on the determined wearing pose. By applying a suitable head-related transfer function, the neck-worn audio device 200 can construct a personalized head-related spatial sound field for different users’ preferences (e.g., specific wearing habits), thereby being able to provide high-quality audio playback effects, optimize users’ spatial audio experience, enhance the sense of immersion and 3D surround, and improve user experience. Thus, in some example embodiments, the neck-worn audio device 200 according to the present disclosure, upon detecting a change in its wearing pose relative to the user’s head, can adjust the applied head-related transfer function accordingly to make it suitable for constructing the desired head-related spatial sound field, thereby being able to ensure that high-quality audio playback effects are provided to the user, optimize users’ spatial audio experience, enhance the sense of immersion and 3D surround, and improve user experience.
[0038] In one example embodiment, the controller 240 in the neck-mounted audio device 200 can also be configured to prompt the user when the determined wearing pose is determined to be an improper wearing pose. For example, the neck-mounted audio device 200 can utilize the speaker 220 to provide a prompt, such as a voice alert, to the user, or when the neck-mounted audio device 200 is used in a split VR display device, the alert prompt can also be provided to the user through the head-mounted display device. In other examples, the neck-mounted audio device can also include a separate warning module (e.g., a buzzer, a vibrator, or an indicator light, etc.), whereby the user can be prompted through the warning module when the wearing pose is determined to be an improper wearing pose. As a non-limiting example, the controller 240 in the neck-mounted audio device 200 can determine whether the determined wearing pose is an improper wearing pose in the following manner: determine coordinates of a location point in the neck-mounted audio device in a head space coordinate system (e.g., the head space coordinate system XYZ shown in FIG. 1) with respect to the user’s head based on the determined wearing pose; determine recommended coordinates of the location point in the neck-mounted audio device in the head space coordinate system based on the recommended wearing pose; and determine the wearing pose to be an improper wearing pose when a difference between the coordinates and the recommended coordinates is greater than a pre-determined difference threshold. Then, when the wearing pose is determined to be the improper wearing pose, the controller 240 can prompt the user in the manner already described in detail above. Further, in this manner, the recommended coordinates corresponding to the location point on the neck-mounted audio device in the pre-determined recommended wearing pose can be pre-determined and stored, and the coordinates corresponding to the location point on the neck-mounted audio device in various pre-determined wearing poses included in the training data can also be pre-determined and stored. In this way, when the neural network model trained with the training data including such coordinate information determines the wearing pose of the neck-mounted audio device on the user’s neck with respect to the head, the coordinates of the location point on the neck-mounted audio device in the head space coordinate system corresponding to the wearing pose are also determined accordingly, thereby enabling determination of whether the wearing pose is an improper wearing pose based on the difference in the coordinates of the location point on the neck-mounted audio device in the determined wearing pose and the recommended wearing pose.
[0039] In one example embodiment, the controller 240 in the neck-worn audio device 200 can also be configured to assist the user to adjust the wearing pose of the neck-worn audio device 200 on the user’s neck when the determined wearing pose is determined to be an improper wearing pose. For example, the controller 240 can compare the determined wearing pose with the recommended wearing pose, determine the difference between the determined wearing pose and the recommended wearing pose, and then can prompt the user to make a corresponding adjustment to the wearing pose of the neck-worn audio device 200 on the user’s neck (e.g., can prompt the user to adjust the neck-worn audio device 200 in which spatial orientation) based on the difference, so that the determined wearing pose can approach the recommended wearing pose. For example, the controller 240 can determine the difference between the determined wearing pose and the recommended wearing pose based on the difference between the coordinates of the positioning point on the neck-worn audio device in the head space coordinate system in the determined wearing pose and the recommended coordinates of the positioning point on the neck-worn audio device in the head space coordinate system in the recommended wearing pose, and then can prompt the user to make a corresponding adjustment to the wearing pose of the neck-worn audio device 200 on the user’s neck according to the difference.
[0040] In the above manner, the neck-worn audio device according to the present disclosure can determine an improper wearing pose, can prompt the user for the improper wearing pose, and can also assist the user to adjust the wearing pose to the recommended wearing pose or to a wearing pose close to the recommended wearing pose, so as to facilitate the construction of a personalized head-related space sound field, thereby making the neck-worn audio device according to the present disclosure more intelligent.
[0041] In one example embodiment, the controller 240 in the neck-worn audio device 200 can also be configured to determine the occlusion condition in the at least three loudspeakers 220 and the at least three microphones 230 based on the sound data obtained from the at least three microphones 230. Referring to FIG. 4, which schematically illustrates a scenario in which the neck-worn audio device according to the present disclosure determines the occlusion condition based on the sound data obtained from the microphones.
[0042] As shown in FIG. 4, the neck-mounted audio device 200 can include four speakers 220 and four microphones 230, each microphone 230 being arranged in a one-to-one corresponding relationship adjacent to one speaker 220, and the four speakers 220 and the four microphones 230 being arranged on the wearing support 210 in such a way that when the neck-mounted audio device 200 is worn on the neck of the user 120, the four speakers 220 are symmetrically arranged on both sides of the head of the user 120. Since the neck-mounted audio device 200 is worn on the neck of the user 120, there can be a shielding of the speakers and / or the microphones by, for example, a collar, hair, etc., which can cause an adverse effect on the sound emission (i.e., emitting a sound signal) of the speakers and the sound reception (i.e., receiving a sound signal) of the microphones. In order to eliminate the difference in the head-related spatial sound field caused by the shielding, the monitoring of the sound field during the wearing process can be used to determine the presence of the shielding, and the user can be warned and reminded of the presence of the shielding.
[0043] With reference back to FIG. 4, in some examples, for the neck-mounted audio device 200 worn on the neck of the user 120, the speakers 220a, 220b can be sounded, and the soundings of the microphones 230c, 230d can be focused on, then the speakers 220c, 220d can be sounded, and the soundings of the microphones 230a, 230b can be focused on. In the above two sounding cases, the sound data recorded by the microphones 230a, 230c are substantially the same (i.e., the frequency band and intensity of the sound signals are substantially the same, and the same applies hereinafter), the sound data recorded by the microphones 230b, 230d are substantially the same, and the signal intensity of the sound signals received by the microphones 230a, 230b, 230c, 230d are all greater than the signal intensity threshold, it can be determined that there is no obstruction for the speakers 220a, 220b, 220c, 220d and the microphones 230a, 230b, 230c, 230d. In some examples, the speakers 220a, 220b, 220c, 220d can be sounded in turn, the microphones 230a, 230b, 230c, 230d can all receive the sound signals, and the sum of the intensities of the sound signals received by each microphone for all the speakers can be determined as the microphone signal overall intensity. If the microphone signal overall intensity of one of the microphones 230a, 230b, 230c, 230d is lower than the microphone signal overall intensities of the other microphones, it can be determined that there is an obstruction for the microphone. In some examples, if the signal intensity of the sound signals received by all the microphones 230a, 230b, 230c, 230d is lower than the signal intensity of the sound signals received from the other speakers when one of the speakers 220a, 220b, 220c, 220d is sounded, it can be determined that there is an obstruction for the speaker. According to the above-illustrated judgment principles, the speakers 220a, 220b, 220c, 220d can be sounded in turn or in groups, and after all the soundings of the microphones 230a, 230b, 230c, 230d are completed, the controller 240 can be used to determine whether there is an obstruction for the speakers 220a, 220b, 220c, 220d and the microphones 230a, 230b, 230c, 230d.
[0044] In one example embodiment, the controller 240 can also be configured to prompt the user to adjust the wearing of the neck-mounted audio device 200 when there is an obstruction in the at least three loudspeakers and / or the at least three microphones. For example, the user can be prompted to remove obstructions such as a collar and / or hair from the loudspeakers and / or the microphones. In addition, in another example embodiment, when the obstruction cannot be completely removed (e.g., obstruction by hair cannot be completely avoided after wearing), the controller 240 can also be configured to adjust the audio sound emitting strategy applied to the at least three loudspeakers when there is an obstruction in the at least three loudspeakers and the at least three microphones, so that the microphone signal overall intensity of each microphone is the same. As an example of adjusting the audio sound emitting strategy, the corresponding loudspeaker among the at least three loudspeakers can be caused to increase the emitted sound pressure to increase the intensity of the sound signal emitted by the corresponding loudspeaker, so that the microphone signal overall intensity of each microphone is the same. It should be understood that the corresponding loudspeaker can refer to, for example, the obstructed loudspeaker in the case of obstruction of the loudspeaker, and the corresponding loudspeaker can refer to, for example, the loudspeaker arranged adjacent to the obstructed microphone in the case of obstruction of the microphone. In this way, the neck-mounted audio device can also ensure that a stable and reliable head-related spatial sound field is established as much as possible in the presence of obstructions.
[0045] In one example embodiment, the controller 240 can also be configured to adjust the audio sound emitting strategy applied to the non-faulty loudspeakers when it is determined that one of the at least three loudspeakers 220 is faulty, so that the neck-mounted audio device 200 establishes a corresponding head-related spatial sound field around the user's head. For example, in the case of sound emission by one of the loudspeakers 220a, 220b, 220c, 220d, the sum of the signal intensities of the sound signals received by all the microphones 230a, 230b, 230c, 230d is determined as the loudspeaker signal overall intensity of the loudspeaker. When the loudspeaker signal overall intensity of one of the loudspeakers 220a, 220b, 220c, 220d is lower than a preset loudspeaker signal overall intensity threshold, it is determined that the loudspeaker is faulty. When a certain loudspeaker is faulty and cannot work normally, the controller 240 can adjust the audio sound emitting strategy applied to the non-faulty loudspeakers according to the status of the currently available loudspeakers, so as to be able to reconstruct an effective spatial sound field without causing the neck-mounted audio device 200 to be in an erroneous spatial sound field state.
[0046] Referring to FIG. 5, a structure of a split VR display device according to one example embodiment of the present disclosure is schematically shown in a block diagram. As shown in FIG. 5, the split VR display device 400 can include a head-mounted display device 410 and a neck-mounted audio device 420 independent of each other, wherein the neck-mounted audio device 420 can be implemented as the neck-mounted audio device 200 and example embodiments thereof described above in connection with FIG. 2. In this way, the split VR display device 400 can track and determine the wearing pose of the neck-mounted audio device on the user's neck in real time only by detecting the head-related spatial sound field constructed by the neck-mounted audio device using an array of radio devices, without using an inertial measurement unit, a pose estimation sensor, a monocular depth camera, etc., thereby greatly reducing the complexity of hardware configuration, reducing the volume and weight of the device, and also reducing the manufacturing cost of the device.
[0047] Referring to FIG. 6, a training data generation method according to one embodiment of the present disclosure is schematically shown in a flowchart. As shown in FIG. 6, the training data generation method 500 includes steps 510, 520 and 530:
[0048] In step 510, a preset wearing pose of the neck-mounted audio device on the user's neck is set;
[0049] In step 520, sound data corresponding to each preset wearing pose is generated;
[0050] In step 530, the whole of the sound data corresponding to all the preset wearing poses is determined as the training data.
[0051] The training data generated by the training data generation method 500 can be used to train a neural network model in a neck-mounted audio device (e.g., the neck-mounted audio device 200 described above in connection with FIG. 2 and its example embodiments) according to the present disclosure. Specifically, at step 510, the training data generation method 500 can pre-set various pre-set wearing poses of the neck-mounted audio device on a user’s neck according to actual needs. In some embodiments, the pre-set various pre-set wearing poses can also include inappropriate wearing poses so that the trained neural network model can determine which wearing poses are inappropriate in use. The differences among these pre-set wearing poses will result in differences among the head-related spatial sound fields constructed by the neck-mounted audio device, and thus, by detecting each head-related spatial sound field, the corresponding wearing pose can be determined. At step 520, for each pre-set wearing pose, a known training audio is used to make all the loudspeakers sound (i.e., emit sound signals) in turn, and all the microphones receive sound (i.e., receive sound signals) to generate corresponding sound data, and thus, the sound data reflects the characteristics of the corresponding head-related spatial sound field (and the corresponding pre-set wearing pose). In one embodiment, step 520 can be specifically implemented as: selecting a pre-set wearing pose; using a training audio to make the at least three loudspeakers sound in turn; in the case that each loudspeaker sounds, the at least three microphones generate corresponding sound data based on the received sound signals, respectively; selecting another pre-set wearing pose, and repeating the above steps of making loudspeakers sound and using microphones to generate sound data until all pre-set wearing poses have been selected.
[0052] Referring to FIG. 7 and in combination with FIG. 6, FIG. 7 further illustrates the details of step 520 in the training data generation method 500 shown in FIG. 5 according to one exemplary embodiment of the present disclosure in the form of a flowchart. As shown in FIG. 7, in this exemplary embodiment, step 520 can be implemented in the following manner: in step 520a, the nth preset wearing pose can be selected, where n is an integer greater than zero, and when initially selected, n can be selected as 1 (i.e., the first preset wearing pose is selected); in step 520b, the mth loudspeaker is made to emit sound using the training audio, where m is an integer greater than zero, and again when initially selected, m can be selected as 1 (i.e., the first loudspeaker is made to emit sound); in step 520c, all the sound receiving devices receive the sound signals and generate corresponding sound data respectively; in step 520d, it is determined whether the value of m is equal to the number M of loudspeakers, if the value of m is not equal to M, it means that there are still loudspeakers that have not emitted sound, so the value of m is incremented by 1 and the process returns to step 520b, if the value of m is equal to M, it means that all the loudspeakers have emitted sound, so the process can proceed to step 520e. In step 520e, it is determined whether the value of n is equal to the number N of preset wearing poses, if the value of n is not equal to N, it means that there are still preset wearing poses for which sound data has not been generated, so the value of n is incremented by 1 and the value of m is set to 1 and the process returns to step 520a, whereby another preset wearing pose is selected and all the loudspeakers are made to emit sound in turn so that all the sound receiving devices receive the sound signals and generate sound data, if the value of n is equal to N, it means that sound data has been generated for all the preset wearing poses, so step 520 can end.
[0053] Continuing to refer to FIG. 6, after sound data has been generated for all the preset wearing poses, the entirety of the sound data corresponding to all the preset wearing poses is determined as the training data. As can be seen, the training data generated in the above manner records the relationship between various preset wearing poses of the neck-mounted audio device on the user's neck and the head-related spatial sound field established by the neck-mounted audio device. Then, any suitable neural network model can be trained using the training data in a suitable manner to obtain a neural network model that can be applied in the neck-mounted audio device according to the present disclosure. Therefore, the neck-mounted audio device according to the present disclosure uses the neural network model trained using the training data to detect the wearing pose of the neck-mounted audio device on the user's neck based on the array of sound receiving devices that detect the head-related spatial sound field established by the neck-mounted audio device in real time, thereby greatly reducing the complexity of the hardware configuration, reducing the size and weight of the device, and also reducing the manufacturing cost of the device.
[0054] Referring to FIG. 8, the structure of a computer device according to one example embodiment of the present disclosure is schematically shown in a block diagram. As shown in FIG. 8, the computer device 600 includes a processor 610 and a memory 620. The memory 620 stores a computer program, and the processor 610 executes the computer program to implement the steps of the training data generation method described above in connection with the example embodiments shown in FIG. 6, FIG. 7.
[0055] The processor 610 can be a single processing unit or a plurality of processing units, all of which can include single or multiple computing units or multiple cores. The processor 610 can be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 610 can be configured to fetch and execute computer-readable program instructions stored in the memory 620 or other computer-readable storage media.
[0056] The memory 620 is an example of computer-readable storage media for storing instructions that can be executed by the processor 610 to implement the various steps described above. By way of example, the memory 620 can include both volatile memory (e.g., RAM) and non-volatile memory (e.g., ROM). In some embodiments, the memory 620 can also include hard disk drives, solid state drives, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CD or DVD), storage arrays, network-attached storage, storage area networks, and the like. Accordingly, the memory 620 can be collectively referred to herein as computer-readable storage or computer-readable storage media, and can be non-transitory media that can store computer-readable, processor-executable computer programs as computer-executable code.
[0057] The present disclosure also relates to a computer-readable storage medium configured to store a computer program, which, when executed by a processor, implements the steps of the training data generation method described above. It should be understood that the computer-readable storage medium should be any suitable storage medium, including but not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or semiconductor media (e.g., solid-state hard drives), or any other non-transitory medium that can be used to store information for access by a processor. The present disclosure does not limit the computer-readable storage medium.
[0058] Furthermore, the present disclosure also relates to a computer program product comprising computer-executable instructions configured to cause a processor to implement the steps of the training data generation method described above when executed on the processor.
[0059] The terminology used in the present disclosure is only for the purpose of describing embodiments of the present disclosure and is not intended to limit the present disclosure. As used in the present disclosure, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and "comprising," when used in this specification, specify the presence of stated features, but do not preclude the presence or addition of one or more other features. As used in the present disclosure, the term "and / or" includes any and all combinations of one or more of the associated listed items. It will be understood that, although the terms "first," "second," "third," etc. can be used herein to describe various features, these features should not be limited by these terms. These terms are only used to distinguish one feature from another.
[0060] Unless otherwise defined, all terms (including technical and scientific terms) used in the present disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or the present specification and will not be interpreted in an idealized or overly formal sense unless expressly so defined in the present disclosure.
[0061] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples, or can omit some technical features from the different embodiments or examples described in the specification, and the embodiments or examples obtained based on such combination, combination or omission are also considered to fall within the scope of the present disclosure.
[0062] The methods described in the present disclosure include one or more steps or actions. The method steps and / or actions need not be performed in the order described in the present disclosure, but can be performed in different orders, for example, they can be performed simultaneously or in reverse order, as long as the principles of the technical solutions described in the present disclosure are not contradicted. In addition, according to actual needs, the steps or actions in the methods described in the present disclosure can be replaced by different steps or actions, or additional steps or actions can also be included.
[0063] Although the present disclosure has been described in detail in conjunction with some exemplary embodiments, it is not limited to the specific forms described in the present disclosure. Instead, the scope of the present disclosure is only limited by the appended claims.
Claims
1. A neck-mounted audio device, comprising: a wearing support for wearing the neck-mounted audio device on a user’s neck; at least three loudspeakers on the wearing support, configured to emit sound signals to establish a head-related spatial sound field around the user’s head; at least three microphones on the wearing support, configured to receive the sound signals and generate corresponding sound data based on the received sound signals; and a controller on the wearing support, configured to determine a wearing pose of the neck-mounted audio device on the user’s neck relative to the head based on the sound data obtained from the at least three microphones, wherein the controller is pre-trained with a neural network model trained with training data recording a relationship between a preset wearing pose of the neck-mounted audio device on the user’s neck and a head-related spatial sound field established by the neck-mounted audio device. The controller is further configured to apply a corresponding head-related transfer function to the neck-mounted audio device based on the determined wearing pose.
2. The neck-worn audio device of claim 1, wherein, The preset wearing pose further comprises a preset recommended wearing pose, and the controller is further configured to:
3. The neck-worn audio device of claim 1, wherein, determine a coordinate of a positioning point in the neck-mounted audio device in a head spatial coordinate system relative to the user’s head based on the wearing pose; determine a recommended coordinate of the positioning point in the head spatial coordinate system based on the recommended wearing pose; determine the wearing pose as an improper wearing pose when a difference between the coordinate and the recommended coordinate is greater than a preset difference threshold; and prompt the user when the wearing pose is determined as the improper wearing pose. The controller is further configured to assist the user to adjust the wearing pose of the neck-mounted audio device on the user’s neck.
4. The neck-worn audio device of claim 3, wherein, The controller is further configured to prompt the user to make a corresponding adjustment to the wearing pose of the neck-mounted audio device on the user’s neck based on the difference between the coordinate and the recommended coordinate.
5. The neck-worn audio device of claim 4, wherein, The controller is further configured to:
6. The neck-worn audio device of claim 1, wherein, determine an occlusion condition in the at least three loudspeakers and the at least three microphones based on the sound data obtained from the at least three microphones; and prompt the user to adjust the wearing of the neck-mounted audio device when there is an occlusion in the at least three loudspeakers and / or the at least three microphones. The controller is further configured to:
7. The neck-worn audio device of claim 6, wherein, make the at least three loudspeakers emit sound signals in turn; make all the microphones receive sound signals, and determine a sum of intensities of sound signals received by each microphone for all the loudspeakers as a microphone signal overall intensity; and determine that a microphone is occluded when the microphone signal overall intensity of the microphone is lower than that of other microphones. 8. The neck-worn audio device of claim 7, wherein, The controller is further configured to adjust an audio sound emitting strategy applied to the at least three loudspeakers so that sounder signal overall intensities of each sounder are equalized when there is an occlusion among the at least three loudspeakers and the at least three sounders.
9. The neck-worn audio device of claim 1, wherein, The controller is further configured to: determine a failure condition of the at least three loudspeakers based on sound data acquired from the at least three sounders; adjust an audio sound emitting strategy applied to the loudspeakers that have not failed when it is determined that one of the at least three loudspeakers has failed so that the neck-mounted audio device establishes a head-related spatial sound field around the user’s head.
10. The neck-worn audio device of claim 9, wherein, The controller is further configured to: make the at least three loudspeakers emit sound signals in turn; make all sounders receive sound signals and determine a sum of intensities of sound signals received by all sounders for one loudspeaker as a loudspeaker signal overall intensity; determine that the one loudspeaker has failed when the loudspeaker signal overall intensity of the one loudspeaker is lower than a preset loudspeaker signal overall intensity threshold. The number of the at least three sounders is equal to the number of the at least three loudspeakers, and each loudspeaker is disposed in a one-to-one corresponding relationship with one sounder.
11. The neck-worn audio device of claim 1, wherein, The number of the at least three loudspeakers is even, and the at least three loudspeakers are configured to be symmetrically arranged on two sides of the user’s head when the neck-mounted audio device is worn on the user’s neck.
12. The neck-worn audio device of claim 11, wherein, The sounders include microphones.
13. The neck-worn audio device of claim 1, wherein, 14. A split-type virtual reality display device, comprising: a head-mounted display device; and the neck-mounted audio device according to any one of claims 1 to 13.
15. A training data generation method, the training data being used to train a neural network model in the neck-mounted audio device according to claim 1, the training data generation method comprising: setting preset wearing poses of the neck-mounted audio device on a user’s neck; generating sound data corresponding to each preset wearing pose; determining a whole of sound data corresponding to all preset wearing poses as the training data. The generating of sound data corresponding to each preset wearing pose comprises: selecting one preset wearing pose; 16. The training data generation method of claim 15, wherein, making the at least three loudspeakers emit sound signals in turn using training audio; in a case where each loudspeaker emits a sound signal, making the at least three sounders respectively generate corresponding sound data based on received sound signals; selecting another preset wearing pose and repeating the above steps of making loudspeakers emit sound signals and making sounders generate sound data until all preset wearing poses have been selected. The processor executes the computer program to implement the steps of the training data generation method according to claim 15 or 16. The computer program, when executed by the processor, implements the steps of the training data generation method according to claim 15 or 16.
17. A computer device comprising a processor, a memory, and a computer program stored on the memory, wherein, 18. A computer readable storage medium having stored thereon a computer program, wherein,