Sound directional frequency conversion method and system, hearing aid equipment, medium and hearing aid robot

By obtaining image information for face recognition and microphone array sound processing, directional frequency conversion is realized, providing a personalized auditory experience for patients with hearing impairments, solving the problem that existing hearing aid devices cannot provide personalized solutions for a variety of hearing-impaired groups, and improving the auditory experience and data processing efficiency.

CN120302223APending Publication Date: 2025-07-11HAINAN QINGWEN YAYIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510430906.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing hearing aids cannot provide personalized hearing solutions for hearing-impaired patients, especially in sound playback locations where special care is not available for multiple hearing-impaired groups.

Method used

By obtaining personnel image information for face recognition, combining microphone array sound for beam synthesis and pronunciation object recognition, using sound recognition results for separation and frequency conversion, matching the current frequency conversion scheme database stored in the cloud, and directional frequency conversion is performed through the local-stop sound conversion algorithm to output directional frequency conversion sound.

Benefits of technology

Provide clear and high-quality auditory experiences for hearing-impaired patients, realize personalized listening solutions, reduce local data storage pressure, improve data processing speed, adapt to more customer groups, and provide convenient and efficient life assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302223A_ABST
    Figure CN120302223A_ABST
Patent Text Reader

Abstract

The invention provides a sound directional frequency conversion method and system, hearing aid equipment, a medium and a hearing aid robot, and the method comprises the steps: carrying out the face recognition of the image information of a person, and obtaining a corresponding face recognition result; performing beam forming on the microphone array sound, and performing pronunciation object recognition and pronunciation object direction recognition according to a beam forming result, or performing scene recognition on external input sound; and performing sound separation based on the sound recognition result, matching a current frequency conversion scheme from a frequency conversion scheme database stored in the cloud according to the face recognition result and the sound separation result, performing directional frequency conversion according to the current frequency conversion scheme by using a sound frequency conversion algorithm stored in the local end, and outputting directional frequency conversion sound to the target person. Through directional frequency conversion of sound, high-quality auditory experience can be provided for hearing impairment patients, the hearing impairment patients can be better integrated into the society and life, and the system has wide application prospects in the fields of old-age care, medical treatment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and sound processing, and particularly to a sound directional frequency conversion method and system, a hearing aid device and medium, and a hearing aid robot. Background Art

[0002] Hearing impairment refers to organic or functional abnormalities in the sound transmission, sound perception, and various levels of nerve centers for comprehensive sound analysis in the auditory system, resulting in varying degrees of hearing loss. Hearing impaired patients mainly include those with congenital diseases, the elderly with hearing decline, those with deafness caused by diseases, and specific groups with hearing loss.

[0003] With the increasing aging of society, the number of hearing impaired patients is also increasing. For hearing impaired patients, they can improve their adaptability by wearing hearing aid devices. However, in some places where sound is played, the sound is set for the general population, and there is no special device to take special care of hearing impaired patients, and it cannot be used for multiple hearing impaired populations to provide personalized hearing solutions. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a sound directional frequency conversion method and system, a hearing aid device and medium, and a hearing aid robot, which are used to solve the problems existing in the prior art.

[0005] To achieve the above purpose and other related purposes, the present invention provides a sound directional frequency conversion method, including the following steps:

[0006] Obtain personnel image information, and perform face recognition on the personnel image information to obtain the corresponding face recognition result; wherein, the personnel image information includes a target person determined in advance or in real time;

[0007] Perform beamforming on the microphone array sound, and perform pronunciation object recognition and pronunciation object direction recognition according to the beamforming result, or perform scene recognition on the external input sound;

[0008] Perform sound separation based on the sound recognition result, wherein the sound recognition result is the pronunciation object recognition result and pronunciation object direction recognition result corresponding to the microphone array sound, or the sound recognition result is the scene recognition result corresponding to the external input sound;

[0009] Match the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud according to the face recognition result and the sound separation result, and perform directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end, and output the directional frequency conversion sound to the target person.

[0010] In an embodiment of the present invention, the process of sound separation based on the sound recognition result includes:

[0011] Converting the microphone array sound into a microphone array sound stream based on the pronunciation object recognition result and the pronunciation object direction recognition result, or converting the external input sound into an external input sound stream based on the scene recognition result;

[0012] Using a pre-trained or real-time trained text recognition model to recognize the text in the microphone array sound stream or the external input sound stream, and performing time marking on the text starting point to obtain corresponding time marking points;

[0013] According to the time marking points, using a pre-trained or real-time trained voiceprint recognition model to perform voiceprint recognition on the microphone array sound or the external input sound, and performing voiceprint time marking to obtain corresponding voiceprint time marking points;

[0014] Starting a new thread, and according to the voiceprint time marking points, using a pre-trained or real-time trained human voice cancellation model to filter the microphone array sound or the external input sound to obtain the corresponding background sound.

[0015] In an embodiment of the present invention, the process of performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end includes:

[0016] According to the current frequency conversion scheme, using the human voice frequency conversion algorithm stored in the local end to change the sound frequency distribution of the human voice audio signal in the target sound, increase the frequency components, increase the sampling points of the human voice audio signal in the target sound, change the frequency of the human voice audio signal in the target sound, reduce the sampling rate of the human voice audio signal in the target sound, and adjust the dynamic range of the human voice audio signal in the target sound; and,

[0017] According to the current frequency conversion scheme, using the environmental frequency conversion algorithm stored in the local end to first perform a fast Fourier transform on the environmental audio signal in the target sound, then perform an inverse fast Fourier transform, and reduce the sampling rate; and,

[0018] According to the current frequency conversion scheme, using the music frequency conversion algorithm stored in the local end to fit the music audio signal in the target sound;

[0019] Wherein, the target sound is the microphone array sound or the external input sound.

[0020] In an embodiment of the present invention, the process of performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end further includes:

[0021] Record the frequency conversion result output by the voice frequency conversion algorithm stored locally as the voice frequency conversion result, and record the frequency conversion result output by the environmental frequency conversion algorithm stored locally as the environmental frequency conversion result, and record the frequency conversion result output by the music frequency conversion algorithm stored locally as the music frequency conversion result;

[0022] Perform digital processing on the voice frequency conversion result, the environmental frequency conversion result, and the music frequency conversion result, and merge the voice frequency conversion result, the environmental frequency conversion result, and the music frequency conversion result into an audio stream based on the digital processing result;

[0023] Perform high-frequency fitting on the low-frequency part in the audio stream, and perform sound zone adjustment according to the high-frequency fitting result, including adjusting the decibel and phase, so that the adjusted sound becomes directional.

[0024] In an embodiment of the present invention, the method further includes:

[0025] Perform pose recognition on the personnel image information to obtain the corresponding pose recognition result;

[0026] Compare the pose recognition result with the standard behavior pose, and give a warning of abnormal human pose when the pose recognition result is inconsistent with the standard behavior pose; and,

[0027] Combine the pose recognition result with the face recognition result, predict the behavior action of the target person at the next moment according to the standard behavior pose, and make an adjustment based on the predicted behavior action at the next moment when outputting the directional frequency conversion sound to the target person.

[0028] In an embodiment of the present invention, before matching the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud according to the face recognition result and the sound separation result, the method further includes:

[0029] Output a test sound to the target person;

[0030] Receive the hearing test data fed back by the target person based on the test sound, and generate a hearing test curve of the target person according to the hearing test data;

[0031] Feed the hearing test curve and the hearing standard curve back to the cloud, so that the cloud generates the frequency sound offset of the target person according to the hearing test curve and the hearing standard curve, forms a frequency conversion scheme based on the frequency sound offset, and adds the frequency conversion scheme to the frequency conversion scheme database.

[0032] The present invention also provides a sound directional frequency conversion system, and the system includes:

[0033] An image acquisition module, configured to acquire personnel image information; wherein, the personnel image information includes a target person determined in advance or in real time;

[0034] A face recognition module, configured to perform face recognition on the personnel image information to obtain a corresponding face recognition result;

[0035] A voice recognition module, configured to perform beamforming on the microphone array voice, and perform pronunciation object recognition and pronunciation object direction recognition according to the beamforming result, or perform scene recognition on the externally input voice;

[0036] A voice separation module, configured to perform voice separation according to the voice recognition result; wherein, the voice recognition result is the pronunciation object recognition result and pronunciation object direction recognition result corresponding to the microphone array voice, or the voice recognition result is the scene recognition result corresponding to the externally input voice;

[0037] A voice directional frequency conversion module, configured to match a current frequency conversion scheme from a frequency conversion scheme database stored in the cloud according to the face recognition result and the voice separation result, and perform directional frequency conversion according to the current frequency conversion scheme by using a voice frequency conversion algorithm stored in the local end, and output a directional frequency conversion voice to the target person.

[0038] The present invention further provides a hearing aid robot, including the voice directional frequency conversion system as described above.

[0039] The present invention further provides a hearing aid device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the voice directional frequency conversion method described in any one of the above.

[0040] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the voice directional frequency conversion method described in any one of the above are implemented.

[0041] As described above, the present invention provides a method and system for sound directional frequency conversion, a hearing aid device and medium, and a hearing aid robot, which have the following beneficial effects: By performing directional frequency conversion on sound, the present invention can provide a clear and high-quality auditory experience for hearing-impaired patients, enabling them to better integrate into society and life. At the same time, by realizing the adjustment of sound directional frequency conversion, the parametric array directional sound system can form a focused sound field, making the sound clearly audible to the listeners within the sound field and rapidly attenuating acoustically outside the sound field, playing the role of not disturbing the people and spreading key information. Moreover, by performing face recognition on images, the present invention can identify and track faces in real time, thereby providing more personalized and high-quality auditory and interactive services for users, making communication more convenient and efficient. In addition, by storing image information, sound, and sound frequency conversion algorithms in the local end, the present invention has extremely high privacy characteristics. At the same time, by storing the frequency conversion scheme in the cloud and through the linkage method between the cloud and the local end, not only can it adapt to more customer groups and provide personalized hearing solutions for more hearing-impaired patients and specific range of listening population, but also it can reduce the local end data storage pressure and improve the data processing speed. Therefore, by combining face tracking technology and sound directional frequency conversion technology, the present invention can follow the movement of specific listeners after face recognition, emit a directional sound field for them, provide more personalized services, provide a clear auditory experience for hearing-impaired patients, and has broad application prospects in the fields of elderly care, medical treatment, etc., and can provide more convenient and efficient life assistance for groups such as the elderly and hearing-impaired patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic flowchart of the sound directional frequency conversion method provided in an embodiment of the present invention;

[0043] Figure 2 It is a schematic diagram of the principle framework of the sound directional frequency conversion method provided in an embodiment of the present invention;

[0044] Figure 3 It is a schematic diagram of the hardware structure of the sound directional frequency conversion system provided in an embodiment of the present invention;

[0045] Figure 4 It is a schematic diagram of the principle of the hearing aid robot performing sound directional frequency conversion provided in an embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of the principle of the motion control system in the hearing aid robot provided in an embodiment of the present invention;

[0047] Figure 6 It is a schematic diagram of the hardware structure of the hearing aid device suitable for implementing one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following describes the implementation manners of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It can be understood that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. In addition, it can be understood that the drawings provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, numbers, and proportions of the components in actual implementation can be arbitrarily changed, and the layout type of the components may also be more complex.

[0049] Figure 1 A flowchart of a method for sound directional frequency conversion is shown. Specifically, in an exemplary embodiment, as Figure 1 shown, this embodiment provides a method for sound directional frequency conversion, including the following steps:

[0050] S110, obtain personnel image information, and perform face recognition on the personnel image information to obtain the corresponding face recognition result; wherein, the personnel image information includes a target person determined in advance or in real time. As an example, in this embodiment or other embodiments, the target person may be a group of people such as the elderly, hearing-impaired patients, etc., such as people with hereditary hearing loss, hyperacusis, etc., and people who are deaf due to diseases (such as long-term exposure to a noisy environment, middle ear or outer ear lesions, etc.). As an example, in this embodiment or other embodiments, the personnel image information may be an image or video stream including the target person captured by a camera. When performing face recognition on the personnel image information, a face recognition model and a face image library stored in the local database can be used to compare the personnel image information obtained in real time to obtain the corresponding face recognition result.

[0051] S120, perform beamforming on the microphone array sound, and perform pronunciation object recognition and pronunciation object direction recognition according to the beamforming result, or perform scene recognition on the external input sound. As an example, in this embodiment or other embodiments, the external input sound may be other external audio actively input, for example, other audio is played through a speaker, and then the sound played by the speaker is used as the external input sound.

[0052] S130, perform voice separation based on the voice recognition result, where the voice recognition result is the recognition result of the pronunciation object corresponding to the microphone array voice and the recognition result of the pronunciation object direction, or the voice recognition result is the scene recognition result corresponding to the external input voice. As an example, in this embodiment or other embodiments, voice separation includes, but is not limited to: text recognition and pronunciation object separation, background voice separation.

[0053] S140, match the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud according to the face recognition result and the voice separation result, perform directional frequency conversion according to the current frequency conversion scheme using the voice frequency conversion algorithm stored in the local end, and output the directional frequency conversion voice to the target person.

[0054] According to the above description, in an exemplary embodiment, the process of voice separation based on the voice recognition result in step S130 includes: converting the microphone array voice into a microphone array voice stream based on the pronunciation object recognition result and the pronunciation object direction recognition result, or converting the external input voice into an external input voice stream based on the scene recognition result; using a pre-trained or real-time trained text recognition model to recognize the text in the microphone array voice stream or the external input voice stream, and performing time marking on the starting point of the text to obtain corresponding time marking points; according to the time marking points, using a pre-trained or real-time trained voiceprint recognition model to perform voiceprint recognition on the microphone array voice and the external input voice, and performing voiceprint time marking to obtain corresponding voiceprint time marking points; starting a new thread, and according to the voiceprint time marking points, using a pre-trained or real-time trained human voice cancellation model to filter the microphone array voice and the external input voice to obtain the corresponding background voice. Therefore, in this embodiment, by converting the audio into a stream mode and then inputting it into three pre-trained or real-time trained neural network models, the first neural network model is a text recognition model, which is used to recognize the text in the voice stream and perform time marking on the starting point of the text; the second neural network model is a voiceprint recognition model, which performs voiceprint recognition on the voice according to the time marking points and marks the voiceprint time marking points. Synchronously start the second thread, and eliminate all human voices through the third neural network model, that is, eliminate all human voices through the human voice cancellation model to obtain the background voice. Among them, the training processes of the human voice cancellation model, the text recognition model, and the voiceprint recognition model can refer to related technologies, and will not be elaborated in this embodiment. As an example, in this embodiment or other embodiments, the text recognition model and the voiceprint recognition model belong to the same data processing thread, and the human voice cancellation model is in a different data processing thread from the text recognition model and the voiceprint recognition model. This embodiment uses two threads, which can process different tasks or data sets simultaneously, so as to make full use of the computing power of the multi-core processor and improve the overall execution efficiency. At the same time, in this embodiment, by adding data processing threads, new functions or services can be easily added without large-scale modification and reconstruction of the existing data processing threads, so as to improve scalability and maintainability and facilitate subsequent upgrade and function expansion.

[0055] According to the above description, in an exemplary embodiment, as Figure 2As shown, the process of step S140 for performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end includes: according to the current frequency conversion scheme, using the human voice frequency conversion algorithm stored in the local end to change the sound frequency distribution of the human voice audio signal in the target sound, increasing the frequency components, adding sampling points to the human voice audio signal in the target sound, changing the frequency of the human voice audio signal in the target sound, reducing the sampling rate of the human voice audio signal in the target sound, and adjusting the dynamic range of the human voice audio signal in the target sound; and, according to the current frequency conversion scheme, using the environmental frequency conversion algorithm stored in the local end to first perform a fast Fourier transform on the environmental audio signal in the target sound, then perform an inverse fast Fourier transform, and reduce the sampling rate; and, according to the current frequency conversion scheme, using the music frequency conversion algorithm stored in the local end to fit the music audio signal in the target sound. Wherein, the target sound is the microphone array sound or the external input sound. Since the sound after frequency conversion generally has serious distortion problems, in this embodiment, according to the current frequency conversion scheme, the human voice frequency conversion algorithm stored in the local end is used to spread spectrum, interpolate, frequency convert, downsample and compress the human voice audio signal in the microphone array sound or the external input sound, so that the sound after frequency conversion can maintain a relatively high degree of restoration. At the same time, by storing the sound frequency conversion algorithm in the local end, it has extremely high privacy characteristics, and then storing the frequency conversion scheme in the cloud, through the linkage method between the cloud and the local end, it can not only adapt to more customer groups, provide personalized hearing solutions for more hearing-impaired patients and specific range of listening populations, but also reduce the local data storage pressure and improve the data processing speed. As an example, in this embodiment or other embodiments, the local end can be the terminal device or application actually used by the hearing-impaired patient, such as a hearing aid device, the APP application in the hearing aid device; or the intelligent robot performing directional frequency conversion hearing aid can be used as the local end.

[0056] According to the above description, in an exemplary embodiment, as Figure 2 shown, the process of step S140 for performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end further includes: recording the frequency conversion result output by using the human voice frequency conversion algorithm stored in the local end as the human voice frequency conversion result, and recording the frequency conversion result output by using the environmental frequency conversion algorithm stored in the local end as the environmental frequency conversion result, and recording the frequency conversion result output by using the music frequency conversion algorithm stored in the local end as the music frequency conversion result; performing digital processing on the human voice frequency conversion result, the environmental frequency conversion result and the music frequency conversion result, and merging the human voice frequency conversion result, the environmental frequency conversion result and the music frequency conversion result into an audio stream based on the digital processing result; performing high-frequency fitting on the low-frequency part in the audio stream, and performing sound zone adjustment according to the high-frequency fitting result, including adjusting the decibel and phase, so that the adjusted sound becomes directional.

[0057] According to the above description, in an exemplary embodiment, before matching the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud based on the face recognition result and the voice separation result, the voice directional frequency conversion method may further include: outputting a test voice to the target person; receiving the hearing test data fed back by the target person based on the test voice, and generating a hearing test curve of the target person according to the hearing test data; feeding back the hearing test curve and the hearing standard curve to the cloud, so that the cloud generates the frequency voice offset of the target person according to the hearing test curve and the hearing standard curve, forms a frequency conversion scheme based on the frequency voice offset, and adds the frequency conversion scheme to the frequency conversion scheme database. As an example, in this embodiment or other embodiments, a hearing aid device may be used to output a test voice to the target person, and the hearing aid device may be an intelligent robot for directional frequency conversion hearing aid, simply referred to as a hearing aid robot. As an example, for example, the hearing aid robot may output a test voice to the target person, and then the hearing aid robot or the local terminal device receives the hearing test data fed back by the target person based on the test voice, generates a hearing test curve of the target person according to the hearing test data, and feeds back the hearing test curve and the hearing standard curve to the cloud after connecting to the network through the hearing aid robot or the local terminal device. The cloud generates the frequency voice offset of the target person according to the uploaded hearing test curve and the hearing standard curve pre-stored in the cloud, forms a frequency conversion scheme based on the frequency voice offset, and adds the frequency conversion scheme to the frequency conversion scheme database. At the same time, after the cloud receives the face recognition result and the voice separation result, it feeds back the formed frequency conversion scheme to the hearing aid robot or the local terminal device, and performs corresponding control through the local terminal device or the hearing aid robot, so that the hearing-impaired patients can enjoy rich sound effects according to their own health conditions. Among them, the frequency of the test voice can be set according to the actual situation. For example, the frequency of the test voice can be set to 100Hz, 200Hz, 300Hz, 500Hz, 800Hz, 1000Hz, 1500Hz, 2000Hz, 3000Hz, 4000Hz, 5000Hz, 6000Hz, 7000Hz, 8000Hz, 10000Hz or 12000Hz.

[0058] According to the above description, in an exemplary embodiment, the sound directional frequency conversion method may further include: performing pose recognition on the personnel image information to obtain a corresponding pose recognition result; comparing the pose recognition result with the standard behavior pose, and when the pose recognition result is inconsistent with the standard behavior pose, giving an early warning of abnormal human poses; and combining the pose recognition result with the face recognition result, predicting the behavior action of the target personnel at the next moment according to the standard behavior pose, and adjusting based on the predicted behavior action at the next moment when outputting the directional frequency conversion sound to the target personnel. In this embodiment, when adjusting based on the predicted behavior action at the next moment, the output volume of the directional frequency conversion sound can be adjusted, the output direction of the directional frequency conversion sound can be adjusted, or other adjustments can be made according to the actual situation, which will not be elaborated in this embodiment. At the same time, in this embodiment, by giving an early warning of abnormal human poses when the pose recognition result is inconsistent with the standard behavior pose, it can be avoided that no one knows when the elderly, hearing-impaired patients and other people have behaviors such as falling and shouting, further strengthening the protection of the elderly, hearing-impaired patients and other people. As an example, in this embodiment or other embodiments, when performing pose recognition on the personnel image information, the pose recognition model and pose image library stored in the local database can be used to compare the personnel image information obtained in real time to obtain the corresponding pose recognition result.

[0059] In summary, the present invention provides a sound directional frequency conversion method. By performing directional frequency conversion on the sound, it can provide a clear and high-quality auditory experience for hearing-impaired patients, enabling them to better integrate into society and life. At the same time, by realizing the adjustment of sound directional frequency conversion, the parametric array directional speaker can form a focused sound field, making the listeners in the sound field clearly audible and the acoustics outside the sound field rapidly attenuate, playing the role of not disturbing the people and spreading key information. Moreover, this method can recognize and track human faces in real time through face recognition of images, thereby providing more personalized and high-quality auditory and interaction services for users and making communication more convenient and efficient. In addition, by storing the image information, sound and sound frequency conversion algorithm in the local end, this method has extremely high privacy characteristics. At the same time, by storing the frequency conversion scheme in the cloud and through the linkage between the cloud and the local end, it can not only adapt to more customer groups, provide personalized hearing solutions for more hearing-impaired patients and specific range of listening people, but also reduce the local data storage pressure and improve the data processing speed. Therefore, by combining the face tracking technology and the sound directional frequency conversion technology, this method can follow the movement of specific listeners after face recognition, emit a directional sound field for them, provide more personalized services, provide a clear auditory experience for hearing-impaired patients, and has broad application prospects in the fields of elderly care, medical treatment, etc., and can provide more convenient and efficient life assistance for groups such as the elderly and hearing-impaired patients.

[0060] In another exemplary embodiment of the present invention, as Figure 3 shown, this embodiment further provides a sound directional frequency conversion system, including:

[0061] An image acquisition module 310, configured to acquire personnel image information; wherein, the personnel image information includes a target person determined in advance or in real time. As an example, in this embodiment or other embodiments, the target person may be a group of people such as the elderly, hearing-impaired patients, etc., for example, people with genetic hearing loss, hyperacusis, etc., people who are deaf due to diseases (such as long-term exposure to a noisy environment, middle ear or outer ear lesions, etc.). As an example, in this embodiment or other embodiments, the personnel image information may be an image or video stream including the target person captured by a camera.

[0062] A face recognition module 320, configured to perform face recognition on the personnel image information to obtain a corresponding face recognition result. As an example, in this embodiment or other embodiments, when performing face recognition on the personnel image information, a face recognition model and a face image library stored in the local database can be used to compare the personnel image information obtained in real time to obtain a corresponding face recognition result.

[0063] A sound recognition module 330, configured to perform beamforming on the microphone array sound, and perform pronunciation object recognition and pronunciation object direction recognition according to the beamforming result, or perform scene recognition on the external input sound. As an example, in this embodiment or other embodiments, the external input sound may be other actively input external audio, for example, other audio is played through a speaker, and then the sound played by the speaker is used as the external input sound.

[0064] A sound separation module 340, configured to perform sound separation according to the sound recognition result; wherein, the sound recognition result is the pronunciation object recognition result and pronunciation object direction recognition result corresponding to the microphone array sound, or the sound recognition result is the scene recognition result corresponding to the external input sound. As an example, in this embodiment or other embodiments, sound separation includes but is not limited to: text recognition and pronunciation object separation, background sound separation.

[0065] A sound directional frequency conversion module 350, configured to match the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud according to the face recognition result and the sound separation result, and perform directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end, and output the directional frequency conversion sound to the target person.

[0066] According to the above description, in an exemplary embodiment, the process of the sound separation module 340 performing sound separation based on the sound recognition result includes: converting the microphone array sound into a microphone array sound stream based on the pronunciation object recognition result and the pronunciation object direction recognition result, or converting the external input sound into an external input sound stream based on the scene recognition result; using a pre-trained or real-time trained text recognition model to recognize the text in the microphone array sound stream or the external input sound stream, and performing time marking on the text starting point to obtain corresponding time marking points; according to the time marking points, using a pre-trained or real-time trained voiceprint recognition model to perform voiceprint recognition on the microphone array sound and the external input sound, and performing voiceprint time marking to obtain corresponding voiceprint time marking points; starting a new thread, and according to the voiceprint time marking points, using a pre-trained or real-time trained human voice cancellation model to filter the microphone array sound and the external input sound to obtain the corresponding background sound. Therefore, in this embodiment, by converting the audio into a stream mode and then inputting it into three pre-trained or real-time trained neural network models, the first neural network model is a text recognition model, which is used to recognize the text in the speech stream and perform time marking on the text starting point; the second neural network model is a voiceprint recognition model, which performs voiceprint recognition on the sound according to the time marking points and marks the voiceprint time marking points. Synchronously start the second thread, and eliminate all human voices through the third neural network model, that is, eliminate all human voices through the human voice cancellation model to obtain the background sound. Among them, the training processes of the human voice cancellation model, the text recognition model, and the voiceprint recognition model can refer to related technologies, and this embodiment will not elaborate on them here. As an example, in this embodiment or other embodiments, the text recognition model and the voiceprint recognition model belong to the same data processing thread, and the human voice cancellation model is in a different data processing thread from the text recognition model and the voiceprint recognition model. This embodiment uses two threads, which can process different tasks or data sets simultaneously, so as to make full use of the computing power of the multi-core processor and improve the overall execution efficiency. At the same time, in this embodiment, by adding data processing threads, new functions or services can be easily added without large-scale modification and reconstruction of the existing data processing threads, so as to improve scalability and maintainability and facilitate subsequent upgrade and function expansion.

[0067] According to the above description, in an exemplary embodiment, as Figure 2As shown in the figure, the process of the sound directional frequency conversion module 350 performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end includes: according to the current frequency conversion scheme, using the human voice frequency conversion algorithm stored in the local end to change the sound frequency distribution of the human voice audio signal in the target sound, increasing the frequency components, adding sampling points to the human voice audio signal in the target sound, changing the frequency of the human voice audio signal in the target sound, reducing the sampling rate of the human voice audio signal in the target sound, and adjusting the dynamic range of the human voice audio signal in the target sound; and, according to the current frequency conversion scheme, using the environmental frequency conversion algorithm stored in the local end to first perform a fast Fourier transform on the environmental audio signal in the target sound, then perform an inverse fast Fourier transform, and reduce the sampling rate; and, according to the current frequency conversion scheme, using the music frequency conversion algorithm stored in the local end to fit the music audio signal in the target sound. Wherein, the target sound is the microphone array sound or the external input sound. Since the sound after frequency conversion generally has serious distortion problems, in this embodiment, according to the current frequency conversion scheme, the human voice frequency conversion algorithm stored in the local end is used to perform spectrum spreading, interpolation, frequency conversion, downsampling and compression on the human voice audio signal in the target sound, so that the sound after frequency conversion can maintain a relatively high degree of restoration as much as possible. At the same time, by storing the sound frequency conversion algorithm in the local end, it has extremely high privacy characteristics, and then storing the frequency conversion scheme in the cloud, through the linkage method between the cloud and the local end, not only can it adapt to more customer groups, provide personalized hearing solutions for more hearing-impaired patients and specific range of listening people, but also can reduce the local data storage pressure and improve the data processing speed. As an example, in this embodiment or other embodiments, the local end can be the terminal device or application program actually used by the hearing-impaired patient, such as a hearing aid device, an APP application program in the hearing aid device; or the intelligent robot performing directional frequency conversion hearing aid can be used as the local end.

[0068] According to the above description, in an exemplary embodiment, as Figure 2 shown in the figure, the process of the sound directional frequency conversion module 350 performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end may further include: recording the frequency conversion result output by using the human voice frequency conversion algorithm stored in the local end as the human voice frequency conversion result, and recording the frequency conversion result output by using the environmental frequency conversion algorithm stored in the local end as the environmental frequency conversion result, and recording the frequency conversion result output by using the music frequency conversion algorithm stored in the local end as the music frequency conversion result; performing digital processing on the human voice frequency conversion result, the environmental frequency conversion result and the music frequency conversion result, and merging the human voice frequency conversion result, the environmental frequency conversion result and the music frequency conversion result into an audio stream based on the digital processing result; performing high-frequency fitting on the low-frequency part in the audio stream, and performing sound zone adjustment according to the high-frequency fitting result, including adjusting the decibel and phase, so that the adjusted sound becomes directional.

[0069] According to the above description, in an exemplary embodiment, before matching the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud based on the face recognition result and the voice separation result, the voice directional frequency conversion module 350 may further include: outputting a test voice to the target person; receiving the hearing test data fed back by the target person based on the test voice, and generating a hearing test curve of the target person according to the hearing test data; feeding back the hearing test curve and the hearing standard curve to the cloud, so that the cloud generates the frequency voice offset of the target person according to the hearing test curve and the hearing standard curve, forms a frequency conversion scheme based on the frequency voice offset, and adds the frequency conversion scheme to the frequency conversion scheme database. As an example, in this embodiment or other embodiments, a hearing aid device may be used to output a test voice to the target person, and the hearing aid device may be an intelligent robot for directional frequency conversion hearing aid, abbreviated as a hearing aid robot. As an example, for example, the hearing aid robot may output a test voice to the target person, and then the hearing aid robot or the local terminal device receives the hearing test data fed back by the target person based on the test voice, generates a hearing test curve of the target person according to the hearing test data, and feeds back the hearing test curve and the hearing standard curve to the cloud after connecting to the network through the hearing aid robot or the local terminal device. The cloud generates the frequency voice offset of the target person according to the uploaded hearing test curve and the hearing standard curve pre-stored in the cloud, forms a frequency conversion scheme based on the frequency voice offset, and adds the frequency conversion scheme to the frequency conversion scheme database. At the same time, after the cloud receives the face recognition result and the voice separation result, it feeds back the formed frequency conversion scheme to the hearing aid robot or the local terminal device, and performs corresponding control through the local terminal device or the hearing aid robot, so that the hearing-impaired patients can enjoy rich sound effects according to their own health conditions. Among them, the frequency of the test voice can be set according to the actual situation. For example, the frequency of the test voice can be set to 100Hz, 200Hz, 300Hz, 500Hz, 800Hz, 1000Hz, 1500Hz, 2000Hz, 3000Hz, 4000Hz, 5000Hz, 6000Hz, 7000Hz, 8000Hz, 10000Hz or 12000Hz.

[0070] According to the above description, in an exemplary embodiment, the sound directional frequency conversion system may further include: an attitude recognition module for performing attitude recognition on the personnel image information to obtain a corresponding attitude recognition result; an early warning module for comparing the attitude recognition result with the standard behavior attitude and giving a human body abnormal attitude warning when the attitude recognition result is inconsistent with the standard behavior attitude; a sound adjustment module for combining the attitude recognition result with the face recognition result, predicting the behavior action of the target person at the next moment according to the standard behavior attitude, and making adjustments based on the predicted behavior action at the next moment when outputting the directional frequency conversion sound to the target person. In this embodiment, when making adjustments based on the predicted behavior action at the next moment, the output volume of the directional frequency conversion sound can be adjusted, the output direction of the directional frequency conversion sound can be adjusted, or other adjustments can be made according to the actual situation, which will not be elaborated herein. At the same time, in this embodiment, by giving a human body abnormal attitude warning when the attitude recognition result is inconsistent with the standard behavior attitude, it can be avoided that no one knows when the elderly, hearing-impaired patients and other people fall, shout and other behaviors, further strengthening the protection of the elderly, hearing-impaired patients and other people. As an example, in this embodiment or other embodiments, when performing attitude recognition on the personnel image information, the attitude recognition model and attitude image library stored in the local database can be used to compare the personnel image information obtained in real time to obtain the corresponding attitude recognition result.

[0071] In summary, the present invention provides a sound directional frequency conversion system. By performing directional frequency conversion on the sound, it can provide a clear and high-quality auditory experience for hearing-impaired patients, enabling them to better integrate into society and life. At the same time, by realizing the directional frequency conversion adjustment of the sound, the parametric array directional speaker can form a focused sound field, making the listeners in the sound field clearly audible and the sound outside the sound field rapidly attenuate, playing the role of not disturbing the people and spreading key information. Moreover, through face recognition of the image, this system can identify and track the face in real time, thereby providing more personalized and high-quality auditory and interaction services for users, making communication more convenient and efficient. In addition, by storing the image information, sound and sound frequency conversion algorithm in the local end, this system has extremely high privacy characteristics. At the same time, by storing the frequency conversion scheme in the cloud and through the linkage mode between the cloud and the local end, it can not only adapt to more customer groups, provide personalized hearing solutions for more hearing-impaired patients and specific range listening people, but also reduce the local data storage pressure and improve the data processing speed. Therefore, by combining the face tracking technology and the sound directional frequency conversion technology, this system can follow the movement of the specific listener after face recognition, emit a directional sound field for him, provide more personalized services, provide a clear auditory experience for hearing-impaired patients, and has broad application prospects in the fields of elderly care, medical treatment, etc., and can provide more convenient and efficient life assistance for groups such as the elderly and hearing-impaired patients.

[0072] It can be understood that the sound directional frequency conversion system provided in the above embodiments and the sound directional frequency conversion method provided in the above embodiments belong to the same concept. The specific manner in which the sound directional frequency conversion method performs operations has been described in detail in the above method embodiments and will not be elaborated here. In practical applications, the sound directional frequency conversion system provided in the above embodiments can, as needed, allocate the above functions to be completed by different functional modules, that is, divide the internal structure of the sound directional frequency conversion system into different functional modules, and then implement all or part of the functions of the corresponding functional modules through the sound directional frequency conversion method described in the above embodiments. Specific limitations are not imposed here either.

[0073] In another exemplary embodiment of the present invention, this embodiment further provides a hearing aid robot, which includes the sound directional frequency conversion system described in the above embodiments. This hearing aid robot can serve as a local end for data interaction with the cloud. In practical applications, the functions of the sound directional frequency conversion system can be allocated to be completed by different functional modules as needed, that is, divide the internal structure of the sound directional frequency conversion system into different functional modules. As an example, as Figure 4 shown, the sound directional frequency conversion system described in the above embodiments can be split into an image recognition system, an intelligent companionship system, a display system, a multi-modal central control processor, an action control system, a sound directional system, and a sound intelligent frequency conversion hearing aid system, etc. in the hearing aid robot body. Among them, the image recognition system can include one or more cameras, a processing unit for face recognition and pose recognition, a storage device for storing personnel image information, face recognition models, and pose recognition models, a communication module for communicating with the main controller and the cloud, etc. The intelligent companionship system can include one or more microphone arrays, artificial intelligence language large models (such as text recognition models), etc. As Figure 5 shown, the action control system can include a main controller that receives data from a distance sensor or lidar. The main controller directly controls the motion chip to drive the actuator to perform the adjustment of the number of frequencies and the adjustment of the frequency bandwidth, so as to dynamically adjust the sound frequency output by the hearing aid according to the user's hearing loss situation or the pose recognition result, ensuring that the user can clearly hear sounds of different frequencies. The sound directional system can include one or more speaker arrays and one or more sound processors, and each sound processor can be configured with a human voice frequency conversion algorithm, an environmental frequency conversion algorithm, or a music frequency conversion algorithm.

[0074] In another exemplary embodiment of the present invention, this embodiment further provides a hearing aid device, which can include a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to cause the hearing aid device to execute Figure 1 the steps of the sound directional frequency conversion method described above. Figure 6A schematic structural diagram of a hearing aid device 1000 is shown. Refer to Figure 6 As shown, the hearing aid device 1000 includes: a processor 1010, a memory 1020, a power supply 1030, a display unit 1040, and an input unit 1060. As an example, in this embodiment or other embodiments, the hearing aid device 1000 may be various electronic devices such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc., may also be an application software or application program installed in the foregoing electronic devices, or may also be an intelligent robot, etc.

[0075] The processor 1010 is the control center of the hearing aid device 1000, connects each component using various interfaces and lines, and executes various functions of the hearing aid device 1000 by running or executing computer programs / instructions stored in the memory 1020, thereby performing overall monitoring of the hearing aid device 1000. In the embodiments of the present invention, when the processor 1010 calls the computer program stored in the memory 1020, it executes the steps of the Figure 1 sound direction frequency conversion method as described. Optionally, the processor 1010 may include one or more processing units; preferably, the processor 1010 may integrate an application processor and a modulation / demodulation processor, where the application processor mainly processes the operating system, user interface, applications, etc., and the modulation / demodulation processor mainly processes wireless communication. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be implemented separately on independent chips. As an example, in this embodiment or other embodiments, the processor 1010 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0076] The memory 1020 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, various applications, etc.; the data storage area may store instruction data created according to the use of the hearing aid device 1000, etc. In addition, the memory 1020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices, etc.

[0077] The hearing device 1000 further includes a power supply 1030 (such as a battery) for powering each component. The power supply can be logically connected to the processor 1010 through a power management system, so as to manage functions such as charging, discharging, and power consumption through the power management system.

[0078] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the hearing device 1000, etc. In the embodiments of the present invention, it is mainly used to display the display interfaces of various applications in the hearing device 1000 and objects such as text and pictures displayed in the display interfaces. The display unit 1040 may include a display panel 1050. The display panel 1050 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0079] The input unit 1060 can be used to receive information such as numbers or characters input by the user. The input unit 1060 may include a touch panel 1070 and other input devices 1080. Among them, the touch panel 1070, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 1070).

[0080] Specifically, the touch panel 1070 can detect the touch operation of the user, detect the signals brought by the touch operation, convert these signals into contact coordinates, send them to the processor 1010, and receive and execute the commands sent by the processor 1010. In addition, the touch panel 1070 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. The other input devices 1080 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0081] Of course, the touch panel 1070 can cover the display panel 1050. After the touch panel 1070 detects a touch operation on or near it, it is transmitted to the processor 1010 to determine the type of touch event. Subsequently, the processor 1010 provides a corresponding visual output on the display panel 1050 according to the type of touch event. Although in Figure 6 the touch panel 1070 and the display panel 1050 are implemented as two independent components to realize the input and output functions of the hearing device 1000, in some embodiments, the touch panel 1070 and the display panel 1050 can be integrated to realize the input and output functions of the hearing device 1000.

[0082] The hearing aid device 1000 may further include one or more sensors, such as a pressure sensor, a gravity acceleration sensor, a proximity light sensor, etc. Of course, according to the needs in specific applications, the above-mentioned hearing aid device 1000 may further include other components such as a camera.

[0083] An embodiment of the present invention also provides a computer-readable storage medium. When the computer program / instructions stored in this storage medium are executed by a processor, the above-mentioned device can execute the steps of the sound direction and frequency conversion method as described in Figure 1 the present invention.

[0084] Those skilled in the art can understand that Figure 6 the above are only examples of the hearing aid device and do not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or different components. For the convenience of description, the above parts are divided into various modules (or units) according to their functions and described separately. Of course, when implementing the present invention, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.

[0085] Those skilled in the art should understand that the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or boxes. Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one box or more boxes.

[0086] It can be understood that when the above embodiments collect, store, use, process, transmit, provide, disclose, delete and other processing of relevant data (such as personnel image information, hearing test data, etc.), it is completed after or with the consent of the user. For example, personnel image information, hearing test data, etc. are obtained by authorization with the user's knowledge and consent; or are actively provided by the user after reading the relevant instructions, or are actively authorized / provided / uploaded by the user when using some or all of the functions described in the above embodiments, or are obtained by other means / ways after or with the consent of the user.

[0087] The above embodiments are only illustrative of the principles and effects of the present invention, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for sound direction and frequency conversion, characterized in that, The method includes: Obtaining personnel image information, performing face recognition on the personnel image information, and obtaining a corresponding face recognition result; wherein, the personnel image information includes a target person determined in advance or in real time; Performing beamforming on the microphone array sound, and performing pronunciation object recognition and pronunciation object direction recognition according to the beamforming result, or performing scene recognition on the externally input sound; Performing sound separation based on the sound recognition result, wherein the sound recognition result is the pronunciation object recognition result and pronunciation object direction recognition result corresponding to the microphone array sound, or the sound recognition result is the scene recognition result corresponding to the externally input sound; Matching the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud according to the face recognition result and the sound separation result, and performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end, and outputting the directional frequency-converted sound to the target person.

2. The sound direction and frequency conversion method according to claim 1, wherein The process of performing sound separation based on the sound recognition result includes: Converting the microphone array sound into a microphone array sound stream based on the pronunciation object recognition result and pronunciation object direction recognition result, or converting the externally input sound into an externally input sound stream based on the scene recognition result; Using a pre-trained or real-time trained text recognition model to recognize the text in the microphone array sound stream or the externally input sound stream, and performing time marking on the starting point of the text to obtain corresponding time marking points; According to the time marking points, using a pre-trained or real-time trained voiceprint recognition model to perform voiceprint recognition on the microphone array sound or the externally input sound, and performing voiceprint time marking to obtain corresponding voiceprint time marking points; Starting a new thread, and according to the voiceprint time marking points, using a pre-trained or real-time trained human voice cancellation model to filter the microphone array sound or the externally input sound to obtain the corresponding background sound.

3. The sound direction and frequency conversion method according to claim 1 or 2, characterized in that, The process of performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end includes: According to the current frequency conversion scheme, using the human voice frequency conversion algorithm stored in the local end to change the sound frequency distribution of the human voice audio signal in the target sound, increase the frequency components, add sampling points to the human voice audio signal in the target sound, change the frequency of the human voice audio signal in the target sound, reduce the sampling rate of the human voice audio signal in the target sound, and adjust the dynamic range of the human voice audio signal in the target sound; and, According to the current frequency conversion scheme, using the environmental frequency conversion algorithm stored in the local end to first perform a fast Fourier transform on the environmental audio signal in the target sound, then perform an inverse fast Fourier transform, and reduce the sampling rate; and, According to the current frequency conversion scheme, using the music frequency conversion algorithm stored in the local end to fit the music audio signal in the target sound; Wherein, the target sound is the microphone array sound or the externally input sound.

4. The sound direction and frequency conversion method according to claim 3, wherein The process of performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored in the local end further includes: The frequency conversion result output by the voice frequency conversion algorithm stored locally is denoted as the voice frequency conversion result, the frequency conversion result output by the environment frequency conversion algorithm stored locally is denoted as the environment frequency conversion result, and the frequency conversion result output by the music frequency conversion algorithm stored locally is denoted as the music frequency conversion result; Perform digital processing on the voice frequency conversion result, the environment frequency conversion result, and the music frequency conversion result, and merge the voice frequency conversion result, the environment frequency conversion result, and the music frequency conversion result into an audio stream based on the digital processing result; Perform high-frequency fitting on the low-frequency part in the audio stream, and perform sound zone adjustment according to the high-frequency fitting result, including adjusting the decibel and phase, so that the adjusted sound becomes directional.

5. The sound direction and frequency conversion method according to claim 1, characterized in that The method further includes: Perform pose recognition on the personnel image information to obtain the corresponding pose recognition result; Compare the pose recognition result with the standard behavior pose, and issue a human abnormal pose warning when the pose recognition result is inconsistent with the standard behavior pose; and, Combine the pose recognition result with the face recognition result, predict the behavior action of the target person at the next moment according to the standard behavior pose, and make adjustments based on the predicted behavior action at the next moment when outputting the directional frequency conversion sound to the target person.

6. The sound direction and frequency conversion method according to claim 1, wherein Before matching the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud according to the face recognition result and the sound separation result, the method further includes: Output a test sound to the target person; Receive the hearing test data fed back by the target person based on the test sound, and generate a hearing test curve of the target person according to the hearing test data; Feed back the hearing test curve and the hearing standard curve to the cloud, so that the cloud generates the frequency sound offset of the target person according to the hearing test curve and the hearing standard curve, forms a frequency conversion scheme based on the frequency sound offset, and adds the frequency conversion scheme to the frequency conversion scheme database.

7. A sound direction and frequency conversion system, characterized in that The system includes: An image acquisition module for acquiring personnel image information; wherein, the personnel image information includes a target person determined in advance or in real time; A face recognition module for performing face recognition on the personnel image information to obtain the corresponding face recognition result; A sound recognition module for performing beamforming on the microphone array sound, and performing pronunciation object recognition and pronunciation object direction recognition according to the beamforming result, or performing scene recognition on the external input sound; A sound separation module for separating sounds according to the sound recognition result; wherein, the sound recognition result is the pronunciation object recognition result and pronunciation object direction recognition result corresponding to the microphone array sound, or the sound recognition result is the scene recognition result corresponding to the external input sound; A sound directional frequency conversion module for matching the current frequency conversion scheme from the frequency conversion scheme database stored in the cloud according to the face recognition result and the sound separation result, performing directional frequency conversion according to the current frequency conversion scheme by using the sound frequency conversion algorithm stored locally, and outputting a directional frequency conversion sound to the target person.

8. A hearing aid robot, characterized in that, It includes a voice direction and frequency conversion system as described in claim 7.

9. A hearing aid device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the voice direction and frequency conversion method described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the voice direction and frequency conversion method described in any one of claims 1 to 6 are implemented.