Sound source localization and identification system for electric power intelligent service and operation method of sound source localization and identification system
By using a sound source localization and recognition system, which employs sound source capture, noise reduction, and visual inspection technologies, the problem of user identification errors caused by noise interference in power service halls has been solved, enabling more accurate user positioning and natural communication.
Patent Information
- Application Number
- CN202511100003.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-12-30
AI Technical Summary
In the power service hall, the noise level is high and the voices are similar to those of users, leading to errors in sound source identification and reducing the naturalness of communication.
A sound source localization and recognition system is adopted, including an input module, a recognition module, and an output module. Through sound source capture, noise reduction, human voice separation, and sound source localization modules, combined with noise processing, echo processing, time delay estimation, spatial spectrum estimation, and visual detection, the user's location can be accurately located.
It improves the accuracy of sound source identification, reduces interaction errors, and increases the naturalness of communication.
Smart Images

Figure CN121237118A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of sound source localization, and in particular to a sound source localization and identification system and its operation method for smart power services. Background Technology
[0002] Virtual digital humans refer to technologies that exist in the non-physical world, created and used by computer means such as computer graphics, graphics rendering, motion capture, deep learning, and speech synthesis, and possess multiple human characteristics. In a narrow sense, digital humans are virtual simulations of the human body using information science, and are a product of the integration of information science and life science. In a broad sense, digital humans refer to the penetration of digital technology into all levels and stages of human anatomy, physics, physiology, and intelligence.
[0003] Virtual digital humans, by integrating artificial intelligence models such as knowledge graphs, speech recognition, and natural language processing, can recognize, understand, and analyze customer problems, provide accurate dialogue responses, and offer personalized services, thereby improving customer service experience and satisfaction. However, power service halls often have a large number of customers conducting business, leading to significant noise levels and the presence of users with similar voices, causing errors in sound source identification and reducing the naturalness of communication. Summary of the Invention
[0004] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a sound source localization and identification system and its operation method for intelligent power services, solving the problems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0006] In a first aspect, embodiments of the present invention provide a sound source localization and identification system for smart power services, comprising: an input module, an identification module, a processing module, and an output module;
[0007] The input module is connected to the recognition module. The input module is used to receive user interaction requests and upload the user interaction requests to the recognition module.
[0008] The recognition module receives the user interaction request, identifies and locates the user's voice location based on the user interaction request, and uploads the identification and location results to the processing module. The recognition module includes a sound source capture module, a noise reduction module, a voice separation module, and a sound source localization module. The capture module is used to obtain the power service keywords of the user interaction request, the noise reduction module reduces noise in the surrounding environment, the voice separation module separates the noise-reduced speech, and the sound source localization module locates the user emitting the voice based on the keywords captured by the capture module.
[0009] One end of the processing module is connected to the identification module, and the other end is connected to the output module. The processing module analyzes and processes the received identification and positioning results for user questions, and the output module answers the user questions based on the analysis and processing results.
[0010] As a preferred embodiment of the sound source localization and identification system for intelligent power services described in this invention, the identification module further includes: a sound source capture module, a noise reduction module, a human voice separation module, and a sound source localization module; the output end of the sound source capture module is connected to the input end of the noise reduction module, the output end of the noise reduction module is connected to the input end of the human voice separation module, and the human voice separation module is connected to the input end of the sound source localization module.
[0011] As a preferred embodiment of the sound source localization and identification system for intelligent power services described in this invention, the noise reduction module includes a noise processing module and an echo processing module.
[0012] The echo processing module is used to cancel the echo components in the sound played by the device and the sound picked up by the microphone in real time. The noise processing module is used to process the ambient noise, enhance the signal in the direction of the target sound source by weighted summation, suppress noise in other directions, and form a beam pointing to the sound source.
[0013] The noise processing module includes dynamic noise detection and noise simulation detection. The dynamic noise detection is used to process environmental noise, and the noise simulation detection is used to process pre-recorded sound.
[0014] The beneficial effect of this preferred technical solution is that adaptive beamforming can dynamically adjust the weights to adapt to complex environments.
[0015] As a preferred embodiment of the sound source localization and identification operation system for intelligent power services described in this invention, the noise reduction module further includes an optimization module at its input end, which is used to optimize the data of the noise reduction module.
[0016] The optimization module includes a noise sample acquisition and training module. The noise sample acquisition module is used to collect various common sounds. The training module uses the acquired noise data to fine-tune the sound source localization and noise reduction model, and optimizes the beamforming parameters based on the reverberation characteristics of the power hall.
[0017] As a preferred embodiment of the sound source localization and identification system for intelligent power services described in this invention, the sound source localization module includes: a time delay estimation module, a spatial spectrum estimation module, and a visual detection module.
[0018] The time delay estimation module is used to calculate the time difference of sound reaching each microphone, calculate the azimuth angle of the sound source using the time difference of sound waves received by the microphone array, and align the timestamps of audio frames and video frames; convert the direction vector of audio positioning into three-dimensional coordinates in the camera coordinate system; establish a fusion model, which combines audio positioning confidence and visual detection confidence, and outputs the final positioning result by weighting the audio positioning confidence and visual detection confidence;
[0019] The spatial spectrum estimation module is used to determine the direction of the sound source by the spatial spectrum peak value. It performs spectral analysis on the signal received by the microphone array and determines the direction of the sound source by the spatial spectrum peak value.
[0020] The visual detection module is used to perform visual detection on the user. The visual detection module captures scene images, locates the sound-emitting object, obtains the target position, and cross-validates it with the sound source localization result.
[0021] As a preferred embodiment of the sound source localization and identification system for intelligent power services described in this invention, the visual detection module further includes: lip movement detection, facial expression detection, and human posture detection.
[0022] The lip movement detection is used to identify the user's mouth movements, the expression detection is used to identify the user's facial expressions, and the body posture detection is used to identify the user's body movements.
[0023] As a preferred embodiment of the sound source localization and identification system for intelligent power services described in this invention, the sound source localization module further includes an analysis module; the analysis module includes frequency domain texture analysis and time domain dynamic feature analysis; the frequency domain texture analysis is used to transform and extract the frequency domain texture features of the sound, construct a sound feature database, pre-store common sound sources, and calculate the similarity through template matching during localization to select the most matching sound source direction.
[0024] The advantages of this preferred technical solution are that it can more accurately identify and locate users by sound source, reduce interaction errors, and increase the naturalness of interaction.
[0025] Secondly, the present invention provides a sound source localization and identification operation method for smart power services, comprising:
[0026] Receive user interaction requests, identify electricity service keywords in the user interaction requests, and obtain the user's electricity service language request;
[0027] The user's voice source is obtained based on the power service language request, and the voice source is preprocessed. Based on the preprocessed language data, human voice separation is performed to identify the user's location of the voice source.
[0028] Receive user questions and requests, analyze and process the questions, and answer user questions based on the analysis and processing results.
[0029] Thirdly, the present invention provides an electronic device, comprising:
[0030] Memory and processor;
[0031] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the sound source localization and identification operation method for smart power services.
[0032] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the sound source localization and identification operation method for smart power services.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention removes noise from the user's speech, records the timbre, speech rate, and audio of the user's speech, and captures the user at the same time, so that the virtual digital human can more accurately identify and locate the user's sound source, reduce interaction errors, and increase the naturalness of the interaction. Attached Figure Description
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0035] Figure 1 This is a schematic diagram of the system architecture of a sound source localization and identification system for smart power services according to an embodiment of the present invention;
[0036] Figure 2 This is a schematic diagram of the identification module structure of a sound source localization and identification system for smart power services according to an embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of an optimized module structure of a sound source localization and identification system for smart power services according to an embodiment of the present invention;
[0038] Figure 4 This is a schematic flowchart of a sound source localization and identification method for smart power services according to an embodiment of the present invention. Detailed Implementation
[0039] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0040] Example 1, referring to Figures 1-3 As one embodiment of the present invention, this embodiment provides a sound source localization and identification operation system for smart power services, comprising:
[0041] Figure 1 This is a schematic diagram of a sound source localization and identification system for smart power services, according to an exemplary embodiment. The sound source localization and identification system for smart power services includes: an input module, an identification module, a processing module, and an output module.
[0042] The input module is connected to the recognition module. The input module is used to receive user interaction requests and upload them to the recognition module.
[0043] The recognition module receives user interaction requests, identifies and locates the user's voice location based on the user interaction requests, and uploads the recognition and location results to the processing module.
[0044] One end of the processing module is connected to the recognition module, and the other end is connected to the output module. The processing module analyzes and processes the received recognition and positioning results to answer the user's questions, and the output module answers the user's questions based on the analysis and processing results.
[0045] It should be noted that the sound source localization and recognition system for intelligent power services of the present invention can remove noise mixed in the user's speech, record the timbre, speech rate, audio and other information in the user's speech, and capture the user at the same time, so that the virtual digital human can more accurately identify and locate the user's sound source, reduce interaction errors and increase the naturalness of the interaction.
[0046] Furthermore, refer to Figure 2 The recognition module includes: a sound source capture module, a noise reduction module, a human voice separation module, and a sound source localization module; the output of the sound source capture module is connected to the input of the noise reduction module, the output of the noise reduction module is connected to the input of the human voice separation module, and the human voice separation module is connected to the input of the sound source localization module.
[0047] The capture module is used to obtain the power service keywords of user interaction needs, the noise reduction module reduces the noise in the surrounding environment, the voice separation module separates the noise-reduced speech, and the sound source localization module locates the user who made the sound based on the keywords captured by the capture module.
[0048] In an optional embodiment, the capture module is used to obtain the power service keywords of user interaction needs, including power service type, power facilities, power consumption scenario, business process and question type.
[0049] For example, the capture module can obtain keywords related to electricity service types, such as electricity bill payment, power outage notice, meter installation, repair application, electricity transfer, charging pile application, and electricity price consultation.
[0050] For example, the capture module obtains keywords for power facilities including transformers, meters, lines, charging piles, distribution boxes, substations, etc.
[0051] For example, the keywords for electricity consumption scenarios obtained by the capture module include residential electricity consumption, industrial electricity consumption, commercial electricity consumption, new energy electricity consumption, and energy-saving renovation.
[0052] For example, the keywords that the capture module obtains for business flows include account opening, account transfer, account cancellation, installation, payment, inquiry, complaint, and refund.
[0053] In another alternative embodiment, the problem types also include fault reporting, policy consultation, and business processing.
[0054] For example, the capture module can obtain keywords for fault reporting, such as power outage and leakage.
[0055] For example, the capture module obtains policy consultation keywords including subsidies and electricity prices.
[0056] For example, the capture module can obtain keywords related to business processing, including process and materials.
[0057] It should be noted that the noise reduction module reduces ambient noise, thereby separating speech.
[0058] Furthermore, the noise reduction module includes: a noise processing module and an echo processing module;
[0059] The echo processing module is used to cancel the echo components in the sound played by the device and the microphone in real time. The noise processing module is used to process the ambient noise, enhance the signal in the direction of the target sound source by weighted summation, suppress noise in other directions, and form a beam pointing to the sound source.
[0060] It should be noted that the noise processing module enhances the signal in the direction of the target sound source by weighted summation, suppresses noise in other directions, and forms a beam pointing towards the sound source. The adaptive beamforming can dynamically adjust the weights to adapt to complex environments.
[0061] Furthermore, the noise processing module includes dynamic noise detection and noise simulation detection. Dynamic noise detection is used to process ambient noise, while noise simulation detection is used to process pre-recorded sound.
[0062] In one optional embodiment, dynamic noise detection monitors the ambient noise intensity in real time. When sudden noise, such as a child crying, occurs, it automatically switches to a strong noise reduction mode, increasing the noise reduction by 5-8 dB. After the noise weakens, it returns to the normal mode. Through the noise classification model ResNet, it identifies common noise types, such as equipment noise, human voice, and broadcast noise, and adopts targeted suppression strategies, such as attenuating broadcast noise in specific frequency bands.
[0063] In an optional embodiment, noise simulation detection is used to automatically identify and remove daily noise samples, including equipment noise, pedestrian noise, and broadcast noise, collected in different areas of the lobby, such as counters, waiting areas, and self-service areas, thereby achieving rapid noise reduction.
[0064] Furthermore, the noise reduction module also includes an optimization module at its input end, which is used to optimize the noise reduction module data.
[0065] The optimization module includes a noise sample acquisition and training module. The noise sample acquisition module is used to collect various common sounds. The training module uses the collected noise data to fine-tune the sound source localization and noise reduction model, and optimizes the beamforming parameters based on the reverberation characteristics of the power hall.
[0066] It should be noted that noise sample collection involved gathering everyday noise samples, including equipment noise, pedestrian noise, and broadcast noise, from different areas of the lobby, such as the counter, waiting area, and self-service area, to establish a scenario-based noise database. The training module used the collected noise data to fine-tune the sound source localization and noise reduction models, and optimized beamforming parameters for the reverberation characteristics of the power lobby.
[0067] In an optional embodiment, the sound source localization module locates the user emitting the sound based on the keywords captured by the capture module. After determining the user's location based on the time and direction of sound reception, the visual detection module detects the user's mouth movements and facial expressions to identify the speaker. Furthermore, the analysis module analyzes the user's speaking speed, audio, and emotions to locate the user, thereby making the sound source localization more accurate. After the user is located, the visual detection module locks onto the user, preventing incorrect judgments due to other sounds during use.
[0068] In another optional embodiment, when multiple users are speaking simultaneously in the lobby and their voices are similar, the system identifies keywords in the users' statements. If multiple users' statements contain keywords, the system determines the user's position by measuring the speed at which each sound is received. At the same time, the system uses a visual detection module to detect the user's position and direction, such as whether they are facing or away from the virtual digital persona. Combined with the sound source, the system makes the sound source localization more accurate.
[0069] Furthermore, the sound source localization module includes: a time delay estimation module, a spatial spectrum estimation module, and a visual detection module;
[0070] The time delay estimation module is used to calculate the time difference of sound arrival at each microphone, calculate the azimuth angle of the sound source using the time difference of sound waves received by the microphone array, and align the timestamps of audio frames and video frames; convert the direction vector of audio positioning into three-dimensional coordinates in the camera coordinate system; establish a fusion model, which combines audio positioning confidence and visual detection confidence, and output the final positioning result by weighting the audio positioning confidence and visual detection confidence.
[0071] Specifically, the time delay estimation module calculates the azimuth angle of the sound source using the time difference between the sound waves received by the microphone array, expressed as:
[0072]
[0073] Where τ is the time delay, d is the microphone spacing, θ is the direction of the sound source, and c is the speed of sound.
[0074] In an optional embodiment, the delay estimation module aligns the timestamps of audio frames and video frames using a hardware clock or software interpolation. The direction vector of the audio localization is converted into three-dimensional coordinates in the camera coordinate system. For example, the direction vector obtained from audio localization is combined with target distance estimation to obtain spatial point P. audio Visual detection yields the target coordinates P. vision The two are mapped to the same coordinate system using a transformation matrix. A fusion model is established, weighting the audio localization confidence and visual detection confidence, and the final localization result is output as follows:
[0075] P final =w1·P audio +w2·P vision
[0076] Where w1 and w2 are the confidence weights for audio localization and visual detection, respectively, and w1+w2=1.
[0077] It should be noted that the confidence weight can be dynamically adjusted according to the environment. For example, when there is a lot of noise, the visual weight is increased, and when there is occlusion, the audio weight is increased.
[0078] For example, the timestamps of approximately 10-30ms audio frames and 25-60fps video frames are aligned using a hardware clock or linear interpolation. The azimuth and elevation angles of the audio localization are converted into three-dimensional coordinates in the camera coordinate system. Spatial points are obtained by calculating visual depth or microphone array delay difference. The target coordinates are obtained by visual detection. The two are mapped to the same coordinate system using rotation or translation matrices. A fusion model of Kalman filtering or Bayesian network is established to weight the delay estimation error and the target classification probability, and the final localization result is output.
[0079] Furthermore, the spatial spectrum estimation module is used to determine the direction of the sound source by using the spatial spectrum peak value. It performs spectral analysis on the signal received by the microphone array and determines the direction of the sound source by using the spatial spectrum peak value.
[0080] In an optional embodiment, the spatial spectrum estimation module uses the MUSIC or ESPRIT algorithm to perform spectral analysis on the signal received by the microphone array, determines the direction of the sound source by the spatial spectrum peak, and, combined with the power hall scenario, preset common user locations as priority detection areas to reduce the amount of computation.
[0081] Furthermore, the visual inspection module is used to perform visual inspection on the user. The visual inspection module captures scene images, locates the sound-emitting object, obtains the target position, and cross-validates it with the sound source localization result.
[0082] In an optional embodiment, the visual detection module captures scene images through a camera, locates the sound-emitting object through target detection technology, obtains the target position through face detection or human pose estimation, and cross-validates the sound source localization results.
[0083] In another optional embodiment, the visual detection module captures scene images through a camera, locates the sound-emitting object through techniques such as posture recognition, obtains the target location through face detection or human posture estimation, and cross-validates the sound source localization results.
[0084] Furthermore, the visual detection module also includes: lip movement detection, facial expression detection, and human posture detection;
[0085] Lip movement detection is used to identify user mouth movements, facial expression detection is used to identify user facial expressions, and body posture detection is used to identify user body movements.
[0086] Furthermore, the sound source localization module also includes an analysis module; the analysis module includes frequency domain texture analysis and time domain dynamic feature analysis; frequency domain texture analysis is used to transform and extract the frequency domain texture features of the sound, build a sound feature database, pre-store common sound sources, and calculate the similarity through template matching during localization to select the most matching sound source direction.
[0087] In an optional embodiment, frequency domain texture analysis can utilize CNN to extract frequency domain texture features of sound, such as formant distribution and harmonic structure. Even if the fundamental frequencies are similar, they can be distinguished by differences in high-frequency overtones, such as the difference in overtone series between male and female voices. A sound feature database is constructed, and common sound sources, such as frequency domain templates of worker A, worker B, printer A, and printer B, are pre-stored. During localization, similarity is calculated by template matching, and the direction of the most matching sound source is selected.
[0088] In another optional embodiment, frequency domain texture analysis can also utilize Wavelet transform to extract frequency domain texture features of sound, such as formant distribution and harmonic structure. Even if the fundamental frequencies are similar, they can be distinguished by differences in high-frequency overtones, such as the difference in overtone series between male and female voices. A sound feature database is constructed, and common sound sources, such as frequency domain templates of worker A, worker B, printer A, and printer B, are pre-stored. During localization, similarity is calculated by template matching, and the direction of the most matching sound source is selected.
[0089] It should be noted that when analyzing human voices, the sound frequency is determined by the speed of vocal cord vibration, while the length, thickness, and tension of the vocal cords are affected by physiological characteristics: adult males have longer vocal cords (about 17-23 mm) and lower vibration frequencies (about 85-150 Hz), resulting in a deep voice; adult females have shorter vocal cords (about 12-17 mm) and higher vibration frequencies (about 165-250 Hz), resulting in a high-pitched voice; children have even shorter vocal cords, with frequencies reaching 200-300 Hz or higher, resulting in a higher pitch. At the same time, the shape and size of the larynx, oral cavity, and nasal cavity affect sound wave resonance, thereby changing the frequency distribution.
[0090] For example, when nasal resonance is enhanced, the high-frequency components are more prominent, and the sound is "sharper"; when chest resonance is dominant, the low-frequency components account for a higher proportion, and the sound is "deeper". When tense or excited, the vocal cord tension increases, the frequency rises, and the sound may become "sharper"; when relaxed or tired, the frequency decreases, and the sound is "deeper"; when angry or shouting, the sound pressure increases, the high-frequency components increase, and the sound is more "piercing".
[0091] The working principle of the sound source localization and identification system for smart power services provided in this embodiment is as follows:
[0092] When in use, after the user states their needs, the virtual digital human receives the user's speech through the input module. The high-resolution recognition module then denoises the user's speech, removing any noise mixed in with the speech and separating the human voice. The sound source localization module records the context, timbre, speech rate, and audio of the user's speech. At the same time, the visual detection module captures and locates the user, enabling one-to-one interaction with the user without being interrupted or mislocated by other voices in the environment.
[0093] This invention provides a sound source localization and recognition system for smart power services. After removing noise from the user's speech through the recognition module, the system records the timbre, speech rate, and audio of the user's speech. At the same time, it works with a visual detection module to capture the user, so that the virtual digital human can more accurately identify and locate the user's sound source, reduce interaction errors, and increase the naturalness of the interaction.
[0094] Example 2, refer to Figure 4 This is one embodiment of the present invention, which differs from the first embodiment in that it provides a sound source localization and identification operation method for smart power services, including:
[0095] S100: Receive user interaction requests, identify electricity service keywords in user interaction requests, and obtain the user's electricity service language request;
[0096] S200: Obtain the user's voice source based on the power service language request, preprocess the voice source, perform human voice separation based on the preprocessed language data, and identify the user's location of the voice source;
[0097] S300: Receives user requests for questions and analyzes and processes the questions, and answers user questions based on the analysis and processing results.
[0098] Example 3: This example also provides an electronic device applicable to the sound source localization and identification operation method for smart power services, including:
[0099] The system includes a memory and a processor. The memory stores computer-executable instructions, and the processor executes these instructions to implement the sound source localization and identification method for smart power services as proposed in the above embodiments.
[0100] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the sound source localization and identification method for smart power services as proposed in the above embodiments.
[0101] The storage medium proposed in this embodiment and the sound source localization and identification operation method for implementing smart power services proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0102] From the above description of the embodiments, those skilled in the art will clearly understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0103] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0104] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0106] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0107] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
[0108] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A sound source localization identification system for electric power smart services, characterized by, The utility model relates to an electric power service interactive system, including: Input module, identification module, processing module and output module; The input module is connected with the identification module, and the input module is used for receiving the user's interactive demand and uploading the user interactive demand to the identification module; The identification module receives the user interactive demand, the identification module is positioned based on the user interactive demand to user sound, and the identification module uploads the identification positioning result to the processing module; the identification module includes sound source capture module, noise reduction module, human voice separation module and sound source positioning module; the capture module is used to obtain the power service keyword of user interactive demand, the noise reduction module carries out noise reduction to the noise in the surrounding environment, the human voice separation module separates the language after noise reduction, and the sound source positioning module positions the user according to the keyword captured by the capture module to the sound emission; One end of the processing module is connected to the identification module, and the other end is connected to the output module, the processing module carries out the analysis processing of user problem to the received identification positioning result, and the output module answers the user question based on the analysis processing result.
2. The sound source localization identification system for electric power smart services of claim 1, wherein, The identification module further includes: the output end of the sound source capture module is connected with the input end of the noise reduction module, the output end of the noise reduction module is connected with the input end of the human voice separation module, and the input end of the human voice separation module is connected with the sound source positioning module.
3. The sound source localization identification system for electric power smart services of claim 2, wherein, The noise reduction module includes: noise processing module and echo processing module; The echo processing module is used for real-time cancellation of echo components in the device playing sound and microphone pickup, and the noise processing module is used for processing environmental noise, enhancing the signal of the target sound source direction by weighted summation, suppressing other direction noise, and forming a beam pointing to the sound source; The noise processing module includes dynamic noise detection and noise simulation detection, the dynamic noise detection is used for processing environmental noise, and the noise simulation detection is used for processing the sound recorded in advance.
4. The sound source localization identification system for electric power smart services of claim 3, wherein, The noise reduction module further includes: an optimization module is arranged at the input end of the noise reduction module, and the optimization module is used for optimizing the data of the noise reduction module; The optimization module includes noise sample collection and training module, the noise sample collection is used for collecting various common sounds; the training module uses the collected noise data to fine-tune the sound source positioning and noise reduction model, and optimizes the beam forming parameters based on the power hall reverberation characteristics.
5. The sound source localization identification system for electric power smart services of claim 4, wherein, The sound source positioning module includes: time delay estimation module, spatial spectrum estimation module and visual detection module; The time delay estimation module is used for calculating the time difference of sound arriving at each microphone, calculating the sound source azimuth angle by using the time difference of microphone array receiving sound wave, and aligning the time stamps of audio frame and video frame; the direction vector of audio positioning is converted into three-dimensional coordinates in the camera coordinate system; a fusion model is established, the fusion model combines the audio positioning confidence and the visual detection confidence, the final positioning result is output by weighting the audio positioning confidence and the visual detection confidence; The spatial spectrum estimation module is used for determining the sound source direction by spatial spectrum peak value, and the signal received by the microphone array is analyzed by spectrum, and the sound source direction is determined by spatial spectrum peak value; The visual detection module is used for visual detection of the user, and the visual detection module captures a scene image, locates a sound-emitting object, obtains a target position, and cross- verifies a sound source positioning result.
6. The sound source localization identification system for electric power smart services of claim 5, wherein, The visual detection module further includes lip movement detection, expression detection, and human body posture detection. The lip movement detection is used for recognizing mouth movements of the user, the expression detection is used for recognizing facial expressions of the user, and the human body posture detection is used for recognizing limb movements of the user.
7. The sound source localization identification system for electric power smart services of claim 6, wherein, The sound source positioning module further includes an analysis module, and the analysis module includes frequency domain texture analysis and time domain dynamic feature analysis.
8. A sound source positioning and identification operation method for electric power intelligent service, applied to the method of any one of claims 1-7, characterized in that, The frequency domain texture analysis is used for transforming and extracting frequency domain texture features of sound, constructing a sound feature database, pre-storing common sound sources, calculating similarity through template matching during positioning, and screening the most matched sound source direction. The method includes the following steps: Receiving a user interaction requirement, identifying a power service keyword in the user interaction requirement, obtaining a power service language request of the user, and obtaining a user sound source based on the power service language request. The method further includes the following steps: Receiving a user question requirement, analyzing and processing the question, and answering the user question based on an analysis and processing result. 9.An electronic device, comprising: a memory and a processor; The memory is used for storing computer executable instructions, and the processor is used for executing the computer executable instructions, which realize the steps of the sound source positioning and identification operation method for power intelligent service in claim 8 when executed by the processor. 10.A computer readable storage medium storing computer executable instructions, which realize the steps of the sound source positioning and identification operation method for power intelligent service in claim 8 when executed by the processor.
Citation Information
Cited By
Human-machine voice precise interaction method based on large model
CN121905189A