Multi-layer natural language interaction method based on deep learning and pattern recognition

By combining a multi-microphone array device and a deep learning model with visual and environmental perception information, sound source signals are collected and separated in real time, solving the problem of poor sound source optimization and collection effect, and realizing efficient multi-layer semantic understanding and personalized interactive services.

CN119721055BActive Publication Date: 2026-03-27BEIJING DIZHIYUAN TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies have poor sound source optimization and acquisition effects in the language interaction process, and fail to perform multi-layer semantic understanding and parsing based on accurate and efficient sound source data, resulting in low efficiency of digital human interaction.

Method used

The system uses a multi-microphone array to collect language interaction information in real time, separates and masks target sound source signals in real time, combines visual perception and environmental perception information to perform parallel multi-user identification and correction, and uses a deep learning model for semantic layering and understanding, including coarse-grained and fine-grained semantic processing.

Benefits of technology

It improves the accuracy and stability of voice acquisition, enhances the digital human's ability to understand user intentions and emotions, optimizes interaction strategies, improves the accuracy and fluency of interaction, and provides more thoughtful personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721055B_ABST
    Figure CN119721055B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to a multilayer natural language interaction method based on deep learning and pattern recognition, which comprises the following steps: S1, collecting language interaction information in real time; S2, positioning a sound source and performing real-time separation and shielding; S3, performing parallel multi-user identification according to visual perception information; S4, adjusting a real-time separation and shielding process of the real-time collection process of the language interaction information; S5, correcting the parallel multi-user identification result; S6, performing model fine-tuning on the parallel multi-user identification; S7, performing information fusion to obtain language interaction fusion data; S8, performing coarse-grained semantic layering on the language interaction fusion data; and S9, performing fine-grained semantic understanding on the language interaction fusion data. The application improves the natural language interaction efficiency in the digital human interaction process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a multi-layer natural language interaction method based on deep learning and pattern recognition. BACKGROUND

[0002] In recent years, with the rapid development of artificial intelligence and deep learning technology, the concept of digital employees has gradually become a reality. In order to realize multi-layer natural language interaction, it is necessary to solve the challenges of multi-modal perception and interaction, including ASR, TTS and visual recognition. Through self-research multi-microphone array and acoustic algorithm, accurate sound can be collected in noisy environment, and dynamic following of multi-person communication can be realized through machine vision algorithm. The rise of AI large model provides the "understanding" thinking ability of digital employees, combined with industry data, can quickly learn and improve service efficiency. Digital employees need to interact with the physical world, which requires them to control existing intelligent systems and reduce the R&D investment of enterprises. The progress of 3D visualization technology enables the realization of realistic image of digital employees, and under the support of AI GC technology, image and video generation, AI driving technology of 3D model are constantly mature. Through the combination of deep learning and pattern recognition, we can develop multi-level natural language interaction solutions, improve the adaptability and business value of digital employees, and open up new opportunities for AI application in business.

[0003] Chinese patent publication No. CN108009285A discloses a forestry ecological environment human-computer interaction method based on natural language processing, belonging to the field of neural networks. The method uses the method of knowledge graph to reason the entity of natural language in the forestry ecological environment, so that the knowledge reasoning is transformed into the problem of processing natural language question through constructing deep neural network, so as to find the corresponding relationship, the knowledge graph deep learning reasoning under representation learning, and the corresponding conclusion is obtained. The present application introduces the concept of knowledge graph in deep learning, and injects the shallow semantic understanding result into the knowledge graph on the basis of constructing the knowledge graph, and obtains the relatively deep semantic understanding through the corresponding knowledge reasoning. But this scheme does not solve the problem of sound source optimization collection in language interaction process, and cannot perform multi-layer semantic understanding analysis based on accurate and efficient sound source data. SUMMARY

[0004] Therefore, the present application provides a multi-layer natural language interaction method based on deep learning and pattern recognition, which overcomes the problem of low natural language interaction efficiency in digital human interaction process caused by poor sound source optimization collection effect in language interaction process and not based on accurate and efficient sound source data for multi-layer semantic understanding analysis in the prior art.

[0005] To achieve the above purpose, the present application provides a multi-layer natural language interaction method based on deep learning and pattern recognition, comprising:

[0006] Step S1, real-time collection of language interaction information by a multi-microphone array device;

[0007] Step S2, positioning of a sound source according to the language interaction information, real-time separation and shielding of the real-time collection process of the language interaction information, and optimization of the multi-microphone array device;

[0008] Step S3, real-time collection of visual perception information, and parallel multi-user recognition according to the visual perception information;

[0009] Step S4, adjustment of the real-time separation and shielding process of the real-time collection process of the language interaction information according to the parallel multi-user recognition result;

[0010] Step S5, real-time collection of environmental perception information, user scene recognition according to the environmental perception information, and correction of the parallel multi-user recognition result according to the user scene recognition result;

[0011] Step S6, judgment of visual fine-tuning conditions according to the parallel multi-user recognition result correction proportion, and model fine-tuning of parallel multi-user recognition according to the visual fine-tuning conditions;

[0012] Step S7, information fusion of language interaction information, visual perception information, and environmental perception information, to obtain language interaction fusion data;

[0013] Step S8, coarse-grained semantic layering of the language interaction fusion data according to a coarse-grained semantic layering method, to obtain intent classification information and context perception information;

[0014] Step S9, after coarse-grained semantic layering, fine-grained semantic understanding of the language interaction fusion data according to a fine-grained semantic understanding method, to obtain sentiment analysis information, tone analysis information, deep context analysis information, and ambiguity resolution information.

[0015] Further, in the step S2, a microphone signal generalized cross-correlation function Rx1 x2(τ) is calculated according to a first microphone signal x1(t) and a second microphone signal x2(t) in the language interaction information, and is set as Ψ x1x2(f) is a weighted function, x1(f) is the Fourier transform of x1(t), x2(f) is the Fourier transform of x2(t), x2*(f) is the conjugate complex of x2(f), e is the base of natural logarithm, j is the imaginary unit, f is the frequency, which represents the identification of different frequency components in the frequency domain when the signal is converted from the time domain to the frequency domain during the analysis of the sound signal received by the microphone, τ is a variable related to the time delay, the peak τmax of the generalized cross-correlation function Rx1x2(τ) of the microphone signal is obtained, and the distance difference △r of the sound source to the first microphone and the second microphone is calculated according to the peak τmax, and △r is set to τmax×c, c is the speed of sound propagation, and the angle θ of the sound source to the plane of the microphone array is calculated, and θ is set to r is the distance of the sound source to the center of the microphone array, and the distance difference △r and the angle θ are taken as the positioning information of the sound source.

[0016] Further, in the step S2, after obtaining the positioning information of the sound source, the effective coefficient D of the sound source is calculated as D=α1×1 / r+α2×T+α3×△F, α1 is the first effective coefficient weight parameter, α2 is the second effective coefficient weight parameter, α3 is the third effective coefficient weight parameter, T is the current speech duration of the sound source, △F is the sum of the differences between the sound source and the reference value on each frequency spectrum feature, and △F is set to ∑ n k=1 |Fk-Fk'|, Fk' is the reference value of each frequency spectrum feature, the sound source corresponding to the maximum effective coefficient Dmax is taken as the target sound source, the real-time separation and shielding process of the real-time collection process of the language interaction information of the target sound source is performed according to the positioning information of the target sound source, and the multi-microphone array device is optimized until the positioning information of the target sound source is the minimum distance difference and the angle is 90°.

[0017] Further, in the step S3, the visual perception information is input into the user recognition deep learning model, and the multi-user recognition result output by the user recognition deep learning model is obtained.

[0018] Further, in the step S4, the adjustment situation of the real-time separation and shielding process of the real-time collection process of the language interaction information is judged according to the parallel multi-user recognition result, wherein:

[0019] When the parallel multi-user recognition result is a single user, it is determined that the real-time separation and shielding process of the real-time collection process of the language interaction information is not adjusted;

[0020] When the parallel multi-user recognition result is multiple users, it is determined to adjust the real-time separation shielding process of the real-time collection process of the language interaction information, the target sound source after adjustment is the sound source of multiple users, and the real-time separation shielding of the real-time collection process of the language interaction information of the target sound source after adjustment is carried out according to the positioning information of the target sound source after adjustment.

[0021] Further, in the step S5, an environment perception feature map is generated according to the environment perception information, and the environment perception feature map is input into a user scene recognition model to obtain a user scene recognition result output by the user scene recognition model, and the parallel multi-user recognition result is corrected according to the user scene recognition result, wherein:

[0022] When the user scene recognition result is a multi-person interaction scene, the parallel multi-user recognition result is not corrected;

[0023] When the user scene recognition result is a single-person interaction scene, the parallel multi-user recognition result is corrected, and the parallel multi-user recognition result of multiple users is corrected to a single user.

[0024] Further, in the step S6, the parallel multi-user recognition result correction proportion A is calculated according to the user scene recognition times L0 and the parallel multi-user recognition result correction times L1, A is set as L1 / L0, the parallel multi-user recognition result correction proportion A is compared with a preset parallel multi-user recognition result correction proportion A0, and the visual fine-tuning condition is judged according to the comparison result, wherein:

[0025] When A≤A0, it is determined that the visual fine-tuning condition does not need to be fine-tuned;

[0026] When A>A0, it is determined that the visual fine-tuning condition needs to be fine-tuned, and the model fine-tuning of the parallel multi-user recognition is carried out.

[0027] Further, in the step S7, the language interaction information, the visual perception information and the environment perception information are time-synchronized, format-unified and normalized to obtain preprocessed language interaction information, preprocessed visual perception information and preprocessed environment perception information, feature extraction is carried out on the preprocessed language interaction information, the preprocessed visual perception information and the preprocessed environment perception information to obtain language interaction information feature extraction data, visual perception information feature extraction data and environment perception information feature extraction data, feature layer fusion is carried out on the language interaction information feature extraction data, the visual perception information feature extraction data and the environment perception information feature extraction data according to a multi-modal convolutional neural network to obtain feature layer fusion data, which is used as language interaction fusion data.

[0028] Further, in the step S8, the language interaction fusion data is subjected to coarse-grained semantic layering according to a coarse-grained semantic layering method to obtain intent classification information and context awareness information, and the coarse-grained semantic layering method comprises:

[0029] Step S100, a neural network model comprising an LSTM layer and a full connection layer is constructed;

[0030] Step S200, the input of the neural network model comprising the LSTM layer and the full connection layer is set as the language interaction fusion data obtained through the previous information fusion step, and the output of the neural network model comprising the LSTM layer and the full connection layer is set as two parts, one part for intent classification and the other part for context awareness;

[0031] Step S300, language interaction fusion data samples are collected;

[0032] Step S400, the language interaction fusion data samples are divided into a training set, a validation set and a test set;

[0033] Step S500, a loss function of the neural network model comprising the LSTM layer and the full connection layer is set;

[0034] Step S600, model parameters of the neural network model comprising the LSTM layer and the full connection layer are updated;

[0035] Step S700, the neural network model comprising the LSTM layer and the full connection layer is outputted;

[0036] Step S800, the language interaction fusion data is inputted into the neural network model comprising the LSTM layer and the full connection layer, and intent classification information and context awareness information outputted by the neural network model comprising the LSTM layer and the full connection layer are obtained.

[0037] Further, in the step S9, after the coarse-grained semantic layering, the language interaction fusion data is subjected to fine-grained semantic understanding according to a fine-grained semantic understanding method to obtain sentiment analysis information, tone analysis information, deep context analysis information and ambiguity resolution information, and the fine-grained semantic understanding method comprises:

[0038] Step S1000, a sentiment dictionary is constructed, and the language interaction fusion data is subjected to fine-grained semantic understanding according to the sentiment dictionary to obtain sentiment analysis information;

[0039] Step S2000, the language interaction fusion data is subjected to fine-grained semantic understanding through feature engineering and a machine learning model to obtain tone analysis information;

[0040] In step S3000, the language interaction fusion data is subjected to fine-grained semantic understanding by combining the attention mechanism with the deep learning model, and deep context analysis information is obtained.

[0041] In step S4000, the language interaction fusion data is subjected to fine-grained semantic understanding by a polysemy representation learning model, and ambiguity resolution information is obtained.

[0042] Compared with the prior art, the method has the beneficial effects that the method provides a basic data source for the whole interaction process through real-time collection of language interaction information in step S1, through the multi-microphone array device, sound wave signals from different directions can be captured, the spatial resolution is enhanced, the digital person can more comprehensively obtain the surrounding voice information, important voice content is avoided to be missed, a foundation is laid for subsequent accurate understanding of user intent, the method helps to accurately identify the sound source direction and distance in a complex environment through sound source positioning in step S2, the physical position is estimated in combination with the environment model, and the sound source position information can be further accurately acquired, so that the digital person can focus on the target sound source, different speakers can be distinguished in a multi-sound source scene, and the voice of the target user is collected and processed in a targeted manner, the method can effectively improve the quality of the voice signal, reduce the influence of interfering sound on the target voice, and let the digital person receive clearer voice instructions or information, dynamically adjust the sensitivity and directivity of the microphone array, so that the collection effect can be automatically optimized according to the sound source position and environmental change, the voice interaction demand in different scenes is adapted, and the accuracy and stability of voice collection are further improved, the method enriches the understanding dimension of the digital person to the user behavior and intent through real-time collection of visual perception information and parallel multi-user identification in step S3, makes the interaction more natural and intuitive, and the method keeps the current collection state in a single user scene through step S4, unnecessary adjustment is avoided.In a multi-user scenario, the method adjusts the collection and separation of multiple user sound sources in time to ensure that the voice of each user can be effectively processed, improving the accuracy and adaptability of voice collection in a multi-user interaction scenario, avoiding voice collection confusion or omission. The method provides context information about the interactive environment for the digital person through step S5, generates an environment perception feature map through humidity, temperature, light, and other sensor data, and combines a user scenario recognition model. The digital person can determine the current environment type, correct the parallel multi-user recognition result according to the scenario recognition result, improve the accuracy of user recognition, avoid misjudgment caused by environmental factors, and make the digital person more accurately recognize the user situation in different scenarios, optimize the interaction strategy and service. The method judges the visual fine-tuning situation by calculating the proportion of the corrected parallel multi-user recognition result and comparing it with the preset value through step S6. This data-based judgment method is scientific and reasonable, and can accurately evaluate the performance of the visual model in actual application. When fine-tuning is needed, the visual model is optimized using transfer learning technology, which can fully utilize the knowledge and experience of existing models, quickly adapt to the specific needs of the digital person language interaction scenario, improve the accuracy of the visual model for parallel multi-user recognition, and thus improve the performance and stability of the entire system in complex scenarios. The method integrates language interaction information, visual perception information, and environmental perception information into an organic whole through step S7 information fusion, forming more comprehensive and rich language interaction fusion data. Combined with voice instructions and user behavior in a specific environment, the user's real needs can be more accurately understood, providing stronger data support for subsequent semantic understanding, improving the accuracy and depth of semantic understanding, and thus improving the digital person's understanding of user intent and the accuracy of interaction response. The method processes the fusion data through step S8 coarse-grained semantic layering to obtain intent classification information, enabling the digital person to quickly identify the general purpose of user voice, such as querying information, issuing instructions, expressing emotions, etc. The method further analyzes the coarse-grained semantic understanding through step S9 fine-grained semantic understanding to obtain emotion analysis information, enabling the digital person to perceive the user's emotional state, such as happiness, sadness, anger, etc., thereby adjusting the interaction method and tone to provide more considerate and personalized services. When detecting that the user's mood is not good, the digital person can respond with a more gentle and comforting tone, which helps the digital person understand the tone characteristics of user speech, such as statements, questions, imperatives, exclamations, etc., better grasp the user's intent and attitude, and respond more appropriately. Deep context analysis information enables the digital person to deeply understand the context meaning behind the voice, accurately understand pronoun reference, implicit information, etc., improve the understanding ability of complex semantics, avoid misunderstanding of user intent, and ambiguity resolution information solves the ambiguity problem in language, ensuring that the digital person accurately understands the user's voice, further improving the accuracy and fluency of interaction, and improving user experience. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 Figure 1 shows a flowchart of the multi-layer natural language interaction method based on deep learning and pattern recognition according to the present embodiment. DETAILED DESCRIPTION

[0044] In order to make the objects and advantages of the present application clearer, the present application will be further described below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0045] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. It should be understood by those skilled in the art that the embodiments are only used to explain the technical principles of the present application and are not used to limit the protection scope of the present application.

[0046] It should be noted that, in the description of the present application, the terms of "upper", "lower", "left", "right", "inner", "outer" and the like indicating the direction or positional relationship are based on the direction or positional relationship shown in the drawings, which is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present application.

[0047] In addition, it should also be noted that, in the description of the present application, unless otherwise explicitly specified and limited, the terms of "mounting", "connecting", "connection" should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0048] Please refer to Figure 1 Figure 1 shows a flowchart of the multi-layer natural language interaction method based on deep learning and pattern recognition according to the present embodiment, which comprises the following steps:

[0049] Step S1, collecting language interaction information in real time through a multi-microphone array device;

[0050] Step S2, positioning a sound source according to the language interaction information, and performing real-time separation and shielding on the real-time collection process of the language interaction information, and optimizing the multi-microphone array device;

[0051] Step S3, collecting visual perception information in real time, and performing parallel multi-user recognition according to the visual perception information;

[0052] Step S4, adjusting the real-time separation and shielding process of the real-time collection process of the language interaction information according to the parallel multi-user recognition result;

[0053] Step S5, real-time collection of environment perception information, user scene recognition according to the environment perception information, and correction of the parallel multi-user recognition result according to the user scene recognition result;

[0054] Step S6, judging the visual fine-tuning situation according to the parallel multi-user recognition result correction proportion, and model fine-tuning of the parallel multi-user recognition according to the visual fine-tuning situation;

[0055] Step S7, information fusion of language interaction information, visual perception information and environment perception information to obtain language interaction fusion data;

[0056] Step S8, coarse-grained semantic layering of the language interaction fusion data according to a coarse-grained semantic layering method to obtain intent classification information and context perception information;

[0057] Step S9, after coarse-grained semantic layering, fine-grained semantic understanding of the language interaction fusion data according to a fine-grained semantic understanding method to obtain sentiment analysis information, tone analysis information, deep context analysis information and ambiguity resolution information.

[0058] Specifically, the method is arranged in a digital human language interaction terminal, through sound source optimization collection of language interaction process, and based on accurate and efficient sound source data, multi-layer semantic understanding and analysis are carried out, and the natural language interaction efficiency in the digital human interaction process is improved, wherein the method provides basic data source for the whole interaction process through real-time collection of language interaction information in step S1, through the multi-microphone array device, sound wave signals from different directions can be captured, the spatial resolution is enhanced, the digital human can more comprehensively obtain the surrounding voice information, important voice content is avoided to be missed, and a basis for subsequent accurate understanding of user intention is laid, the method through sound source positioning in step S2 is helpful to accurately identify the sound source direction and distance in a complex environment, the physical position is estimated combined with the environment model, and the sound source position information can be further accurately acquired, so that the digital human can focus on the target sound source, different speakers can be distinguished in a multi-sound source scene, and the voice of the target user is collected and processed in a targeted manner, the method through real-time separation and shielding of target sound source signals and shielding of background noise can effectively improve the quality of voice signals, reduce the influence of interference sound on target voice, let the digital human receive clearer voice instructions or information, dynamically adjust the sensitivity and directivity of the microphone array, so that it can automatically optimize the collection effect according to the sound source position and environmental change, adapt to the voice interaction demand in different scenes, and further improve the accuracy and stability of voice collection, the method through step S3 real-time collection of visual perception information and parallel multi-user identification enriches the understanding dimension of the digital human to user behavior and intention, and makes the interaction more natural and intuitive, the method through step S4 keeps the current collection state in a single user scene, unnecessary adjustment is avoided.In a multi-user scenario, the method adjusts the collection and separation of multiple user sound sources in time to ensure that the voice of each user can be effectively processed, improving the accuracy and adaptability of voice collection in a multi-user interaction scenario, avoiding voice collection confusion or omission. The method provides context information about the interactive environment for the digital person through step S5, generates an environment perception feature map through humidity, temperature, light, and other sensor data, and combines a user scenario recognition model. The digital person can determine the current environment type, correct the parallel multi-user recognition result according to the scenario recognition result, improve the accuracy of user recognition, avoid misjudgment caused by environmental factors, and make the digital person more accurately recognize the user situation in different scenarios, optimize the interaction strategy and service. The method judges the visual fine-tuning situation by calculating the correction proportion of the parallel multi-user recognition result and comparing it with the preset value through step S6. This data-based judgment method is scientific and reasonable, and can accurately evaluate the performance of the visual model in actual application. When fine-tuning is needed, the visual model is optimized using transfer learning technology, which can fully utilize the knowledge and experience of existing models, quickly adapt to the specific needs of the digital person language interaction scenario, improve the accuracy of the visual model for parallel multi-user recognition, and thus improve the performance and stability of the entire system in complex scenarios. The method integrates language interaction information, visual perception information, and environmental perception information into an organic whole through step S7 information fusion, forming more comprehensive and rich language interaction fusion data. Combined with voice instructions and user behavior in a specific environment, the user's real needs can be more accurately understood, providing stronger data support for subsequent semantic understanding, improving the accuracy and depth of semantic understanding, and thus improving the digital person's understanding of user intent and the accuracy of interaction response. The method processes the fusion data through step S8 coarse-grained semantic layering to obtain intent classification information, enabling the digital person to quickly identify the general purpose of user voice, such as querying information, issuing instructions, expressing emotions, etc. The method further analyzes the coarse-grained semantic understanding through step S9 fine-grained semantic understanding to obtain emotion analysis information, enabling the digital person to perceive the user's emotional state, such as happiness, sadness, anger, etc., thereby adjusting the interaction method and tone to provide more considerate and personalized services. When detecting that the user's mood is not good, the digital person can respond with a more gentle and comforting tone, which helps the digital person understand the tone characteristics of the user's speech, such as statements, questions, imperatives, exclamations, etc., better grasp the user's intent and attitude, and respond more appropriately. Deep context analysis information enables the digital person to deeply understand the context meaning behind the voice, accurately understand pronoun reference, implicit information, etc., improve the understanding ability of complex semantics, avoid misunderstanding of user intent, and ambiguity resolution information solves the ambiguity problem in language, ensuring that the digital person accurately understands the user's voice, further improving the accuracy and fluency of interaction, and improving user experience.

[0059] Specifically, in the step S1, the language interaction information includes microphone signals of a plurality of microphone arrays, a current speech duration of a sound source, and a spectrum feature of the sound source, wherein the microphone signals of the plurality of microphone arrays refer to speech signals collected by each microphone in the plurality of microphone arrays.

[0060] Specifically, in the step S1, the configuration of the plurality of microphone arrays is used to capture sound wave signals from different directions, enhance spatial resolution, support subsequent sound source positioning and noise suppression, realize real-time capture of speech signals through fast sampling and data transmission technology, and store data in a cache to support subsequent fast processing.

[0061] Specifically, in the step S2, a microphone signal generalized cross-correlation function Rx1x2(τ) is calculated according to the first microphone signal x1(t) and the second microphone signal x2(t) in the language interaction information, and is set as Ψ x1x2 (f) is a weighting function, x1(f) is a Fourier transform of x1(t), x2(f) is a Fourier transform of x2(t), x2*(f) is a conjugate complex of x2(f), e is a base of a natural logarithm, j is an imaginary unit, f is a frequency, which represents an identification of different frequency components in a frequency domain when a sound signal received by a microphone is analyzed and the signal is converted from a time domain to the frequency domain, τ is a variable related to a time delay, a peak τmax of the microphone signal generalized cross-correlation function Rx1x2(τ) is obtained, and a distance difference △r of a sound source to a first microphone and a second microphone is calculated according to the peak τmax, and is set as △r=τmax×c, c is a sound propagation speed, and an angle θ of the sound source to a microphone array plane is calculated, and is set as r is a distance of the sound source to a center of the microphone array, and the distance difference △r and the angle θ are taken as positioning information of the sound source.

[0062] Specifically, the first microphone signal refers to the microphone signal collected by a randomly specified single microphone in the multi-microphone array device, and the second microphone signal refers to the microphone signal collected by another randomly specified single microphone in the multi-microphone array device. The embodiment does not limit the random specification manner of the multi-microphone array device, and a person skilled in the art can freely set it according to the actual situation, for example, the microphones in the multi-microphone array device can be digitally marked, and two random numbers corresponding to the microphones are selected by a random number generator to serve as the first microphone and the second microphone, respectively. The microphone signal collected by the first microphone is taken as the first microphone signal, and the microphone signal collected by the second microphone is taken as the second microphone signal. The sound source refers to the user sound source in front of the digital human language interaction terminal for voice interaction. The microphone array plane refers to the plane formed after the connection of the microphones in the multi-microphone array device. The microphone array center refers to the center of the microphone array plane. The embodiment does not limit the acquisition method of the distance from the sound source to the microphone array center, and a person skilled in the art can freely set it according to the actual situation, as long as the acquisition requirement of the distance from the sound source to the microphone array center is met, for example, the distance from the sound source to the microphone array center can be collected by a sonar sensor.

[0063] Specifically, in the step S2, after obtaining the positioning information of the sound source, the effective coefficient D of the sound source is calculated as D=α1×1 / r+α2×T+α3×△F, where α1 is the first effective coefficient weight parameter, α2 is the second effective coefficient weight parameter, α3 is the third effective coefficient weight parameter, T is the current speech duration of the sound source, and △F is the sum of the differences between the sound source and the reference value on each frequency spectrum feature, that is, △F=∑ n k=1 |Fk-Fk'|, Fk' is the reference value of each frequency spectrum feature. The sound source corresponding to the maximum effective coefficient Dmax is taken as the target sound source. The real-time collection process of the language interaction information of the target sound source is separated and shielded in real time according to the positioning information of the target sound source. The multi-microphone array device is optimized until the positioning information of the target sound source has the minimum distance difference and the angle of 90°.

[0064] Specifically, the embodiment does not limit the setting manner of the effective coefficient weight parameter, and a person skilled in the art can freely set according to the demand, as long as the demand for identifying the target sound source is met. For example, the first effective coefficient weight parameter can be set as 0.5, the second effective coefficient weight parameter can be set as 0.3, and the third effective coefficient weight parameter can be set as 0.2. The current voice duration of the sound source refers to the duration of the same sound source continuously outputting voice in front of the digital human language interaction terminal at the current moment. The reference value of each spectral feature refers to the preset reference value of the energy distribution of the voice signal at different frequency bands. The embodiment does not limit the specific parameters of the reference value of each spectral feature, and a person skilled in the art can freely set according to the actual situation, as long as the reflection demand of the energy distribution is met. For example, the reference value of each spectral feature can be preset according to the different setting scenes of the digital human language interaction terminal. The maximum effective coefficient refers to the maximum value of the effective coefficient of each sound source that performs voice interaction in front of the digital human language interaction terminal. The positioning information of the target sound source refers to the distance difference Δr and the included angle θ of the target sound source. The real-time separation shielding refers to that the digital human language interaction terminal only accepts the language interaction information from the target sound source. The embodiment does not limit the specific manner of the real-time separation shielding, and a person skilled in the art can freely set according to the actual situation, as long as the demand for separately receiving the language interaction information of the target sound source is met. For example, the real-time separation shielding can be performed through blind source separation technology. The embodiment does not limit the optimization manner of the multi-microphone array device, and a person skilled in the art can freely set according to the actual situation, as long as the optimization demand is met. For example, the built-in controller in the multi-microphone array device can be used to control each microphone, and each microphone can be moved through a slide rail until the positioning information of the target sound source is the minimum distance difference and the included angle is 90°.

[0065] Specifically, in the step S3, the visual perception information is input into a user identification deep learning model, and a multi-user identification result output by the user identification deep learning model is obtained.

[0066] Specifically, the visual perception information refers to image information in front of the digital human language interaction terminal, including visible light image data, infrared image data and laser radar data, the user recognition deep learning model refers to a neural network model for obtaining the parallel multi-user recognition result after analyzing the visual perception information, the basic framework of the user recognition deep learning model is set as a convolutional neural network model in this embodiment, historical visual perception information-historical multi-user recognition result is used as a recognition model construction data set, 70% of the recognition model construction data set is used as a recognition model training set, 30% of the recognition model construction data set is used as a recognition model validation set, the convolutional neural network model is trained according to the recognition model training set, a trained convolutional neural network model is obtained, the trained convolutional neural network model is verified according to the recognition model validation set, and the trained convolutional neural network model is output as the user recognition deep learning model after the accuracy reaches 98%. The construction parameters of the convolutional neural network model are not specifically limited in this embodiment, and a person skilled in the art can freely set them according to the actual situation, as long as the multi-user recognition requirement is met. For example, the loss function of the convolutional neural network model can be set as a cross-entropy function, the parallel multi-user recognition result is obtained by the user recognition deep learning model according to the visual perception information, and indicates whether the user in front of the digital human language interaction terminal is a single user or multiple users, including a single user and multiple users.

[0067] Specifically, in the step S4, the adjustment of the real-time separation shielding process of the real-time collection process of the language interaction information is judged according to the parallel multi-user recognition result, wherein:

[0068] When the parallel multi-user recognition result is a single user, it is determined that the real-time separation shielding process of the real-time collection process of the language interaction information is not adjusted;

[0069] When the parallel multi-user recognition result is multiple users, it is determined that the real-time separation shielding process of the real-time collection process of the language interaction information is adjusted, the target sound source after adjustment is the sound source of multiple users, and the real-time collection process of the language interaction information of the target sound source after adjustment is separated and shielded according to the positioning information of the target sound source after adjustment.

[0070] Specifically, in the step S5, an environment perception feature map is generated according to the environment perception information, the environment perception feature map is input into the user scene recognition model, the user scene recognition result output by the user scene recognition model is obtained, and the parallel multi-user recognition result is corrected according to the user scene recognition result, wherein:

[0071] When the user scene recognition result is a multi-person interaction scene, the parallel multi-user recognition result is not corrected;

[0072] When the user scene recognition result is a single-person interaction scene, the parallel multi-user recognition result is corrected, and the parallel multi-user recognition result of multiple users is corrected to a single user.

[0073] Specifically, the environment perception information refers to information collected by Internet of Things sensors such as humidity sensors, temperature sensors, and illumination sensors on the interactive environment of the digital human language interaction terminal, and the environment perception feature map refers to a data feature map related to the interactive environment of the digital human language interaction terminal generated according to the environment perception information. The present embodiment does not limit the manner of generating the environment perception feature map, and those skilled in the art can freely set it according to the actual situation, as long as the requirement of converting the text features of the environment perception information into a picture is met. For example, the environment perception feature map can be generated by direct mapping based on sensor data. When the environment perception information is collected by various sensors such as temperature sensors, humidity sensors, illumination sensors, gas sensors, etc., the sensor data can be directly mapped to the pixel values of the image according to certain rules, thereby generating an environment data feature map. Taking a temperature sensor as an example, assuming that multiple temperature sensors are arranged to form a sensor network for monitoring the temperature at different positions on the digital human language interaction terminal. First, the size and resolution of the feature map are determined, such as generating a 50x50 pixel feature map to represent the indoor temperature distribution. Then, the indoor space is divided according to a division rule, such as a uniform grid, and mapped to the pixels of the feature map. For example, the entire indoor space is divided into 2500 small areas, corresponding to 50x50 pixels, and each small area corresponds to the monitoring range of a sensor. The data collected by each temperature sensor is normalized so that its value falls within an appropriate range, such as 0 to 255, to match the pixel value range of the image. Assuming that the temperature value collected by a certain temperature sensor is 20°C, the normalized value is 100. The normalized temperature value is assigned to the pixel corresponding to the corresponding small area, thereby generating an environment data feature map representing the indoor temperature distribution. Different pixel gray values or color values in the map reflect the temperature conditions at different positions. The user scene recognition model refers to a deep learning model that takes the environment perception feature map as input and outputs a user scene recognition result. The present embodiment does not limit the construction method of the user scene recognition model, and those skilled in the art can freely set it according to the actual situation, as long as the requirement of recognizing the interactive environment of the digital human language interaction terminal is met. For example, the user scene recognition model can be a convolutional neural network model. The user scene recognition result refers to the recognition result of the interactive environment of the digital human language interaction terminal obtained from the environment perception feature map, including a multi-person interactive scene and a single-person interactive scene. The multi-person interactive scene refers to a situation where multiple people interact with the interactive object in the interactive environment of the digital human language interaction terminal, such as schools, hospitals, and subway stations. The single-person interactive scene refers to a situation where there is no multi-person interaction with the interactive object in the interactive environment of the digital human language interaction terminal, and only single-person interaction is possible, such as a private office.

[0074] Specifically, in the step S6, the parallel multi-user recognition result correction proportion A is calculated according to the user scene recognition times L0 and the parallel multi-user recognition result correction times L1, A is set as L1 / L0, the parallel multi-user recognition result correction proportion A is compared with the preset parallel multi-user recognition result correction proportion A0, and the visual fine-tuning situation is judged according to the comparison result, wherein:

[0075] When A≤A0, it is determined that the visual fine-tuning situation does not need fine-tuning;

[0076] When A>A0, it is determined that the visual fine-tuning situation needs fine-tuning, and the model fine-tuning of the parallel multi-user recognition is performed.

[0077] Specifically, the user scene recognition times refer to the total number of user scene recognitions in the step S5 according to the environmental perception information in the fine-tuning monitoring period, the parallel multi-user recognition result correction times refer to the total number of corrections of the parallel multi-user recognition result in the step S5 according to the user scene recognition result in the fine-tuning monitoring period, the fine-tuning monitoring period refers to a preset period for monitoring the correction situation, the preset parallel multi-user recognition result correction proportion refers to a parameter reflecting the correction situation, and the visual fine-tuning situation refers to the possibility of judging whether to perform model fine-tuning on the parallel multi-user recognition according to the correction situation, including no need for fine-tuning and need for fine-tuning. In this embodiment, the model fine-tuning of the parallel multi-user recognition is performed by using the transfer learning technology. The basic idea of transfer learning is to use the model parameters trained on other related tasks or data sets to initialize the model of the current task, and then further train on the data set of the current task. Assuming that there is a convolutional neural network model pre-trained on a large-scale general image data set, the model has good feature extraction capability, but is not accurate enough for parallel multi-user recognition in the digital human language interaction scene. First, the parameters of part of the layers of the pre-trained model are kept unchanged as the initialization parameters of the new model, such as the previous convolutional layers. These layers have learned general image features such as edges and textures. The visual perception information collected in the digital human language interaction scene and the corresponding parallel multi-user recognition result labels are used to construct a new training data set, including visible light image data, infrared image data, and laser radar data, etc. Then, the last few layers of the model are retrained, such as fully connected layers, etc. The parameters of the layers are adjusted so that the model can better adapt to the parallel multi-user recognition task in the digital human language interaction scene. In the training process, according to the characteristics of the new data set and the task requirements, a suitable loss function such as cross-entropy function and an optimization algorithm such as stochastic gradient descent are selected to update the model parameters.

[0078] Specifically, in the step S7, the language interaction information, the visual perception information and the environmental perception information are time-synchronized, format-unified and normalized to obtain pre-processed language interaction information, pre-processed visual perception information and pre-processed environmental perception information, feature extraction is performed on the pre-processed language interaction information, the pre-processed visual perception information and the pre-processed environmental perception information to obtain language interaction information feature extraction data, visual perception information feature extraction data and environmental perception information feature extraction data, and feature layer fusion is performed on the language interaction information feature extraction data, the visual perception information feature extraction data and the environmental perception information feature extraction data according to a multi-modal convolutional neural network to obtain feature layer fusion data as language interaction fusion data.

[0079] Specifically, in the time synchronization of the language interaction information, the visual perception information and the environmental perception information, the timestamp technology is used to mark the time when the language interaction information, the visual perception information and the environmental perception information are collected, and then the language interaction information, the visual perception information and the environmental perception information collected in the same time window are associated, for example, when the multi-microphone array device collects language interaction information, the time stamp of the collection time is recorded at the same time; when the camera collects visual perception information, the corresponding time stamp is also marked; when various sensors collect environmental perception information, the time stamp is also recorded, by comparing these time stamps, the data in the same or similar time range is regarded as multi-modal data at the same time for fusion processing, in the format unification and normalization of the language interaction information, the visual perception information and the environmental perception information, for the image pixel value (usually between 0-255) in the visual perception information and the temperature value (may be in a larger actual temperature range) in the environmental perception information, they are normalized to 0-1 or -1 to 1 through linear mapping, in the feature extraction of the preprocessed language interaction information, the preprocessed visual perception information and the preprocessed environmental perception information, the Fourier transform is used to obtain the frequency spectrum of the voice in the preprocessed language interaction information, and the frequency component and energy distribution therein are analyzed to extract the spectral features; the prosody features are calculated by the length, pitch change of the voice signal in the preprocessed language interaction information; the semantic features are obtained by using the morphological and syntactic analysis of the text after the voice in the preprocessed language interaction information is transcribed by using the natural language processing tool; the face feature point coordinates in the preprocessed visual perception information are obtained by using the face recognition technology in the computer vision algorithm, the human body posture angle is determined by using the pose estimation algorithm, the thermal imaging features are determined by analyzing the temperature gradient of the infrared image in the preprocessed visual perception information, the three-dimensional model of the environment is constructed according to the point cloud data of the laser radar in the preprocessed visual perception information, and the geometric features of the object are extracted;The average value and variance of the data collected by the humidity sensor in the pre-processed environment perception information within a preset time are calculated, the temperature fluctuation range is obtained by analyzing the difference between the maximum value and the minimum value of the temperature sensor data in the pre-processed environment perception information in different time periods, and the light direction is determined according to the measurement values of the light sensor in different directions in the pre-processed environment perception information. When performing feature layer fusion, each branch of the multi-modal convolutional neural network is designed to process the language interaction information feature extraction data, the visual perception information feature extraction data, and the environment perception information feature extraction data, respectively, and then information fusion and interaction are performed in the subsequent layers of the multi-modal convolutional neural network. The construction method of the multi-modal convolutional neural network is not limited in the embodiment, and a person skilled in the art can freely set it according to the actual situation, as long as the data fusion requirement is met. For example, when constructing the multi-modal convolutional neural network, the number of nodes of the input layer is determined according to the feature dimension of the fused data, multiple hidden layers can be set, the complex relationship in the multi-modal data is learned by adjusting the number of neurons and the activation function of each layer, and the number of nodes and the output format of the output layer are determined according to the target of the language interaction fusion data, such as generating intent classification information and context perception information.

[0080] Specifically, in the step S8, the language interaction fusion data is subjected to coarse-grained semantic layering according to a coarse-grained semantic layering method to obtain intent classification information and context perception information. The coarse-grained semantic layering method includes the following steps.

[0081] In step S100, a neural network model including an LSTM layer and a fully connected layer is constructed. The LSTM layer is used to model the time sequence information in the language interaction fusion data, and the fully connected layer is used to map the output of the LSTM layer to the dimensions of intent classification and context perception.

[0082] In step S200, the input of the neural network model including the LSTM layer and the fully connected layer is set to the language interaction fusion data obtained through the previous information fusion step, and the output of the neural network model including the LSTM layer and the fully connected layer is set to two parts. One part is used for intent classification, and the output category number of the intent classification is determined according to the pre-defined intent categories, including query information, issuing instructions, expressing emotions, etc. The output node number of the neural network model including the LSTM layer and the fully connected layer can be set to the category number, and the probability distribution of each intent classification is obtained through the Softmax activation function. The other part is used for context perception, and the output is a feature vector related to the context, such as a vector representing the current dialogue topic, dialogue history key information, etc. The dimension of the vector is determined according to the specific context modeling requirement.

[0083] Step S300, collect language interaction fusion data samples, and artificially mark the language interaction fusion data samples to obtain an intent category to which each sample artificially marked belongs and context information artificially marked related to each sample; the context information related to each sample includes a dialogue topic, a previously mentioned key entity, and the like;

[0084] Step S400, divide the language interaction fusion data samples into a training set, a validation set, and a test set; the training set is used for training of the neural network model containing an LSTM layer and a full connection layer, and accounts for 80% of the language interaction fusion data samples; the validation set is used for evaluating performance of the neural network model containing the LSTM layer and the full connection layer in a training process, and accounts for 10% of the language interaction fusion data samples; and the test set is used for finally evaluating generalization ability of the neural network model containing the LSTM layer and the full connection layer on unseen data, and accounts for 10% of the language interaction fusion data samples;

[0085] Step S500, set a loss function of the neural network model containing the LSTM layer and the full connection layer; for an intent classification part of the neural network model containing the LSTM layer and the full connection layer, a cross-entropy loss function is adopted, and a formula is c is a number of classes of intent classification, y i is an element (0 or 1) in a one-hot encoding vector corresponding to a class of a real intent classification of a sample, and y i' is an element in a class probability distribution vector of an intent classification predicted by the model; for a context perception part, a mean square error (MSE) loss function can be adopted, and a formula is n is a dimension of a context perception vector, b i is an element in a real context perception vector of a sample, and b i' is an element in a context perception vector predicted by the neural network model containing the LSTM layer and the full connection layer;

[0086] Step S600, update model parameters of the neural network model containing the LSTM layer and the full connection layer; a stochastic gradient descent (SGD) and a variant (Adam) optimization algorithm are used to update the model parameters;

[0087] Step S700, outputting the neural network model comprising the LSTM layer and the fully connected layer; inputting the training set batch into the neural network model comprising the LSTM layer and the fully connected layer, calculating a loss function value, then updating the parameters of the neural network model comprising the LSTM layer and the fully connected layer by reverse propagation of gradients through an optimization algorithm; after the end of each training cycle, evaluating the performance of the neural network model comprising the LSTM layer and the fully connected layer using a validation set, the performance of the neural network model comprising the LSTM layer and the fully connected layer including the accuracy of intent classification, the similarity index of the context awareness vector and the real context information; if the performance does not improve or there are signs of overfitting, such as the validation set loss starting to rise, the training is stopped in advance, and the model parameters are adjusted, such as increasing the regularization strength, adjusting the learning rate, etc., and the training is continued until the model reaches good performance on the validation set, then the model is output as the neural network model comprising the LSTM layer and the fully connected layer.

[0088] Step S800, inputting the language interaction fusion data into the neural network model comprising the LSTM layer and the fully connected layer, and obtaining the intent classification information and the context awareness information output by the neural network model comprising the LSTM layer and the fully connected layer.

[0089] Specifically, in the step S9, after the coarse-grained semantic layering, the language interaction fusion data is subjected to fine-grained semantic understanding according to a fine-grained semantic understanding method, to obtain sentiment analysis information, tone analysis information, deep context analysis information and ambiguity resolution information, the fine-grained semantic understanding method comprising:

[0090] Step S1000, constructing a sentiment dictionary, and performing fine-grained semantic understanding on the language interaction fusion data according to the sentiment dictionary to obtain sentiment analysis information;

[0091] Step S2000, performing fine-grained semantic understanding on the language interaction fusion data through feature engineering and a machine learning model to obtain tone analysis information;

[0092] Step S3000, performing fine-grained semantic understanding on the language interaction fusion data through a combination of an attention mechanism and a deep learning model to obtain deep context analysis information;

[0093] Step S4000, performing fine-grained semantic understanding on the language interaction fusion data through a polysemy representation learning model to obtain ambiguity resolution information.

[0094] Specifically, the sentiment dictionary refers to a collection of rich sentiment words and their sentiment tendencies such as positive, negative, neutral, and intensity. It constructs a custom sentiment dictionary according to specific fields and application scenarios to better adapt to the language characteristics of digital human interaction. When performing fine-grained semantic understanding on the language interaction fusion data according to the sentiment dictionary to obtain sentiment analysis information, the text in the language interaction fusion data is matched with the words in the sentiment dictionary. For each matched sentiment word, a corresponding score is assigned according to its sentiment tendency and intensity in the dictionary, for example, "happy" is assigned a positive score of +3, and "sad" is assigned a negative score of -3. Then, the sentiment tendency score of the text is obtained by summing the scores of all sentiment words in the text. According to the score, the sentiment category is judged, such as a score greater than 0 indicating positive sentiment, a score less than 0 indicating negative sentiment, and a score close to 0 indicating neutral sentiment, thereby obtaining sentiment analysis information.In the fine-grained semantic understanding of the language interaction fusion data by feature engineering and machine learning model, features related to tone are extracted from the language interaction fusion data, such as the use of punctuation marks (the frequency of exclamation marks, question marks, etc.), the part of speech of words (such as the proportion and type of verbs, adjectives, adverbs, etc.), sentence structure (the proportion of declarative sentences, interrogative sentences, imperative sentences, exclamatory sentences, etc.), tone variation features (if there is voice tone information, the rising, falling, and flat features can be analyzed), etc. The machine learning model, such as Naive Bayes, is trained using the features related to tone. The language interaction data labeled with different tone types (such as declarative tone, interrogative tone, imperative tone, exclamatory tone, etc.) is used as the training set. The relationship between features and tone types is learned by the machine learning model. The features of the language interaction fusion data are input into the trained machine learning model, and the tone type is predicted by the machine learning model to obtain tone analysis information. In the fine-grained semantic understanding of the language interaction fusion data by combining attention mechanism and deep learning model, the attention mechanism is introduced into the Transformer deep learning model, which can focus on the important parts of the text related to the context. The attention mechanism assigns different weights to different parts of the input text according to their relevance to the current task, highlighting the key information in the text. Through training on a large amount of language interaction data, the Transformer deep learning model learns how to understand the deep context based on the context and attention weight. For example, in a conversation about tourism, when referring to "that place", the Transformer deep learning model focuses on the previously mentioned tourism destination-related information through the attention mechanism, accurately understanding the specific content of "that place", and further obtaining deep context analysis information to grasp the context of the entire conversation. In the fine-grained semantic understanding of the language interaction fusion data by the polysemy representation learning model, a deep learning model is used to learn multiple semantic representations of words and sentences. For example, through training on a large-scale corpus, the polysemy representation learning model can learn different vector representations of "bank" in different contexts. In practical applications, according to the context of the language interaction fusion data, the polysemy representation learning model selects the semantic representation that best fits the context. The polysemy representation learning model selects the semantic representation with the highest probability or strongest correlation by calculating the relevance or probability of different semantic representations and the context, thereby realizing ambiguity resolution and obtaining ambiguity resolution information.

[0095] The technical scheme of the present application has been described in combination with the preferred embodiments shown in the drawings, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical schemes after the changes or replacements will all fall within the protection scope of the present application.

Claims

1. A multi-layer natural language interaction method based on deep learning and pattern recognition, characterized in that, include: Step S1: Real-time acquisition of language interaction information using a multi-microphone array device; Step S2: Locate the sound source based on the language interaction information, perform real-time separation and shielding of the real-time acquisition process of the language interaction information, and optimize the multi-microphone array device; Step S3: Real-time acquisition of visual perception information, and parallel multi-user identification based on the visual perception information; Step S4: Adjust the real-time separation and masking process of the real-time acquisition process of the language interaction information based on the parallel multi-user recognition results; Step S5: Collect environmental perception information in real time, identify user scenarios based on the environmental perception information, and correct the parallel multi-user identification results based on the user scenario identification results. Step S6: Determine the visual fine-tuning situation based on the correction ratio of the parallel multi-user recognition results, and fine-tune the model of parallel multi-user recognition based on the visual fine-tuning situation. Step S7: Information fusion is performed on language interaction information, visual perception information and environmental perception information to obtain language interaction fusion data; Step S8: Perform coarse-grained semantic layering on the language interaction fusion data according to the coarse-grained semantic layering method to obtain intent classification information and context-aware information; Step S9: After performing coarse-grained semantic layering, fine-grained semantic understanding is performed on the language interaction fusion data according to the fine-grained semantic understanding method to obtain sentiment analysis information, tone analysis information, deep context analysis information, and ambiguity resolution information.

2. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 1, characterized in that, In step S2, the generalized cross-correlation function Rx1x2(τ) of the microphone signals is calculated based on the first microphone signal x1(t) and the second microphone signal x2(t) in the language interaction information, and the setting is... Ψ x1x2 (f) is the weighting function, x1(f) is the Fourier transform of x1(t), x2(f) is the Fourier transform of x2(t), x2*(f) is the complex conjugate of x2(f), e is the base of the natural logarithm, j is the imaginary unit, f is the frequency, which represents the identifier of different frequency components in the frequency domain when the signal is converted from the time domain to the frequency domain during the analysis of the sound signal received by the microphone, τ is a time delay-related variable, the peak value τmax of the generalized cross-correlation function Rx1x2(τ) of the microphone signal is obtained, and the distance difference Δr between the sound source and the first and second microphones is calculated based on the peak value τmax, setting Δr = τmax × c, c is the sound propagation speed, and the angle θ between the sound source and the microphone array plane is calculated, setting r is the distance from the sound source to the center of the microphone array, and the distance difference Δr and the included angle θ are used as the positioning information of the sound source.

3. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 2, characterized in that, In step S2, after obtaining the location information of the sound source, the effective coefficient D of the sound source is calculated as follows: D = α1 × 1 / r + α2 × T + α3 × ΔF, where α1 is the first effective coefficient weighting parameter, α2 is the second effective coefficient weighting parameter, α3 is the third effective coefficient weighting parameter, T is the current speech duration of the sound source, and ΔF refers to the sum of the differences between the sound source and the reference value in each spectral feature. Fk' is the reference value of each spectral feature. The sound source corresponding to the maximum effective coefficient Dmax is taken as the target sound source. The real-time acquisition process of the language interaction information of the target sound source is separated and shielded in real time according to the positioning information of the target sound source. The multi-microphone array device is optimized until the positioning information of the target sound source is the minimum distance difference and the included angle is 90°.

4. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 3, characterized in that, In step S3, the visual perception information is input into the user recognition deep learning model, and the multi-user recognition result output by the user recognition deep learning model is obtained.

5. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 4, characterized in that, In step S4, the adjustment of the real-time separation and masking process of the real-time acquisition process of the language interaction information is judged based on the parallel multi-user recognition results, wherein: When the parallel multi-user identification result is a single user, it is determined that the real-time separation and masking process of the real-time acquisition process of the language interaction information will not be adjusted. When the parallel multi-user identification result is multiple users, it is determined that the real-time separation and masking process of the real-time acquisition process of the language interaction information is adjusted. The adjusted target sound source is the sound source of multiple users. Based on the positioning information of the adjusted target sound source, the real-time separation and masking process of the language interaction information of the adjusted target sound source is performed.

6. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 5, characterized in that, In step S5, an environmental perception feature map is generated based on the environmental perception information, and the environmental perception feature map is input into the user scene recognition model to obtain the user scene recognition result output by the user scene recognition model. The parallel multi-user recognition result is then corrected based on the user scene recognition result, wherein: When the user scene identification result is a multi-user interaction scene, the parallel multi-user identification result is not corrected; When the user scene recognition result is a single-person interaction scene, the parallel multi-user recognition result is corrected to correct the parallel multi-user recognition result of multiple users to a single user.

7. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 6, characterized in that, In step S6, the parallel multi-user recognition result correction ratio A is calculated based on the number of user scene recognition attempts L0 and the number of parallel multi-user recognition result correction attempts L1, and A = L1 / L0 is set. The parallel multi-user recognition result correction ratio A is compared with the preset parallel multi-user recognition result correction ratio A0, and the visual fine-tuning situation is judged based on the comparison result, wherein: When A≤A0, the visual fine-tuning situation is determined to be that no fine-tuning is needed; When A > A0, the visual fine-tuning situation is determined to be in need of fine-tuning, and model fine-tuning is performed on the parallel multi-user recognition.

8. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 1, characterized in that, In step S7, the language interaction information, visual perception information, and environmental perception information are synchronized in time, formatted, and normalized to obtain preprocessed language interaction information, preprocessed visual perception information, and preprocessed environmental perception information. Feature extraction is performed on the preprocessed language interaction information, preprocessed visual perception information, and preprocessed environmental perception information to obtain language interaction information feature extraction data, visual perception information feature extraction data, and environmental perception information feature extraction data. Feature layer fusion is performed on the language interaction information feature extraction data, visual perception information feature extraction data, and environmental perception information feature extraction data according to a multimodal convolutional neural network to obtain feature layer fused data, which is used as language interaction fused data.

9. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 8, characterized in that, In step S8, the language interaction fusion data is subjected to coarse-grained semantic layering according to a coarse-grained semantic layering method to obtain intent classification information and context-aware information. The coarse-grained semantic layering method includes: Step S100: Construct a neural network model containing LSTM layers and fully connected layers; Step S200: Set the input of the neural network model containing LSTM layers and fully connected layers to the language interaction fusion data obtained through the previous information fusion step, and set the output of the neural network model containing LSTM layers and fully connected layers to two parts, one part for intent classification and the other part for context awareness. Step S300: Collect language interaction fusion data samples; Step S400: Divide the language interaction fusion data samples into a training set, a validation set, and a test set; Step S500: Set the loss function for the neural network model containing LSTM layers and fully connected layers; Step S600: Update the model parameters of the neural network model containing LSTM layers and fully connected layers; Step S700: Output the neural network model containing LSTM layers and fully connected layers; Step S800: Input the language interaction fusion data into the neural network model containing LSTM layers and fully connected layers, and obtain the intent classification information and context-aware information output by the neural network model containing LSTM layers and fully connected layers.

10. The multi-layer natural language interaction method based on deep learning and pattern recognition according to claim 9, characterized in that, In step S9, after performing coarse-grained semantic layering, fine-grained semantic understanding is performed on the language interaction fusion data according to a fine-grained semantic understanding method to obtain sentiment analysis information, tone analysis information, deep context analysis information, and disambiguation information. The fine-grained semantic understanding method includes: Step S1000: Construct an emotion dictionary and perform fine-grained semantic understanding on the language interaction fusion data based on the emotion dictionary to obtain emotion analysis information; Step S2000: Fine-grained semantic understanding of the language interaction fusion data is performed through feature engineering and machine learning models to obtain tone analysis information; Step S3000: Fine-grained semantic understanding of the language interaction fusion data is performed by combining the attention mechanism with a deep learning model to obtain deep context analysis information; Step S4000: Fine-grained semantic understanding of the language interaction fusion data is performed through a polysemous representation learning model to obtain ambiguity resolution information.

Citation Information

Patent Citations

  • Forestry ecology environment human-computer interaction method based on natural language processing

    CN108009285A

  • Group behavior identification method based on multi-feature fusion

    CN106529467A

  • Control system and method of artificial intelligence drive robot

    CN108297098A