Multifunctional integrated service terminal and data processing method applied by same

Through the combination of a multimodal fusion framework and an intelligent decision-making engine, the problem that traditional multi-functional all-in-one machines are difficult to understand user intentions in complex noise environments is solved, efficient and accurate data processing and office service response are achieved, and user experience and office efficiency are improved.

CN120277620AActive Publication Date: 2025-07-08GUANGZHOU SUNRISE ELECTRONICS TECH

Patent Information

Application Number
CN202510764003.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Traditional multi-functional all-in-one machines are difficult to accurately understand user intentions in complex noise environments, resulting in poor user experience and inefficient office efficiency.

Method used

Using a multimodal fusion framework and an intelligent decision-making engine, combining environment perception algorithms and intention recognition models, an integrated control center is built through high-precision microphone arrays, multi-touch screens and biometric sensors to optimize data processing capabilities in complex noise environments.

Benefits of technology

It improves the recording quality and user identity verification accuracy in complex noise environments, enhances the intuitiveness of human-computer interaction and system security, and improves office efficiency and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277620A_ABST
    Figure CN120277620A_ABST
Patent Text Reader

Abstract

The invention discloses a multifunctional integrated service terminal and a data processing method applied by the same, and relates to the technical field of intelligent office equipment.The method comprises the steps that the cooperative relation between a multi-mode interaction module and a business processing module is determined, and the mapping relation from the business processing module to different scene service output is determined; an integrated control center is constructed based on a multi-modal fusion framework corresponding to the cooperative relation and an intelligent decision engine corresponding to the mapping relation, the multi-modal fusion framework comprises an environment perception algorithm, and the intelligent decision engine comprises an intention recognition model; obtaining a training data set containing interaction data in a complex noise environment; training an integrated control center based on the training data set, and combining an environment perception algorithm and an intention recognition model for joint optimization; inputting a to-be-served multi-mode input signal into the integrated control center, and outputting a directional sound recording file, a business handling instruction and an office service response; the office efficiency of the multifunctional all-in-one machine product is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent office equipment, and particularly relates to a multi-functional integrated service terminal and a data processing method applied thereto. Background Art

[0002] With the development of information technology, the single-function service devices in the traditional office environment can gradually no longer meet the needs of efficient multitask processing. Especially for scenarios that require audio recording, identity authentication, and intelligent assistance, the scattered use of separate devices such as computers, high-speed document scanners, and voice recorders not only occupies space but also makes it difficult to work collaboratively, reducing work efficiency and service experience.

[0003] Currently, there are all-in-one machine products integrating multiple functions in the market, such as desktop all-in-one machines integrating functions such as cameras and touch screens. However, they usually lack professional audio processing capabilities and intelligent human-computer interaction designs. Especially in terms of accurate role separation recording and intention recognition in complex noise environments, they perform poorly and cannot well understand the intentions of customers in noisy environments, resulting in poor user experience and low office efficiency. Therefore, improvements are needed. Summary of the Invention

[0004] In order to improve the office efficiency of multi-functional all-in-one machine products and enhance the user experience, this application provides a multi-functional integrated service terminal and a data processing method applied thereto.

[0005] In the first aspect, the invention objective of this application is achieved by adopting the following technical solutions: A multi-functional integrated service terminal, comprising: Determine the collaboration relationship between the multi-modal interaction module and the service processing module, and determine the mapping relationship from the service processing module to service outputs in different scenarios; Based on the multi-modal fusion framework corresponding to the collaboration relationship and the intelligent decision-making engine corresponding to the mapping relationship, construct an integrated control center, where the multi-modal fusion framework includes an environment perception algorithm, and the intelligent decision-making engine includes an intention recognition model; Obtain a training data set containing interaction data in a complex noise environment, where the training data set includes several types of multi-source heterogeneous signals; Based on the training data set, train the integrated control center, and jointly optimize the environment perception algorithm and the intention recognition model during the training process; Input the multi-modal input signal to be served into the trained integrated control center, and output a directional recording file, a service handling instruction, and an office service response.

[0006] By adopting the above technical solutions, several types of multi-source heterogeneous signals, including voice, touch, biometric features, etc., cover heterogeneous business government affairs data involved in office scenarios (such as bank business halls); based on different business service scenarios and actual business interaction scenarios, the multi-functional integrated service terminal of the present application is provided with a multi-modal fusion framework and an intelligent decision-making engine to improve the clarity and / or recognizability of different types of multi-source heterogeneous signals, and is applicable to scenarios that require accurate recording of conversations and office content, such as bank counter services, legal consultations, etc.; the intent recognition model trained based on a large-scale language model can understand the needs of customers in a complex noise environment and automatically guide the subsequent business process, which is conducive to improving service efficiency and service quality; by jointly training the environmental perception algorithm and the intent recognition model, the intent resolution of the intent recognition model in complex semantic scenarios is improved. Through the deep coordination of multi-modal data interaction, seamless fusion of multi-modal signals such as voice, touch, and face is supported, and the interaction response speed of personalized services such as data interaction is increased. In summary, the integrated control center of the present application that integrates environmental perception optimization, intent recognition model, and multi-modal processing capabilities is conducive to improving the office efficiency of multi-functional all-in-one products and enhancing the user experience.

[0007] In a preferred example of the present application: The multi-modal interaction module includes: A directional recording sub-module integrating a high-precision microphone array and a deep learning noise reduction algorithm; A multi-touch screen supporting touch operations and gesture recognition; An identity authentication unit configured with a biometric sensor.

[0008] By adopting the above technical solutions, the directional recording sub-module can effectively capture and process sound signals in the environment by using a high-precision microphone array and a deep learning noise reduction algorithm, and can accurately record the user's voice even in a noisy environment, improving the recording quality; the multi-touch screen provides an intuitive and convenient human-computer interaction method, which is conducive to enhancing the user experience; the identity authentication unit can ensure the security of the system through biometric technologies (such as fingerprint, face recognition, etc.) for identity verification.

[0009] In a preferred example of the present application: The business processing module includes: An intent recognition engine provided with an adaptive learning algorithm and an environment adaptive algorithm, supporting natural language understanding and multi-round conversations, dynamically updating the intent recognition database based on the adaptive learning algorithm and historical interaction data, and adjusting the signal processing strategies for environmental noise signals and interactive voice signals in combination with the environment adaptive algorithm and the obtained environmental noise detection data; A modular service interface for loading extended functions; A terminal-cloud collaborative architecture for realizing real-time collaboration between local processed data and cloud big data.

[0010] By adopting the above technical solution, the ability to dynamically update the intent recognition database using the adaptive learning algorithm and historical interaction data enables the system to continuously optimize its understanding ability according to the user's habits. Combining with the environment adaptive algorithm to adjust the processing strategy of voice signals in different environments further improves the stability and response speed of the system in various environments; the modular service interface is used to load extended functions to adapt to diverse functional and service response requirements.

[0011] In the second aspect, the invention object of the present application is realized by adopting the following technical solution: A data processing method applied to a multi-functional integrated service terminal, the method comprising: Determine the cooperation relationship between the multi-modal interaction module and the service processing module, and determine the mapping relationship from the service processing module to different service scenarios; Based on the multi-modal fusion framework corresponding to the cooperation relationship and the intelligent decision-making engine corresponding to the mapping relationship, construct an integrated control center for controlling the multi-functional integrated service terminal, wherein the multi-modal fusion framework includes an environment perception algorithm, and the intelligent decision-making engine includes an intent recognition model; Obtain a training data set, the training data set including multi-source heterogeneous signals collected in a complex noise environment; Use the training data set to train the integrated control center, and jointly optimize the environment perception algorithm and the intent recognition model, so that the integrated control center can adapt to complex operating environments and accurately recognize user intents; Input the multi-modal input signal to be processed into the integrated control center to obtain a directional recording file, a service handling instruction, and an office service response.

[0012] By adopting the above technical solution, the integrated control center constructed using the multi-modal fusion framework and the intelligent decision-making engine realizes seamless connection from data collection to service provision. It can not only efficiently process multi-source heterogeneous data, but also make accurate service responses; by training with a data set containing a complex noise environment, the system can better cope with real-world challenges in actual applications, which is beneficial to improving the complex scenario adaptation ability and service reliability of the multi-functional integrated service terminal.

[0013] In a preferred example of the present application: the training of the integrated control center using the training data set includes: Based on the training data set, determine the training multi-source heterogeneous signals, environmental perception results, and intention recognition results; the multi-source heterogeneous signals and the environmental perception results satisfy the data collaboration relationship from diverse original data sources to multi-functional service requirements, and the environmental perception results and the intention recognition results satisfy the data mapping relationship from the quality assessment of the original data source to the feedback of multi-functional services; Input the multi-source heterogeneous signals into the integrated control center. The multi-modal fusion framework outputs test environmental perception results based on the multi-source heterogeneous signals, and the intelligent decision-making engine outputs test intention recognition results based on the test environmental perception results; Based on the environmental perception results and the test environmental perception results, determine the environmental perception loss function of the environmental perception algorithm; Based on the intention recognition results and the test intention recognition results, determine the intention recognition loss function of the intention recognition model; Based on the environmental perception loss function and the intention recognition loss function, determine the total loss function; Optimize the total loss function to train and optimize the integrated control center.

[0014] By adopting the above technical solution, using a training data set containing multi-source heterogeneous signals collected in a complex noise environment, learn how to extract valuable information from diverse original data sources and convert it into data that meets multi-functional service requirements; by inputting the multi-source heterogeneous signals into the integrated control center for testing, verify the performance of the integrated control center in actual operation, that is, whether it can accurately generate corresponding environmental perception results and intention recognition results according to the input multi-source heterogeneous signals, which helps to evaluate the performance of the system, and by calculating the respective loss functions of the environmental perception algorithm and the intention recognition model and combining them into a total loss function to quantify the error level in the current state of the system, and then adjust the model parameters by optimizing the total loss function, thereby gradually reducing the error and improving the accuracy and efficiency of the system.

[0015] In a preferred example of the present application: the intelligent decision-making engine outputs test intention recognition results based on the test environmental perception results, including: Obtain the signal features of the test environmental perception results and perform preprocessing, input the preprocessed signal features into the intention recognition model of the intelligent decision-making engine, identify the specific needs and service preferences of the user, and obtain the initial user intention result; Based on the initial user intention result, combined with the preset business logic and service process, generate the test intention recognition result.

[0016] By adopting the above technical solution, after preprocessing the test environment perception results, an intelligent decision-making engine is used to identify the specific needs and service preferences of users, so as to realize personalized services and improve the system's ability to understand user intentions. This application not only considers the direct needs of users, but also combines business logic and service processes to ensure that the provided services meet both user expectations and established operating specifications, which helps to enhance the user experience while also ensuring the consistency and reliability of the services.

[0017] In a preferred example of this application: The method for judging whether the integrated control center training is completed includes: Input the multi-source heterogeneous signals into the integrated control center, and output the verified environment perception results based on the multi-modal fusion framework; Perform error calculation based on the environment perception results and the verified environment perception results to determine the environment perception error; perform numerical comparison based on the environment perception error and a preset environment perception threshold, and judge whether the training of the environment perception algorithm is completed based on the comparison result; and / or, Perform error calculation based on the intention recognition results and the test intention recognition results to determine the intention recognition error; Perform numerical comparison based on the intention recognition error and a preset intention recognition threshold, and judge whether the training of the intention recognition model is completed based on the comparison result.

[0018] By adopting the above technical solution, this application can scientifically evaluate the state of system training by comparing the environment perception error with the preset environment perception threshold and the intention recognition error with the preset intention recognition threshold. If the error is lower than the set threshold, it is considered that the training has achieved the expected goal; otherwise, continue to adjust until the conditions are met. At the same time, by independently evaluating the environment perception and intention recognition respectively, it is possible to more carefully understand the performance of each part and facilitate targeted optimization.

[0019] In a preferred example of this application: The steps for obtaining the training data set include: Obtain an original data set containing the original multi-source heterogeneous signals collected in a complex noise environment; Perform data preprocessing on the original data set to obtain the training multi-source heterogeneous signals; Based on the training multi-source heterogeneous signals, use a simulation environment to generate several environment perception results under different environmental conditions as the training environment perception results; According to business requirements and service scenarios, define corresponding service feedbacks and user intentions, and construct intention recognition results; Combine the training multi-source heterogeneous signals, the training environment perception results and the intention recognition results to construct a complete training data set.

[0020] By adopting the above technical solutions, this application expands the diversity of the dataset through a simulation environment generation algorithm, such as scenarios like financial counter whistling and hospital equipment beeping. By generating training data of various noise types, it is beneficial to improve the noise intensity coverage range of the training dataset and the spatio-temporal alignment accuracy of multi-source signals. For example, through dynamic noise injection and signal enhancement techniques, the data quality is improved. For instance, the voice signal is denoised by Wave-U-Net, and the SNR in a 50dB noise environment is increased from 12dB to 28dB; the touch signal generates 200% touch trajectory variation samples based on GAN to cover user behavior deviations; the biometric signal simulates hand movements through an elastic deformation algorithm, reducing the false recognition rate of fingerprint recognition by 37%.

[0021] In a preferred example of this application: generating a test intent recognition result based on the initial user intent result, in combination with a preset business logic and service process, includes: Analyze the key information in the initial user intent result and match it with the preset service process; According to the matching result, call the relevant service module or service instruction set to form a specific business handling instruction or office service response; Compare the formed business handling instruction or office service response with the initial user intent result; Output the final test intent recognition result based on the comparison information.

[0022] By adopting the above technical solutions, the business scenario adaptability of the test intent recognition result is improved. Through a dynamic service process matching mechanism, by analyzing the initial user intent result and matching the preset service process, the system can more efficiently call the relevant service module or service instruction set to form a specific business handling instruction or office service response, which can not only speed up the service response speed but also improve the accuracy and satisfaction of the service.

[0023] In a preferred example of this application: optimizing the total loss function includes: optimizing the total loss function through an Adam optimizer.

[0024] By adopting the above technical solutions, optimizing the total loss function through an Adam optimizer can effectively improve the model training efficiency and convergence speed.

[0025] In summary, this application includes at least one of the following beneficial technical effects: 1. By determining the collaboration relationship between the multimodal interaction module and the service processing module and establishing a mapping relationship to different scenario service outputs, more efficient and accurate service responses can be achieved; the environmental perception algorithm included in the multimodal fusion framework can process interaction data collected from a complex noise environment, including various types of multi-source heterogeneous signals; the intent recognition model in the intelligent decision-making engine is optimized by combining multi-source heterogeneous signals in the training dataset, and can more accurately understand and predict the needs and intents of users. 2. Using the simulation environment to generate perception results under different environmental conditions as part of the training helps improve the system's perception and adaptation capabilities to different working environments, and is beneficial to providing stable services in unpredictable or ever-changing application scenarios. 3. Defining service feedback and user intents according to business requirements and service scenarios, and constructing a training dataset by combining multi-source heterogeneous signals and environmental perception results can significantly improve the accuracy of intent recognition. Description of the Drawings

[0026] Figure 1 is a framework diagram of a multifunctional integrated service terminal in an embodiment of the present application; Figure 2 is a flowchart of a data processing method applied to a multifunctional integrated service terminal in an embodiment of the present application. Detailed Embodiments

[0027] The present application will be further described in detail below with reference to the accompanying drawings.

[0028] In one embodiment, as Figure 1 shown, the present application discloses a multifunctional integrated service terminal, which includes a multimodal interaction module, a service processing module, and an integrated control center; determine the collaboration relationship between the multimodal interaction module and the service processing module, and determine the mapping relationship of the service processing module to different scenario service outputs; based on the multimodal fusion framework corresponding to the collaboration relationship and the intelligent decision-making engine corresponding to the mapping relationship, construct an integrated control center, where the multimodal fusion framework includes an environmental perception algorithm, and the intelligent decision-making engine includes an intent recognition model; obtain a training dataset containing interaction data in a complex noise environment, and the training dataset includes several types of multi-source heterogeneous signals; train the integrated control center based on the training dataset, and jointly optimize the environmental perception algorithm and the intent recognition model during the training process; input the multimodal input signal to be served into the trained integrated control center, and output a directional recording file, a service handling instruction, and an office service response.

[0029] Specifically, the multi-modal interaction module includes a directional recording sub-module integrating a high-precision microphone array and a deep learning noise reduction algorithm, a multi-touch screen supporting touch operation and gesture recognition, and an identity authentication unit configured with a biometric sensor; the high-precision microphone array includes a six-microphone circular array (spacing 8 cm), a MEMS digital microphone (signal-to-noise ratio 65 dB), and a prefabricated DSP chip (TIC6748) to achieve real-time reverberation suppression; the identity authentication unit includes a multi-modal biometric authentication method of fingerprint + vein dual-factor authentication.

[0030] The service processing module includes an intention recognition engine, a modular service interface, and an edge-cloud collaboration architecture. The intention recognition engine is provided with an adaptive learning algorithm and an environment adaptation algorithm, supports natural language understanding and multi-round conversations, dynamically updates the intention recognition database based on the adaptive learning algorithm and historical interaction data, and adjusts the signal processing strategies for environmental noise signals and interactive voice signals in combination with the environment adaptation algorithm and the obtained environmental noise detection data; the modular service interface is used to load extended functions, and the modular service interface allows flexible addition or removal of specific service functions, improving the scalability of the system; the edge-cloud collaboration architecture is used to achieve real-time collaboration between local processed data and cloud big data.

[0031] Each module in the above-mentioned multi-functional integrated service terminal can be implemented in whole or in part by software, hardware, and their combinations; each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0032] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the above-mentioned multi-functional integrated service terminal is divided into different functional units or modules to complete all or part of the functions described above.

[0033] In one embodiment, as Figure 2 shown, a data processing method applied to a multi-functional integrated service terminal is provided. The data processing method applied to a multi-functional integrated service terminal is applied to a multi-functional integrated service terminal. The data processing method applied to a multi-functional integrated service terminal specifically includes the following steps: S1: Determine the collaboration relationship between the multi-modal interaction module and the service processing module, and determine the mapping relationship of the service processing module to the outputs of different service scenarios.

[0034] In this embodiment, the determination of the collaboration mode includes clarifying the data interfaces between multiple modality interaction modules, such as how voice input triggers the response of the service processing module, and defining a unified standardized data interaction format; the determination of the mapping relationship of the service scenarios includes defining the output formats (such as the JSON structure of service instructions, the storage path of recording files) for different service scenarios (such as business handling, office services), and establishing the association rules between the environmental perception results and the service scenarios; by defining the collaboration relationship and the mapping relationship, the types and formats of information that need to be transmitted between the multi-modal interaction module and the service processing module are clarified, as well as the expected service output modes under different service scenarios.

[0035] Specifically, a complete data stream from the original signal quality to the service feedback is established, including: Input layer: Multi-source heterogeneous signals (including noise pollution); Processing layer: Environmental perception (signal quality assessment) → Intent recognition (service instruction understanding); Output layer: Structured service response (such as financial transfer instructions, medical registration operations); This application enables the environmental perception results to directly affect the decision weights of the intent recognition model. For example, in a noisy environment, the recognition confidence of voice signals is automatically enhanced.

[0036] Exemplarily, the high-precision microphone array adopts an 8-microphone linear array, supports beamforming technology, and the directional accuracy is ±5°; the multi-touch screen supports 10-point touch, the sampling rate ≥120Hz, and the resolution ≥4K; the biometric sensor includes a fingerprint recognition module (false recognition rate <0.001%) and an infrared camera (for face recognition); the intent recognition engine is deployed on an edge computing device (such as NVIDIA Jetson) and supports real-time inference; the end-cloud collaboration architecture can process critical instructions locally (such as emergency calls) and complex tasks in the cloud (such as big data analysis).

[0037] S2: Based on the multi-modal fusion framework corresponding to the collaboration relationship and the intelligent decision-making engine corresponding to the mapping relationship, an integrated control center for controlling the multi-functional integrated service terminal is constructed, where the multi-modal fusion framework includes an environmental perception algorithm, and the intelligent decision-making engine includes an intent recognition model.

[0038] In this embodiment, the intelligent decision-making engine is a deep learning model; the multi-modal fusion framework is responsible for integrating heterogeneous signals (such as audio, video, touch data, etc.) input by the multi-modal interaction module, preprocessing and feature extraction of the signals through environmental perception algorithms, and finally outputting a unified environmental perception result (such as noise level, user location, operation intention, etc.). Multi-source data fusion refers to integrating features of different modalities (such as speech feature vectors, touch coordinates) into a unified environmental perception result through a feature fusion layer (such as an attention mechanism); the intention recognition model understands the user's voice commands through natural language processing (NLP), or recognizes the user's operation intention through a touch trajectory analysis model (such as LSTM).

[0039] Specifically, the environmental perception algorithm uses a deep learning noise reduction model (such as a DNN-based denoiser), inputs noisy audio, and outputs clean speech; the sensor fusion algorithm fuses the direction of the microphone array and the accelerometer data through Kalman filtering to locate the user's position.

[0040] S3: Obtain a training data set, which contains multi-source heterogeneous signals collected in a complex noise environment.

[0041] In this embodiment, the training data set contains multi-source heterogeneous signals collected in a complex noise environment (such as high background noise, multi-user interference), such as speech (in wav format), touch (pressure / position data), biometric features (fingerprint / vein images), environmental sensors (noise decibel values), image data, and other types of data sources in a noisy environment, with different sampling rates, data formats, and physical dimensions.

[0042] S4: Use the training data set to train the integrated control center, and jointly optimize the environmental perception algorithm and the intention recognition model, so that the integrated control center can adapt to a complex operation environment and accurately recognize the user's intention.

[0043] In this embodiment, training the integrated control center using the training data set includes: S41: Based on the training data set, determine the training multi-source heterogeneous signals, environmental perception results, and intention recognition results; the multi-source heterogeneous signals and the environmental perception results satisfy the data coordination relationship from diverse original data sources to multi-functional service requirements, and the environmental perception results and the intention recognition results satisfy the data mapping relationship from the quality assessment of the original data source to the multi-functional service feedback.

[0044] Specifically, multi-modal data collection includes environmental voice signals collected through the voice module, touch pressure values collected through the touch module (capacitive touch screen), biometric signals collected through the biometric module, and environmental monitoring signals collected through the environmental monitoring module; then, each data in the training dataset is labeled with relationships, and each data standard includes the following dimensions: (1) Modal combination, such as "voice + touch + biometrics"; (2) Business intent, such as ["cross-border transfer", "medical insurance registration", "provident fund query"] (3) Environmental status: including noise type, SNR value, and modal availability, such as ["voice interference", "mechanical vibration", "electromagnetic noise"], modal availability: {"voice": 1, "touch": 0, "biometrics": 1} (4) Timestamp, accurate to the millisecond level.

[0045] Through the time series matching algorithm, align the end moment of the voice command with the start moment of the touch operation to complete signal alignment; and splice the 13-dimensional voice MFCC coefficients (Mel Frequency Cepstrum Coefficient) with the touch pressure curve (time series data), and input them into the graph neural network to learn the dependencies between modalities to complete feature-level fusion. Among them, the graph neural network training uses GraphConvolutionalNetwork (GCN) to construct a modality relationship graph, where the nodes represent each modality data, and the edge weights represent the collaboration intensity, to learn the complementary relationships between modalities in the input data and obtain the feature-level collaboration coefficients; and construct a rule engine to define conflict scenario handling strategies (such as when the voice command conflicts with the touch operation, biometric verification is preferred) to achieve instruction decision-level association.

[0046] In this embodiment, voice signal processing is performed at the software layer of the multi-modal fusion framework, specifically including: Voice processing, perform real-time Wave-U-Net noise reduction, and output clean voice signals and noise suppression coefficients; Touch processing, extract sliding trajectory features (such as acceleration, curvature) through a CNN1D network to generate touch behavior codes; Biometric processing, after preprocessing the fingerprint image through OpenCV (graying, binarization), input it into a lightweight CNN model to extract feature vectors; Next, reference values for automated tool standard environment perception are adopted, such as noise level, user location, and operation intention, and corresponding service scenario outputs are labeled according to the multi-modal signals input by the user; the Transformer architecture is used to fuse multi-modal features, a cross-attention mechanism is set up to focus on key modalities, and data collaboration relationships and mapping relationships are established. The data collaboration relationship includes a logical association between multi-source heterogeneous signals and the results of environment perception. For example, the voice signal in a noisy environment needs to correspond to the "high-noise environment" label. The data mapping relationship includes defining the mapping rules from the results of environment perception to service feedback. For example, if the result of environment perception is "user authentication passed", it is mapped to "allow access to sensitive services".

[0047] S42: Input the multi-source heterogeneous signals into the integrated control center. The multi-modal fusion framework outputs the test environment perception results based on the multi-source heterogeneous signals, and the intelligent decision-making engine outputs the test intention recognition results based on the test environment perception results.

[0048] In this embodiment, the test environment perception results include noise level, user intention category, and operation priority; the test intention recognition results include structured service instructions and / or service responses.

[0049] Specifically, the content executed by the environment perception algorithm includes: Audio processing, using a deep learning noise reduction model (such as DNN-based denoiser) to remove background noise and extract clear speech features; Touch analysis, analyzing the touch trajectory through an LSTM network to identify the operation intention (such as "long press" corresponding to "confirm"); Sensor fusion, combining the direction of the microphone array and the accelerometer data to locate the user's position (such as "the user is located on the left side of the device").

[0050] S43: Based on the results of environment perception and the test environment perception results, determine the environment perception loss function of the environment perception algorithm.

[0051] Specifically, the environment perception loss function selects the mean square error or cross-entropy loss function based on the type of the results of environment perception.

[0052] When the type of the results of environment perception is the first environment perception result based on continuous values, the mean square error function is used for calculation, and the environment perception loss function The calculation formula is: , where N is the number of samples, referring to the total number of samples of the results of environment perception in the training dataset; is the true result of environment perception, is the predicted result of environment perception.

[0053] When the type of the environmental perception result is the second environmental perception result of a classification task, it is calculated using the cross-entropy loss function, and the formula is: , where is the true label. (For example, when the environmental perception result is "high noise", the label corresponding to category c is 1, and the others are 0); is the probability distribution predicted by the model (the prediction probability of the multi-modal fusion framework for category c); C is the total number of categories (such as the number of classifications of the environmental perception result, such as "low noise", "medium noise", "high noise").

[0054] S44: Based on the intention recognition result and the test intention recognition result, determine the intention recognition loss function of the intention recognition model.

[0055] Specifically, the intention recognition loss function is the cross-entropy loss function , and the calculation formula is: Among them; M is the total number of intention categories, is the output vector of the intention recognition model for the j-th sample, in the form of (p1, p2,..., pK), representing the probabilities belonging to K intention categories; Convert the output vector of the intention recognition model into a probability distribution.

[0056] In this embodiment, the environmental perception loss function and the intention recognition loss function satisfy the following model parameter consistency constraint function: Among them, is the set of all weight parameters of the environmental perception model; is the environmental perception weight value of the k parameter; is the set of all weight parameters of the intention recognition model; is the intention recognition weight value of the k parameter; K is the shared parameter index, including the shared parameters in the environmental perception loss function and the intention recognition loss function; is the L2 norm.

[0057] S45: Based on the environmental perception loss function and the intention recognition loss function, determine the total loss function.

[0058] Specifically, the total loss function is a weighted combination of the environmental perception and intention recognition losses; in this application, the parameter consistency constraint term (parameter consistency constraint function) is used to force the environmental perception module and the intention recognition module to share key feature representations, solving the problem of "recognition errors caused by perception bias" in the traditional solution, so as to design a loss function system that simultaneously optimizes the environmental perception accuracy and the intention recognition accuracy.

[0059] S46: Optimize the total loss function to train and optimize the integrated control center.

[0060] Specifically, it combines a training method with dynamic weight adjustment: for example, at the initial stage of training, it focuses on the environmental perception loss, and gradually increases the weight of the intent recognition loss in the later stage. It uses the AdamW optimizer and sets a learning rate decay strategy (1e-4 → 5e-5); the specific training optimization process includes forward propagation, backward propagation, and iterative training. Among them, forward propagation includes inputting multi-source heterogeneous signals and calculating the test environmental perception result and the test intent recognition result; backward propagation is to calculate the gradient according to the total loss function and update the parameters of the multi-modal fusion framework and the intelligent decision-making engine; iterative training is to repeat the above process until convergence (such as training for 100 rounds, or the loss of the validation set no longer decreases).

[0061] S5: Input the multi-modal input signal to be processed into the integrated control center to obtain a directional recording file, a service handling instruction, and an office service response.

[0062] In this embodiment, the directional recording file is, for example, a noise-reduced voice instruction; the service handling instruction is, for example, "submit a reimbursement application"; the office service response is, for example, automatically sending an email confirmation.

[0063] Specifically, the method for determining whether the training of the integrated control center is completed includes: S51: Input the multi-source heterogeneous signal into the integrated control center and output the verification environmental perception result based on the multi-modal fusion framework.

[0064] Specifically, the environmental perception result includes the noise type and the signal quality score.

[0065] S52: Calculate the error based on the environmental perception result and the verification environmental perception result to determine the environmental perception error; compare the environmental perception error with the preset environmental perception threshold, and judge whether the training of the environmental perception algorithm is completed based on the comparison result.

[0066] Specifically, the environmental perception threshold is a preset fault tolerance range. For example, the noise type classification error rate threshold is ≤5%, the SNR prediction error threshold is ±3dB, and the modal availability judgment error rate threshold is ≤2%.

[0067] And / or S53: Calculate the error based on the intent recognition result and the test intent recognition result to determine the intent recognition error.

[0068] S54: Compare the intent recognition error with the preset intent recognition threshold, and judge whether the training of the intent recognition model is completed based on the comparison result.

[0069] Specifically, the intent recognition threshold includes an intent classification accuracy threshold (such as ≥95%) and an intent confidence threshold (such as ≥0.85), and only the intent with the highest confidence is considered.

[0070] In one embodiment, in step S42, the intelligent decision-making engine outputs a test intention recognition result based on the test environment perception result, including: S421: Obtain the signal features of the test environment perception result and perform preprocessing. Input the preprocessed signal features into the intention recognition model of the intelligent decision-making engine to identify the specific needs and service preferences of the user, and obtain the initial user intention result.

[0071] In this embodiment, the signal preprocessing includes capturing the voice frequency domain features of the voice signal and improving the voice signal-to-noise ratio, performing wavelet decomposition and morphological filtering on the touch signal, and performing normalization processing, binary processing, etc. on the biometric signal.

[0072] Specifically, input the preprocessed multi-modal features into the intention recognition model, and focus on the key time slices through the self-attention mechanism (such as the verb position in the voice command), and generate the probability distribution information containing the user intention. For example, in the office scenario of a bank branch, generate the probability distribution information containing intentions such as "transfer" and "query" as the initial user intention result.

[0073] S422: Generate a test intention recognition result based on the initial user intention result, combined with the preset business logic and service process.

[0074] In this embodiment, the initial user intention result is fused with the business logic rules. For example, trigger the business logic engine. If the initial user intention result is: "transfer" AND the environmental noise type = "voice interference", then biometric verification is required, and the intention confidence of the current initial user intention recognition result is improved. Then, call the corresponding financial domain knowledge graph to verify the intention legality (such as "transfer" requires sufficient account balance) and associate the relevant service process (such as transfer needs to synchronously call the identity authentication and amount verification module). This application dynamically adjusts the model parameters in real time to keep the intention recognition accuracy above 92% in a 65dB noise environment, and solves the problem of the increase in the intention recognition error rate caused by the environmental perception deviation through parameter consistency constraints.

[0075] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0076] The foregoing embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A multi-functional integrated service terminal, characterized in that, Including: Determine the collaborative relationship between the multimodal interaction module and the service processing module, and determine the mapping relationship from the service processing module to the outputs of different scenario services; Based on the multimodal fusion framework corresponding to the collaborative relationship and the intelligent decision-making engine corresponding to the mapping relationship, construct an integrated control center, where the multimodal fusion framework includes an environment perception algorithm, and the intelligent decision-making engine includes an intention recognition model; Obtain a training data set containing interaction data in a complex noise environment, where the training data set includes several types of multi-source heterogeneous signals; Based on the training data set, train the integrated control center, and jointly optimize the environment perception algorithm and the intention recognition model during the training process; Input the multimodal input signal to be served into the trained integrated control center, and output a directional recording file, a service handling instruction, and an office service response.

2. The multifunctional integrated service terminal according to claim 1, characterized in that, The multimodal interaction module includes: A directional recording sub-module integrating a high-precision microphone array and a deep learning noise reduction algorithm; A multi-touch screen supporting touch operations and gesture recognition; An identity authentication unit configured with a biometric sensor.

3. The multifunctional integrated service terminal according to claim 1, characterized in that, The service processing module includes: An intention recognition engine provided with an adaptive learning algorithm and an environment adaptive algorithm, supporting natural language understanding and multi-round conversations, dynamically updating the intention recognition database based on the adaptive learning algorithm and historical interaction data, and combining the environment adaptive algorithm and the obtained environment noise detection data to adjust the signal processing strategies for environmental noise signals and interaction voice signals; A modular service interface for loading extended functions; A terminal-cloud collaborative architecture to achieve real-time collaboration between local processed data and cloud big data.

4. A data processing method applied to a multi-functional integrated service terminal, characterized in that, The method includes: Determine the collaborative relationship between the multimodal interaction module and the service processing module, and determine the mapping relationship from the service processing module to the outputs of different service scenarios; Based on the multimodal fusion framework corresponding to the collaborative relationship and the intelligent decision-making engine corresponding to the mapping relationship, construct an integrated control center for controlling the multifunctional integrated service terminal, where the multimodal fusion framework contains an environment perception algorithm, and the intelligent decision-making engine contains an intention recognition model; Obtain a training data set, where the training data set contains multi-source heterogeneous signals collected in a complex noise environment; Use the training data set to train the integrated control center, and jointly optimize the environment perception algorithm and the intention recognition model, so that the integrated control center can adapt to a complex operation environment and accurately recognize user intentions; Input the multimodal input signal to be processed into the integrated control center to obtain a directional recording file, a service handling instruction, and an office service response.

5. A data processing method applied to a multi-functional integrated service terminal according to claim 4, characterized in that The training the integrated control center using the training data set includes: Based on the training data set, determine the training multi-source heterogeneous signals, environmental perception results, and intention recognition results; the multi-source heterogeneous signals and the environmental perception results satisfy the data collaboration relationship from diverse original data sources to multi-functional service requirements, and the environmental perception results and the intention recognition results satisfy the data mapping relationship from the quality assessment of the original data source to the feedback of multi-functional services; Input the multi-source heterogeneous signals into the integrated control center. The multi-modal fusion framework outputs test environmental perception results based on the multi-source heterogeneous signals, and the intelligent decision-making engine outputs test intention recognition results based on the test environmental perception results; Based on the environmental perception results and the test environmental perception results, determine the environmental perception loss function of the environmental perception algorithm; Based on the intention recognition results and the test intention recognition results, determine the intention recognition loss function of the intention recognition model; Based on the environmental perception loss function and the intention recognition loss function, determine the total loss function; Optimize the total loss function to train and optimize the integrated control center.

6. A data processing method applied to a multi-functional integrated service terminal according to claim 5, characterized in that, The intelligent decision-making engine outputs test intention recognition results based on the test environmental perception results, including: Obtain the signal features of the test environmental perception results and perform preprocessing. Input the preprocessed signal features into the intention recognition model of the intelligent decision-making engine to identify the specific needs and service preferences of the user, and obtain the initial user intention result; Based on the initial user intention result, combined with the preset business logic and service process, generate the test intention recognition result.

7. A data processing method applied to a multi-functional integrated service terminal according to claim 6, characterized in that, The method for judging whether the training of the integrated control center is completed includes: Input the multi-source heterogeneous signals into the integrated control center, and output the verification environmental perception results based on the multi-modal fusion framework; Perform error calculation based on the environmental perception results and the verification environmental perception results to determine the environmental perception error; perform numerical comparison based on the environmental perception error and the preset environmental perception threshold, and judge whether the training of the environmental perception algorithm is completed based on the comparison result; and / or, Perform error calculation based on the intention recognition results and the test intention recognition results to determine the intention recognition error; Perform numerical comparison based on the intention recognition error and the preset intention recognition threshold, and judge whether the training of the intention recognition model is completed based on the comparison result.

8. The data processing method applied to a multi-functional integrated service terminal according to claim 5, characterized in that The step of obtaining the training data set includes: Obtain the original data set containing the original multi-source heterogeneous signals collected in a complex noise environment; Perform data preprocessing on the original data set to obtain the training multi-source heterogeneous signals; Based on the training multi-source heterogeneous signals, use the simulation environment to generate several environmental perception results under different environmental conditions as the training environmental perception results; According to the business requirements and service scenarios, define the corresponding service feedback and user intentions, and construct the intention recognition results; Combine the training multi-source heterogeneous signals, the training environmental perception results, and the intention recognition results to construct a complete training data set.

9. The data processing method applied to a multi-functional integrated service terminal according to claim 6, characterized in that, The generating the test intention recognition result based on the initial user intention result, combined with the preset business logic and service process, includes: Analyze the key information in the initial user intention result and match it with the preset service process; According to the matching result, call the relevant service module or service instruction set to form specific business handling instructions or office service responses; Compare the formed business handling instructions or office service responses with the initial user intention result; Output the final test intention recognition result based on the comparison information.

10. A data processing method applied to a multi-functional integrated service terminal according to claim 5, characterized in that, The optimization of the total loss function includes: optimizing the total loss function through the Adam optimizer.

Citation Information

Patent Citations

  • Human-computer interaction method, device and equipment and storage medium

    CN116737883A

  • Neural rehabilitation training device

    CN117809812A

  • Dynamic data adaptive desensitization method and device based on artificial intelligence

    CN119128990A

  • Multi-mode synchronous fusion speech recognition system

    CN119296523A

  • Method for training end-to-end automatic driving strategy

    CN119622456A

Cited By

  • Smart phone and method

    CN120856822A

  • Method and system for optimizing anti-interference performance of touch screen

    CN121092009A

  • Desktop type intelligent digital assistant terminal

    CN121900861A