A multi-functional integrated service terminal and its application data processing method
By combining a multimodal fusion framework and an intelligent decision engine, the problem of inaccurate intent recognition in traditional multifunction printers under complex and noisy environments is solved, enabling efficient office work and personalized services in noisy environments.
Patent Information
- Application Number
- CN202510764003.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Traditional multifunction printers struggle to accurately identify user intent in noisy environments, resulting in poor user experience and low work efficiency.
By employing a multimodal fusion framework and intelligent decision engine, combined with environmental perception algorithms and intent recognition models, and optimizing the integrated control center through training datasets, seamless fusion and accurate recognition of multi-source heterogeneous signals are achieved.
It improves the office efficiency and user experience of multifunction printers in complex and noisy environments, and supports efficient personalized service response through deep collaboration of multimodal data interaction.
Smart Images

Figure CN120277620B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent office equipment technology, and in particular to a data processing method for a multi-functional integrated service terminal and its application. Background Technology
[0002] With the development of information technology, single-function service devices in traditional office environments are gradually unable to meet the needs of efficient multitasking. In particular, for scenarios that require audio recording, identity authentication, and intelligent assistance, the separate use of individual devices such as computers, document scanners, and voice recorders not only takes up space but also makes it difficult to work together, reducing work efficiency and service experience.
[0003] Currently, there are all-in-one products on the market that integrate multiple functions, such as desktop all-in-ones that integrate cameras and touch screens. However, they usually lack professional audio processing capabilities and intelligent human-computer interaction design. In particular, they are not good at accurate role separation recording and intent recognition in complex noisy environments. They cannot understand customer intent well in noisy environments, resulting in poor user experience and low office efficiency. Therefore, they need to be improved. Summary of the Invention
[0004] To improve the office efficiency and enhance the user experience of multifunction printers, this application provides a data processing method for a multifunction integrated service terminal and its application.
[0005] Firstly, the objective of this invention is achieved through the following technical solution:
[0006] A multi-functional integrated service terminal, comprising:
[0007] Determine the collaborative relationship between the multimodal interaction module and the business processing module, and determine the mapping relationship between the business processing module and the service output of different scenarios;
[0008] Based on the multimodal fusion framework corresponding to the aforementioned collaborative relationship and the intelligent decision engine corresponding to the aforementioned mapping relationship, an integrated control center is constructed, wherein the multimodal fusion framework includes an environmental perception algorithm and the intelligent decision engine includes an intent recognition model;
[0009] Obtain a training dataset containing interactive data under complex noise environments, wherein the training dataset includes several types of multi-source heterogeneous signals;
[0010] Based on the training dataset, the integrated control center is trained, and the environmental perception algorithm and the intention recognition model are jointly optimized during the training process.
[0011] The multimodal input signals to be served are input into the integrated control center after training, and the center outputs directional recording files, business processing instructions and office service responses.
[0012] By adopting the above technical solutions, several types of multi-source heterogeneous signals, including voice, touch, and biometrics, cover heterogeneous business and government data involved in office scenarios (such as bank branches). The multi-functional integrated service terminal of this application, based on different business service scenarios and actual business interaction scenarios, sets up a multi-modal fusion framework and intelligent decision engine to improve the clarity and / or recognizability of different types of multi-source heterogeneous signals. It is suitable for scenarios requiring accurate recording of dialogue and office content, such as bank counter services and legal consultations. The intent recognition model, trained based on a large-scale language model, can understand customer needs in complex noisy environments and automatically guide subsequent business processes, which is beneficial to improving service efficiency and quality. Jointly training the environmental perception algorithm and the intent recognition model improves the intent resolution of the intent recognition model in complex semantic scenarios. Through deep collaboration of multi-modal data interaction, it supports seamless fusion of multi-modal signals such as voice, touch, and face, improving the interactive response speed of personalized services such as data interaction. In summary, this application's integrated control center, which integrates environmental perception optimization, intent recognition model, and multi-modal processing capabilities, is beneficial to improving the office efficiency of multi-functional all-in-one products and enhancing user experience.
[0013] In a preferred embodiment, the multimodal interaction module includes:
[0014] A directional recording submodule integrating a high-precision microphone array and a deep learning noise reduction algorithm;
[0015] Multi-touch screen that supports touch operation and gesture recognition;
[0016] An identity authentication unit equipped with a biometric sensor.
[0017] By adopting the above technical solutions, the directional recording submodule can effectively capture and process sound signals in the environment by using a high-precision microphone array and deep learning noise reduction algorithm, and can accurately record the user's voice even in noisy environments, thus improving the recording quality; the multi-touch screen provides an intuitive and convenient human-computer interaction method, which is conducive to improving the user experience; the identity authentication unit uses biometric technology (such as fingerprint, facial recognition, etc.) to verify identity, which can ensure the security of the system.
[0018] In a preferred embodiment of this application, the business processing module includes:
[0019] An intent recognition engine equipped with adaptive learning and environment adaptive algorithms supports natural language understanding and multi-turn dialogue. It dynamically updates the intent recognition database based on the adaptive learning algorithm and historical interaction data, and adjusts the signal processing strategies for environmental noise signals and interactive voice signals by combining the environment adaptive algorithm and the acquired environmental noise detection data.
[0020] Modular service interfaces are used to load extended functionalities;
[0021] The edge-cloud collaborative architecture enables real-time collaboration between locally processed data and cloud-based big data.
[0022] By adopting the above technical solutions and utilizing the ability of adaptive learning algorithms and historical interaction data to dynamically update the intent recognition database, the system can continuously optimize its understanding capabilities based on user habits. Combined with environmental adaptive algorithms, it adjusts the processing strategies for speech signals in different environments, further improving the system's stability and response speed in various environments. Modular service interfaces are used to load extended functions to adapt to diverse functional and service response needs.
[0023] Secondly, the objective of this invention is achieved through the following technical solution:
[0024] A data processing method applied to a multi-functional integrated service terminal, the method comprising:
[0025] Determine the collaborative relationship between the multimodal interaction module and the business processing module, and determine the mapping relationship between the business processing module and the output of different service scenarios;
[0026] Based on the multimodal fusion framework corresponding to the collaborative relationship and the intelligent decision engine corresponding to the mapping relationship, an integrated control center for controlling the multifunctional integrated service terminal is constructed, wherein the multimodal fusion framework includes an environmental perception algorithm and the intelligent decision engine includes an intent recognition model.
[0027] Obtain a training dataset, which contains multi-source heterogeneous signals collected in a complex noise environment;
[0028] The integrated control center is trained using the training dataset, and the environmental perception algorithm and the intent recognition model are jointly optimized, so that the integrated control center can adapt to complex operating environments and accurately recognize user intents.
[0029] The multimodal input signals to be processed are input into the integrated control center to obtain directional recording files, business processing instructions, and office service responses.
[0030] By adopting the above technical solutions, an integrated control center built using a multimodal fusion framework and intelligent decision engine is achieved, realizing a seamless connection from data acquisition to service provision. It can not only efficiently process multi-source heterogeneous data, but also make accurate service responses. By training on datasets containing complex noise environments, the system can better cope with real-world challenges in practical applications, which is conducive to improving the adaptability of multifunctional integrated service terminals to complex scenarios and service reliability.
[0031] In a preferred embodiment, the step of training the integrated control center using the training dataset includes:
[0032] Based on the training dataset, training multi-source heterogeneous signals, environmental perception results, and intent recognition results are determined; the multi-source heterogeneous signals and the environmental perception results satisfy a data collaboration relationship from diverse original data sources to multi-functional service requirements, and the environmental perception results and the intent recognition results satisfy a data mapping relationship from original data source quality assessment to multi-functional service feedback;
[0033] The multi-source heterogeneous signals are input to the integrated control center, the multi-modal fusion framework outputs test environment perception results based on the multi-source heterogeneous signals, and the intelligent decision engine outputs test intent recognition results based on the test environment perception results.
[0034] Based on the environmental perception results and the test environmental perception results, the environmental perception loss function of the environmental perception algorithm is determined.
[0035] Based on the intent recognition results and the test intent recognition results, the intent recognition loss function of the intent recognition model is determined;
[0036] Based on the environmental perception loss function and the intent recognition loss function, the total loss function is determined;
[0037] The total loss function is optimized to train and optimize the integrated control center.
[0038] By adopting the above technical solution and utilizing a training dataset containing multi-source heterogeneous signals collected under complex noise environments, the system learns how to extract valuable information from diverse raw data sources and transform it into data that meets the needs of multi-functional services. Testing with the multi-source heterogeneous signals input into the integrated control center verifies its performance in actual operation, specifically its ability to accurately generate corresponding environmental perception and intent recognition results based on the input multi-source heterogeneous signals. This helps evaluate the system's performance. Furthermore, by calculating the loss functions of the environmental perception algorithm and the intent recognition model and combining them into a total loss function, the error level of the system in its current state is quantified. Then, by optimizing the total loss function, the model parameters are adjusted to gradually reduce errors and improve the system's accuracy and efficiency.
[0039] In a preferred embodiment of this application: the intelligent decision engine outputs a test intent recognition result based on the test environment perception result, including:
[0040] The signal features of the test environment perception results are obtained and preprocessed. The preprocessed signal features are then input into the intent recognition model of the intelligent decision engine to identify the user's specific needs and service preferences, thereby obtaining the initial user intent results.
[0041] Based on the initial user intent result and combined with the preset business logic and service process, a test intent recognition result is generated.
[0042] By adopting the above technical solution, after preprocessing the test environment perception results, the intelligent decision engine is used to identify the user's specific needs and service preferences, so as to realize personalized services and improve the system's ability to understand user intentions. This application not only considers the user's direct needs, but also combines business logic and service processes to ensure that the services provided meet the user's expectations and comply with established operating procedures. This helps to improve the user experience while also ensuring the consistency and reliability of the services.
[0043] In a preferred embodiment, this application provides a method for determining whether the training of the integrated control center is complete, comprising:
[0044] The multi-source heterogeneous signals are input into the integrated control center, and the environmental perception results are output based on the multimodal fusion framework.
[0045] Error calculation is performed based on the environmental perception results and the verified environmental perception results to determine the environmental perception error; a numerical comparison is performed based on the environmental perception error and a preset environmental perception threshold, and the training of the environmental perception algorithm is determined based on the comparison result.
[0046] And / or,
[0047] Based on the intent recognition result and the test intent recognition result, an error calculation is performed to determine the intent recognition error;
[0048] A numerical comparison is made between the intent recognition error and the preset intent recognition threshold, and the training of the intent recognition model is determined based on the comparison result.
[0049] By adopting the above technical solution, this application can scientifically evaluate the training status of the system by comparing the environmental perception error with the preset environmental perception threshold and the intention recognition error with the preset intention recognition threshold. If the error is lower than the set threshold, the training is considered to have achieved the expected goal; otherwise, adjustments are made until the conditions are met. At the same time, by conducting independent evaluations of environmental perception and intention recognition, the performance of each part can be understood in more detail, which is convenient for targeted optimization.
[0050] In a preferred example of this application: the step of obtaining the training dataset includes:
[0051] Obtain the raw dataset containing raw multi-source heterogeneous signals collected under complex noise environments;
[0052] The original dataset is preprocessed to obtain training multi-source heterogeneous signals;
[0053] Based on the training multi-source heterogeneous signals, several environmental perception results under different environmental conditions are generated using the simulated environment as training environmental perception results.
[0054] Based on business needs and service scenarios, define corresponding service feedback and user intent, and construct intent recognition results;
[0055] By combining the training multi-source heterogeneous signals, the training environment perception results, and the intent recognition results, a complete training dataset is constructed.
[0056] By adopting the above technical solutions, this application expands the diversity of the dataset through simulated environment generation algorithms, such as scenarios like financial counter whistling and hospital equipment buzzing. By generating training data with multiple noise types, it helps to improve the noise intensity coverage of the training dataset and the spatiotemporal alignment accuracy of multi-source signals. For example, through dynamic noise injection and signal enhancement techniques, data quality is improved. For instance, speech signals are denoised using Wave-U-Net, increasing the SNR from 12dB to 28dB in a 50dB noise environment; touch signals are generated based on GAN to produce 200% of the touch trajectory variation samples, covering user behavior deviations; and biometric signals are simulated using an elastic deformation algorithm to reduce the fingerprint recognition false recognition rate by 37%.
[0057] In a preferred embodiment of this application, the step of generating a test intent recognition result based on the initial user intent result and in conjunction with preset business logic and service processes includes:
[0058] Analyze the key information in the initial user intent results and match it with the preset service process;
[0059] Based on the matching results, relevant service modules or service instruction sets are invoked to generate specific business processing instructions or office service responses.
[0060] Compare the generated business processing instructions or office service responses with the initial user intent results;
[0061] The final test intent recognition result is output based on the comparison information.
[0062] By adopting the above technical solutions, the business scenario adaptability of the test intent recognition results is improved. Through the dynamic service process matching mechanism, by analyzing the initial user intent results and matching them with the preset service process, the system can more efficiently call relevant service modules or service instruction sets to form specific business processing instructions or office service responses. This not only speeds up service response but also improves service accuracy and satisfaction.
[0063] In a preferred embodiment of this application, optimizing the total loss function includes optimizing the total loss function using the Adam optimizer.
[0064] By adopting the above technical solution and using the Adam optimizer to optimize the total loss function, the training efficiency and convergence speed of the model can be effectively improved.
[0065] In summary, this application includes at least one of the following beneficial technical effects:
[0066] 1. By determining the collaborative relationship between the multimodal interaction module and the business processing module, and establishing a mapping relationship to service outputs in different scenarios, more efficient and accurate service responses can be achieved; the environmental perception algorithm included in the multimodal fusion framework can process interaction data collected from complex noisy environments, including various types of multi-source heterogeneous signals; the intent recognition model in the intelligent decision engine is optimized by combining multi-source heterogeneous signals in the training dataset, enabling it to more accurately understand and predict user needs and intents;
[0067] 2. Using simulated environments to generate perception results under different environmental conditions as part of the training helps improve the system's perception and adaptability to different working environments, which is beneficial for providing stable services in unpredictable or ever-changing application scenarios.
[0068] 3. By defining service feedback and user intent based on business needs and service scenarios, and constructing a training dataset by combining multi-source heterogeneous signals and environmental perception results, the accuracy of intent recognition can be significantly improved. Attached Figure Description
[0069] Figure 1 This is a framework diagram of a multi-functional integrated service terminal according to one embodiment of this application;
[0070] Figure 2 This is a flowchart of a data processing method applied to a multi-functional integrated service terminal according to an embodiment of this application. Detailed Implementation
[0071] The present application will be further described in detail below with reference to the accompanying drawings.
[0072] In one embodiment, such as Figure 1 As shown, this application discloses a multi-functional integrated service terminal, which includes a multimodal interaction module, a business processing module, and an integrated control center. The application determines the collaborative relationship between the multimodal interaction module and the business processing module, and establishes a mapping relationship between the business processing module and service outputs in different scenarios. Based on a multimodal fusion framework corresponding to the collaborative relationship and an intelligent decision engine corresponding to the mapping relationship, an integrated control center is constructed. The multimodal fusion framework includes an environmental perception algorithm, and the intelligent decision engine includes an intent recognition model. A training dataset containing interaction data under complex noise conditions is acquired, including several types of multi-source heterogeneous signals. The integrated control center is trained based on the training dataset, and jointly optimized during the training process using the environmental perception algorithm and the intent recognition model. The multimodal input signals to be served are input into the trained integrated control center, which then outputs directional recording files, business processing instructions, and office service responses.
[0073] Specifically, the multimodal interaction module includes a directional recording submodule integrating a high-precision microphone array and a deep learning noise reduction algorithm, a multi-touch screen supporting touch operation and gesture recognition, and an identity authentication unit equipped with a biometric sensor; the high-precision microphone array includes a six-microphone ring array (8cm spacing), a MEMS digital microphone (65dB signal-to-noise ratio), and a pre-built DSP chip (TIC6748) to achieve real-time reverberation suppression; the identity authentication unit includes a multimodal biometric authentication method of fingerprint + vein dual-factor authentication.
[0074] The business processing module includes an intent recognition engine, modular service interfaces, and an edge-cloud collaborative architecture. The intent recognition engine is equipped with adaptive learning and environment adaptation algorithms, supporting natural language understanding and multi-turn dialogue. It dynamically updates the intent recognition database based on the adaptive learning algorithm and historical interaction data. Combined with the environment adaptation algorithm and acquired environmental noise detection data, it adjusts the signal processing strategies for environmental noise signals and interactive voice signals. The modular service interface is used to load extended functions, allowing for the flexible addition or removal of specific service functions, thus improving the system's scalability. The edge-cloud collaborative architecture enables real-time collaboration between locally processed data and cloud-based big data.
[0075] Each module in the aforementioned multifunctional integrated service terminal can be implemented entirely or partially through software, hardware, or a combination thereof; each module can be embedded in the processor of the computer device in hardware form or independent of it, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0076] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the multifunctional integrated service terminal can be divided into different functional units or modules to complete all or part of the functions described above.
[0077] In one embodiment, such as Figure 2 As shown, a data processing method for a multi-functional integrated service terminal is provided. This method is applied to a multi-functional integrated service terminal and specifically includes the following steps:
[0078] S1: Determine the collaborative relationship between the multimodal interaction module and the business processing module, and determine the mapping relationship between the business processing module and the output of different service scenarios.
[0079] In this embodiment, determining the collaboration method includes clarifying the data interfaces between multiple modal interaction modules, such as how voice input affects the response of the business processing module, and defining a standardized data interaction format; determining the mapping relationship of service scenarios includes defining output formats (such as the JSON structure of business instructions, the storage path of recording files) for different service scenarios (such as business processing, office services), and the association rules between the monitoring environment perception results and service scenarios; by defining collaboration and mapping relationships, the types and formats of information that need to be transmitted between the multimodal interaction modules and the business processing modules, as well as the expected service output modes under different service scenarios, are clarified.
[0080] Specifically, establishing a complete data flow from raw signal quality to service feedback includes:
[0081] Input layer: multi-source heterogeneous signals (including noise contamination);
[0082] Processing layer: Environmental perception (signal quality assessment) → Intent recognition (service instruction understanding);
[0083] Output layer: Structured service responses (such as financial transfer instructions and medical registration operations);
[0084] This application enables environmental perception results to directly influence the decision weights of the intent recognition model, such as automatically enhancing the recognition confidence of speech signals in noisy environments.
[0085] For example, the high-precision microphone array uses an 8-microphone linear array, supports beamforming technology, and has a directional accuracy of ±5°; the multi-touch screen supports 10-point touch, a sampling rate of ≥120Hz, and a resolution of ≥4K; the biometric sensors include a fingerprint recognition module (false recognition rate <0.001%) and an infrared camera (for facial recognition); the intent recognition engine is deployed on an edge computing device (such as NVIDIA Jetson) and supports real-time inference; the edge-cloud collaborative architecture can process key instructions locally (such as emergency calls) and process complex tasks in the cloud (such as big data analysis).
[0086] S2: Based on a multimodal fusion framework corresponding to collaborative relationships and an intelligent decision engine corresponding to mapping relationships, an integrated control hub for controlling multifunctional integrated service terminals is constructed. The multimodal fusion framework includes an environmental perception algorithm, and the intelligent decision engine includes an intent recognition model.
[0087] In this embodiment, the intelligent decision engine is a deep learning model; the multimodal fusion framework is responsible for integrating heterogeneous signals (such as audio, video, touch data, etc.) input from the multimodal interaction module, preprocessing and extracting features from the signals through environmental perception algorithms, and finally outputting a unified environmental perception result (such as noise level, user position, operation intention, etc.). Multi-source data fusion refers to integrating features of different modalities (such as voice feature vectors, touch coordinates) into a unified environmental perception result through a feature fusion layer (such as attention mechanism); the intention recognition model understands the user's voice commands through natural language processing (NLP), or recognizes the user's operation intention through a touch trajectory analysis model (such as LSTM).
[0088] Specifically, the environmental perception algorithm uses a deep learning noise reduction model (such as a DNN-based denoiser) to take noisy audio as input and output clean speech; the sensor fusion algorithm uses Kalman filtering to fuse microphone array orientation and accelerometer data to locate the user's position.
[0089] S3: Obtain the training dataset, which contains multi-source heterogeneous signals collected in a complex noisy environment.
[0090] In this embodiment, the training dataset contains multi-source heterogeneous signals collected in complex noisy environments (such as high background noise and multi-user interference), such as voice (wav format), touch (pressure / position data), biometrics (fingerprint / vein images), environmental sensors (noise decibel values), image data, and other data sources with different sampling rates, data formats, and physical dimensions.
[0091] S4: The integrated control center is trained using the training dataset, and the environmental perception algorithm and intent recognition model are jointly optimized to enable the integrated control center to adapt to complex operating environments and accurately identify user intent.
[0092] In this embodiment, the integrated control center is trained using a training dataset, including:
[0093] S41: Based on the training dataset, determine the training multi-source heterogeneous signals, environmental perception results, and intent recognition results; the multi-source heterogeneous signals and environmental perception results satisfy the data collaboration relationship from diverse original data sources to multi-functional service requirements, and the environmental perception results and intent recognition results satisfy the data mapping relationship from original data source quality assessment to multi-functional service feedback.
[0094] Specifically, multimodal data acquisition includes environmental speech signals acquired through the speech module, touch pressure values acquired through the touch module (capacitive touchscreen), biometric signals collected through the biometric module, and environmental monitoring signals collected through the environmental monitoring module. Next, each data point in the training dataset is labeled with relationships, and each data point's criteria include the following dimensions:
[0095] (1) Modal combination, such as “voice + touch + biometrics”;
[0096] (2) Business intent, such as ["cross-border transfer", "medical insurance registration", "housing provident fund inquiry"].
[0097] (3) Environmental conditions: including noise type, SNR value and modal availability, such as ["voice interference", "mechanical vibration", "electromagnetic noise"], modal availability: {"voice": 1, "touch": 0, "biometric": 1}
[0098] (4) Timestamp, accurate to the millisecond level.
[0099] A time-series matching algorithm is used to align the end time of voice commands with the start time of touch operations to achieve signal alignment. The 13-dimensional voice MFCC coefficients (Mel Frequency Cepstrum Coefficients) are then concatenated with the touch pressure curve (time-series data) and input into a graph neural network to learn inter-modal dependencies, achieving feature-level fusion. The graph neural network training uses a Graph Convolutional Network (GCN) to construct a modal relationship graph, where nodes represent each modality's data and edge weights represent the cooperation strength, learning the complementary relationships between modalities in the input data to obtain feature-level cooperation coefficients. A rule engine is also constructed to define conflict scenario handling strategies (e.g., prioritizing biometric verification when voice commands and touch operations contradict each other) to achieve command decision-level association.
[0100] In this embodiment, speech signal processing is performed at the software layer of the multimodal fusion framework, specifically including:
[0101] Speech processing, real-time Wave-U-Net noise reduction, outputting clean speech signal and noise suppression coefficient;
[0102] Touch processing involves extracting sliding trajectory features (such as acceleration and curvature) through a CNN1D network to generate touch behavior codes.
[0103] Biometric processing: After fingerprint images are preprocessed using OpenCV (grayscale conversion, binarization), they are input into a lightweight CNN model to extract feature vectors.
[0104] Next, automated tools are used to standardize environmental perception reference values, such as noise levels, user location, and operational intent. Based on the multimodal signals input by the user, corresponding service scenario outputs are labeled. A Transformer architecture is used to fuse multimodal features, and a cross-attention mechanism is set to focus on key modalities, establishing data collaboration and mapping relationships. The data collaboration relationship includes a logical association between multi-source heterogeneous signals and environmental perception results. For example, voice signals in noisy environments need to correspond to the "high-noise environment" label. The data mapping relationship includes defining the mapping rules from environmental perception results to service feedback. For example, if the environmental perception result is "user authentication passed," it is mapped to "allowed access to sensitive services."
[0105] S42: Input multi-source heterogeneous signals to the integrated control center. The multi-modal fusion framework outputs test environment perception results based on the multi-source heterogeneous signals. The intelligent decision engine outputs test intent recognition results based on the test environment perception results.
[0106] In this embodiment, the test environment perception results include noise level, user intent category, and operation priority; the test intent recognition results include structured service instructions and / or service responses.
[0107] Specifically, the environmental perception algorithm performs the following tasks:
[0108] Audio processing uses deep learning noise reduction models (such as DNN-based denoiser) to remove background noise and extract clear speech features;
[0109] Touch analysis uses an LSTM network to analyze touch trajectories and identify the intended operation (e.g., "long press" corresponds to "confirm").
[0110] Sensor fusion combines microphone array orientation and accelerometer data to pinpoint the user's location (e.g., "the user is on the left side of the device").
[0111] S43: Based on the environmental perception results and the test environmental perception results, determine the environmental perception loss function of the environmental perception algorithm.
[0112] Specifically, the environmental perception loss function is selected based on the type of environmental perception result, using either mean squared error or cross-entropy loss function.
[0113] When the environmental perception result is a first environmental perception result based on continuous values, the mean square error function is used to calculate the environmental perception loss function. The calculation formula is: Where N is the number of samples, referring to the total number of environmental perception results samples in the training dataset; The results are based on real-world environmental perception. This refers to the predicted environmental perception results.
[0114] When the environmental perception result is a second environmental perception result of a classification task, the cross-entropy loss function is used for calculation, and the formula is: ,in The label is the true label (for example, when the environmental perception result is "high noise", the label for category c is 1, and for others it is 0). C represents the probability distribution predicted by the model (the predicted probability of category c by the multimodal fusion framework); C represents the total number of categories (such as the number of categories in the environmental perception result, such as "low noise", "medium noise" and "high noise").
[0115] S44: Based on the intent recognition results and the test intent recognition results, determine the intent recognition loss function of the intent recognition model.
[0116] Specifically, the intention recognition loss function is the cross-entropy loss function. The calculation formula is:
[0117] Where M represents the total number of intent categories. Let p1 be the output vector of the intent recognition model for the j-th sample, in the form (p1, p2, ..., pK), representing the probability of belonging to one of the K intent categories. Convert the output vector of the intent recognition model into a probability distribution.
[0118] In this embodiment, the environment perception loss function and the intent recognition loss function satisfy the following model parameter consistency constraint function:
[0119] in, This is the set of all weight parameters for the environmental perception model. The environment perception weight value for parameter k; This is the set of all weight parameters for the intent recognition model; K represents the intent recognition weight value for parameter k; K is the shared parameter index, which includes shared parameters from the environment perception loss function and the intent recognition loss function. It is an L2 norm.
[0120] S45: Determine the total loss function based on the environment perception loss function and the intent recognition loss function.
[0121] Specifically, the total loss function is a weighted combination of the environmental perception and intent recognition losses. This application forces the environmental perception module and the intent recognition module to share key feature representations through parameter consistency constraint terms (parameter consistency constraint function), thereby solving the problem of "perception bias leading to recognition error" in traditional schemes, and designing a loss function system that simultaneously optimizes the accuracy of environmental perception and the accuracy of intent recognition.
[0122] S46: Optimize the total loss function to train and optimize the integrated control center.
[0123] Specifically, the training method is adjusted by combining dynamic weights: for example, focusing on environmental perception loss in the early stage of training, and gradually increasing the weight of intent recognition loss in the later stage, using the AdamW optimizer, and setting a learning rate decay strategy (1e-4→5e-5); the specific training optimization process includes forward propagation, back propagation and iterative training, where forward propagation includes inputting multi-source heterogeneous signals, calculating the test environmental perception result and the test intent recognition result; back propagation is to calculate the gradient according to the total loss function and update the parameters of the multimodal fusion framework and intelligent decision engine; iterative training is to repeat the above process until convergence (e.g., training for 100 rounds, or the validation set loss no longer decreases).
[0124] S5: Input the multimodal input signal to be processed into the integrated control center to obtain directional recording files, business processing instructions and office service responses.
[0125] In this embodiment, directional audio recordings include noise-reduced voice commands; business processing commands include "submit reimbursement application"; and office service responses include automatically sent email confirmations.
[0126] Specifically, methods for determining whether the training of the integrated control center is complete include:
[0127] S51: Input multi-source heterogeneous signals into the integrated control center and output the verification environmental perception results based on the multi-modal fusion framework.
[0128] Specifically, the environmental perception results include noise type and signal quality score.
[0129] S52: Calculate the error based on the environmental perception results and the verification environmental perception results to determine the environmental perception error; compare the environmental perception error with the preset environmental perception threshold, and determine whether the training of the environmental perception algorithm is complete based on the comparison results.
[0130] Specifically, the environmental perception threshold is to achieve a set fault tolerance range, such as a noise type classification error rate threshold of ≤5%, an SNR prediction error threshold of ±3dB, and a modal availability judgment error rate threshold of ≤2%.
[0131] And / or,
[0132] S53: Calculate the error based on the intent recognition result and the test intent recognition result to determine the intent recognition error.
[0133] S54: Compare numerical values based on the intent recognition error and the preset intent recognition threshold, and determine whether the training of the intent recognition model is complete based on the comparison results.
[0134] Specifically, the intent recognition threshold includes an intent classification accuracy threshold (e.g., ≥95%) and an intent confidence threshold (e.g., ≥0.85), and only the intent with the highest confidence is considered.
[0135] In one embodiment, in step S42, the intelligent decision engine outputs a test intent recognition result based on the test environment perception result, including:
[0136] S421: Acquire signal features of the test environment perception results and preprocess them. Input the preprocessed signal features into the intent recognition model of the intelligent decision engine to identify the user's specific needs and service preferences, and obtain the initial user intent results.
[0137] In this embodiment, signal preprocessing includes capturing speech domain features and improving speech signal-to-noise ratio for speech signals, performing wavelet decomposition and morphological filtering on touch signals, and normalizing and binarizing biometric signals.
[0138] Specifically, the preprocessed multimodal features are input into the intent recognition model, which uses a self-attention mechanism to focus on key time slices (such as the position of verbs in voice commands) to generate probability distribution information containing user intent. For example, in an office scenario in a bank branch, probability distribution information containing intents such as "transfer" and "query" is generated as the initial user intent result.
[0139] S422: Generate test intent recognition results based on the initial user intent results and in combination with preset business logic and service processes.
[0140] In this embodiment, the initial user intent result is fused with business logic rules. For example, if the business logic engine is triggered and the initial user intent result is "transfer" AND environmental noise type = "voice interference", then biometric verification is mandatory, and the intent confidence of the current initial user intent recognition result is increased. Then, the corresponding financial domain knowledge graph is called to verify the legality of the intent (e.g., "transfer" requires sufficient account balance) and associate related service processes (e.g., transfer requires simultaneous calling of identity authentication and amount verification modules). This application dynamically adjusts the model parameters in real time so that the intent recognition accuracy remains above 92% in a 65dB noise environment. Through parameter consistency constraints, the problem of increased intent recognition error rate caused by environmental perception deviation is solved.
[0141] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0142] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A data processing method applied to a multi-functional integrated service terminal, characterized in that, The method is applied to a multi-functional integrated service terminal, and the method includes: Determine the collaborative relationship between the multimodal interaction module and the business processing module, and determine the mapping relationship between the business processing module and the output of different service scenarios; Based on the multimodal fusion framework corresponding to the collaborative relationship and the intelligent decision engine corresponding to the mapping relationship, an integrated control center for controlling the multifunctional integrated service terminal is constructed, wherein the multimodal fusion framework includes an environmental perception algorithm and the intelligent decision engine includes an intent recognition model. Obtain a training dataset, which contains multi-source heterogeneous signals collected in a complex noise environment; The integrated control center is trained using the training dataset, and the environmental perception algorithm and the intent recognition model are jointly optimized, so that the integrated control center can adapt to complex operating environments and accurately recognize user intents. The multimodal input signal to be processed is input to the integrated control center to obtain directional recording files, business processing instructions and office service responses; The process of training the integrated control center using the training dataset includes: Based on the training dataset, training multi-source heterogeneous signals, environmental perception results, and intent recognition results are determined; the multi-source heterogeneous signals and the environmental perception results satisfy a data collaboration relationship from diverse original data sources to multi-functional service requirements, and the environmental perception results and the intent recognition results satisfy a data mapping relationship from original data source quality assessment to multi-functional service feedback; The multi-source heterogeneous signals are input to the integrated control center, the multi-modal fusion framework outputs test environment perception results based on the multi-source heterogeneous signals, and the intelligent decision engine outputs test intent recognition results based on the test environment perception results. Based on the environmental perception results and the test environmental perception results, the environmental perception loss function of the environmental perception algorithm is determined. Based on the intent recognition results and the test intent recognition results, the intent recognition loss function of the intent recognition model is determined; Based on the environmental perception loss function and the intent recognition loss function, the total loss function is determined; Optimize the total loss function to train and optimize the integrated control center; When the environmental perception result is a first environmental perception result based on continuous values, the mean square error function is used to calculate the environmental perception loss function Loss. Env The calculation formula is: Where N is the number of samples, which refers to the total number of environmental perception results samples in the training dataset; The results are based on real-world environmental perception. For the predicted environmental perception results; When the environmental perception result is a second environmental perception result of a classification task, the cross-entropy loss function is used for calculation, and the formula is: Where t ic For real labels; p ic denoted as , where is the probability distribution predicted by the model; C represents the total number of categories.
2. The data processing method applied to a multi-functional integrated service terminal according to claim 1, characterized in that, The intelligent decision engine outputs test intent recognition results based on the test environment perception results, including: The signal features of the test environment perception results are obtained and preprocessed. The preprocessed signal features are then input into the intent recognition model of the intelligent decision engine to identify the user's specific needs and service preferences, thereby obtaining the initial user intent results. Based on the initial user intent result and combined with the preset business logic and service process, a test intent recognition result is generated.
3. The data processing method applied to a multi-functional integrated service terminal according to claim 2, characterized in that, The method for determining whether the training of the integrated control center is complete includes: The multi-source heterogeneous signals are input into the integrated control center, and the environmental perception results are output based on the multimodal fusion framework. Error calculation is performed based on the environmental perception results and the verified environmental perception results to determine the environmental perception error; a numerical comparison is performed based on the environmental perception error and a preset environmental perception threshold, and the training of the environmental perception algorithm is determined based on the comparison result. And / or, Based on the intent recognition result and the test intent recognition result, an error calculation is performed to determine the intent recognition error; A numerical comparison is made between the intent recognition error and the preset intent recognition threshold, and the training of the intent recognition model is determined based on the comparison result.
4. The data processing method applied to a multi-functional integrated service terminal according to claim 1, characterized in that, The steps for obtaining the training dataset include: Obtain the raw dataset containing raw multi-source heterogeneous signals collected under complex noise environments; The original dataset is preprocessed to obtain training multi-source heterogeneous signals; Based on the training multi-source heterogeneous signals, several environmental perception results under different environmental conditions are generated using the simulated environment as training environmental perception results. Based on business needs and service scenarios, define corresponding service feedback and user intent, and construct intent recognition results; By combining the training multi-source heterogeneous signals, the training environment perception results, and the intent recognition results, a complete training dataset is constructed.
5. The data processing method applied to a multi-functional integrated service terminal according to claim 2, characterized in that, The process of generating test intent recognition results based on the initial user intent result and in conjunction with preset business logic and service processes includes: Analyze the key information in the initial user intent results and match it with the preset service process; Based on the matching results, relevant service modules or service instruction sets are invoked to generate specific business processing instructions or office service responses. Compare the generated business processing instructions or office service responses with the initial user intent results; The final test intent recognition result is output based on the comparison information.
6. The data processing method applied to a multi-functional integrated service terminal according to claim 1, characterized in that, The optimization of the total loss function includes: optimizing the total loss function using the Adam optimizer.
Citation Information
Patent Citations
Human-computer interaction method, device and equipment and storage medium
CN116737883A
Multi-mode synchronous fusion speech recognition system
CN119296523A