Business handling method and device based on self-service terminal, self-service terminal and storage medium

By employing multimodal recognition technology in bank self-service terminals, combined with voice and touchscreen guidance, operation command recognition and emotion monitoring are achieved, solving the problems of low interaction efficiency and insufficient security of self-service terminals in special scenarios and improving the user experience.

CN121963358APending Publication Date: 2026-05-01INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2026-01-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Bank self-service terminals suffer from low interaction efficiency in special scenarios, making it difficult to meet diverse user needs and lacking security, resulting in a poor user experience.

Method used

Multimodal recognition technology is employed, including voice guidance strategies, touchscreen text guidance strategies, operation command recognition, and emotion monitoring. Guidance strategies are adjusted based on the user's age and emotion to improve business processing efficiency and security.

Benefits of technology

It improves the efficiency and security of business processing in special scenarios and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963358A_ABST
    Figure CN121963358A_ABST
Patent Text Reader

Abstract

The invention discloses a service handling method and device based on a self-service terminal, the self-service terminal and a storage medium. The method comprises the steps of receiving a business handling request and triggering verification of a target user identity; after the identity verification is passed, generating guidance information according to a first type of guidance strategy matched with the age of the user; the first type of guiding strategies comprises a voice guiding strategy and a touch screen character guiding strategy; performing operation instruction recognition and emotion monitoring on the target user in the process that the target user handles the business according to the guidance information; the operation instruction recognition comprises voice instruction recognition and gesture instruction recognition; and when it is monitored that the negative emotion exists in the process that the target user issues the operation instruction, determining a second type of guiding strategy according to the current operation instruction of the target user and the current emotion type, and guiding the target user to handle the business according to the second type of guiding strategy. According to the technical scheme, the business handling efficiency in a special scene is improved, the business handling safety is improved, and thus the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A business processing method, device, self-service terminal, and storage medium based on a self-service terminal. Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a business processing method, apparatus, self-service terminal, and storage medium based on a self-service terminal. Background Technology

[0002] Currently, bank self-service terminals mostly use a combination of touchscreens and buttons for human-computer interaction. In special business scenarios (such as multi-user concurrent interaction during peak hours, interaction scenarios for the elderly, and business scenarios with high security requirements), self-service terminals suffer from low interaction efficiency, leading to low business processing efficiency and failing to meet the diverse needs of users. Summary of the Invention

[0003] This invention provides a business processing method, device, self-service terminal, and storage medium based on a self-service terminal, in order to improve the efficiency and security of business processing in special scenarios, thereby enhancing the user experience.

[0004] According to one aspect of the present invention, a business processing method based on a self-service terminal is provided, the method comprising:

[0005] Upon receiving a business processing request from a target user, the system triggers the authentication process for the target user.

[0006] Upon successful identity verification, guidance information is generated according to a first type of guidance strategy that matches the target user's age to guide the target user in handling the business; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy.

[0007] During the process of the target user handling business according to the guidance information, the operation instructions and emotions of the target user are recognized based on the current business handling node; the operation instruction recognition includes voice instruction recognition and gesture instruction recognition.

[0008] If negative emotions are detected during the issuance of operation instructions by the target user, a second type of guidance strategy is determined based on the target user's current operation instructions and current emotion category, and the target user is guided to complete the business according to the second type of guidance strategy.

[0009] According to another aspect of the present invention, a service processing device based on a self-service terminal is provided, the device comprising:

[0010] The identity verification triggering module is used to receive a business processing request from a target user and trigger the identity verification of the target user.

[0011] The first guidance module is used to generate guidance information to guide the target user to conduct business in accordance with a first type of guidance strategy that matches the age of the target user after the identity verification is passed; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy.

[0012] The instruction recognition and emotion monitoring module is used to recognize operation instructions and monitor emotions of the target user based on the current business processing node during the process of the target user handling business according to the guidance information; the operation instruction recognition includes voice instruction recognition and gesture instruction recognition.

[0013] The second guidance module is used to determine a second type of guidance strategy based on the target user's current operation command and current emotion category when negative emotions are detected during the target user's issuance of operation instructions, and guide the target user to complete the business according to the second type of guidance strategy.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the self-service terminal-based business processing method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the self-service terminal-based business processing method according to any embodiment of the present invention.

[0019] The technical solution of this invention, by receiving a business processing request from a target user, triggers identity verification of the target user; if the identity verification is successful, guidance information is generated according to a first type of guidance strategy matching the target user's age to guide the target user in processing the business; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy; during the target user's business processing according to the guidance information, operation instructions and emotion monitoring are performed on the target user based on the current business processing node; operation instruction recognition includes voice instruction recognition and gesture instruction recognition; if negative emotions are detected during the target user's issuance of operation instructions, a second type of guidance strategy is determined based on the target user's current operation instructions and current emotion category, and the target user is guided to process the business according to the second type of guidance strategy. This technical means of using multimodal recognition and executing corresponding guidance strategies based on the recognition results solves the problems of low business processing efficiency and insufficient security leading to poor user experience in special application scenarios of existing technologies, improves business processing efficiency and security in special scenarios, and thus enhances the user experience.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1a is a flowchart of a business processing method based on a self-service terminal provided in Embodiment 1 of the present invention;

[0023] Figure 1b is a structural schematic diagram of a self-service terminal provided in this embodiment;

[0024] Figure 2 is a schematic diagram of a business processing device based on a self-service terminal provided in Embodiment 2 of the present invention;

[0025] Figure 3 is a schematic diagram of the structure of a self-service terminal that implements the business processing method based on a self-service terminal according to an embodiment of the present invention. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] Example 1

[0029] Figure 1a is a flowchart of a business processing method based on a self-service terminal according to Embodiment 1 of the present invention. This embodiment is applicable to situations where users process business using a bank self-service terminal. The method can be executed by a business processing device based on the self-service terminal, which can be implemented in hardware and / or software and can be configured in the central processing system of the self-service terminal. As shown in Figure 1a, the method includes:

[0030] S110: Receive the business processing request from the target user and trigger the authentication of the target user.

[0031] In this embodiment, service requests can be submitted via the touchscreen interface of the self-service terminal (e.g., clicking the "Check Balance" button on the touchscreen) or via the voice recognition module of the self-service terminal (e.g., giving a voice command to "Check Balance" in front of the self-service terminal). Upon receiving a service request from a target user, the self-service terminal can promptly trigger verification of the target user's identity.

[0032] In one optional implementation, triggering the authentication of the target user may include: triggering the acquisition of a real-time facial image of the target user; analyzing the real-time facial image using a pre-trained facial recognition model to determine the degree of facial occlusion of the target user; dynamically adjusting the marginal parameters of the loss function of the pre-trained facial recognition model during training based on the degree of facial occlusion of the training samples; extracting target facial features from the real-time facial image based on the degree of facial occlusion; comparing the target facial features with facial features in a pre-built database, and authenticating the target user based on the comparison results.

[0033] In this embodiment, real-time facial images of the target user can be acquired, and a pre-trained facial recognition model can be called to analyze the facial occlusion in the real-time facial images. If facial occlusion exists, the degree of occlusion can be determined, and target features can be extracted based on the degree of occlusion. These target features are then compared with facial features in a pre-built database to verify the target user's identity. In this embodiment, by analyzing the degree of facial occlusion, a balance can be struck between "recognition strictness" and "user experience," improving the user identity verification pass rate while reducing the false recognition rate. For example, if the occlusion area... 20% can have its marginal parameter adjusted to 0.3, resulting in a smaller shading area. At times, the marginal parameter can be adjusted to 0.2, etc.

[0034] In this embodiment, the self-service terminal can also prompt the target user to verify their identity by combining facial recognition with document verification via a touch screen. This combination of facial recognition and document verification can increase security for business transactions.

[0035] S120. If the identity verification is successful, generate guidance information according to the first type of guidance strategy that matches the target user's age to guide the target user to handle business; the first type of guidance strategy includes voice guidance strategy and touch screen text guidance strategy.

[0036] In this embodiment, the self-service terminal can guide users through transactions according to a working mode matched to the target user's age. For example, if the user is under 60 years old, they can be guided through transactions using a regular working mode; if the user is over 60 years old, they can be guided through transactions using a senior citizen mode. Guidance methods can include voice guidance and touchscreen text guidance.

[0037] In one optional implementation, generating guidance information according to a first type of guidance strategy matching the target user's age to guide the target user in handling business may include: determining the target user's age information based on the target user's real-time facial image; if the age information is greater than 60 years old, generating target voice guidance information or target touchscreen text guidance information to guide the target user in handling business; the speech rate of the target voice guidance information is lower than a preset normal speech rate, and the target voice guidance information is generated through a preset speech synthesis model; the preset speech synthesis model is obtained based on business domain corpus after at least one round of training, and the preset speech synthesis model includes a scene-adaptive prosody module; the text size of the touchscreen text guidance information is larger than a preset normal text size.

[0038] In this embodiment, if the target user is detected to be an elderly person, guidance information suitable for the elderly can be generated for the target user based on the current business processing node.

[0039] For example, a preset speech synthesis model can generate slow-paced (e.g., 30% slower than the preset normal speaking speed) voice guidance information to guide business transactions. The preset speech synthesis model can also generate voice guidance information based on dialects, which helps prevent elderly users from having difficulty hearing instructions. The speech synthesis model in this embodiment can be trained on a large amount of banking scenario speech data (including technical terms, operation prompts, etc.), ensuring the accuracy of pronunciation of technical terms in the voice guidance information. The speech synthesis model can also adapt to different scenarios by adjusting the pronunciation rhythm of the generated voice guidance information. For example, a brisk tone can be used when announcing "Operation successful," a calm tone when announcing "Insufficient account balance," and a pause can be added for emphasis when announcing "Please confirm the limit," improving the user interaction experience. The speech synthesis model in this embodiment can be an end-to-end speech synthesis model.

[0040] For example, text guidance information can be displayed on the touchscreen interface (the text guidance information can be 20% larger than the preset normal text), and the touchscreen interface can be simplified to retain only the core functions.

[0041] S130. During the process of the target user handling business according to the guidance information, the operation instructions of the target user are identified and emotions are monitored based on the current business handling node; the operation instruction recognition includes voice instruction recognition and gesture instruction recognition.

[0042] In this embodiment, when the target user conducts business according to the target voice guidance information and / or the target touch screen text guidance information, the self-service terminal can recognize the target user's operation instructions and perform emotion monitoring on the target user during the issuance of operation instructions.

[0043] In one optional implementation, identifying the target user's operation command based on the current business processing node may include:

[0044] When the operation command is recognized as a voice command, a pre-trained voice recognition module is used to process the current voice information and identify the target user's current voice command. The pre-trained voice recognition module includes an adaptive noise suppression module, a bidirectional temporal module, and an attention mechanism. The adaptive noise suppression module extracts specified noise spectrum features from the current voice information and generates inverse noise that matches the specified noise spectrum features to suppress noise in the current voice information. The bidirectional temporal module captures the forward and backward temporal sequences of the current voice information to accurately segment the current voice information. The attention mechanism focuses on key information in the current voice information. When the operation command is recognized as a gesture command, a pre-trained gesture recognition model is used to process the current gesture information and identify the target user's current gesture command. The pre-trained gesture recognition model includes a convolutional block attention module and a detection module. The convolutional block attention module focuses on the skin texture features of the hand through channel attention and locates the hand region through spatial attention to achieve hand region localization. The detection module sets anchor boxes of specified sizes to detect small targets on the hand.

[0045] In this embodiment, if the target user inputs a command via voice, the self-service terminal can collect the current voice information and use a pre-trained speech recognition module to recognize the command. Specifically, depthwise separable convolution can be used for feature extraction, reducing the computational power consumption of the self-service terminal (adapting to the limited hardware resources of the self-service terminal); an adaptive noise suppression module can be used to extract the spectral features of common noises in banking scenarios (such as queuing sounds and printer sounds) and generate reverse noise signals to cancel interference, greatly improving the accuracy of speech recognition in noisy environments; a bidirectional temporal module can also be used to simultaneously capture the "forward temporal sequence (such as the word 'account' after 'transfer')" and the "backward temporal sequence (such as the word 'transfer' before 'account')" of the current voice information, avoiding errors in command segmentation (such as misrecognizing "transfer 50,000" as "transfer 5"); a speech segment attention mechanism can also be used to focus on key information segments such as "quota" and "account number" to further reduce speech recognition errors.

[0046] In this embodiment, if the target user inputs a command using gestures, the self-service terminal can collect the current gesture information and use a pre-trained gesture recognition module to recognize the command. Specifically, this embodiment can focus the channel attention of the convolutional block attention module on the skin texture features of the hand, and locate the hand region through spatial attention (excluding background interference, such as terminal screen patterns), thereby improving the accuracy of hand region positioning and solving the positioning deviation problem when the user's hand is close to the screen; it can also set "16" for commonly used gestures in the detection module (such as swiping left and right to turn pages). 16, 32 32, 64 The 64” anchor frame size is suitable for small targets such as hands, solving the problem of recognition failure caused by unclear gesture features such as users wearing gloves or making small gesture movements.

[0047] Based on the above optional implementation methods, when the target user is older than 60 years old, the gesture capture time window of the pre-trained gesture recognition model can be extended; the extended pre-trained gesture recognition model can be used to process the current gesture information and recognize the target user's current gesture command.

[0048] This embodiment can extend the gesture capture time window of the pre-trained gesture recognition model and optimize the recognition of slow gestures when the target user is an elderly user, so as to adapt to the problem that the elderly user's gesture movements are small and slow.

[0049] In one optional implementation, monitoring the target user's emotions based on the target user's current business processing node may include: extracting key facial region features from the current face image using a pre-trained emotion recognition model; comparing the key facial region features with a pre-built emotion feature library; and determining the current emotion category based on the comparison results to achieve emotion monitoring.

[0050] In this embodiment, facial image features can be extracted using the convolutional backbone network of a pre-trained emotion recognition model. An attention mechanism is used to focus on key facial areas (such as eyebrows and corners of the mouth). The extracted features are compared with a pre-built emotion feature library to determine the current emotion category of the target user. The self-service terminal can then determine the assistance operation for the target user based on the current emotion category.

[0051] S140. If negative emotions are detected during the process of a target user issuing an operation instruction, a second type of guidance strategy is determined based on the target user's current operation instruction and current emotion category, and the target user is guided to complete the business according to the second type of guidance strategy.

[0052] In this embodiment, if the target user has negative emotions (such as anxiety or confusion) when issuing operation instructions to handle business, the self-service terminal can provide corresponding guidance to the target user based on the emotion category to help the user.

[0053] In one optional implementation, determining the second type of guidance strategy based on the target user's current operation and current mood category may include: triggering secondary verification of the target user's current operation instruction based on the target user's current operation and current mood category; triggering operation process prompts for the target user's current operation; or triggering remote assistance to assist the target user in handling business.

[0054] For example, if the user is anxious (worried about making a mistake), the self-service terminal can automatically display safety prompts (such as "Secondary confirmation required, beware of operational errors") and provide access to human customer service; if the user is confused (unfamiliar with the operation or does not know how to operate), the self-service terminal can automatically pop up a step-by-step video guide (such as "How to insert a card") and read out the operation steps in slow speech through a voice synthesis module.

[0055] To enable those skilled in the art to better understand the self-service terminal of this embodiment, Figure 1b is a structural schematic diagram of a self-service terminal provided in this embodiment. The self-service terminal may include a voice interaction module, a gesture recognition module, a face recognition module, a touchscreen interaction module, and a central processing module. The voice interaction module, gesture recognition module, face recognition module, and touchscreen interaction module are all communicatively connected to the central processing module and work collaboratively under the coordination of the central processing module.

[0056] 1. Voice Interaction Module: This module may include a microphone, a speech recognition unit, a speech synthesis unit, and a speaker. The microphone is used to collect the user's voice commands; the speech recognition unit can convert the collected voice signals into text information and further parse them into operation signals, improving the accuracy of speech recognition in noisy bank environments; the speech synthesis unit can convert the feedback information from the self-service terminal (such as operation results, prompts, etc.) into natural and fluent voice signals; the speaker is used to output the voice signals generated by the speech synthesis unit.

[0057] 2. Gesture Recognition Module: This module can consist of a camera and a gesture analysis unit. The camera captures the user's hand gestures; the gesture analysis unit processes and recognizes the captured gesture images. It quickly locates the hand area, then uses a key point extraction algorithm to extract key feature points such as finger joints, thereby recognizing the user's preset gestures (e.g., waving to confirm, clenching a fist to cancel, swiping left or right to turn pages, etc.) and converting them into corresponding operation commands. The recognition response time can be controlled within 300ms.

[0058] 3. Face Recognition Module: This module may include a high-definition camera and a face comparison unit. The high-definition camera captures the user's facial information; the face comparison unit compares the captured facial information with the user's identity information in the database to complete identity verification. Simultaneously, this module can also monitor the user's facial expressions in real time during interaction using facial expression analysis algorithms, helping to determine the user's operational intentions or emotional state. For example, if the user shows confusion, the terminal can automatically provide more detailed operation guidance.

[0059] 4. Touchscreen interaction module: It can retain the traditional touchscreen operation mode and serve as a supplement to multimodal interaction. Users can complete input and selection operations by touching the screen when needed.

[0060] 5. Central Processing Module: As the core of the self-service terminal, this module employs an attention-based modal fusion algorithm. It receives operation signals and instructions from various interaction modules, analyzes and processes them, and then sends control commands to the terminal's execution mechanisms (such as the card processing module). Simultaneously, it coordinates the work of each interaction module, enabling intelligent switching between modalities. For example, when the voice interaction module detects unclear user instructions, it can automatically activate the touchscreen interaction module, displaying options for the user to choose from. When the gesture recognition module does not detect user gestures, it can prompt the user to use voice or the touchscreen for operation.

[0061] The technical solution of this embodiment, by receiving a business processing request from a target user, triggers identity verification of the target user; if the identity verification is successful, guidance information is generated according to a first type of guidance strategy matching the target user's age to guide the target user in processing the business; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy; during the target user's business processing according to the guidance information, operation instructions and emotion monitoring are performed on the target user based on the current business processing node; operation instruction recognition includes voice instruction recognition and gesture instruction recognition; if negative emotions are detected during the target user's issuance of operation instructions, a second type of guidance strategy is determined based on the target user's current operation instructions and current emotion category, and the target user is guided to process the business according to the second type of guidance strategy. By adopting multimodal recognition and executing the corresponding guidance strategy based on the recognition results, this technical means solves the problem of low business processing efficiency and insufficient security in special application scenarios, resulting in poor user experience, and improves business processing efficiency and security in special scenarios, thereby enhancing the user experience.

[0062] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to maintain the security of user personal information and network security.

[0063] Example 2

[0064] Figure 2 is a schematic diagram of a business processing device based on a self-service terminal provided in Embodiment 2 of the present invention. As shown in Figure 2, the device includes: an identity verification triggering module 210, a first guidance module 220, an instruction recognition and emotion monitoring module 230, and a second guidance module 240. Wherein:

[0065] The identity verification triggering module 210 is used to receive a business processing request from a target user and trigger identity verification for the target user.

[0066] The first guidance module 220 is used to generate guidance information to guide the target user to handle business according to a first type of guidance strategy that matches the age of the target user after the identity verification is passed; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy.

[0067] The instruction recognition and emotion monitoring module 230 is used to perform operation instruction recognition and emotion monitoring on the target user based on the current business processing node during the process of the target user handling business according to the guidance information; the operation instruction recognition includes voice instruction recognition and gesture instruction recognition.

[0068] The second guidance module 240 is used to determine a second type of guidance strategy based on the target user's current operation command and current emotion category when negative emotions are detected during the target user's issuance of operation instructions, and guide the target user to handle business according to the second type of guidance strategy.

[0069] The technical solution of this invention, by receiving a business processing request from a target user, triggers identity verification of the target user; if the identity verification is successful, guidance information is generated according to a first type of guidance strategy matching the target user's age to guide the target user in processing the business; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy; during the target user's business processing according to the guidance information, operation instructions and emotion monitoring are performed on the target user based on the current business processing node; operation instruction recognition includes voice instruction recognition and gesture instruction recognition; if negative emotions are detected during the target user's issuance of operation instructions, a second type of guidance strategy is determined based on the target user's current operation instructions and current emotion category, and the target user is guided to process the business according to the second type of guidance strategy. This technical means of using multimodal recognition and executing corresponding guidance strategies based on the recognition results solves the problems of low business processing efficiency and insufficient security leading to poor user experience in special application scenarios of existing technologies, improves business processing efficiency and security in special scenarios, and thus enhances the user experience.

[0070] Optionally, the authentication triggering module 210 can be used for:

[0071] Trigger the acquisition of the target user's real-time facial image;

[0072] The pre-trained face recognition model is used to analyze the real-time face image to determine the degree of facial occlusion of the target user; the pre-trained face recognition model dynamically adjusts the marginal parameters of the loss function according to the degree of facial occlusion of the training samples during the training process;

[0073] Extract target facial features from the real-time facial image based on the degree of facial occlusion;

[0074] The target facial features are compared with facial features in a pre-built database, and the identity of the target user is verified based on the comparison results.

[0075] Optionally, the first guidance module 220 can be used for:

[0076] The age information of the target user is determined based on the real-time facial image of the target user;

[0077] If the age information is greater than 60 years old, generate target voice guidance information or target touch screen text guidance information to guide the target user to handle business.

[0078] The speech rate of the target speech guidance information is lower than the preset normal speech rate. The target speech guidance information is generated by a preset speech synthesis model. The preset speech synthesis model is obtained by training at least one round of training based on business domain corpus. The preset speech synthesis model includes a scene-adaptive prosody module.

[0079] The text size of the touchscreen text guidance information is larger than the preset normal text size.

[0080] Optional, the instruction recognition and emotion monitoring module 230 can be used for:

[0081] When the operation command is identified as a voice command, a pre-trained voice recognition module processes the current voice information to identify the target user's current voice command. The pre-trained voice recognition module includes an adaptive noise suppression module, a bidirectional timing module, and an attention mechanism. The adaptive noise suppression module extracts specified noise spectrum features from the current voice information and generates inverse noise matching the specified noise spectrum features to suppress noise in the current voice information. The bidirectional timing module captures the forward and backward timing sequences of the current voice information to accurately segment the current voice information. The attention mechanism focuses on key information in the current voice information.

[0082] When the operation command is recognized as the gesture command, the current gesture information is processed using a pre-trained gesture recognition model to identify the current gesture command of the target user. The pre-trained gesture recognition model includes a convolutional block attention module and a detection module. The convolutional block attention module focuses on the skin texture features of the hand through channel attention and locates the hand region through spatial attention to achieve hand region localization. The detection module sets a specified size anchor box to detect small targets on the hand.

[0083] Optionally, the instruction recognition and emotion monitoring module 230 can also be used for:

[0084] The pre-trained emotion recognition model is used to extract key facial region features from the current face image;

[0085] The facial key region features are compared with a pre-built emotion feature library, and the current emotion category is determined based on the comparison results to achieve emotion monitoring.

[0086] Optionally, the self-service terminal-based business processing device further includes an elderly person's gesture recognition module, used for:

[0087] If the target user is older than 60 years old, extend the gesture capture time window of the pre-trained gesture recognition model;

[0088] The extended pre-trained gesture recognition model is used to process the current gesture information and identify the current gesture command of the target user.

[0089] Optionally, the second guidance module 240 can be used for:

[0090] A secondary verification of the target user's current action instruction is triggered based on the target user's current action and current emotion category;

[0091] Trigger an operation flow prompt for the target user's current operation; or...

[0092] Trigger remote assistance mode to help the target user conduct business.

[0093] The self-service terminal-based business processing device provided in the embodiments of the present invention can execute the self-service terminal-based business processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0094] Example 3

[0095] Figure 3 shows a schematic diagram of the structure of a self-service terminal 300 that can be used to implement an embodiment of the present invention. Electronic devices are intended to represent various forms of digital computers or various forms of mobile devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.

[0096] As shown in Figure 3, the self-service terminal 300 includes at least one processor 301 and a memory, such as a read-only memory (ROM) 302 and a random access memory (RAM) 303, communicatively connected to the processor 301. The memory stores computer programs executable by the processor. The processor 301 can perform various appropriate actions and processes based on the computer program stored in the ROM 302 or loaded into the RAM 303 from storage unit 308. The RAM 303 can also store various programs and data required for the operation of the self-service terminal 300. The processor 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0097] Multiple components in the self-service terminal 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard and mouse; an output unit 307, such as various types of displays and speakers; a storage unit 308, such as a disk and optical disk; and a communication unit 309, such as a network card, modem, or wireless transceiver. The communication unit 309 allows the self-service terminal 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0098] Processor 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 301 performs the various methods and processes described above, such as business processing methods based on self-service terminals.

[0099] In some embodiments, the self-service terminal-based service processing method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the self-service terminal 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by processor 301, one or more steps of the self-service terminal-based service processing method described above can be performed. Alternatively, in other embodiments, processor 301 can be configured to perform the self-service terminal-based service processing method by any other suitable means (e.g., by means of firmware).

[0100] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0101] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0102] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0104] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0105] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0106] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0107] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A business processing method based on a self-service terminal, characterized in that, include: Upon receiving a business processing request from a target user, the system triggers the authentication process for the target user. Upon successful identity verification, guidance information is generated according to a first type of guidance strategy that matches the target user's age to guide the target user in handling the business; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy. During the process of the target user handling business according to the guidance information, the operation instructions and emotions of the target user are recognized based on the current business handling node; the operation instruction recognition includes voice instruction recognition and gesture instruction recognition. If negative emotions are detected during the issuance of operation instructions by the target user, a second type of guidance strategy is determined based on the target user's current operation instructions and current emotion category, and the target user is guided to complete the business according to the second type of guidance strategy.

2. The method according to claim 1, characterized in that, Triggering authentication of the target user includes: triggering the acquisition of a real-time facial image of the target user; analyzing the real-time facial image using a pre-trained facial recognition model to determine the degree of facial occlusion of the target user; dynamically adjusting the marginal parameters of the loss function of the pre-trained facial recognition model according to the degree of facial occlusion of the training samples during training; extracting target facial features from the real-time facial image according to the degree of facial occlusion; comparing the target facial features with facial features in a pre-built database, and authenticating the target user based on the comparison results.

3. The method according to claim 1, characterized in that, Generating guidance information according to a first type of guidance strategy matching the target user's age to guide the target user in handling business includes: determining the target user's age information based on the target user's real-time facial image; if the age information is greater than 60 years old, generating target voice guidance information or target touchscreen text guidance information to guide the target user in handling business; the speech rate of the target voice guidance information is lower than a preset normal speech rate, and the target voice guidance information is generated through a preset speech synthesis model; the preset speech synthesis model is obtained based on business domain corpus after at least one round of training, and the preset speech synthesis model includes a scene-adaptive prosody module; the text size of the touchscreen text guidance information is larger than a preset normal text size.

4. The method according to claim 1, characterized in that, Based on the current business processing node, the operation command recognition of the target user includes: when the operation command is recognized as a voice command, processing the current voice information using a pre-trained voice recognition module to recognize the target user's current voice command; the pre-trained voice recognition module includes an adaptive noise suppression module, a bidirectional temporal module, and an attention mechanism; the adaptive noise suppression module is used to extract specified noise spectrum features from the current voice information and generate inverse noise matching the specified noise spectrum features to suppress noise in the current voice information; the bidirectional temporal module is used to capture the forward and backward temporal sequences of the current voice information to accurately segment the current voice information; the attention mechanism is used to focus on key information in the current voice information; when the operation command is recognized as a gesture command, processing the current gesture information using a pre-trained gesture recognition model to recognize the target user's current gesture command; the pre-trained gesture recognition model includes a convolutional block attention module and a detection module; the convolutional block attention module focuses on hand skin texture features through channel attention and locates the hand region through spatial attention to achieve hand region localization; the detection module sets a specified size anchor box to detect small hand targets.

5. The method according to claim 1, characterized in that, Based on the target user's current business processing node, emotion monitoring of the target user includes: extracting key facial region features from the current face image using a pre-trained emotion recognition model; comparing the key facial region features with a pre-built emotion feature library; and determining the current emotion category based on the comparison results to achieve emotion monitoring.

6. The method according to claim 4, characterized in that, Also includes: If the target user is older than 60 years old, extend the gesture capture time window of the pre-trained gesture recognition model; The extended pre-trained gesture recognition model is used to process the current gesture information and identify the current gesture command of the target user.

7. The method according to claim 1, characterized in that, The second type of guidance strategy is determined based on the target user's current operation and current emotion category, including: triggering secondary verification of the target user's current operation instruction based on the target user's current operation and current emotion category; triggering operation process prompts for the target user's current operation; or triggering remote assistance mode to assist the target user in handling business.

8. A business processing device based on a self-service terminal, characterized in that, include: The identity verification triggering module is used to receive a business processing request from a target user and trigger the identity verification of the target user. The first guidance module is used to generate guidance information to guide the target user to conduct business in accordance with a first type of guidance strategy that matches the age of the target user after the identity verification is passed; the first type of guidance strategy includes a voice guidance strategy and a touch screen text guidance strategy. The instruction recognition and emotion monitoring module is used to recognize the operation instructions and monitor the emotions of the target user based on the current business processing node during the process of the target user handling business according to the guidance information. The operation command recognition includes voice command recognition and gesture command recognition; The second guidance module is used to determine a second type of guidance strategy based on the target user's current operation command and current emotion category when negative emotions are detected during the target user's issuance of operation instructions, and guide the target user to complete the business according to the second type of guidance strategy.

9. A self-service terminal, characterized in that, The self-service terminal includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a business processing method based on a self-service terminal according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute a business processing method based on a self-service terminal as described in any one of claims 1-7.