A payment processing method based on a desktop robot and a desktop robot

By using desktop robots to acquire user data, match payment methods, and provide various interactive guidance, the problem of the lack of intelligence and personalization in existing payment terminal devices is solved, thereby improving the payment experience and emotional satisfaction.

CN122434513APending Publication Date: 2026-07-21ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2026-04-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing commercial POS scenarios, payment terminal devices only provide one-way information output and lack intelligent and personalized payment guidance and service experience.

Method used

By acquiring visual and voice data through desktop robots, user profile characteristics are determined, and target payment methods are matched accordingly. Various interactive methods such as body movements, lighting effects, voice broadcasts, screen displays, and projections are used to guide users to make payments.

Benefits of technology

It improves the intelligence of payment interaction, reduces the time cost for users to choose payment methods, enhances the naturalness and emotional experience of the payment process, and alleviates users' waiting anxiety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434513A_ABST
    Figure CN122434513A_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a payment processing method based on a desktop robot, a desktop robot and a computing device. The scheme comprises: at least one perception data in visual data and voice data collected by the desktop robot; determining a user portrait feature based on the perception data; determining a target payment method matched with the user portrait feature from a plurality of payment methods supported by the desktop robot; the desktop robot presents interaction information corresponding to the target payment method, so that the user interacts with the desktop robot and adopts the target payment method for payment. Improve user payment experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a payment processing method based on a desktop robot. This specification also relates to a desktop robot, a payment processing device based on a desktop robot, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the rapid development of mobile payment technology and the intelligent upgrading of commercial retail scenarios, payment terminal devices have been widely used in various commercial venues, such as retail stores, restaurants, cafes, and cosmetic counters. Currently, commercial checkout scenarios mainly use devices such as barcode scanners, QR code stands, NFC tags, and POS machines to complete payment interactions. These devices only provide one-way information output.

[0003] How to achieve more intelligent and personalized payment guidance and service experience is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide a desktop robot-based payment processing method, desktop robot, apparatus, device, and computer-readable medium to improve the user's payment experience.

[0005] According to a first aspect of one or more embodiments of this specification, a payment processing method based on a desktop robot is provided, comprising: acquiring perception data; the perception data including at least one of visual data and voice data collected by the desktop robot; Based on the perceived data, user profile features are determined; From the various payment methods supported by the desktop robot, determine the target payment method that matches the user profile characteristics; The desktop robot is controlled to present interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method; the interactive information includes at least one form of interactive information such as body movements, light display, voice broadcast, screen display and projection.

[0006] According to a second aspect of one or more embodiments of this specification, a desktop robot is provided, including: a head module, a body module, and a limb module; The head module and the limb module are respectively connected to the body module; At least one of the head module, the body module, and the limb module includes an information sensing unit; the information sensing unit is used to collect at least one type of sensing information, namely visual data and voice data. At least one of the head module, the body module, and the limb module is used to present interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method; the interactive information includes at least one form of interactive information such as body movements, light presentation, voice broadcast, screen display, and projection; the target payment method is a payment method that matches the user profile characteristics from a variety of payment methods supported by the desktop robot, and the user profile characteristics are determined based on the perception data.

[0007] According to a third aspect of one or more embodiments of this specification, a desktop robot-based payment processing apparatus is provided, comprising: The data acquisition module is used to acquire sensory data; the sensory data includes at least one of visual data and voice data collected by the desktop robot.

[0008] The feature determination module is used to determine user profile features based on the perceived data.

[0009] The payment method determination module is used to determine the target payment method that matches the user profile characteristics from among the various payment methods supported by the desktop robot.

[0010] The information control module is used to control the desktop robot to present interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method; the interactive information includes at least one form of interactive information such as body movements, light display, voice broadcast, screen display and projection.

[0011] According to a fourth aspect of one or more embodiments of this specification, a computing device, a memory, and a processor are provided; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of a method for payment processing based on a desktop robot.

[0012] According to a fifth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions, which, when executed by a processor, implement the steps of a method for payment processing based on a desktop robot.

[0013] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of a method for payment processing based on a desktop robot.

[0014] One or more embodiments of this specification can achieve at least the following beneficial effects: by acquiring at least one type of perceptual data, including visual and voice data, user profile features are determined based on the perceptual data, and a matching target payment method is determined from multiple payment methods accordingly. This can reduce the time cost for users to choose a payment method or the waiting time for merchants to select a payment method, improve the naturalness and acceptability of payment interactions, and enhance the level of intelligence, thereby improving the user experience.

[0015] On the other hand, desktop robots can present interactive information corresponding to the target payment method. Through the coordinated cooperation of multi-channel feedback, the payment process is no longer a mechanical output of information, but an interactive process with emotional expression. For example, when the payment is successful, the robot can perform a celebratory action accompanied by light effects; when the payment fails, it can perform a reassuring action and guide the user to try again, effectively alleviating the user's waiting anxiety and enhancing the emotional experience. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram illustrating an application scenario of a desktop robot-based payment processing method provided in one embodiment of this specification. Figure 2 This is a flowchart illustrating a payment processing method based on a desktop robot provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the structure of a desktop robot provided in one embodiment of this specification; Figure 4 This is a flowchart illustrating an interaction method for a desktop robot provided in one embodiment of this specification. Figure 5 This is a schematic diagram of the structure of a payment processing device based on a desktop robot provided in one embodiment of this specification; Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0019] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0020] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “an,” “an,” “the,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification includes any or all possible combinations of one or more associated listed items.

[0021] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded.

[0022] Although the terms "first," "second," etc., may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of this specification. Ordinal numbers such as "first," "second," etc., do not necessarily indicate order; often they are used to facilitate the distinction of objects. For example, "first server" and "second server" usually refer to two servers. To distinguish these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0023] The word "if" can be interpreted as "when," "when," or "in response to a determination," depending on the context.

[0024] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0025] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0026] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. For example, in locations where desktop robots are deployed, video capture prompts may be displayed by pasting or showing; or, when a user becomes a registered user of the terminal application or processes business, authorization prompts may be displayed through terms and conditions, allowing data collection and use based on user authorization; or, authorization prompts may be displayed on the desktop robot's display interface. In practical applications, authorization prompts may be presented to users in one or more ways, and the specific methods are not specifically limited.

[0027] The following explains the terms and concepts used in one or more embodiments of this specification.

[0028] Desktop robots are small, intelligent, and interactive devices that can be placed on desktops, workbenches, or other similar locations to assist in business processes. They typically possess one or more of the following functions: speech recognition, emotion simulation, facial expression display, motion feedback, and environmental awareness. A desktop robot can include hardware components (such as the main structure, drive unit, and sensors) and software components (such as control algorithms and human-computer interaction interfaces). Alternatively, desktop robots may also have network connectivity to interact with servers.

[0029] QR code payment is a wireless payment method based on an account system. Merchants generate QR codes containing transaction information for users to scan, or users can complete payment by scanning their payment code. The main types of QR code payments include active scanning and passive scanning.

[0030] Active Scan Mode: Also known as the active scan payment method, this refers to a payment method where the payer scans a code to complete the payment. Specifically, the payer can use their mobile device to read the barcode displayed by the recipient to complete the payment. For example, the payer uses the "scan" function in their mobile application to scan the merchant's QR code, with the payer acting as the active scanner.

[0031] Scanned Payment Mode: Also known as the scanned payment method, this refers to a payment method where the recipient scans a barcode displayed on the payer's mobile device to complete the payment. Specifically, it involves the payee reading the barcode displayed on the payer's mobile device. For example, the payer displays a payment code through a mobile application, and the merchant's device scans this code to complete the transaction.

[0032] Near Field Payment: This refers to a payment method where users complete on-site transactions with a payment terminal using near-field wireless communication technologies (such as NFC, Bluetooth, UWB, etc.).

[0033] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0034] Figure 1 This is a schematic diagram illustrating an application scenario of a desktop robot-based payment processing method provided in one embodiment of this specification.

[0035] like Figure 1 As shown, the desktop robot 100 can be positioned at locations such as cash registers and self-service checkout machines. It can collect environmental information, such as images from a camera and sounds from a microphone. If the desktop robot 100 determines that a user 200 is nearby, it can analyze the user's profile and present interactive information corresponding to the payment method matched to that user, enabling the user to make a payment. For example, if it is predicted that the user will most likely use a merchant's QR code for payment, the desktop robot can display the merchant's QR code and prompt the user to scan it. If it is predicted that the user will most likely use near-field communication (NFC) payment, the desktop robot can light up the NFC touch area or prompt the user to touch it through gestures or voice prompts. The user 200 can interact with the desktop robot using a user terminal to complete the payment. For example, the user can touch the desktop robot with their user terminal to make a near-field payment; or the user can scan the QR code displayed by the desktop robot using their user terminal; or the user can display a payment code on their user terminal, which the desktop robot can scan to make the payment.

[0036] In this way, desktop robots can provide payment methods tailored to different users, making them suitable for various users and allowing them to use their preferred payment methods. This minimizes the need for users to switch payment methods when needed, thus improving the user experience. It also simplifies the process for merchants to switch payment methods based on user preferences or change the point-of-sale (POS) devices, improving payment processing efficiency.

[0037] In practical applications, desktop robots can be equipped with business processing logic. For example, they can analyze collected images or sounds to determine user profile characteristics and predict the payment methods the user might use, then control the desktop robot to present corresponding interactive information. Alternatively, the desktop robot can connect to a server, which can also contain business processing logic. The server can analyze images or sounds and determine payment methods, then issue execution commands to the desktop robot, which will then present corresponding interactive information based on these commands. Alternatively, some business processing logic can be executed in the desktop robot, while other logic can be executed on the server, with the server and desktop robot working together to perform business processing.

[0038] The user terminal can be one or more of the following: smartphone, laptop, tablet, IoT device, portable wearable device, or immersive image display device. Specifically, IoT devices can be one or more of the following: smart speaker, smart TV, smart air conditioner, or smart in-vehicle device. Portable wearable devices can be one or more of the following: smartwatch, smart bracelet, or head-mounted device. Immersive image display devices include, but are not limited to, augmented reality (AR) devices and virtual reality (VR) devices.

[0039] Users can represent registered or active users of a terminal device or application. When a user registers or uses the terminal device or application, the terminal device or application can display terms and conditions regarding the collection, use, or processing of user information. Users can choose to agree to authorization or disagree, or, after authorization, can exit through a preset path. The payment methods provided by the desktop robot can be any payment method supported by the terminal device or application. After collecting data containing a user's facial or voiceprint features, the desktop robot can locally determine or request the server whether the user is a registered or active user of the terminal device or application, or whether the user is an authorized user. If the user is a registered or active user, or an authorized user, the desktop robot can match the corresponding payment method and display relevant information. If the user is not a registered or active user, or an authorized user, the desktop robot may not match a corresponding payment method for that user, may not build a user profile, may not store the user's biometric data, and may delete the collected image or sound information; the desktop robot can then present the default payment method.

[0040] A server can be a standalone physical server, a server cluster consisting of multiple physical servers, or a distributed file system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0041] The server can connect to one or more desktop robots via a local area network (LAN), a wide area network (WAN), the internet, or other types of data networks. Similarly, if user terminals can also connect to the server, the service can also connect to one or more user terminals via a LAN, WAN, internet, or other types of data networks.

[0042] This specification provides a payment processing method based on a desktop robot, and also relates to a desktop robot, a payment processing device based on a desktop robot, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0043] Figure 2 This is a flowchart illustrating a payment processing method based on a desktop robot, as provided in one embodiment of this specification.

[0044] From a programming perspective, the entity executing the process can be a program hosted on an application server or a desktop robot. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.

[0045] like Figure 2 As shown, the process may include the following steps.

[0046] Step 202: Acquire sensory data.

[0047] The perception data may include at least one of visual data and voice data collected by the desktop robot. Alternatively, the perception data may also include short-range communication signal data, such as radio frequency field signals, Bluetooth signals, etc. Or, the perception data may include infrared perception data, pressure perception data, etc., used to detect whether a user is touching the desktop robot. The perception data can be a collection of data collected by the desktop robot through various sensors that reflects user characteristics and / or environmental conditions.

[0048] Visual data can represent image or video data acquired through image acquisition devices (such as cameras, depth cameras, etc.). Visual data can include, but is not limited to, RGB images, infrared images, depth images, video streams, etc.

[0049] Speech data can represent audio signals acquired through sound acquisition devices (such as microphones, microphone arrays, etc.). Speech data can include, but is not limited to, raw audio waveforms, pre-processed audio features, and text after speech recognition. It can contain human voices, ambient sounds, etc., or it can preserve human voices after denoising ambient sounds.

[0050] Desktop robots can be intelligent interactive devices deployed in desktop settings (such as cash registers, service counters, and counters). Desktop robots may have a sensing module for acquiring sensory data.

[0051] In one implementation, the desktop robot may have an image acquisition module for acquiring visual data. For example, the desktop robot may have a camera, which can be one or more combinations of an RGB camera, a depth camera, and an infrared camera. The camera can be fixed or rotatable. Alternatively, the desktop robot may have an image sensor, such as a CMOS sensor or a CCD sensor, for converting optical signals into electrical signals. Exemplarily, the desktop robot can activate the camera, set acquisition parameters (such as resolution, frame rate, exposure time, etc.) or use default acquisition parameters; the camera acquires images or video streams of the current scene; the acquired images are preprocessed (such as noise reduction, enhancement, cropping, etc.); and the preprocessed visual data is stored in a buffer for use in subsequent steps.

[0052] In one implementation, the desktop robot may have a sound acquisition module for collecting voice data. The desktop robot may have a microphone, which can be a single microphone or a microphone array (such as a 2-microphone, 4-microphone, or 6-microphone circular array); it may also have an audio processor for amplifying, filtering, and performing analog-to-digital conversion on the audio signal; and it may have a sound source localization module for calculating the azimuth angle of the sound source, supporting user direction tracking. For example, the desktop robot may activate the microphone array, set acquisition parameters (such as sampling rate, bit depth, gain, etc.) or use default acquisition parameters; the microphone array may acquire sound signals from the current environment; then, the acquired audio may be preprocessed (such as filtering, noise reduction, echo cancellation, etc.); and the preprocessed voice data may be stored in a buffer for use in subsequent steps.

[0053] In practical applications, desktop robots can also acquire visual and voice data simultaneously. For example, a desktop robot can activate its camera and microphone array at the same time to collect multimodal data, which can then be used to determine user profile characteristics. Alternatively, visual and voice data can be combined with other data, such as infrared sensor data, Bluetooth signals, NFC signals, and other multimodal data, to determine user profile characteristics or whether a user is interacting with the desktop robot.

[0054] In one implementation, the desktop robot can connect to external devices, and the sensing data can be data collected by the external devices. For example, the desktop robot establishes a connection with external sensors via a wireless communication module (such as Wi-Fi, Bluetooth, etc.) or a wired communication module (such as a USB interface); the external sensors collect visual data and / or voice data; and send the collected data to the robot.

[0055] Step 204: Determine user profile features based on the perceived data.

[0056] User profile features can represent a set of labels or parameters describing a user's attributes and state. A user profile can be an abstract representation of a user's multi-dimensional characteristics, and features can be specific attribute items in the profile. User profile features are an abstraction and refinement of raw data, enabling desktop robots to understand a user's basic attributes and current state. User profile features can include, but are not limited to, one or more of the following: age features, gender features, emotional features, identity features, behavioral features, historical transaction features, product interest features, and environmental features.

[0057] For example, deep learning networks can be used to integrate features such as facial structure analysis, clothing style, voice tone, behavioral habits, historical payment data, and gender and age recognition to generate multi-dimensional user profile feature tags, including customer identity, age group, gender, preferences, purchase frequency, and product interests.

[0058] To ensure data security and reduce the risk of leakage, an edge computing architecture can be adopted. The extraction of user profile features can be completed on the desktop robot's local processor (such as an NPU or CPU). Raw visual data (such as facial images) and voice data are cleared from memory after feature extraction, retaining only irreversible feature vectors. These feature vectors cannot be reconstructed from the original image or voice, thus achieving data anonymization. Alternatively, the user profile feature extraction process can also be performed on a server. The desktop robot can provide the collected data to the server after encryption or anonymization processing, where the server will perform feature extraction.

[0059] Step 206: Determine the target payment method that matches the user profile characteristics from the various payment methods supported by the desktop robot.

[0060] Multiple payment methods can refer to the set of payment channels or methods that the desktop robot can support and process. "Multiple" indicates two or more methods. Payment methods can include, but are not limited to, contactless payment, scanned payment, near-field payment, facial recognition payment, and bank card payment. Near-field payment can refer to payment via contactless communication; scanned payment can refer to payment by presenting a payment code; contactless payment can refer to payment by scanning a payment code; facial recognition payment can refer to payment by scanning a face using a merchant's device; and bank card payment can refer to payment using a bank card, such as via a magnetic stripe or NFC tag.

[0061] The target payment method can refer to the payment method that is finally determined and recommended to the user after matching and selection.

[0062] In practical applications, different user groups may have varying levels of acceptance and operational skills regarding payment methods. For example, elderly users may be unfamiliar with QR code scanning and are more comfortable with being scanned for payment; younger users, prioritizing efficiency, are more comfortable with near-field communication (NFC) payments. By matching user profile characteristics with payment methods, the most suitable payment options can be provided to users, reducing operational difficulties.

[0063] One implementation method is to determine the target payment method based on rule matching. Specifically, a payment recommendation rule table can be preset in the server or desktop robot. This table can set corresponding payment methods for one or more dimensions of user profile features. For example, elderly people can correspond to the payment method of being scanned, children to the payment method of being initiated, and young and middle-aged people to the payment method of near-field payment. Alternatively, user terminals of model A or brand B can correspond to the payment method of being scanned, and user terminals of model B or brand C to the payment method of being initiated. Then, the user profile features are compared with the payment recommendation rule table; the payment method corresponding to the rule that matches successfully is taken as the target payment method. If multiple rules match successfully, they can be selected according to priority, such as prioritizing those with higher historical preferences; or, weights can be set for each dimension of user profile features, and the target payment method can be determined by these weights.

[0064] One implementation method is to determine the target payment method based on a machine learning model. Specifically, a payment recommendation neural network model can be pre-trained, with user profile features as input and payment method probabilities as output. After the desktop robot acquires the perception data, it can input the extracted user profile features into the trained payment recommendation neural network model. The model outputs the recommendation probabilities of each payment method and selects the payment method with the highest probability as the target payment method.

[0065] One implementation method is to determine the target payment method based on historical preferences. Specifically, user identity information can be determined based on perceived data, and then the user's historical payment information can be retrieved based on that identity information to identify the user's frequently used payment methods, which are then designated as the target payment method. If the user's historical payment information is unavailable or insufficient, rule matching or model recommendation can also be used.

[0066] As one implementation method, the target payment method can be dynamically determined based on the real-time scenario. Specifically, current scenario parameters, such as the number of people in the queue, network status, and time, can be obtained. These parameters are then input into a recommendation algorithm to determine the target payment method. For example, if there are many people in the queue, payment methods with shorter processing times, such as NFC-based near-field payment, are prioritized. Alternatively, if the merchant's network is poor, payment methods that require scanning are prioritized, utilizing the user's terminal to trigger the payment process and avoiding long waiting times for the merchant to process the transaction.

[0067] In practical applications, one or more specific implementation methods can be selected for determining the target payment method. If multiple methods are selected and multiple payment methods are initially determined, the final recommended target payment method can be determined based on the weights corresponding to each specific implementation method.

[0068] By identifying a target payment method from multiple payment options that matches the user's profile characteristics, intelligent payment method recommendations can be achieved, reducing the time users spend manually selecting and judging. Furthermore, identifying the target payment method ensures that subsequent interactive information is presented in a targeted manner, avoiding invalid or redundant interactive content and improving the overall efficiency of the payment process.

[0069] Step 208: Control the desktop robot to present interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method.

[0070] The interactive information includes at least one form of interactive information such as body movements, lighting, voice broadcasting, screen display, and projection.

[0071] The main control system of a desktop robot can send instructions to its actuators to perform specific actions or display specific content. These instructions can be generated by the main control system or sent from a server to the desktop robot.

[0072] Interactive information can represent the information conveyed during the interaction between the desktop robot and the user. This interactive information can be any one or more combinations of body movements, lighting effects, voice announcements, screen displays, and projections. Different payment methods can correspond to different interactive information. After determining the target payment method for the user, the desktop robot can present the interactive information corresponding to that target payment method.

[0073] Optionally, the interactive information includes at least one of the following: operational guidance information used to guide the user to use the target payment method before the payment process is executed; prompt information used to indicate that payment is in progress during the payment process; and prompt information indicating the payment result after the payment process is completed.

[0074] The period before payment process execution can be considered the time stage before the payment process officially begins. During this time, the user has not yet initiated the payment operation, and the server has not yet performed resource transfers or other operations for the current transaction. For example, the user has not yet scanned the merchant's payment code, the desktop robot has not yet scanned the payment code provided by the user's terminal, the user's terminal has not yet touched the desktop robot, or the desktop robot has not yet established short-range communication (NFC, etc.) with the user's terminal.

[0075] Operation guidance information can represent information content used to guide users on how to perform payment operations, and may include, but is not limited to: payment method descriptions, operation step guidance, QR code display, NFC sensing area indication, etc.

[0076] The payment process can be a stage in which the payment process is in progress but has not yet been completed. For example, the server has received the processing request for the current transaction and has started to execute steps such as risk detection and resource transfer, but has not yet received the payment processing result.

[0077] The notification message used to indicate that a payment is in progress can be a message that informs the user that the transaction has not yet been completed, thus conveying the current status to the user. For example, it may display words such as "Payment in progress" or "Please wait".

[0078] The completion of the payment process indicates that the server has completed the resource transfer process for the current transaction. This could mean the required resources have been deducted from the payer's account or added to the payee's account, confirming the payment result. The payment result can indicate either successful or failed payment.

[0079] The notification message indicating the payment result can be a message conveying whether the payment was successful or not, and is used to communicate the outcome to the user. This may include text or symbols indicating whether the payment was successful or failed.

[0080] The payment process is a sequential process, and users' information needs differ at each stage. Before payment, users need to know "how to operate," during payment, users need to know the "current status," and after payment, users need to know "what the result is." Designing interactive information in stages can accurately match users' information needs at each stage. In addition, staged interactive information can reduce users' uncertainty and anxiety. For example, during the payment process, users often worry about whether the payment was successful or whether they need to wait; displaying the "paying" status can alleviate user anxiety. Providing timely feedback on the result after payment allows users to clarify their next action, such as leaving or retrying.

[0081] In practical applications, desktop robots can display one or more of the following information based on the payment process: operation guidance information, payment process prompts, and payment result prompts. Alternatively, they can dynamically select the information type based on user profiles, such as analyzing user profile characteristics to determine if the user is a repeat customer or their emotional state. If the user is a repeat customer, pre-payment guidance information can be omitted, and only payment result prompts or prompts during the payment process and payment result prompts can be displayed, simplifying the interaction. If the user is a new customer, all three types of information can be displayed for complete guidance. Furthermore, if the user is impatient, payment process prompts can be skipped, shortening the interaction time.

[0082] By categorizing interactive information according to the time stages of the payment process, the information is presented in a clear chronological order and logical hierarchy, allowing users to judge the current payment progress based on the information and reducing confusion and misoperation.

[0083] A desktop robot can be a desktop robot with movable parts such as a head, arms, legs, and body. Limb movements can represent the actions performed by the robot through motor-driven movable parts, and can include, but are not limited to: head turning, shaking or nodding, arm raising, rotating or opening and closing, leg movement, body tilting, etc.

[0084] Desktop robots can have light signals emitted by one or more LEDs or other light sources. The lighting presentation can represent the light signals emitted by the LEDs or other light sources, and may include, but is not limited to: color changes, brightness changes, flashing patterns, marquee effects, etc.

[0085] Desktop robots can have voice broadcasting modules such as speakers. Voice broadcasting can represent audio content played through speakers, etc., and can include, but is not limited to: pre-recorded audio, text-to-speech, sound effects, etc.

[0086] Desktop robots can have one or more displays. The screen display can show content displayed through the display (such as LCD, OLED, etc.), and can include, but is not limited to, text, images, animations, videos, etc.

[0087] Desktop robots can have components such as projectors. Projection can represent content projected onto external surfaces (such as desktops or walls) via a projector, and can include, but is not limited to, QR codes, operation instructions, game interfaces, etc.

[0088] For example, if the target payment method is determined to be QR code payment, the user needs to scan the merchant's payment code. The desktop robot can project the QR code onto the table below the robot, such as in front of the robot's feet. The user can scan the projected QR code to make the payment. After payment, the desktop robot can also project a message such as "Payment successful." Alternatively, while projecting the QR code, it can also announce "Please scan the projected QR code," and the lights can display a blue breathing effect. The desktop robot's arm can point to the projection area. After successful payment, the lights can turn green, the voice can announce "Payment successful," and the desktop robot's arm can perform a high-five gesture. Furthermore, it can present differentiated interactive information based on user profiles. For example, it can adjust the font or volume according to the user's age group; if the user is determined to be elderly, the projected font or the volume of the voice message can be increased. Or, it can adjust the tone of voice according to the user's gender; for example, a gentle tone can be used for female users, and a humorous tone for male users. Or, it can adjust the interaction rhythm according to the user's mood; if the user is determined to be impatient, the voice duration can be shortened.

[0089] By controlling a desktop robot to present interactive information corresponding to the target payment method, payment decisions can be transformed into user-understandable and actionable instructions. Furthermore, the coordinated use of multi-channel interactive information improves the reliability of information transmission and enriches the user experience.

[0090] The payment technologies described in one or more embodiments of this specification may include, for example, Near Field Communication (NFC), Wi-Fi, 3G / 4G / 5G, POS machine card swiping technology, QR code scanning technology, barcode scanning technology, Bluetooth, infrared, Short Message Service (SMS), Multimedia Message Service (MMS), etc.

[0091] In one or more embodiments of this specification, after receiving a request, the server can generate a QR code and return it to the client. In some embodiments of this specification, the server may return a code value generated by the server based on the received request. The client can map the code value returned by the server to the corresponding QR code and render and display the QR code. Alternatively, the server may directly generate a QR code image based on the received request and return the generated QR code image to the client for the client to display. Furthermore, depending on actual usage needs, the QR code generation process includes, but is not limited to, the above explanation, and the embodiments in this specification do not impose specific limitations. The aforementioned client can be a client in a user terminal, a client in a desktop robot, or a client in a merchant device (such as a cash register).

[0092] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps and does not represent the only possible execution order. The order of some steps may be adjusted according to actual needs, or some steps may be omitted. When the claims involve method steps, changes in the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.

[0093] Figure 2 The method described in the article automatically determines user profile characteristics through sensory data and recommends target payment methods accordingly, reducing the number of manual selection and judgment steps for users and improving the intelligence of payment interaction. Through the coordinated use of one or more channels of interactive information, such as body language, lighting, voice announcements, screen displays, and projection, the payment process is no longer a mechanical information output, but an interactive process with emotional expression, effectively alleviating user waiting anxiety and enhancing the emotional experience.

[0094] based on Figure 2 In addition to the method described herein, this specification also provides some implementation methods of the method, which will be described below.

[0095] To provide more accurate services to users, the method in one or more embodiments of this specification may further include: determining the target direction of the user based on the perceived data; and controlling the desktop robot to move toward the target direction.

[0096] The target direction of the user represents the user's spatial orientation relative to the desktop robot. Once the user's direction is determined, the desktop robot can be controlled to face that direction, allowing for better presentation of interactive information. The target direction can be represented by angle values ​​(such as azimuth θ), coordinate values ​​(such as x / y coordinates), or orientation markers (such as left front, directly front, right front).

[0097] In practical applications, the desktop robot's head and / or torso can be controlled to face the user's target direction. The desktop robot may have a rotatable head structure and may be equipped with components such as a camera, display screen, microphone, and near-field communication module. The desktop robot's torso can represent its main body structure and may house the main control module, battery, and actuators. The desktop robot's torso rotation can be achieved through the movement of its legs or feet, or through the rotation of its base. The desktop robot may have anthropomorphic leg structures or a horizontally rotatable base to achieve torso rotation.

[0098] The desktop robot can first adjust its posture, such as facing the user's target direction, and then present the interactive information corresponding to the target payment method. Alternatively, the desktop robot can gradually adjust its posture until it faces the user's target direction while presenting the interactive information. The process of adjusting the desktop robot's posture and analyzing and determining the interactive information to be presented can be executed sequentially or in parallel. No specific limitations are made here.

[0099] When a user is in different positions, if the desktop robot maintains a fixed orientation, the user may not be able to see the screen content clearly, hear the voice prompts, or experience eye contact. By determining the user's orientation and controlling the desktop robot to face the user, effective communication of interactive information can be ensured. Furthermore, when the desktop robot actively turns towards the user, the user feels that the robot is looking at them; this anthropomorphic behavior can establish an emotional connection, enhancing the naturalness and comfort of the interactive experience. In addition, accurate user orientation positioning ensures that the camera is aimed at the user's face, improving the accuracy and efficiency of facial recognition and feature analysis.

[0100] As one implementation method, the target direction can be determined based on visual data. For example, the camera in the desktop robot can capture images of the current scene; a face detection algorithm can be used to identify the face region from the captured image and obtain the center pixel coordinates of the face bounding box. For example, if the image resolution is 640×360, the center pixel coordinates of the face bounding box are (240, 180); then, based on the center pixel coordinates of the face bounding box and the camera's field of view parameters, the azimuth angle of the user relative to the desktop robot can be determined; the desktop robot can also be rotated according to this azimuth angle so that the camera's optical axis is aligned with the user's direction.

[0101] As another implementation, the target direction can be determined based on voice data. For example, a microphone array in a desktop robot can collect ambient sound; calculate the time delay difference (TDOA) using a sound source localization algorithm (such as GCC-PHAT); calculate the sound source azimuth angle based on the time delay difference and the microphone spacing; and determine this azimuth angle as the target direction of the user. The desktop robot can also be rotated based on this azimuth angle so that its front (e.g., face) faces the user.

[0102] As another implementation, the target direction can be determined based on the fusion of visual and voice data. Optionally, determining the target direction of the user based on the perceived data may include: using the microphone array in the desktop robot to calculate the azimuth angle of the user through a sound source localization algorithm; based on the azimuth angle, controlling the camera of the desktop robot to point towards the user area to perform face detection; and determining the target direction of the user based on the face center coordinates obtained from the face detection.

[0103] A microphone array is a sound acquisition component consisting of multiple microphones arranged in a specific geometric structure. Microphone arrays can be circular arrays, linear arrays, planar arrays, etc.

[0104] Sound source localization algorithms refer to algorithms that calculate the spatial location of a sound source by processing multi-channel audio signals collected by a microphone array. Sound source localization algorithms may include, but are not limited to: Generalized Cross-Correlation with Phase Transform (GCC-PHAT), Multiple Signal Classification (MUSIC), Beamforming, etc.

[0105] Azimuth represents the angle of a sound source on a horizontal plane relative to the front of the desktop robot. It can be expressed in degrees or radians and can range from -180° to +180° or from 0° to 360°.

[0106] Pointing the camera towards the user area indicates controlling the camera's orientation, aligning its optical axis with the area in the user's direction. The user area can be a search range determined based on the azimuth of the sound source, and does not have to be a precise location.

[0107] Face detection can be described as the process of identifying facial regions in an image using face detection algorithms. Face center coordinates represent the position of the face bounding box center in the image's pixel coordinate system, and can be expressed in pixel coordinates.

[0108] A two-stage fusion strategy combining initial sound source localization and fine visual localization achieves low-latency, highly robust user orientation tracking. Sound source localization, as prior information, narrows the visual search area, reducing face detection latency and significantly improving tracking response speed. Visual detection, as a fine localization method, reduces orientation errors, meeting the accuracy requirements for robot eye-tracking.

[0109] Optionally, the azimuth angle can be weighted and fused with the face center coordinates obtained through face detection to determine the target direction of the user.

[0110] As an example, the user's target direction can be determined through a serial two-stage fusion method. Specifically, the microphone array can acquire multi-channel audio signals; calculate the sound source azimuth angle θ_mic using the GCC-PHAT algorithm or other algorithms; then control the camera to rotate to the θ_mic direction; perform face detection within the range of θ_mic ± a preset offset angle (such as 15°, 20°, etc.), and output the face center pixel coordinates (x, y); convert these pixel coordinates (x, y) into a visual azimuth angle θ_cam; and then perform a weighted fusion of θ_mic and θ_cam to obtain the target direction θ_target = α·θ_mic + (1-α)·θ_cam. Here, α is the weight corresponding to the sound source azimuth angle θ_mic. The specific weight value can be set according to actual needs.

[0111] As another example, the user's target direction can be determined through parallel two-level fusion. Specifically, the microphone array and camera can start acquiring data simultaneously; sound source localization and full-image face detection can be performed in parallel; the consistency between the sound source azimuth and the visual azimuth can be compared; if the deviation between the two is less than a threshold (e.g., 20°), the fused output is obtained; if the deviation is too large, the one with higher confidence is selected; and the target direction is output.

[0112] As another example, the target direction for a user can be determined through dynamic weight fusion. Specifically, the confidence level C_mic (e.g., signal-to-noise ratio, signal strength, etc.) for sound source localization can be calculated using voice information collected by a microphone array; the confidence level C_cam (e.g., bounding box confidence, face clarity, etc.) for face detection can also be calculated using image information collected by a camera. The fusion weights are then dynamically adjusted based on the confidence level, with higher confidence levels corresponding to larger weights. For example, the weight α for sound source localization can be determined using the formula α=C_mic / (C_mic+C_cam). The target direction is then determined based on the weights, such as the target direction θ_target=α·θ_mic+(1-α)·θ_cam. This dynamic weight fusion approach can automatically adjust the audio-visual weights according to environmental conditions (e.g., lighting, noise), improving the system's adaptability in complex scenarios.

[0113] As another example, visual-assisted sound source correction can be used to determine the user's target direction. Specifically, an initial azimuth angle θ_mic can be calculated through sound source localization; then, the camera is controlled to point in the direction of θ_mic for face detection; if face detection is successful, the visual azimuth angle θ_cam can be used as the target direction; if face detection fails (e.g., due to insufficient light, face occlusion, etc.), the sound source azimuth angle θ_mic is used as the target direction.

[0114] Successful face detection indicates that the face detection algorithm outputs results that meet preset validity conditions, including but not limited to one or more of the following conditions: detection confidence greater than or equal to a first threshold (e.g., 0.7), face bounding box size greater than or equal to a second threshold (e.g., width 50 pixels), face quality score greater than or equal to a third threshold (e.g., 0.6), face detection in N consecutive frames (e.g., 3 frames), deviation of face position from sound source azimuth angle not exceeding a fourth threshold (e.g., 20°), and detection of only one face. All or some of these conditions may be met, and the specific combination of conditions can be adjusted according to the actual application scenario.

[0115] In one or more embodiments of this specification, user profile features may include at least one of user person profile features, user device profile features, and user environment features; the user person profile features include at least one of age features, gender features, emotional features, voice features, behavioral posture features, and historical transaction features; the user device profile features include at least one of device type features and device function features; the user environment features include at least one of queue length features and ambient light intensity features.

[0116] Among them, user profile features can represent a set of features describing the user's attributes and state, including physiological, psychological, and behavioral attributes. User device profile features can represent a set of features describing the attributes and functions of the payment device used by the user, including the features of the hardware device held by the user (such as a mobile phone, watch, etc.). User environment features can represent the contextual features of the user's current scene and physical environment.

[0117] Age features can represent a user's age group (e.g., child, youth, elderly). Age features can be discrete labels (e.g., <18, 18-30, 30-50, >50; or child, youth, elderly, etc.) or continuous values ​​(e.g., 25 years old). Age features can be used to adapt to the operational capabilities and payment preferences of users in different age groups. Optionally, if the user's estimated age is less than the first age, or if the user is estimated to be a child, the target payment method can be the primary scan payment method, indicating payment via scanning a QR code; if the estimated age is greater than the second age, or if the user is estimated to be elderly, the target payment method can be the secondary scan payment method, indicating payment via presenting a payment code; if the estimated age is between the first and second ages, or if the user is estimated to be a youth, the target payment method can be the near-field payment method, indicating payment via contactless communication; where the first age is less than the second age. The specific values ​​for the first and second ages can be determined by statistical analysis of historical transactions, or they can be set based on expert experience. For example, the first age can be 18 years old, the second age can be 50 years old, or other values ​​can be used. For example, based on facial image data and / or voice data obtained by a desktop robot, an age estimation model can be used to estimate the user's age characteristics.

[0118] Gender characteristics can be attributes describing a user's gender. These characteristics can be binary labels (e.g., male / female) or multi-category labels (e.g., male / female / unknown). Gender characteristics can be used to adjust interaction tone, visual style, etc. For example, if the user is female, the desktop robot's voice tone can be gentler; the lighting color can prefer warm tones (e.g., pink, orange), etc. For instance, based on facial image data and / or voice data acquired by the desktop robot, a gender recognition model can be used to identify the user's gender characteristics.

[0119] Emotional characteristics can be attributes that describe a user's current emotional state. These characteristics can include, but are not limited to, calmness, pleasure, impatience, confusion, and anxiety. Emotional characteristics can be used to adjust the pace and content of interactions. For example, if a user is impatient, shorter payment methods can be prioritized, such as near-field communication; voice prompts can be shortened or skipped; lights can flash rapidly as cues; and body language can be simplified. For instance, facial image data and / or voice data acquired by a desktop robot can be used to identify the user's emotional characteristics using an emotion recognition model.

[0120] Voice features can represent a set of features describing a user's voice attributes. Voice features may include, but are not limited to, speech rate, timbre, volume, language type, and accent. Voice features can be used to assist in identity recognition and interaction adaptation. For example, voice features can help infer a user's age and thus determine the appropriate payment method; similarly, voice features can be used to analyze whether a user is in a hurry. If a user expresses a need to save time, the system can prioritize recommending near-field payment methods that take less time, adjust the speech rate to match the user's needs, or skip some prompts. For instance, based on voice data acquired by a desktop robot, a voice recognition model can be used to identify the user's voice features, such as speech rate and timbre.

[0121] Behavioral posture features can represent attributes describing a user's body movements and postures. These features can include, but are not limited to, holding objects, arm posture, body orientation, and hand gestures. Behavioral posture features can be used to analyze a user's willingness to pay, such as whether the user wants to pay or which payment method they prefer. For example, if a user is holding a phone with the screen off or displaying the main interface, it can be inferred that the user wants to use near-field payment; similarly, if a user is holding a phone with the screen lit up and displaying a payment code, it can be inferred that the user wants to use a QR code payment method. Furthermore, if a user is carrying shopping bags, holding a child, or has difficulty operating the device, it can be inferred that the user wants to use facial recognition payment. For instance, visual data acquired by a desktop robot can be used with a behavioral analysis model to identify user behavioral characteristics.

[0122] Historical transaction features can represent attributes describing a user's historical payment behavior and transaction records. User identity is matched using facial recognition or voiceprint recognition, and historical transaction data is queried from the cloud server. Historical transaction features may include, but are not limited to: frequently used payment methods, purchase frequency, purchase amount, preferred product categories, membership level, etc. Historical transaction features can be used to respect user habits and provide personalized recommendations. If a user frequently uses a payment method that involves scanning a payment code, that payment method can be prioritized, and the desktop robot can display the corresponding interactive information for that payment method.

[0123] Device type characteristics can represent attributes describing the type of hardware device a user is using for payment. Device type characteristics can include, but are not limited to: smartphones, smartwatches, smart rings, tablets, feature phones, etc. Device type characteristics can be used to recommend suitable payment methods. If a user is using a smartwatch, it can be inferred that the target payment method is near-field payment, and the desktop robot can display the corresponding interactive information for near-field payment.

[0124] Device functional characteristics can represent a set of functional attributes that describe the user's payment device supports. Device functional characteristics may include, but are not limited to: Bluetooth functionality, NFC functionality, camera functionality, and a display screen. Device functional characteristics can be used to determine the device's payment capabilities and the payment methods it supports. For example, if the device does not support NFC, near-field payment can be excluded, and the target payment method can be either active scanning or passive scanning.

[0125] Queue size can represent the number of users currently waiting to pay, specifically the number of users waiting to pay within the desktop robot's field of vision or at the checkout counter, and can be used to determine the urgency level. For example, face detection can be performed using captured image information to count the number of faces in the current frame or consecutive frames, thereby determining the queue size. The desktop robot's camera component can count the number of faces within its field of vision in real time.

[0126] For example, when the number of people in the queue exceeds a preset threshold (e.g., 5 or 7 people), it can be determined that the current situation is a peak congestion scenario. The desktop robot can automatically adjust its recommendation strategy, prioritizing the display of the payment method with the shortest processing time (e.g., NFC contactless payment). It can also guide users to operate quickly by performing pointing gestures using its body module, and simplify the voice broadcast content or skip marketing information to reduce user waiting time and improve passage efficiency. For instance, if there are many people in the queue, it can prioritize recommending contactless payment methods with faster payment efficiency, and can also use interactive methods such as faster speech, simplified guidance, and skipping marketing to reduce user waiting time. If there are few people in the queue, it can prioritize recommending the use of main scan payment or scan-based payment, or it can display some benefits information before, after, or during payment. Or, if there is only one user waiting to settle, the desktop robot can display some information for the user to claim benefits, such as deducting points through interactive games or claiming coupons for payment, and then present information related to the payment method.

[0127] Ambient light intensity characteristics can represent the brightness of light in the surrounding environment and can be used to adjust the presentation. For example, ambient light conditions can be determined using an ambient light sensor, or the ambient light conditions can be estimated using the average brightness value of a captured image.

[0128] For example, if the lighting is strong, payment methods that require users to show their payment code or contactless payment methods can be prioritized. The screen or projector brightness of the desktop robot can be adjusted to a higher level to enhance visual contrast, make colors vibrant, and ensure that content is clearly visible without glare interference.

[0129] If the lighting is moderate, the probability of recommending each payment method can be the same based on the light intensity dimension. Other dimension feature information can be combined to determine the target payment method to be recommended. The screen or projection brightness of the desktop robot can be adjusted to a moderate state to maintain a visually comfortable state.

[0130] If the lighting is weak, such as at night or in a dark area, based on the dimension of light intensity, the probability of contactless payment methods that are not affected by light can be increased, and users should be given priority to use contactless payment methods; the brightness of the desktop robot's screen or projector can be reduced, warm colors can be used, and the voice broadcast volume can be increased to reduce glare.

[0131] The characteristics of the user's environment can also include network status information, such as the network status of the merchant where the desktop robot is located. If the merchant's network status is poor, the main payment method should be recommended first, and the payment process should be triggered by the user's terminal to avoid the user waiting for a long time for the merchant to process the payment.

[0132] In practical applications, user profile features can include multi-dimensional features, and the target payment method can be determined by fusing multiple profile features. For example, weights can be assigned to each dimension of the feature, and the target payment method can be determined by calculating the combined weights. Alternatively, multi-dimensional features can be provided to a recommendation model, which can then combine these features to determine the recommended target payment method.

[0133] In one or more embodiments of this specification, the desktop robot may present different interactive information for different payment methods. Optionally, controlling the desktop robot to present interactive information corresponding to the target payment method may include: if the target payment method indicates that the user pays by scanning a QR code, then controlling the desktop robot to display the merchant's QR code image on the screen or by projection; or, if the target payment method indicates that the user pays by presenting a payment code, then controlling the desktop robot to display a prompt message prompting the user to present the payment code on the screen or by projection; or, if the target payment method indicates that the user pays through contactless communication, then controlling the desktop robot to display a prompt message prompting the user to touch the desktop robot to make the payment on the screen or by projection.

[0134] These methods include: 1) Payment via scanning a QR code, also known as active scanning payment, where the user scans the merchant's QR code using their mobile phone's camera. 2) Payment by presenting a payment code, also known as passive scanning payment, where the user displays the code on their mobile phone or other mobile device, which is then scanned by the merchant's scanning device, such as a desktop robot's camera. 3) Payment via contactless communication, also known as near-field communication (NFC) payment, which utilizes short-range communication technologies such as NFC, Bluetooth, and UWB.

[0135] Desktop robots can be equipped with components such as screens and / or projectors, which can project different interactive information. For active scanning payment methods, the desktop robot can display the merchant's QR code image on the screen or projector, or it can display prompts for the user to scan the payment code, such as "Please scan the QR code." For passive scanning payment methods, the desktop robot can display prompts for the user to present their payment code on the screen or projector, such as "Please present the QR code." For near-field communication (NFC) payment methods, the desktop robot can display prompts for the user to touch the robot to make the payment, such as "Tap me," "Tap my head to pay," or "Tap my body to pay."

[0136] In practical applications, a desktop robot can include either a screen component or a projector component, or both. Optionally, if the desktop robot includes both a screen component and a projector component, the content displayed on the screen and the content projected on the projector can be the same or different. For example, for the main QR code payment method, to facilitate user scanning, the desktop robot's screen can display a prompt message such as "Please scan the QR code," while the projector can project the merchant's payment code; or, both the screen and the projector can display the merchant's payment code. Furthermore, for different payment methods, the desktop robot's screen can display the same content, such as a welcome message, advertisement, or mini-game, or it can display a welcome message, advertisement, or mini-game tailored to user preferences; while the projector can display different interactive information for different payment methods. Similarly, for different payment methods, the desktop robot's projector can display the same content, such as a welcome message, advertisement, or mini-game, or it can display a welcome message, advertisement, or mini-game tailored to user preferences; while the screen can display different interactive information for different payment methods. There are no restrictions on the specific display method or content here.

[0137] Optionally, if the desktop robot includes a screen component and a projector component, the component to be used can be selected based on the ambient light intensity. For example, if the light intensity is greater than or equal to a first preset intensity, such as greater than 500 lumens, the projected content may be unclear, so screen display can be selected first; if the light intensity is less than or equal to a second preset intensity, such as less than 200 lumens, the screen may reflect light, so projection can be selected first.

[0138] By mapping payment methods to interactive information, users can clearly understand the operation to be performed, reducing operational errors and payment failures caused by unclear guidance. Furthermore, clear interactive guidance reduces user thinking and hesitation time, shortening payment processing time. Additionally, the differentiated interactive content design for different payment methods demonstrates the desktop robot's professional understanding of the payment process, enhancing user trust.

[0139] Desktop robots can also be humanoid robots, capable of displaying different body movements. As one implementation, controlling the desktop robot to display interactive information corresponding to the target payment method can include: if the target payment method indicates that the user is paying by scanning a QR code, then controlling the desktop robot's arm to point towards the area displaying the merchant's QR code image; or, if the target payment method indicates that the user is paying by presenting a payment code, then controlling the desktop robot to perform a head-raising motion, capturing the payment code presented by the user through an image acquisition module located on its head; or, if the target payment method indicates that the user is paying via contactless communication, then controlling the desktop robot's arm to point towards the desktop robot's contactless communication area.

[0140] The area displaying the merchant's QR code image can represent the physical location where the QR code image is displayed, such as the area on the screen or the desktop area projected onto.

[0141] The contactless communication area can refer to the sensing area on a desktop robot that supports short-range communication such as NFC, and can be located on the top of the head, chest, or arm.

[0142] Desktop robots can be driven by arm motors to point the ends of their arms or fingers toward a specific direction or area. For example, the arms can rotate 45° to point toward a projected QR code area; or the arms can be raised vertically to indicate that the user should touch the top of the desktop robot.

[0143] The head-raising motion indicates that the desktop robot rotates its head upwards at a certain angle, driven by a head motor. The image acquisition module located on the head refers to a camera or sensor module installed on the robot's head. Users can place the payment code image displayed on their mobile phones or other user terminals into the acquisition area of ​​the image acquisition module so that the desktop robot can acquire the user's payment code for payment.

[0144] When the user initiates a payment scan, their arm points towards the QR code area; when the user is being scanned, they look up to adjust the camera angle; and during near-field payments, their arm points towards the sensor area. These gestures provide clear functional guidance rather than simply emotional expression. Pointing with the arm is more intuitive than text or voice prompts, allowing users to quickly understand the operation location and reducing cognitive load. Raising the head during a payment scan ensures the camera is aligned with the user's phone height, optimizing the scanning angle and improving QR code recognition success rate.

[0145] The desktop robot can also provide different interactive information at different stages of the payment process. For example, before payment, the desktop robot can present pre-payment guidance information. For instance, if the target payment method is determined to be a primary scan-to-pay method based on the user's profile characteristics, the desktop robot can project a QR code image of the merchant's payment code and also announce prompts such as "Please scan the projected QR code," while the lights can display a blue breathing effect. If the target payment method is determined to be a near-field payment method, the desktop robot can project a prompt asking the user to tap the screen, and also announce prompts such as "Please tap the halo area above your head with your phone," and can also perform a gesture of raising both hands vertically, while the lights can display a blue breathing effect. If the target payment method is determined to be a passive scan-to-pay method, the desktop robot can project a prompt asking the user to show the QR code, and can also announce prompts, while the lights can display a blue breathing effect.

[0146] After a user interacts with the desktop robot for payment, such as after the user presents their payment code to the robot for scanning, after the user taps the robot to obtain NFC tag information, or after the user scans the payment code displayed on the robot to confirm the transaction, the desktop robot can also display interactive information. The interactive information differs depending on the payment method, or some interactive information may be the same. For example, for the main scanning payment method, the desktop robot can project a QR code image and announce "Payment in progress, please wait a moment." The robot's arm can also rotate 45° to point at the QR code and display a marquee light effect. For near-field communication (NFC) payment, the desktop robot can project a tap prompt and announce "Payment in progress, please wait a moment." The robot can also perform actions such as nodding and swinging its arm back, and display a marquee light effect. For the scanned payment method, the desktop robot can project a prompt asking the user to present their QR code and announce "Payment in progress, please wait a moment." The robot can also perform actions such as looking up, and display a marquee light effect.

[0147] After the current transaction with the user is completed, the desktop robot can also present feedback information after payment. For example, after successful payment, it can provide one or more of the following: personalized blessing animation, voice thanks, body celebration, projected content display, lighting, and brand incentive push notifications. If the payment fails, it can trigger reassurance and guidance for retry. For example, the desktop robot can keep the projected information unchanged and announce the payment success or failure via voice. Alternatively, if the payment is successful, it can perform a hand opening and closing and shoulder rotation similar to a high-five. Alternatively, if the payment fails, it can perform a head tilt or arm drooping action, and can also display or play a retry prompt. It can also present different lighting effects than before and during payment. For example, if the payment is successful, it can display a green light for 3 seconds, and if the payment fails, it can display a red light for 3 seconds.

[0148] It is understandable that the above interactive information is only an example, and specific interactive content can be set according to actual needs, without specific limitations here.

[0149] To improve device utilization and to draw users' attention to the payment-enabled desktop robot, thereby increasing user engagement and providing richer services, the desktop robot can display information that attracts users or disseminates public welfare knowledge or other information when not processing payment transactions. As one implementation method, the method in one or more embodiments of this specification may further include: if no user interacting with the desktop robot is detected within a first preset time period, the desktop robot displays non-payment interaction information.

[0150] The first preset duration can represent a preset time threshold used to determine whether the desktop robot is idle. For example, 3 minutes, 5 minutes, 8 minutes, 10 minutes, etc. The first preset duration can be a fixed value or a threshold that dynamically changes based on environmental factors. For example, the first preset duration can be dynamically adjusted based on at least one of the following factors: the current time period, including peak or off-peak hours; the date type, including weekdays, weekends, or holidays; the store type, including fast food restaurants, full-service restaurants, or retail stores; ambient light intensity; ambient noise level; the number of users currently present; the robot's remaining battery power; promotional activity status; queuing status, etc. For example, during peak lunch hours (e.g., 11:30-13:30), the first preset duration can be set to 2 minutes to quickly enter the customer attraction state and attract queuing customers; during off-peak afternoon hours (e.g., 14:30-17:00), the first preset duration can be set to 10 minutes to avoid frequent switching and disturbing a small number of customers. Furthermore, when the robot's battery power is below a threshold, such as 20%, the first preset duration can be extended to 15 minutes to reduce unnecessary interactions and save battery power. For example, when a promotional activity is detected in the store, the first preset duration is shortened to 1 minute to increase the frequency of exposure of the traffic-driving information; or, when the camera detects that the number of people queuing in front of the cashier exceeds the preset number, such as 5 people, the first preset duration can be shortened to 40 seconds in order to quickly attract customers' attention and improve turnover efficiency.

[0151] If no user interaction is detected, it means that the desktop robot's perception module has not recognized the target object or target event within a specified time. For example, no face appears in the captured image, or no user voice information is collected.

[0152] The specific content of non-payment interactive information may include at least one form of information such as: body movement sequences (e.g., dance moves), lighting effects (e.g., color changes, flashing patterns), voice content (e.g., greetings, music), screen display content (e.g., animations, images), and projected content (e.g., interactive game interfaces). Specific information content can be personalized promotions, brand interactions, games, event invitations, and other multimedia content.

[0153] By presenting non-payment interactive information during idle periods, the utilization rate of desktop robot devices was improved, and the attention of potential users was attracted. At the same time, diverse interactive content was provided without interfering with the payment process, enriching the user experience.

[0154] As one implementation, the desktop robot presenting non-payment interaction information may include: if the number of times a face is detected within the visual range of the desktop robot within a second preset time period after the first preset time period is greater than or equal to a preset number of times, then the desktop robot performs a preset dance action; the second preset time period is less than the first preset time period.

[0155] The second preset duration represents the time window used to count the number of face detections after the first preset duration ends. The second preset duration can be used to determine the current pedestrian traffic level and serves as the sampling period for face detection. A shorter second preset duration allows for quick assessment of pedestrian traffic, enabling users to promptly see the robot's interactive feedback and improving the responsiveness and naturalness of the interaction; a longer first preset duration can be used to confirm idle status.

[0156] The visual range refers to the spatial area in which the image acquisition modules of a desktop robot, such as its camera, can effectively detect human faces. The visual range can be determined by factors such as the camera's field of view (FOV), installation height, and detection algorithm.

[0157] The number of detected faces can represent the cumulative number of times the face detection algorithm successfully identified faces within a second preset time period. The number of detections can be the number of unique faces, which is the number of people after deduplication; or it can be the number of detection frames, which is the number of image frames in which faces were detected.

[0158] The preset number of times represents a threshold for detecting faces in high-traffic areas. When the number of detected faces is greater than or equal to the preset number, it is considered a high-traffic area, and the dance action can be triggered. The preset number of times can be a pre-configured parameter, such as 3, 5, or 10 times. The specific value of the preset number of times can be a fixed value or a dynamically adjusted threshold. For example, the preset number of times can be dynamically adjusted according to time periods (such as peak and off-peak periods); the preset number of times can be reduced during peak periods (e.g., set to 3 times) and increased during off-peak periods (e.g., set to 7 times).

[0159] Preset dance choreography can represent pre-configured dance sequence stored in the robot's memory. Dance choreography is a special form of physical movement, involving coordinated movements of multiple joints, and possessing rhythm and visual appeal. An example of dance choreography (30-second sequence): 0-5s: alternating left and right arm movements performing a wave motion ±30°; 5-15s: leg movements, with leg rotation ±15°; 15-25s: a combination of nodding and shaking the head; 25-30s: a fixed pose with both arms raised. This is just one example; it can be customized according to actual needs, and is not limited here.

[0160] By counting the number of faces within a second preset time period after the first preset time period, and determining whether to execute a dance move based on the comparison between the number of faces and the preset number, an intelligent interactive strategy is implemented in high-traffic scenarios. From the user's perspective, when multiple users pass by the robot simultaneously, the dance moves can provide entertainment for waiting users, alleviating the boredom of the waiting process; the visual appeal of the dance moves can attract users' attention, giving users a relaxed and pleasant experience while waiting to pay or passing by; and by executing the dance moves during idle periods, it avoids disturbing users who are already paying, ensuring the smoothness and privacy of the payment process.

[0161] As one implementation, the desktop robot presenting non-payment interactive information may include: if the number of times a face is detected within the visual range of the desktop robot within a second preset time period is less than the preset number, then the desktop robot launches an interactive game; if a user is detected participating in the interactive game, then the desktop robot performs a preset physical action.

[0162] Deploying interactive games can mean that the desktop robot presents game content available for user participation through projection, screen display, or other means. Interactive games can take various forms, including but not limited to game dot displays, touch games, motion-sensing games, and question-and-answer games. User participation in interactive games can mean that users interact with the game content through touch, clicks, gestures, voice, or other means. Preset body movements can refer to pre-configured sequences of body movements stored in the robot's memory, which can include movements of movable parts such as head rotation, arm raising, and body tilting.

[0163] For example, in a light-dot touch game, the projection module can be controlled to project multiple light dots onto the desktop, such as projecting 5 random light dots (e.g., 5cm in diameter); then, it can detect whether the user touches the light dots, such as by using a camera to detect the click location through background subtraction; if a touch is detected, a preset body action feedback is executed, such as the desktop robot's arm performing a high-five action if the click is successful.

[0164] For example, in a shape elimination game, the projection module can be controlled to project multiple shapes (such as stars, circles, squares, etc.) onto the desktop; then, it can detect whether the user clicks on a shape; if the clicked shape disappears, the robot performs a body movement feedback; if all shapes are eliminated, a celebration action is played.

[0165] For example, in motion-tracking games, the screen can be controlled to display the game interface (e.g., the user's gestures control the movement of objects on the screen); the user's gestures can be detected through a camera or other sensing units; the game state can be updated based on the gestures, and the robot can perform limb movements as feedback.

[0166] For example, in a question-and-answer interactive game, the question can be read aloud (e.g., "How's the weather today?"); the user's voice response can be detected; and then different physical actions can be performed to provide feedback based on the response.

[0167] By launching an interactive game when the number of facial recognitions falls below a preset threshold within a second preset time period, and executing preset physical actions upon detecting user participation, a deep interaction strategy was implemented in low-traffic scenarios. From the user's perspective, interactive games enhance the autonomy and enjoyment of the interaction; when users participate, the robot performs physical actions as feedback, making users feel noticed and responded to, thus increasing the pleasure of the interaction. Furthermore, executing game interactions during low-traffic periods allows users ample time to participate, preventing interruptions due to queuing pressure and ensuring the integrity and comfort of the interaction.

[0168] Considering that ambient light affects the visibility of interactive information, and to reduce unnecessary resource waste, the method in one or more embodiments of this specification may further include: acquiring ambient light information. The desktop robot presenting non-payment interactive information may include: if the light intensity of the ambient light information is greater than or equal to a preset intensity, then the desktop robot presents non-payment interactive information; if the light intensity of the ambient light information is less than the preset intensity, then the desktop robot enters a sleep state.

[0169] Ambient lighting information represents the light intensity data of the environment surrounding a desktop robot, and can include parameters such as light intensity, light direction, and light color temperature. The desktop robot can collect lighting information through sensors or receive it from external devices.

[0170] Illumination intensity represents the brightness value of ambient light, reflecting the brightness of the environment. Preset intensity represents a pre-configured illumination intensity threshold used to determine whether the environment is suitable for displaying non-payment interactive information. Preset intensity is a pre-set parameter, such as 50 Lux or 100 Lux.

[0171] Entering sleep mode indicates that the desktop robot is running at reduced power or pausing non-core functions. It is a low-power mode, not a complete shutdown, and the desktop robot can still monitor wake-up conditions (such as payment triggers, sound triggers, etc.). Optionally, the desktop robot can also be automatically woken up after entering sleep mode. For example, the light sensor can periodically detect if the light intensity is greater than a preset intensity, and then resume normal interaction; the payment module can continuously monitor and wake up and initiate the payment process if a payment request is detected, such as an NFC radio frequency signal emitted by a card reader device; the microphone array can monitor and wake up if user voice or a loud sound is detected; the touch sensor can monitor and wake up if a user touches the robot; the internal clock can trigger wake-up, such as waking up once at a preset time (e.g., every hour); or the communication module can wake up if it receives a remote command from the cloud or an administrator.

[0172] For example, ambient light information can be detected using a light sensor. A desktop robot could have a light sensor (such as a photodiode or ambient light sensor) to collect the current ambient light intensity value; then compare it with a preset intensity. If the light intensity is greater than or equal to the preset intensity, a non-payment interaction information presentation process is executed; if the light intensity is less than the preset intensity, the desktop robot enters a sleep state. Alternatively, during sleep, it can be periodically woken up to detect changes in light intensity.

[0173] For example, ambient lighting information can be determined through camera image analysis. A desktop robot can use a camera to capture environmental images; then analyze the image brightness values, converting them into estimated light intensity values; then compare these estimated light intensity values ​​with a preset intensity; and based on the comparison result, decide whether to present interactive information or enter a sleep state. Light sensors and cameras can be installed on the top of the desktop robot, such as on its head, to avoid the robot itself blocking ambient light and ensure detection accuracy. Alternatively, they can be installed on the front of the desktop robot, such as near the robot's face or screen, to detect lighting conditions in the user's field of vision. They can also be installed on the sides of the desktop robot, such as on both sides of the robot's body, to detect ambient lighting from multiple directions. Alternatively, they can be distributed across multiple parts of the desktop robot, such as placing multiple light sensors at multiple locations on the robot, and taking the average or maximum value as the ambient light intensity. For example, time periods can also be used to assist in the judgment. For example, it can obtain the current time information (such as 14:00, 23:00, etc.); it can also obtain the ambient light intensity value; if the current time is daytime and the light intensity is greater than or equal to the preset intensity, interactive information is displayed; if the current time is nighttime or the light intensity is less than the preset intensity, it enters a sleep state. Time conditions can be used as an auxiliary verification of light conditions.

[0174] For example, the preset intensity value can also be dynamically adjusted. For instance, the preset intensity can be dynamically adjusted according to the scene type (e.g., 50 Lux indoors in a shopping mall, 200 Lux outdoors, etc.), and the scene type can be manually configured or automatically identified.

[0175] By acquiring ambient lighting information and comparing the light intensity with a preset intensity, the system determines whether to present non-payment interactive information or enter a sleep state, achieving intelligent interactive control based on environmental conditions. From the user's perspective, presenting interactive information in well-lit conditions ensures that users can clearly see the projected content, dance movements, etc., avoiding confusion caused by insufficient light making the interactive content invisible; the sleep state avoids unnecessary light pollution or noise interference during unoccupied periods (such as at night), providing a quiet and comfortable experience for the surrounding environment. In addition, light intensity detection provides an objective environmental judgment basis for the presentation of interactive information, avoiding blind execution of interactive actions; the sleep state can reduce robot power consumption, extend equipment lifespan, and improve system reliability.

[0176] Desktop robots can interact with user terminals to process transactions as payment devices. For example, a desktop robot can act as a cash register to perform processes such as payment collection; or, a desktop robot can also act as an auxiliary payment device to work with a cash register to complete the transaction process.

[0177] In one implementation, the desktop robot can communicate with the merchant's POS device via wired or wireless means. The method in one or more embodiments of this specification may further include: if the desktop robot collects an image of a payment code presented by a user, it sends the payment code image or the payment code number obtained by parsing the payment code image to the merchant's POS device, so that the merchant's POS device can execute the payment processing procedure; or, if the desktop robot sends tag information for triggering the payment process to the user terminal via contactless communication, the desktop robot sends payment identification information for determining the payment account of the payer user on the user terminal to the merchant's POS device, so that the merchant's POS device can execute the payment processing procedure.

[0178] Merchant POS equipment refers to the equipment used by merchants to complete payment transactions, such as POS machines, cash registers, POS system servers, self-service checkout devices, and smart vending machines. User terminals refer to payment devices held by users, such as smartphones, smartwatches, and tablets.

[0179] Desktop robots can serve as auxiliary devices to help cashiers complete transactions; they can also be called auxiliary payment devices or auxiliary transaction devices. This allows merchants who previously did not support transactions based on short-range communication to now have the capability to conduct transactions via short-range communication. For the merchant, no modifications are required to their existing equipment, enabling them to conduct transactions via short-range communication at a low cost, or even without any additional cost.

[0180] Wired communication connections refer to connections that transmit data via physical cables, such as Ethernet cables, USB cables, and serial cables. Wireless communication connections refer to connections that transmit data via wireless signals, such as Wi-Fi, Bluetooth, ZigBee, and 4G / 5G.

[0181] A payment code image can represent an image such as a QR code or barcode displayed on a user's terminal screen, which can be used to identify the user's payment account. The payment code number can represent the numerical or string information obtained after parsing the payment code image (such as QR code decoding). The payment code number is the essential content of the payment code image, with a smaller data size and higher transmission efficiency.

[0182] The tag information may contain information used to trigger the transaction process. The tag information can be a string conforming to communication rules or protocols, such as a link string or other types of strings. As one implementation, the tag information may include the identification information of the payment application; after obtaining this information, the user terminal can launch the corresponding payment application. The tag information may also include business information indicating the payment transaction, such as a transaction token. After obtaining this information, the user terminal or the launched payment application can execute the corresponding business processing flow, such as executing a preset process to send confirmation information to the server, sending other instructions to the server, or accessing the server to obtain information issued by the server. The tag information may also include merchant-related information, such as merchant ID, merchant device ID, etc., or transaction-related information, such as information indicating the transaction amount.

[0183] Tag information can be pre-sent by the server to the desktop robot, for example, before the desktop robot receives the amount to be paid from the merchant's device or before the merchant's device performs any settlement-related operations (e.g., the merchant's device is in standby mode). Alternatively, the tag information can be sent by the server to the desktop robot after the merchant's device determines the amount to be paid; or it can be partial or template information containing tag information within the desktop robot, which can be concatenated or assembled according to local preset rules before the desktop robot interacts with the user terminal or based on user terminal interactions. Tag information can be reusable, with multiple transactions corresponding to the same tag information. In this case, once a transaction is completed, the correspondence between that transaction and the tag information can be released or invalidated, thus not affecting the execution of subsequent transactions. Alternatively, tag information can be non-reusable, with one transaction corresponding to one tag information. For each interaction, the desktop robot can send new tag information to the user terminal, with different transactions corresponding to different tag information. The server can also record the correspondence between tag information and the desktop robot, or the correspondence between the desktop robot and the merchant's device.

[0184] The confirmation message can be automatically sent by the user terminal after obtaining the tag information, based on user authorization. For example, after obtaining the tag information, the user terminal can launch the payment application corresponding to that tag information. The payment application can then parse at least part of the information in the tag information and automatically send a confirmation message to the server.

[0185] Alternatively, the confirmation message can be sent based on the user's confirmation action. For example, after obtaining the tag information, the user terminal can interpret the content of the tag information. For instance, if the tag information contains an identifier or string that triggers the display of the payment confirmation page, the user terminal can display a payment confirmation page containing information such as the amount to be paid. This page can contain operation controls. If the control is a confirmation control, after the user performs an operation on the control, the user terminal can send a confirmation message to the server.

[0186] Alternatively, the tag information obtained by the user terminal may contain the identification information of a payment application with payment function. After obtaining the tag information, the user terminal may display a prompt page to launch the payment application. This page may contain operation controls. If the operation control is a control that indicates agreement to launch and authorize the payment application, the user terminal may automatically launch the payment application after the user performs an operation on the control. The launched payment application may send confirmation information to the server.

[0187] The confirmation message may also include user identification information (such as UID), digital signature, timestamp, etc. In practice, once the transaction corresponding to the confirmation message has been processed, the server can delete the confirmation message, mark it as invalid, or mark it as processed. The server will not execute any further transactions based on this confirmation message.

[0188] Payment identification information can be a string of a preset length. For example, it can contain one or more characters such as numbers, symbols, etc., with a preset number of digits. Payment identification information can trigger the merchant's POS device to execute the payment process, such as the POS device triggering the acquiring process after obtaining the payment identification information. The acquiring process can also represent the payment collection process, which can request the server to execute a resource transfer process for the current transaction, such as deducting funds from the payer's account on the user's terminal and adding funds to the payee's account on the merchant's device.

[0189] The payment identification information may contain specific characters that conform to the payment communication protocol, such as a string starting with "28". After obtaining the payment identification information, the merchant's POS device can determine that the payment identification information is used for payment and can execute the corresponding acquiring process. As one implementation method, the payment identification information may have the same starting character as the payment code numbers, and / or, the number of characters in the payment identification information may be the same as the number of characters in the payment code numbers.

[0190] The starting character can represent the first character or a character at a predetermined position starting from the first character. For example, the first character of the payment identifier and the payment code numbers are the same; or the first two, three, or four characters of the payment identifier and the payment code numbers are the same. For instance, both the payment identifier and the payment code numbers are strings starting with "28", or both are strings starting with "13", etc.

[0191] The number of characters in a payment code can refer to the total number of characters it contains, such as 18 characters, 17 characters, etc. Alternatively, it can also refer to the number of characters in various formats, such as the number of numeric characters, alphabetic characters, or symbolic characters.

[0192] For example, the number of characters in the payment identifier information can be the same as the number of characters in the payment code. For instance, the payment identifier information can be 18 characters, 17 characters, or 13 characters, etc. Alternatively, the payment identifier information can be an 18-, 17-, or 13-character string of pure numbers. Or, the payment identifier information can contain a portion that consists of 5 consecutive numbers, or a portion that consists of 3 consecutive letters, etc.

[0193] For scan-to-pay methods, for example, a desktop robot can guide the user to present their payment code, such as displaying a prompt message like "Please present your payment code" on the screen or projection. Then, the desktop robot's camera captures an image of the user's payment code and sends the captured image to the merchant's POS device, or it parses the payment code image to obtain the payment code number and sends the payment code number to the merchant's POS device. After receiving the payment code image or payment code number, the merchant's POS device can call the payment interface to execute the transaction processing. After processing the transaction, the server can also provide feedback to at least one of the POS device, the desktop robot, and the user's terminal.

[0194] For near-field payment methods, for example, a desktop robot can guide the user close to the NFC sensing area. As an NFC passive device, the desktop robot's NFC module can respond to the radio frequency signal emitted by the user terminal and send tag information back to the user terminal, triggering the payment process. After the user terminal completes payment authorization or confirmation, the server can send payment identification information to the desktop robot, or the user terminal can send payment identification information to the desktop robot. The desktop robot can then send the payment identification information to the merchant's POS device, or the server can send the payment identification information to the merchant's POS device. After obtaining the payment identification information, the merchant's POS device can call the payment interface to execute the acquiring process. After processing the transaction, the server can also send result information back to at least one of the POS device, the desktop robot, and the user terminal.

[0195] For the main payment method, for example, the desktop robot can pre-store the merchant's payment code information, or the merchant's POS device can send the payment code to the desktop robot after settling the payment for the goods. The desktop robot can project or display the merchant's payment code so that the user's terminal can scan the code to make the payment.

[0196] As one implementation method, after determining that the target payment method is a certain payment method, the desktop robot can start the relevant process or module of that payment method, while the functional modules of other payment methods can be in a non-working or standby state. For example, if the target payment method is determined to be a near-field payment method such as NFC, the desktop robot may not project or display the merchant's payment code; or if the target payment method is determined to be a scan-based or active scan-based payment method, the near-field communication module of the desktop robot may be in a non-working or standby state.

[0197] As another implementation method, after determining that the target payment method is a specific payment method, the desktop robot can initiate the relevant processes or modules for that payment method, and the functional modules for other payment methods can also be operational. For example, if the target payment method is determined to be a scan-based or active scan payment method, the desktop robot's near-field communication (NFC) module can be operational. This way, if the user does not use the scan-based or active scan payment method but instead uses a tap-to-pay NFC method, the desktop robot can still respond normally to the user's terminal interaction, and the user can successfully complete the payment. As another example, if the target payment method is active scan payment, the desktop robot can project or display the merchant's payment code image. Simultaneously, the desktop robot's image acquisition module (such as a camera) and / or the NFC module can also be operational. If the user presents the payment code image shown on their terminal within the desktop robot's image acquisition range, the desktop robot can successfully acquire the payment code image and execute subsequent transaction processing. Alternatively, the user can complete the payment by interacting with the desktop robot via NFC. For example, if the target payment method is determined to be the scan-to-pay method, the desktop robot can project and display a prompt message for the user to present the QR code. It can also display the merchant's QR code image on the screen so that the user can also make a transaction by scanning the merchant's QR code. And / or, the near-field communication module can also be in working order so that the user can also make a payment by using the near-field payment method.

[0198] To ensure data security, information transmission between different terminals can be encrypted or the data can be anonymized before transmission.

[0199] Desktop robots can be used to interact with users, such as guiding them to show payment codes or providing NFC touch guidance. POS devices can then execute the payment process, allowing users to enjoy a smooth and natural payment experience. Furthermore, the communication connection between the robot and the POS device can adapt to different merchant deployment environments without requiring major modifications to the merchant's existing POS equipment, making it more practical.

[0200] Based on the same idea, one or more embodiments of this specification also provide a desktop robot capable of performing the methods in the foregoing embodiments.

[0201] Figure 3 This is a schematic diagram of the structure of a desktop robot provided in one embodiment of this specification. Figure 3 As shown, the desktop robot may include a head module 302, a body module 304, and a limb module 306; wherein the head module 302 and the limb module 306 are respectively connected to the body module 304; at least one of the head module 302, body module 304, and limb module 306 includes an information sensing unit; the information sensing unit is used to collect at least one type of perceptual information, including visual data and voice data. At least one of the head module 302, body module 304, and limb module 306 presents interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method; the interactive information includes at least one form of interactive information, including limb movements, light presentation, voice broadcast, screen display, and projection projection; the target payment method is a payment method determined from multiple payment methods supported by the desktop robot that matches the user profile characteristics, and the user profile characteristics are determined based on the perceptual data.

[0202] The head module can represent the upper structural module of a desktop robot and may include components such as a facial display, camera, and microphone. In one implementation, the head module may include a camera component 3022. This camera component can be used to collect visual data, or, if the target payment method is a payment method where the user pays by presenting a payment code, the camera can be used to capture an image of the payment code presented by the user's terminal.

[0203] Alternatively, the desktop robot can be equipped with multiple camera components. The camera component that collects visual data of the surrounding environment can be different from the camera component that collects the image of the payment code displayed by the user terminal. The acquisition range of one or more cameras on the desktop robot can cover the foot area of ​​the desktop robot, so that the desktop robot can detect whether it is close to the edge of the table during movement (such as performing dance steps or interactive games) to prevent the desktop robot from falling; or it can detect the movement path of the desktop robot, or plan the movement path to avoid obstacles or reach the target location.

[0204] A camera assembly can refer to a collection of optical devices used to acquire image or video data, and may include one or more cameras or auxiliary optical elements (such as lenses, filters, fill lights, etc.).

[0205] To better interact with the user, the head module 302 can optionally perform head movements. If the target payment method is a payment method where the user pays by showing a payment code, the head module will perform a head-up movement so that the camera component faces diagonally upwards to capture the payment code shown by the user.

[0206] The head module can be movable, achieving movements such as rotation, tilting, and nodding through actuators such as motors and servos. A head-tilting motion indicates that the head module rotates upwards, causing the camera assembly's shooting direction to change from horizontal or downward to obliquely upward. "Obliquely upward" can indicate that the angle between the camera's optical axis and the horizontal plane is acute, such as 15° or 30°.

[0207] The desktop robot tilts its head up so that the camera is angled slightly upwards, allowing for a more accurate view of the user's hand holding the phone. Users don't need to lower the phone or adjust their posture, reducing the burden on the user.

[0208] In practical applications, the head-raising angle can be dynamically adjusted based on the user's height. For example, face tracking can be used to make the camera face the direction of the user's face as much as possible.

[0209] In practical applications, the camera component may not be located in the head module, but rather in the body module.

[0210] The body module can represent the central structural module of a desktop robot, and may contain core components such as the main control processor, battery, and communication module. It is the main structure of the robot, and the head and limb modules are connected to it. Connection can represent the physical connection relationship between modules, and may include at least one of the following connection methods: mechanical connection (such as screws, clips), electrical connection (such as ribbon cables, connectors), etc.

[0211] In one implementation, the body module may include a projector 3042 for projecting projection information corresponding to the target payment method; the projection information includes at least one of the following: operation guidance information used to guide the user to use the target payment method before the payment process is executed, prompt information used to indicate that payment is in progress during the payment process, and prompt information indicating the payment result after the payment process is completed.

[0212] A projector can refer to an optical device that projects images or video content onto an external surface (such as a desktop or floor). Projection information can represent the visual content presented by the projector, including images, text, animations, QR codes, etc. For example, if the target payment method indicates that the user pays by scanning a QR code, the projection information may include the merchant's QR code image; or, if the target payment method indicates that the user pays by presenting a payment code, the projection information may include a prompt message asking the user to present the payment code; or, if the target payment method indicates that the user pays via contactless communication, the projection information may include a prompt message asking the user to touch a desktop robot to make the payment.

[0213] The projector can be angled downwards to project information onto the surface where the desktop robot is located. Alternatively, the projector can be angled horizontally or upwards to project information onto a projection surface near the desktop robot.

[0214] The projector can be located on the front of the body module, projecting information onto the front of the desktop robot and its feet. Alternatively, the projector can be located on the back or side of the body module, projecting information onto a preset projection surface. The projector can also be located on the head module.

[0215] A limb module can represent a movable structural module of a desktop robot, typically including components such as arms, hands, legs, and feet capable of performing limb movements, used to present limb movement interaction information. In one implementation, the limb module may include a movable arm component, which can be used to perform limb movements corresponding to the target payment method. For example, if the target payment method indicates that the user pays by scanning a QR code, the arm component can move its arm towards the area displaying the merchant's QR code image; or, if the target payment method indicates that the user pays via contactless communication, the arm component can move its arm towards the contactless communication area of ​​the desktop robot.

[0216] A movable arm assembly can represent a robotic arm structure capable of movement, which can achieve actions such as rotation, extension, and pointing through actuators such as motors and servos. The arm assembly can represent a mechanical structure that simulates a human arm, and may include segmented structures such as the upper arm, forearm, and hand, each segment capable of independent or coordinated movement; alternatively, it can be a non-segmented mechanism, such as including a single, integrated arm structure, or a single, integrated arm structure combined with a hand structure without fingers. The specific style of the arm assembly is not limited here.

[0217] In one implementation, the limb module may include a movable leg component, which may also perform limb movements corresponding to the target payment method.

[0218] The desktop robot in one or more embodiments of this specification can be a small robot that can be placed on a desktop. Its height can be less than or equal to 50 cm, for example, it can be 8 cm, 15 cm, 22 cm, 27 cm, 35 cm, etc. It can occupy a small area, so as to minimize the placement requirements and minimize the impact on existing devices on the desktop.

[0219] like Figure 3 As shown, the desktop robot can be located on base 308, which provides the robot's mobility, and the projection can be displayed on the base. Alternatively, the base can also have wireless charging functionality to charge the desktop robot. The desktop robot can also be detached from the base and placed directly on a desktop or other supporting surface, where the projection can be displayed.

[0220] Desktop robots can also provide non-payment interactive information such as interactive games and dance movements. The specific process for determining the target payment method and the presentation of interactive information can be found in the descriptions of the aforementioned embodiments, and will not be repeated here.

[0221] The above is an illustrative scheme of a desktop robot according to this embodiment. It should be noted that the technical solution of this desktop robot and the technical solution of the payment processing method based on the desktop robot described above belong to the same concept. For details not described in detail in the technical solution of this desktop robot, please refer to the description of the technical solution of the payment processing method based on the desktop robot described above.

[0222] Figure 4 This is a flowchart illustrating an interaction method for a desktop robot provided in one embodiment of this specification. Figure 4As shown, after the desktop robot is activated, it continuously senses environmental information through sensing components such as cameras and microphones. Once it detects interactive signals such as human voices or faces, it can construct a user profile using the collected visual and voice data. This allows it to intelligently recommend a target payment method that matches the user's characteristics from various payment options. Before payment, it provides operational guidance information through projectors, display screens, voice, and body movements. During payment, it provides emotional support through voice broadcasts, screen displays, projection displays, or body movements to reduce user anxiety while waiting. After payment, it can present differentiated results of successful or failed payments. After completion, it can return to a continuous sensing state, or maintain a continuous sensing state during the payment process. If no interaction is detected within a preset time, it can enter an idle mode. The desktop robot can then perform non-payment interactive activities such as dance moves or projection interactive games, maintaining a continuous sensing state during this process. If a user wants to pay, the non-interactive information presented by the desktop robot can be terminated promptly. Alternatively, returning to a continuous sensing state after dance moves or projection interactive games improves user experience and device utilization. For example, in near-field payment scenarios using short-range communication technologies such as NFC, during the pre-payment guidance phase, the desktop robot can project the text "Tap to pay," and the halo around the NFC area on its head can light up, flash, or display a blue breathing light effect. The robot can also announce, "Please take out your phone and tap the ring area on your head," and raise its arms vertically towards the NFC area. The screen can also display an animated NFC icon. During the payment process, after detecting the NFC signal or sending tag information to the user's terminal, the robot can announce, "Payment in progress, please wait," and display a marquee effect. Body movements can include nodding the head and swinging the arms back (simulating a waiting posture), and the screen can display an animated progress bar. In the post-payment feedback phase, if a successful payment is received, the robot can announce, "Payment successful, thank you for your patronage," and the light can remain solid green for 3 seconds. The arms can perform a high-five, the screen can display a smiling and blinking animation, and the projection can display brand logo animations, coupon QR codes, and other information. If payment fails, the system can announce messages such as "Payment failed, please try again" via voice, flash yellow lights for 3 seconds, and display gestures such as a slightly lowered head and an arm making a "please" gesture. The screen can also display a concerned animated expression, and the projector can show instructions on how to retry the operation.

[0223] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0224] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.

[0225] Figure 5 This is a schematic diagram of a desktop robot-based payment processing device provided in one embodiment of this specification.

[0226] like Figure 5 As shown, the device may include: The data acquisition module 502 is used to acquire perception data; the perception data includes at least one of visual data and voice data collected by the desktop robot.

[0227] The feature determination module 504 is used to determine user profile features based on the perceived data.

[0228] The payment method determination module 506 is used to determine the target payment method that matches the user profile characteristics from the multiple payment methods supported by the desktop robot.

[0229] The interactive information control module 508 is used to control the desktop robot to present interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method; the interactive information includes at least one form of interactive information such as body movements, light presentation, voice broadcast, screen display and projection.

[0230] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.

[0231] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0232] The above is an illustrative scheme of a desktop robot-based payment processing device according to this embodiment. It should be noted that the technical solution of this device belongs to the same concept as the technical solution of the desktop robot-based payment processing method described above. Details not described in detail in the technical solution of this device can be found in the description of the technical solution of the desktop robot-based payment processing method described above.

[0233] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.

[0234] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification.

[0235] The computing device 600 includes: Memory 610 and processor 620; The memory 610 is used to store computer programs / instructions, and the processor 620 is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor 620, they implement the steps of the above-described payment processing method based on desktop robots.

[0236] Specifically, the components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and the database 650 is used to store data.

[0237] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0238] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0239] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.

[0240] When the processor 620 executes the computer instructions, it implements the steps of the above-described desktop robot-based payment processing method.

[0241] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the payment processing method based on desktop robots described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the payment processing method based on desktop robots described above.

[0242] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the payment processing method based on a desktop robot as described above.

[0243] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the desktop robot-based payment processing method described above. Details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the desktop robot-based payment processing method described above.

[0244] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described desktop robot-based payment processing method.

[0245] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the payment processing method based on desktop robots described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the payment processing method based on desktop robots described above.

[0246] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the embodiments of desktop robots, devices, equipment, media, and products, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The desktop robots, devices, equipment, media, and products provided in the embodiments of this specification correspond to the methods, and therefore the desktop robots, devices, equipment, media, and products also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding desktop robots, devices, equipment, media, and products will not be repeated here.

[0247] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0248] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0249] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0250] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0251] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0252] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, the invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0253] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0254] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0255] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0256] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0257] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0258] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0259] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0260] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A payment processing method based on a desktop robot, comprising: Acquire sensory data; the sensory data includes at least one of visual data and voice data collected by the desktop robot; Based on the perceived data, user profile features are determined; From the various payment methods supported by the desktop robot, determine the target payment method that matches the user profile characteristics; The desktop robot is controlled to present interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method; the interactive information includes at least one form of interactive information such as body movements, light display, voice broadcast, screen display and projection.

2. The method according to claim 1, further comprising: Based on the perceived data, the user's target direction is determined; Control the desktop robot to move towards the target direction.

3. The method according to claim 2, wherein determining the target direction of the user based on the perceived data includes: Using the microphone array in the desktop robot, the azimuth angle of the user is calculated through a sound source localization algorithm; Based on the azimuth angle, the desktop robot's camera is controlled to point towards the user area for face detection; Based on the face center coordinates obtained from face detection, the target direction of the user is determined.

4. The method according to claim 1, wherein the interactive information includes at least one of the following: operation guidance information for guiding the user to make payment using the target payment method before the payment process is executed, prompt information for indicating that payment is in progress during the payment process, and prompt information indicating the payment result after the payment process is completed.

5. The method according to claim 1, wherein the user profile features include at least one of user person profile features, user device profile features, and user environment features; the user person profile features include at least one of age features, gender features, emotional features, voice features, behavioral posture features, and historical transaction features; the user device profile features include at least one of device type features and device function features; and the user environment features include at least one of queue length features and ambient light intensity features.

6. The method according to claim 1, wherein determining a target payment method that matches the user profile characteristics from among multiple payment methods supported by the desktop robot, includes: Based on the user profile features, a target payment method that corresponds to the user profile features is found from a preset payment recommendation rule table; The preset payment recommendation rule table records various payment methods and the profile characteristics of users who use various payment methods.

7. The method according to claim 1, wherein the payment method includes at least one of the following: payment via contactless communication, payment by presenting a payment code, and payment by scanning a payment code.

8. The method according to claim 7, wherein controlling the desktop robot to present interactive information corresponding to the target payment method includes: If the target payment method is a method in which the user pays by scanning a QR code, then the desktop robot is controlled to display the merchant's QR code image on the screen or by projection. Alternatively, if the target payment method is a method in which the user makes payment by presenting a payment code, then the desktop robot is controlled to display a prompt message prompting the user to present the payment code via screen display or projection. Alternatively, if the target payment method is a payment method that indicates contactless communication, then the desktop robot is controlled to display a prompt message on the screen or projected to encourage the user to touch the desktop robot to make the payment.

9. The method according to claim 7, wherein controlling the desktop robot to present interactive information corresponding to the target payment method includes: If the target payment method is a method in which the user pays by scanning a QR code, then control the arm of the desktop robot to point to the area displaying the merchant's QR code image; Alternatively, if the target payment method is a payment method where the user pays by presenting a payment code, then the desktop robot is controlled to perform a head-raising action, and the payment code presented by the user is captured by the image acquisition module located on the head. Alternatively, if the target payment method is a payment method that indicates payment via contactless communication, then the arm of the desktop robot is controlled to point towards the contactless communication area of ​​the desktop robot.

10. The method according to claim 1, further comprising: If no user interacting with the desktop robot is detected within the first preset time period, the desktop robot will display non-payment interaction information.

11. The method according to claim 10, wherein the desktop robot presents non-payment interaction information, including: If the number of times a face is detected within the visual range of the desktop robot is greater than or equal to a preset number within a second preset time period following the first preset time period, then the desktop robot performs a preset dance move; the second preset time period is less than the first preset time period.

12. The method according to claim 11, wherein the desktop robot presents non-payment interaction information, including: If the number of times a face is detected within the visual range of the desktop robot within the second preset time period is less than the preset number, then the desktop robot will launch an interactive game. If a user is detected participating in the interactive game, the desktop robot will perform a preset physical action.

13. The method according to claim 10, further comprising: Obtain ambient lighting information; The desktop robot presents non-payment interactive information, including: If the light intensity of the ambient light information is greater than or equal to the preset intensity, the desktop robot will display non-payment interaction information. If the ambient light intensity is less than a preset intensity, the desktop robot enters a sleep state.

14. The method according to any one of claims 1 to 13, wherein the desktop robot and the merchant's POS device are connected via wired or wireless means, and the method further comprises: If the desktop robot collects the payment code image presented by the user, it sends the payment code image or the payment code number obtained by parsing the payment code image to the merchant's POS device so that the merchant's POS device can execute the payment collection process. Alternatively, if the desktop robot sends tag information to the user terminal to trigger the payment process via contactless communication, the desktop robot will send payment identification information to the merchant's POS device to determine the payment account of the payer user on the user terminal, so that the merchant's POS device can execute the payment process.

15. A desktop robot, comprising: Head module, body module, and limb module; The head module and the limb module are respectively connected to the body module; At least one of the head module, the body module, and the limb module includes an information sensing unit; the information sensing unit is used to collect at least one type of sensing information, namely visual data and voice data. At least one of the head module, the body module, and the limb module is used to present interactive information corresponding to the target payment method, so that the user can interact with the desktop robot to make payment using the target payment method; the interactive information includes at least one form of interactive information such as body movements, light presentation, voice broadcast, screen display, and projection; the target payment method is a payment method that matches the user profile characteristics from a variety of payment methods supported by the desktop robot, and the user profile characteristics are determined based on perception data.

16. The desktop robot according to claim 15, wherein the head module includes a camera component; the camera component is used to collect the visual data, or, if the target payment method is a method indicating that the user makes payment by presenting a payment code, the camera is used to collect an image of the payment code presented by the user terminal.

17. The desktop robot according to claim 16, wherein the head module is capable of performing head movements; if the target payment method is a method indicating that the user makes payment by presenting a payment code, the head module performs a head-raising movement so that the camera assembly faces diagonally upward in order to capture the payment code presented by the user.

18. The desktop robot according to claim 15, wherein the body module includes a projector for projecting projection information corresponding to the target payment method; the projection information includes at least one of the following: operation guidance information for guiding the user to use the target payment method to make payment before the payment process is executed, prompt information for indicating that payment is in progress during the payment process, and prompt information indicating the payment result after the payment process is completed.

19. The desktop robot of claim 15, wherein the limb module includes a movable arm assembly for performing limb movements corresponding to the target payment method.

20. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 14.